Skip to main content

Command Palette

Search for a command to run...

1. Major data objects

Updated
•5 min read•View as Markdown
1. Major data objects

What is Schema ?

A schema is the blueprint of a database that defines its structure and organization, detailing tables, fields, relationships, indexes, and constraints. It serves as a framework, showing how data is stored and managed, and enables users and applications to interact with it effectively. The schema includes tables, which organize data into rows and columns, with each table representing an entity, such as "Customers" or "Orders." Fields, or columns, specify the type of data (like text, date, or integer) in each column. Relationships illustrate connections between tables, such as one-to-one, one-to-many, or many-to-many, while constraints ensure data integrity by enforcing rules like unique identifiers (PRIMARY KEY), references to other tables (FOREIGN KEY), and mandatory values (NOT NULL). Indexes improve database performance by speeding up data retrieval. Schemas come in several types: logical (abstract structure), physical (actual storage), and specialized types like star or snowflake schemas for data warehouses, which organize data in central fact tables connected to dimension tables. For example, in an e-commerce database, a schema could connect a "Customer" table containing fields like "CustomerID" and "Name" to an "Order" table with fields like "OrderID" and "OrderDate," linked by a foreign key relationship. By defining these elements, schemas provide a structured roadmap for efficient data organization and access.

What is Table ?

In a database, a table is a structured set of data organized into rows and columns, where each row represents a single record, and each column represents a specific attribute of that record. Tables store data about a particular entity, such as "Customers," "Products," or "Orders," and are a fundamental building block of relational databases.

Each column in a table has a specific data type, such as text, integer, or date, which defines the kind of information it can hold, like a customer’s name, email, or address. Rows, often called records, contain individual entries for that entity, with each value corresponding to one of the table’s columns. Tables also frequently include constraints to ensure data accuracy, like PRIMARY KEY, which uniquely identifies each record, or FOREIGN KEY, which links records between tables. By organizing data into tables, relational databases make it easy to store, retrieve, and manage structured information efficiently.

What are the major data storage methods ?

Data storage methods vary widely based on the type of data, the storage purpose, and the access requirements. Here are some common data storage methods:

  1. File Storage: Data is stored in files, usually on a file system, which can be local (like on hard drives) or cloud-based (such as AWS S3 or Google Cloud Storage). It’s commonly used for unstructured data like documents, images, and videos, and enables easy retrieval and sharing.

  2. Database Storage:

    • Relational Databases: Store structured data in tables with rows and columns, using a schema to define data relationships. Examples include MySQL, PostgreSQL, and SQL Server.

    • NoSQL Databases: Store unstructured or semi-structured data without rigid schemas, using formats like key-value pairs, documents, or wide-column stores. Examples include MongoDB, Cassandra, and DynamoDB.

    • In-Memory Databases: Store data in RAM for faster access, suitable for real-time applications. Examples are Redis and Memcached.

  3. Data Warehouses: Centralized storage solutions for large volumes of historical data from various sources, optimized for analytical queries. Data warehouses are used for business intelligence and analytics, with examples like Amazon Redshift, Google BigQuery, and Snowflake.

  4. Data Lakes: Store raw, unstructured, or semi-structured data in a central repository, typically on cloud storage. Data lakes support large-scale storage and analysis, enabling flexibility to process diverse data types. Common platforms include Azure Data Lake, AWS S3, and Hadoop-based systems.

  5. Block Storage: Data is stored in fixed-sized chunks or blocks, each with a unique address, making it ideal for databases and applications requiring high performance and flexibility. Common in cloud environments, with examples like Amazon EBS and Google Persistent Disk.

  6. Object Storage: Organizes data into objects (files with metadata) in a flat address space, ideal for unstructured data, and widely used in cloud environments. Examples include AWS S3, Azure Blob Storage, and Google Cloud Storage.

  7. Tape Storage: A magnetic storage method used for long-term archival and backup purposes, with a low cost per byte but slower retrieval speed. Often used in compliance and data retention requirements for rarely accessed data.

  8. Hybrid Storage Solutions: Combines on-premise and cloud storage, offering flexibility and scalability while meeting compliance requirements. These systems allow companies to balance the speed of local storage with the scalability of the cloud.

These methods help businesses manage data according to requirements like scalability, speed, cost, and flexibility. Choosing the right storage method depends on factors such as the data structure, access frequency, and performance needs.

Data processing methods

Data processing methods vary in terms of speed, use case, and data volume. Here’s an overview of the main types—batch, real-time, and near real-time:

  1. Batch Processing
  • Definition: Batch processing involves collecting and processing data in large groups (batches) at specified intervals.

  • When It's Used: Suitable for handling high volumes of data where immediate processing isn’t critical, like end-of-day reports, payroll, and billing.

  • Examples: ETL (Extract, Transform, Load) jobs in data warehouses, daily transaction processing, and report generation.

  • Advantages: Efficient for large datasets, cost-effective, and minimizes resource demand since it processes data in bulk.

  • Challenges: Not suitable for scenarios requiring immediate data insights, as processing happens at scheduled intervals.

  1. Real-time Processing
  • Definition: Real-time processing involves processing data instantly as it is generated or received, with virtually no delay.

  • When It's Used: Critical for applications where immediate response is essential, like financial trading systems, fraud detection, online gaming, and IoT sensors.

  • Examples: Real-time payment processing, stock market trading platforms, and autonomous vehicle systems.

  • Advantages: Provides instant feedback and up-to-the-moment insights, enabling immediate action based on data.

  • Challenges: High computational and storage requirements, making it resource-intensive and costly. Requires robust systems to handle peak loads continuously.

  1. Near real-time Processing
  • Definition: Near real-time processing processes data almost as soon as it’s received, but with a slight delay, usually seconds to minutes.

  • When It's Used: Suitable for applications that need timely data but can tolerate slight latency, like customer recommendations, social media monitoring, and fraud alerts.

  • Examples: Social media feed updates, personalized e-commerce recommendations, and monitoring systems in healthcare.

  • Advantages: Balances timely data insights with fewer resources than real-time systems, suitable for semi-urgent insights and decision-making.

  • Challenges: May not be suitable for applications needing truly instantaneous processing, as latency is still present, albeit minimal.

Each processing method fits specific requirements based on factors like data volume, response time, and cost considerations.