Modular Blocks Minimize Network I/O in Distributed Database
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing distributed databases face challenges in scaling to hundreds of servers, managing frequent server additions and removals, handling network and server failures, and minimizing network I/O during large table joins, especially when using commodity hardware and cloud infrastructure.
Innovation Solution
The system employs modular blocks of 5G bytes or less, each with an associated log file, managed by a master node that distributes and replicates data across worker nodes, allowing for efficient data transfer, independent block operations, and flexible replication strategies to handle node failures and additions without impacting performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is distributed across hundreds of servers in a distributed database, then scalability and availability are improved, but network I/O increases and system complexity increases
Solution Approach 1:
The distributed database is segmented into modular blocks of 5G bytes or less, where each block is a self-contained unit with associated metadata and log files. This segmentation allows independent management, transfer, and replication of small data units across worker nodes, reducing the network I/O required for operations compared to moving larger data partitions.
2Loss of energy
If modular blocks of 5G bytes or less are used with associated log files, then data transfer efficiency is improved and network I/O is minimized, but device complexity increases
Solution Approach 1:
Each modular block is merged with its associated log file and metadata to create a self-contained data unit. This combining eliminates the need for separate management of data and its accompanying information, simplifying operations like transfer and replication despite the fine-grained modular structure.
3Adaptability or versatility
If frequent server additions and removals are handled in a distributed database, then adaptability is improved, but query execution time increases and performance degrades
Solution Approach 1:
The system performs preliminary actions by maintaining ready-to-transfer modular blocks and log files on worker nodes before server additions or removals occur. When nodes are added or removed, pre-prepared data blocks can be immediately assigned or transferred without requiring complex real-time data shuffling, thus maintaining query performance during dynamic changes.
4Reliability
If data is replicated across multiple worker nodes for high availability, then reliability is improved, but network I/O and storage requirements increase
Solution Approach 1:
Data replication is performed at the modular block level rather than at the partition level. Each worker node stores copies of specific modular blocks, allowing selective replication of only the necessary small data units. This segmented approach reduces the total network I/O and storage requirements compared to replicating entire data partitions across multiple nodes.
Data Source
AI summary
A method implemented by a computer includes receiving a segment of data that has a time dimension, where the time dimension of the segment of data is bounded by a start time stamp and an end time stamp. The segment of data is added to an append-only database table of a distributed database. The addition operation imposes an inherent data order based upon the start time stamp and end time stamp without the manual definition off database table partition in the distributed database.


