Dynamic Boundary Determination for Parallel Sort Runs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current database systems face inefficiencies in parallel processing of large data sets due to fixed division points in disk sort techniques, which fail to adapt to data distribution and resource availability, limiting scalability and parallelism.
Innovation Solution
The method involves sorting records based on key values, gathering metadata to determine dynamic boundary values for disjoint subsets, and outputting sorted data in parallel, allowing for adaptive division and efficient resource utilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If fixed division points are used in disk sort techniques, then the system structure is simple and easy to implement, but the parallel processing efficiency is limited and cannot adapt to data distribution
Solution Approach 1:
The patent implements dynamic division points that adapt to actual data distribution and resource availability. The system determines optimal division points based on metadata about data characteristics and current system state, allowing the parallel processing configuration to change dynamically rather than being fixed, thereby improving parallel efficiency without requiring overly complex static structures
Solution Approach 2:
The system changes the parameter of division points from fixed values to dynamically determined values based on data distribution analysis. By gathering metadata and computing optimal boundary values, the system adjusts the division parameters to match actual data characteristics, resolving the contradiction between simple implementation and processing efficiency
2Adaptability or versatility
If division points are specified at query optimization time, then the decision-making process is simple, but the system cannot correct for uneven distribution of rows among disjoint sets
Solution Approach 1:
The system implements feedback by gathering metadata about actual data distribution and using this information to determine optimal division points. This feedback loop allows the system to adapt to uneven data distribution by adjusting division points based on observed data characteristics rather than relying on static pre-determined values
Solution Approach 2:
The system performs preliminary gathering of metadata about data distribution before determining division points. This preliminary action provides the necessary information about data characteristics, enabling the system to make informed decisions about optimal division points that account for actual data distribution patterns
3Adaptability or versatility
If key-based binning is used to separate records into bins, then the implementation is straightforward, but the system suffers from inability to adapt to circumstances including contents of sort and processing resources
Solution Approach 1:
The patent replaces static key-based binning with dynamic division points that adapt to both data contents and processing resources. The system determines division points based on metadata about data distribution and current system state, allowing the binning strategy to be dynamic rather than fixed, thereby improving adaptability without requiring overly complex resource management mechanisms
Data Source
AI summary
A system, method, and computer program product are provided for sorting a set of records in a sort run. As the records are sorted, metadata regarding the sort run is gathered, and subsequently used to determine bounds of two or more disjoint subsets of the sort run. This enables the parallelization of several tasks over the sort run data using efficient, dynamic bounds determination, such as the outputting of sorted data from the disjoint subsets in parallel.


