Data Partitioning System for Live Dataset Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Transferring large data sets between computing systems is challenging due to resource limitations, making manual partitioning difficult, especially for live or actively modified datasets, which can result in uneven system performance, backups, slowdowns, and crashes.
Innovation Solution
A data partitioning and transferring system (DPS) automatically determines optimal partition sizes and numbers based on available resources and dataset characteristics, using unique columns and profiling results to ensure equal or roughly equal partition sizes, thereby maximizing resource utilization and preventing system interruptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If manual partitioning is attempted for large datasets, then data can be divided into smaller units, but the process becomes difficult or impossible to accomplish successfully
Solution Approach 1:
The system enables self-service automatic partitioning where the database management system autonomously analyzes dataset characteristics, determines optimal partitioning strategies, and executes the partitioning without administrator intervention. This resolves the contradiction by making the system self-sufficient for handling large data partitioning tasks.
Solution Approach 2:
The patent replaces manual mechanical partitioning operations with automated computational algorithms that analyze data patterns, calculate optimal partition sizes, and execute divisions. This substitution eliminates the difficulty of manual partitioning while effectively managing large datasets.
2Ease of operation
If data is transferred as a single block, then transfer simplicity is maintained, but resource limitations prevent successful transfer
Solution Approach 1:
The system automatically segments large datasets into smaller partitions based on resource constraints and data characteristics. This segmentation enables transfer operations to proceed successfully by breaking down the overwhelming single-block transfer into manageable units while maintaining operational simplicity through automation.
Solution Approach 2:
The partitioning strategy is dynamically adjusted based on available resources, data patterns, and transfer requirements. The system adapts partition sizes and numbers automatically, providing both simplicity and resource efficiency without requiring manual configuration.
3Quantity of substance
If manual partitioning is performed on live datasets, then data can be divided, but system performance degrades with backups and slowdowns
Solution Approach 1:
The system performs preliminary analysis of live datasets to determine optimal partitioning strategies before execution. It identifies data patterns, hotspots, and access patterns in advance, allowing partitioning to be performed with minimal disruption to ongoing operations and maintaining system reliability.
Solution Approach 2:
The system incorporates feedback mechanisms that monitor system performance during partitioning operations on live datasets. It adjusts partitioning strategies in real-time based on observed performance metrics, preventing degradation and ensuring reliable operation throughout the partitioning process.
4Ease of operation
If uneven partitions are created, then manual partitioning is simpler, but system performance becomes inconsistent
Solution Approach 1:
The system automatically adjusts partition parameters including size, number, and distribution based on data characteristics and system resources. This dynamic parameter optimization ensures uniform partition sizes that provide consistent system performance while eliminating the need for manual intervention.
Data Source
AI summary
Disclosed herein are system, method, and computer program product embodiments for providing data partitioning and transferring operations. An embodiment operates by determining a partition size and a number of partitions for an initial data set to be transferred from a first location to a second location. A uniqueness factor for at least a subset of the columns of the dataset is determined, and a set of unique columns is identified from the initial data set based on the uniqueness factor. Based on the partition size, a set of values from the row records from the set of unique columns is identified. Based on the identified set of values, the initial data set is partitioned into the number of partitions. One of transmitting or receiving at least one of the partitions is performed.


