Adaptive Cloud Data Transfer via File Segmentation and Parallel Paths
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for transferring large data sets face challenges in achieving optimal speed and cost efficiency, particularly as data sizes exceed 100 GBytes, due to fixed bandwidth limitations in terrestrial and cloud networks.
Innovation Solution
A system and method that dynamically select folders, determine file types and sizes, and adjust transfer parameters, including parallel transport instances and bit-exact compression, to optimize data transfer speed and cost, utilizing machine learning models to adapt and improve transfer efficiency over time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If fixed bandwidth network paths are used for data transfer, then network infrastructure is simple and cost-effective, but transfer speed is limited and cannot meet large data set requirements
Solution Approach 1:
The patent segments large data sets into smaller files and uses multiple parallel transport instances to transfer them simultaneously over heterogeneous network paths. This segmentation allows the system to overcome fixed bandwidth limitations by distributing data across multiple channels, thereby increasing overall transfer speed without requiring a single high-capacity network connection.
Solution Approach 2:
The system dynamically adjusts the number of parallel transport instances based on measured network conditions and transfer performance. By continuously monitoring transfer speeds and adapting the level of parallelization, the system optimizes data transfer performance in real-time without requiring complex pre-configured network infrastructure.
2Productivity
If multiple parallel transport instances are used to increase transfer speed, then data transfer speed improves, but system complexity and resource consumption increase
Solution Approach 1:
The system automatically measures transfer speeds, determines optimal file handling strategies, and adjusts parallel transport instance configuration without external intervention. This self-service capability allows the system to manage its own complexity by adapting to network conditions and file characteristics, maintaining high productivity while minimizing the need for external system management.
Solution Approach 2:
The system implements continuous feedback loops where transfer performance is measured and used to adjust subsequent transfer parameters. By monitoring actual transfer speeds and comparing them against targets, the system dynamically optimizes the number of parallel instances and file handling approaches, ensuring high throughput while adapting to changing conditions.
3Speed
If large files are transferred as single units, then transfer overhead is minimized, but transfer speed is limited by single path bandwidth
Solution Approach 1:
Large files are automatically segmented into smaller parts that can be transferred in parallel over multiple network paths. This segmentation enables the system to overcome single-path bandwidth limitations while the automated process minimizes the time lost to file preparation by using efficient file splitting algorithms and concurrent transfer initiation.
4Productivity
If small files are transferred individually, then each file can be optimized separately, but total transfer time increases due to overhead
Solution Approach 1:
Multiple small files are merged into aggregated transfer units that can be moved efficiently across the network. This merging reduces the overhead associated with individual file transfers while maintaining the ability to optimize transfer parameters. The system balances file aggregation with parallel transport to maximize throughput without excessive preparation time.
Data Source
AI summary
An system and method of adaptively transporting large data assets to meet desired transfer speed and cost, includes selecting folders containing files having the data, receiving input of the desired transfer speed and cost, selecting a destination memory device, reading a file header of the files for the selected folders to obtain header information, determining a file type and file size based on the header information, determining a large file size threshold based on the desired transfer speed and cost and previously measured transfer speed. When the file size of a file is above the large file size threshold, dividing the file into two or more file parts in accordance with the file type, transferring the one or more files, including any of the file parts, using a number of parallel transport machine instances, measuring a total speed of the transfer, and storing the measured total speed for use as the previously measured transfer speed.


