Elastic Data Mover Utility Selective Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems for moving data across multiple server clusters are inefficient, time-consuming, and require significant computing resources, often leading to errors and suboptimal utilization of processing and storage resources.
Innovation Solution
A method and system that utilize an elastic data mover utility (EDMU) to selectively move data between clusters, allowing for efficient data transfer by generating jobs based on user input and authenticating users, thereby optimizing resource usage and reducing the need for full data replication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data is moved in its entirety from one server cluster to another server cluster, then complete data transfer is achieved, but computing resources and time are significantly consumed
Solution Approach 1:
The patent extracts and moves only the necessary portions of data rather than transferring entire datasets. The system identifies specific data elements that need to be moved between clusters and transfers only those, leaving the rest behind. This selective extraction approach maintains data transfer completeness for required information while dramatically reducing the volume of data moved, thereby improving transfer efficiency and reducing resource consumption.
Solution Approach 2:
The patent segments data into identifiable units or elements that can be selectively moved. By breaking down large datasets into discrete data elements with specific identifiers, the system can choose which segments to transfer based on requirements. This segmentation enables partial data movement, resolving the contradiction between transferring complete necessary data and minimizing overall transfer volume.
2Reliability
If existing systems move data between clusters, then data transfer is achieved, but the systems are time-consuming and require significant computing resources
Solution Approach 1:
The system extracts only the specific data elements that need to be moved rather than transferring entire datasets. By identifying and moving only necessary data portions, the transfer time is significantly reduced while still achieving the required data transfer capability. This approach maintains reliability for the needed data while eliminating unnecessary transfer time for irrelevant data.
3Reliability
If full data replication is performed between clusters, then data availability is ensured, but resource utilization becomes suboptimal
Solution Approach 1:
The patent extracts and moves only specific data elements that are necessary for operational requirements rather than replicating entire datasets. This selective approach ensures data availability for required information while avoiding the energy waste of copying and storing unnecessary data in target clusters. Resources are optimized by eliminating redundant data replication.
Solution Approach 2:
The system segments data into discrete elements and selectively replicates only those segments that are necessary. This segmentation enables the system to maintain data availability for critical information while avoiding the energy consumption associated with full dataset replication. The selective segment-based approach optimizes resource utilization by matching replication activity to actual needs.
Data Source
AI summary
A method, apparatus, computer-readable medium, and/or system described herein may be used to efficiently store, move, and/or process data across a plurality of computing clusters. For example, a computing device may receive an indication of one or more data storage locations within a first cluster of servers and/or an indication of one or more data storage locations within a second cluster of servers. The computing device may generate a data file comprising the indication of the one or more data storage locations within the first cluster of servers and/or the indication of one or more data storage locations within the second cluster of servers. Based on the generated data file, the computing device may generate a job to move data stored at the one or more data storage locations within the first cluster of servers to the one or more data storage locations within the second cluster of servers. Based on the job, the computing device may transmit, e.g., to the first cluster of servers and/or the second cluster of servers, instructions to move data stored at the one or more data storage locations within the first cluster of servers to the one or more data storage locations within the second cluster of servers.


