ML-Based Distributed Data Archival for Efficient Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data archiving processes in distributed data networks require transferring entire databases repeatedly, leading to inefficient use of resources, time, and potential replication errors, as they do not differentiate between changed and unchanged data.
Innovation Solution
A system utilizing machine learning algorithms to divide data into blocks, assign unique identifiers, characterize changes, and categorize them for targeted transfer via a distributed ledger, ensuring only changed data is replicated, thereby optimizing resource usage and reducing replication errors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If entire databases are transferred for archiving, then data completeness is ensured, but resource consumption and time increase significantly
Solution Approach 1:
The patent segments data into individual records or subsets that can be selectively transferred. Instead of moving entire databases, the system identifies and transfers only the changed or affected records, reducing resource consumption while maintaining data completeness for the archived portions.
Solution Approach 2:
The patent extracts only the necessary data elements for archiving by identifying changed records through comparison mechanisms. This extraction approach allows the system to transfer only essential data while leaving unchanged data in the original database, thereby reducing resource consumption.
2Reliability
If entire databases are transferred for archiving, then data completeness is ensured, but transfer time increases significantly
Solution Approach 1:
The patent segments data into individual records or subsets that can be selectively transferred. Instead of moving entire databases, the system identifies and transfers only the changed or affected records, reducing transfer time while maintaining data completeness for the archived portions.
Solution Approach 2:
The patent performs preliminary comparison actions to identify changed records before the actual transfer. By pre-comparing data versions and identifying only the differences, the system prepares a targeted list of records to transfer, significantly reducing transfer time while ensuring data completeness.
3Reliability
If recurring archiving is performed, then data backup is maintained, but redundant data transfer occurs
Solution Approach 1:
The patent extracts only the necessary data elements for archiving by identifying changed records through comparison mechanisms. This extraction approach allows the system to transfer only essential data while leaving unchanged data in the original database, thereby reducing redundant data transfer.
Solution Approach 2:
The patent implements feedback mechanisms that track which records have been archived and their current states. This feedback information is used to determine what data needs to be transferred in subsequent archiving operations, preventing redundant transfers of unchanged data while maintaining backup reliability.
4Measurement precision
If machine learning algorithms are used to identify changes, then transfer precision is improved, but system complexity increases
Solution Approach 1:
The patent uses machine learning models that are copied or replicated across different system components to maintain consistency in change detection. By distributing the ML models rather than requiring a single complex centralized system, the patent reduces overall system complexity while maintaining high change detection precision.
Data Source
AI summary
Embodiments of the invention are directed to a system, method, or computer program product for an approach to electronic data archival in a distributed data network. The system allows for replicating and transmitting data for archival purposes from a source to a destination using a machine learning algorithm. The machine learning algorithm selectively distributes only data which has changed within a database, and thus prevents the necessity of repetitively archiving an entire database. The system categorizes changed data based on the characteristics of the data, and thereafter distributes the data via a distributed data network.


