Token-Based Data Segment Representation for Backup Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data storage systems face inefficiencies in resource utilization and scalability during backup and recovery operations, particularly in managing and restoring data across multiple hosts and storage devices, with existing techniques failing to effectively reduce backup and recovery time while ensuring data integrity.
Innovation Solution
A method involving the use of tokens, such as hash values, and a unique identifier to represent data segments, allowing for efficient data synchronization and restoration by comparing current and previous data states, and utilizing digital signatures for verification and data operations, enabling decentralized verification and restoration without server interaction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional backup and recovery methods are used, then data can be stored and retrieved, but resource utilization is inefficient and backup/recovery time is excessive
Solution Approach 1:
The patent segments data into data segments and represents them using tokens (hash values) and unique identifiers. Instead of backing up entire datasets, the system segments data and stores only unique segments with their corresponding tokens, significantly reducing backup time and improving recovery speed by enabling parallel processing and selective restoration of individual segments.
Solution Approach 2:
The patent creates simplified representations (tokens and identifiers) of data segments instead of storing complete copies. These tokens serve as efficient proxies that can be rapidly transmitted and verified, reducing the time required for backup operations while maintaining the ability to reconstruct original data when needed.
2Adaptability or versatility
If data is stored across multiple hosts and storage devices, then data availability is improved, but resource utilization becomes inefficient and system scalability is limited
Solution Approach 1:
The patent creates a universal token-based representation system that can be applied across multiple hosts and storage devices uniformly. The tokens and unique identifiers serve as universal keys that work across different storage locations, enabling efficient resource utilization and easy scalability without requiring host-specific or device-specific backup mechanisms.
Solution Approach 2:
By storing lightweight token representations instead of complete data copies across multiple hosts, the system achieves scalable distribution with minimal resource overhead. Each host can verify data integrity using tokens without requiring full data replication, improving resource utilization while maintaining system scalability.
3Reliability
If data integrity verification is performed during backup operations, then data reliability is ensured, but additional processing time and computational resources are consumed
Solution Approach 1:
The patent extracts the essential verification information into separate tokens (hash values) and unique identifiers that are stored alongside the data segments. This separation allows integrity verification to be performed by simply comparing tokens without requiring complex processing of the actual data, ensuring data reliability while minimizing additional computational complexity.
4Reliability
If complete data sets are restored during recovery operations, then full data availability is achieved, but restoration time is excessive and network resources are overwhelmed
Solution Approach 1:
The patent enables selective restoration by segmenting data into individual data segments, each identified by unique tokens and identifiers. During recovery operations, only the specific segments that need to be restored can be retrieved using their tokens, rather than restoring complete datasets. This significantly reduces restoration time and network resource consumption while maintaining data completeness for the required segments.
Data Source
AI summary
Described are techniques for representing a data segment comprising. A list of one or more tokens representing one or more data portions included in the data segment is received. A unique identifier uniquely identifying said data segment from other data segments is received. A signature value determined in accordance with said list of tokens and said unique identifier is received. The list of tokens, said unique identifier, and said signature value are stored as information corresponding to said data segment.


