Token-Based Data Segment Representation for Backup Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data storage systems face inefficiencies in resource utilization and scalability during backup and recovery operations, particularly in managing and restoring data across multiple hosts and storage devices, with existing techniques failing to effectively reduce backup and recovery time while ensuring data integrity.

Innovation Solution

A method involving the use of tokens, such as hash values, and a unique identifier to represent data segments, allowing for efficient data synchronization and restoration by comparing current and previous data states, and utilizing digital signatures for verification and data operations, enabling decentralized verification and restoration without server interaction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional backup and recovery methods are used, then data can be stored and retrieved, but resource utilization is inefficient and backup/recovery time is excessive

Engineering Contradiction:
Improvebackup and recovery speedVSAvoidbackup and recovery time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent segments data into data segments and represents them using tokens (hash values) and unique identifiers. Instead of backing up entire datasets, the system segments data and stores only unique segments with their corresponding tokens, significantly reducing backup time and improving recovery speed by enabling parallel processing and selective restoration of individual segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates simplified representations (tokens and identifiers) of data segments instead of storing complete copies. These tokens serve as efficient proxies that can be rapidly transmitted and verified, reducing the time required for backup operations while maintaining the ability to reconstruct original data when needed.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If data is stored across multiple hosts and storage devices, then data availability is improved, but resource utilization becomes inefficient and system scalability is limited

Engineering Contradiction:
Improvesystem scalabilityVSAvoidresource utilization efficiency
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent creates a universal token-based representation system that can be applied across multiple hosts and storage devices uniformly. The tokens and unique identifiers serve as universal keys that work across different storage locations, enabling efficient resource utilization and easy scalability without requiring host-specific or device-specific backup mechanisms.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

By storing lightweight token representations instead of complete data copies across multiple hosts, the system achieves scalable distribution with minimal resource overhead. Each host can verify data integrity using tokens without requiring full data replication, improving resource utilization while maintaining system scalability.

Inventive Principle:
Principle #26Copying

3Reliability

If data integrity verification is performed during backup operations, then data reliability is ensured, but additional processing time and computational resources are consumed

Engineering Contradiction:
Improvedata integrityVSAvoidverification processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential verification information into separate tokens (hash values) and unique identifiers that are stored alongside the data segments. This separation allows integrity verification to be performed by simply comparing tokens without requiring complex processing of the actual data, ensuring data reliability while minimizing additional computational complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

4Reliability

If complete data sets are restored during recovery operations, then full data availability is achieved, but restoration time is excessive and network resources are overwhelmed

Engineering Contradiction:
Improvedata completenessVSAvoidrestoration time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent enables selective restoration by segmenting data into individual data segments, each identified by unique tokens and identifiers. During recovery operations, only the specific segments that need to be restored can be retrieved using their tokens, rather than restoring complete datasets. This significantly reduces restoration time and network resource consumption while maintaining data completeness for the required segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8082231B1Techniques using identifiers and signatures with data operations
Publication Date: 2011.12.20 EMC IP HLDG CO LLC
  • US8082231B1 patent drawing
  • US8082231B1 patent drawing
  • US8082231B1 patent drawing

AI summary

Described are techniques for representing a data segment comprising. A list of one or more tokens representing one or more data portions included in the data segment is received. A unique identifier uniquely identifying said data segment from other data segments is received. A signature value determined in accordance with said list of tokens and said unique identifier is received. The list of tokens, said unique identifier, and said signature value are stored as information corresponding to said data segment.