Deduplicated Data Copies Using Temporal Hash Structures

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems require multiple point solutions for managing the lifecycle of application data, leading to complex and expensive infrastructure with redundant data copies and inefficient data movement across storage repositories.

Innovation Solution

The Data Management Virtualization System leverages a unified engine to manage data protection across various storage repositories by tracking changes over time, using deduplication and compression algorithms, and abstracting physical storage resources into virtualized pools, enabling efficient data movement and reduction of redundant copies.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple point solutions are deployed to manage data lifecycle, then data protection coverage is improved, but infrastructure complexity and cost increase

Engineering Contradiction:
Improvedata protection coverageVSAvoidinfrastructure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate data management functions (backup, replication, archiving, disaster recovery) into a single unified data management system. This consolidation maintains comprehensive data protection coverage while reducing infrastructure complexity by eliminating redundant components and simplifying system architecture.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified data management system performs multiple data protection functions simultaneously - backup, replication, archiving, and disaster recovery - through a single multi-functional platform. This universal approach improves data protection coverage while avoiding the complexity of deploying separate specialized solutions for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple copies of data are created for different storage repositories, then data protection and availability are improved, but storage capacity requirements increase

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent implements nested deduplication where duplicate data blocks are identified and stored only once, with multiple references pointing to the same stored copy. This nesting approach maintains data availability across multiple storage repositories while dramatically reducing total storage capacity requirements by eliminating redundant data copies.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

Instead of creating full physical copies of data for each storage repository, the system creates logical references or pointers to the same underlying data blocks. This virtual copying mechanism maintains data availability and accessibility while minimizing actual storage capacity consumption.

Inventive Principle:
Principle #26Copying

3Productivity

If data is moved multiple times across storage repositories, then data lifecycle management is improved, but network bandwidth consumption increases

Engineering Contradiction:
Improvedata lifecycle management efficiencyVSAvoidnetwork bandwidth
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent extracts and stores only the unique portions of data that change during lifecycle transitions, rather than moving complete data sets. When data moves between storage repositories, only the differential changes are transmitted across the network, significantly reducing bandwidth consumption while maintaining efficient lifecycle management.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs preliminary deduplication and change identification before data movement operations. By pre-processing data to identify what actually needs to be transferred, the system avoids unnecessary network traffic while maintaining productive data lifecycle management across repositories.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9384207B2System and method for creating deduplicated copies of data by tracking temporal relationships among copies using higher-level hash structures
Publication Date: 2016.07.05 GOOGLE LLC
  • US9384207B2 patent drawing
  • US9384207B2 patent drawing
  • US9384207B2 patent drawing

AI summary

Systems and methods are disclosed for forming deduplicated images of a data object that changes over time using difference information between temporal states of the data object. The method includes organizing the content of the data object for a first temporal state as a plurality of content segments and storing the content segments in a data store; creating an organized arrangement of hash structures to represent the data object in its first temporal state; receiving difference information for the data object; forming at least one hash signature for the changed content; and storing the changed content that is unique in the data store as content segments. The method also includes determining, subsequent to receiving the changed content at the deduplicating content store, whether the changed content should be stored by searching for the hash signature for the changed higher-level hash structure in the global cache of the deduplicating content store.