Content Addressable Deduplication Using Inline Non-Lossy Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data management systems require multiple point solutions for different stages of a data lifecycle, leading to complex and expensive infrastructure with redundant data copies and inefficient data movement, which is exacerbated by server virtualization and the gap between emerging compute models and current data management implementations.

Innovation Solution

A Data Management Virtualization System that leverages Service Level Agreements to unify data protection across storage repositories, using deduplication and compression algorithms to reduce data redundancy, and abstracting physical storage resources into virtualized pools for efficient data movement and storage, allowing for non-uniform frequency and retention specifications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple point solutions are deployed for different data lifecycle stages, then data protection and management requirements are met, but infrastructure complexity and cost increase

Engineering Contradiction:
Improvedata protectionVSAvoidinfrastructure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple data management functions (backup, archiving, disaster recovery, compliance) into a single unified platform that handles all data lifecycle stages. This consolidation eliminates the need for separate point solutions while maintaining comprehensive data protection capabilities through integrated services.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The data management platform is designed to perform multiple functions simultaneously - creating backups, maintaining compliance records, enabling disaster recovery, and managing archiving all through one system. This multi-functional approach reduces infrastructure complexity while meeting diverse data protection requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple copies of data are created and moved to individual storage repositories, then data protection requirements are satisfied, but redundant data copies and inefficient data movement occur

Engineering Contradiction:
Improvedata protectionVSAvoidredundant data copies
Core Design Contradiction:
ReliabilityVSLoss of substance

Solution Approach 1:

The patent employs intelligent copying mechanisms that create data copies only when necessary and only the required portions of data. The system uses differential copying and selective replication to minimize redundant data copies while ensuring adequate protection across backup, archiving, and disaster recovery scenarios.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system automatically manages data retention by discarding obsolete copies and recovering only the necessary data for current protection needs. Through intelligent retention policies and lifecycle management, the system eliminates unnecessary redundant copies while maintaining required data protection levels.

Inventive Principle:
Principle #34Discarding and recovering

3Reliability

If data is copied frequently to various storage locations, then data protection and availability are improved, but storage capacity and network bandwidth are consumed

Engineering Contradiction:
Improvedata availabilityVSAvoidstorage capacity
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent implements different data handling strategies for different storage locations based on their specific requirements. Critical data receives frequent copying to nearby storage for quick recovery, while less critical data is archived with longer retention. This localized quality approach optimizes both availability and storage consumption.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts copying frequencies and retention periods based on data importance, access patterns, and storage resource availability. This dynamic lifecycle management ensures adequate data availability while minimizing storage capacity consumption through adaptive policies.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8299944B2System and method for creating deduplicated copies of data storing non-lossy encodings of data directly in a content addressable store
Publication Date: 2012.10.30 GOOGLE LLC
  • US8299944B2 patent drawing
  • US8299944B2 patent drawing
  • US8299944B2 patent drawing

AI summary

Systems and methods are disclosed for storing deduplicated images in which a portion of the image is stored in encoded form directly in a hash table, the method comprising: organizing unique content of each data object as a plurality of content segments and storing the content segments in a data store; receiving content to be included in the deduplicated image of the data object; determining if the received content may be encoded using a predefined non-lossy encoding technique and in which the encoded value would fit within the field for containing a hash signature; if so, placing the encoding in the field and marking the hash structure to indicate that the field contains encoded content; otherwise, generating a hash signature for the received content and placing the hash signature in the field and placing the received content in a corresponding content segment if it is unique.