Cloud Storage Gateway with Unified Namespace for Multi-Site Deduplication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current storage management systems struggle to efficiently manage and scale data storage in cloud environments, particularly in handling massive volumes of unstructured data and performing value-added storage operations such as deduplication, content indexing, and policy-driven storage across multiple cloud storage sites.

Innovation Solution

The system employs a data storage enterprise architecture that includes content indexing, containerized deduplication, and policy-driven storage, allowing for efficient data management and scaling across multiple cloud storage sites. This architecture supports wide-area network data transfer and integrates with cloud gateways for block-level and sub-object-level data migration and restoration.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional storage management systems are used to manage cloud storage, then basic storage operations can be performed, but the systems cannot efficiently handle massive volumes of unstructured data and perform value-added storage operations such as deduplication, content indexing, and policy-driven storage across multiple cloud storage sites

Engineering Contradiction:
Improvecapability to perform value-added storage operationsVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is divided into multiple independent cloud storage sites that can be distributed across different locations. Each site operates semi-autonomously but can collaborate through the unified namespace, allowing the system to scale horizontally while maintaining manageable complexity at each individual site.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The cloud storage gateway implements a unified namespace that provides universal access to data across multiple cloud storage sites. This single interface handles diverse operations including storage, retrieval, deduplication, content indexing, and policy-driven management, eliminating the need for separate systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If data is stored across multiple cloud storage sites, then storage capacity and reliability are improved, but data transfer costs and latency increase

Engineering Contradiction:
Improvedata storage reliabilityVSAvoiddata transfer latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs data deduplication and content indexing in advance before data needs to be retrieved. By pre-processing data and creating indexes locally at each cloud storage site, the system eliminates the need to transfer entire data sets across wide area networks when data is accessed, significantly reducing latency while maintaining multi-site redundancy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The cloud storage gateway acts as an intermediary that manages data placement and retrieval across multiple cloud storage sites. It intelligently routes data operations to the appropriate site, caching frequently accessed data locally and using the unified namespace to transparently handle cross-site data access, thereby reducing the impact of wide area network latency.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If data is replicated across multiple cloud storage sites, then data availability and fault tolerance are improved, but storage costs increase

Engineering Contradiction:
Improvedata availabilityVSAvoidtotal data storage volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

Instead of replicating entire data sets across multiple cloud storage sites, the system creates and stores only unique data copies. The deduplication engine identifies and eliminates redundant data, storing only distinct data objects while maintaining references to them across multiple sites. This provides fault tolerance through redundancy of references rather than redundant copies of all data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system dynamically changes the replication parameter from full data replication to selective replication of only unique data objects. By monitoring data uniqueness and replication status, the system adjusts the amount of data stored at each site, maintaining adequate availability while minimizing total storage volume across the distributed system.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If traditional storage systems are used, then implementation is straightforward, but they cannot scale to handle massive volumes of unstructured data

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidstorage architecture complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system transitions from traditional hierarchical storage architecture to a distributed cloud-based architecture with a unified namespace. This dimensional change allows the system to scale horizontally by adding more cloud storage sites rather than vertically expanding single systems, enabling handling of massive data volumes while maintaining manageable operational complexity through standardized interfaces.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12321592B2Data object store and server for a cloud storage environment
Publication Date: 2025.06.03 COMMVAULT SYSTEMS INC
  • US12321592B2 patent drawing
  • US12321592B2 patent drawing
  • US12321592B2 patent drawing

AI summary

Data storage operations, including content-indexing, containerized deduplication, and policy-driven storage, are performed within a cloud environment. The systems support a variety of clients and cloud storage sites that may connect to the system in a cloud environment that requires data transfer over wide area networks, such as the Internet, which may have appreciable latency and/or packet loss, using various network protocols, including HTTP and FTP. Methods are disclosed for content indexing data stored within a cloud environment to facilitate later searching, including collaborative searching. Methods are also disclosed for performing containerized deduplication to reduce the strain on a system namespace, effectuate cost savings, etc. Methods are disclosed for identifying suitable storage locations, including suitable cloud storage sites, for data files subject to a storage policy. Further, systems and methods for providing a cloud gateway and a scalable data object store within a cloud environment are disclosed, along with other features.