Tiered Replication Groups for Global Data Stores

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

As the number of geographic regions in data systems increases, replication latency and resource demands grow, leading to performance, consistency, and data integrity issues that fail to meet user needs.

Innovation Solution

Implementing a geographically distributed data store with tiered replication, where data is organized into replication groups with local replication states, allowing asynchronous replication within groups and optimizing robustness, region isolation, performance, and protection against data loss by configuring replica groups independently of the number of replica regions and nodes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If data is replicated across more geographic regions, then data availability and accessibility are improved, but replication latency and resource demands increase

Engineering Contradiction:
Improvedata availabilityVSAvoidreplication latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent divides the distributed data store into multiple replication groups, where each group contains a primary region and one or more replica regions. This segmentation allows independent management and optimization of replication within each group, reducing the overall replication latency by limiting the scope of synchronous replication operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces replication coordinators as intermediary components that manage replication operations between primary and replica regions. These coordinators act as mediators that can optimize replication timing and sequencing, thereby reducing replication latency while maintaining data consistency across geographic regions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If data is replicated across more geographic regions, then data availability is improved, but resource demands on storage systems increase

Engineering Contradiction:
Improvedata availabilityVSAvoidresource demands
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

By organizing the distributed data store into multiple replication groups with independent primary and replica regions, the patent enables resource demands to be distributed and managed at the group level rather than system-wide, reducing the overall resource burden on storage systems.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent allows each replication group to independently determine its own replication configuration and timing, enabling partial replication strategies that balance data availability requirements with available storage resources, rather than requiring full replication across all regions simultaneously.

Inventive Principle:
Principle #16Partial or excessive action

3Stability of the object's composition

If replication is performed across all replica regions, then data consistency is improved, but performance and data integrity issues increase

Engineering Contradiction:
Improvedata consistencyVSAvoidperformance
Core Design Contradiction:
Stability of the object's compositionVSProductivity

Solution Approach 1:

The patent segments the replication process into independent replication groups, where data consistency is maintained within each group through coordinated replication. This segmentation allows performance optimization within groups without compromising overall data consistency, as each group can be managed independently with appropriate replication strategies.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12147310B1Group replication for highly global workloads
Publication Date: 2024.11.19 AMAZON TECH INC
  • US12147310B1 patent drawing
  • US12147310B1 patent drawing
  • US12147310B1 patent drawing

AI summary

A geographically distributed data store including a number of geographically distributed regions may be implemented using replication groups that include multiple regions configured according to replication criteria. First tier replication of particular changes to data stored in the distributed data store may be performed in compliance with the replication criteria, where management of replication state is performed with respect to replication across the replication groups. Independent of the first tier replication, individual replication groups may implement second tier replication of changes to data where management of replication state is performed with respect to replication within the particular replication group. Replication group configuration may be determined using the replication criteria which may include thresholds for replication resource utilization, replication latency and utilization of data change logs.