Remote Control Plane Failover Across Distributed Data Centers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The management of large-scale computing resources in data centers has become increasingly complicated due to the increased scale and scope of data centers, and existing technologies do not effectively address the challenges of remote management and failover of control planes in distributed computing environments.

Innovation Solution

A remote control plane system with automated failover capabilities is implemented, allowing management of hardware resources from a geographically distinct location, with a control plane monitoring service that monitors availability and status, enabling seamless failover to a different control plane in case of impairment or performance issues.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the control plane is collocated with the data plane in the same area, then management is simpler and monitoring is easier, but the system lacks fault isolation capability and cannot achieve remote failover

Engineering Contradiction:
Improvefault isolation capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the control plane and data plane into separate geographic areas. The control plane is deployed in a first area while the data plane resources are in a second area, creating physical separation that enables fault isolation. This segmentation allows the control plane to survive data plane failures and vice versa, directly resolving the contradiction between reliability through fault isolation and system architecture complexity.

Inventive Principle:
Principle #1Segmentation

2Reliability

If the control plane manages hardware resources in the same area, then response time is faster and latency is reduced, but the system cannot achieve remote management and geographic redundancy

Engineering Contradiction:
Improvegeographic redundancyVSAvoidremote management capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces an intermediary mechanism in the form of a control plane that remotely manages data plane resources across geographic boundaries. The control plane acts as a mediator between operators and the distributed hardware resources, enabling remote management while maintaining coordination. This intermediary approach resolves the contradiction by providing remote management capability and geographic redundancy without requiring direct collocation.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a single control plane manages all hardware resources, then the system is simpler to administer, but the system lacks automated failover capability and is vulnerable to single points of failure

Engineering Contradiction:
Improveautomated failover capabilityVSAvoidcontrol plane architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-configuring multiple control planes in different geographic areas before failures occur. Each control plane is prepared in advance to potentially assume management of hardware resources, with established communication channels and coordination protocols. This preliminary preparation enables automated failover when failures occur, resolving the contradiction between reliability through failover capability and control plane architecture complexity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3980892B1Remote control planes with automated failover
Publication Date: 2025.10.29 AMAZON TECH INC
  • EP3980892B1 patent drawingFigure 1
  • EP3980892B1 patent drawingFigure 2
  • EP3980892B1 patent drawingFigure 3

AI summary

Techniques for automated failover of remote control planes are described. A method of automated failover of remote control planes include determining failover event associated with a first control plane has occurred, the first control plane associated with a first area of a provider network, identifying a second control plane associated with a second area of the provider network, and failing over the first area of the provider network from the first control plane to the second control plane, wherein the child area updates one or more references to endpoints of the first control plane to be references to endpoints of the second control plane.