Simulated Network Outages for Backup Scaling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computing service providers face challenges in ensuring that secondary networks can scale sufficiently to handle redirected workloads during primary network failures, often resulting in instability and downtime due to underutilization and inadequate resource allocation.

Innovation Solution

Implementing a system that simulates outages in secondary networks to regularly redirect and scale resources, allowing for better preparation and handling of potential primary network failures by periodically redirecting events and monitoring performance across multiple regions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If computing resources are dynamically allocated or deallocated based on usage over time, then resource optimization is improved, but the secondary network cannot scale sufficiently fast enough to accommodate redirected workloads during primary network failures

Engineering Contradiction:
Improvecomputing resourcesVSAvoidscaling speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The system performs preliminary actions by proactively simulating outages and redirecting workloads to secondary networks before actual failures occur. This allows the secondary network to pre-scale its computing resources, ensuring it is ready to handle redirected workloads at full capacity when needed, rather than reacting slowly to actual failure events.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements periodic action by conducting simulated outages at regular intervals to stress-test and scale the secondary network. This periodic stress-testing ensures the secondary network maintains its scaling capability over time, preventing resource optimization from causing degradation of failover readiness.

Inventive Principle:
Principle #19Periodic action

2Reliability

If a secondary network is maintained as backup, then service reliability is improved, but the secondary network remains underutilized and cannot scale sufficiently when needed

Engineering Contradiction:
Improveservice reliabilityVSAvoidnetwork utilization
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system ensures continuity of useful action by continuously utilizing the secondary network through simulated outages, even when the primary network is operational. This continuous stress-testing and workload redirection maintains the secondary network's scaling capability and readiness, transforming it from a dormant backup to an actively prepared failover target.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system applies parameter changes by dynamically adjusting the operational state of the secondary network between normal and stressed conditions. During simulated outages, the secondary network's resource allocation and scaling parameters are modified to reflect actual failure scenarios, ensuring it can transition smoothly to handling all workloads when needed.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11841780B1Simulated network outages to manage backup network scaling
Publication Date: 2023.12.12 AMAZON TECH INC
  • US11841780B1 patent drawing
  • US11841780B1 patent drawing
  • US11841780B1 patent drawing

AI summary

A system is configured to simulate outages of network resources. The system is configured to provide a control plane for computing resources of a provider network. The control plane is configured to cause simulated outages of a primary region of the plurality of regions selected to host the plurality of different computing resources. During the simulated outages, the control plane moves respective workloads of the plurality of different computing resources to be performed in the one or more secondary networks and tracks a performance of the one or more secondary regions hosting the moved respective workloads of the plurality of computing resources. After completing individual ones of the simulated outages of the first network, the control plane moves the respective workloads of the plurality of different computing resources back to the primary region.