Hardware Fault Injection Service for Cloud Resilience Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Chaos engineering in virtualized cloud environments is limited, as cloud providers restrict tests on shared-tenant hosts and only allows simulation of virtualized computing resource availability, not physical resource changes, hindering comprehensive resilience testing.

Innovation Solution

A test service simulates hardware-based faults by disabling resources or introducing errors on host machines and network devices, monitoring recovery processes, and recording results for analysis, enabling more robust resilience testing across virtual and physical layers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If chaos engineering tests are performed in a virtualized environment, then virtualized computing resource availability can be tested, but physical resource changes cannot be simulated

Engineering Contradiction:
Improveresilience testing capabilityVSAvoidtesting scope
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces a test service as an intermediary component that bridges the gap between virtualized environment constraints and physical resource testing requirements. The test service communicates with both the virtual machine manager and hardware components, enabling indirect access to physical resources through controlled interfaces while maintaining virtualization layer abstraction

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The testing capability is segmented into multiple independent components: virtual resource testing, physical resource testing, and integrated testing. This allows the system to selectively enable physical resource fault injection only when appropriate, while maintaining virtualized environment safety constraints, thus expanding versatility without compromising reliability

Inventive Principle:
Principle #1Segmentation

2Productivity

If tests are performed on shared-tenant hosts, then resource utilization can be improved, but cloud providers limit the types of tests that can be performed

Engineering Contradiction:
Improvehost utilizationVSAvoidtest type variety
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by enabling different types of tests on different portions of the host system. Physical resource fault injection can be targeted to specific hardware components or isolated regions, allowing comprehensive testing on shared hosts while maintaining security and stability for critical shared resources

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The testing system dynamically adjusts the type and scope of tests based on host configuration, tenant requirements, and system state. This allows the system to maximize test variety on shared-tenant hosts while adapting to cloud provider constraints in real-time, thus maintaining both productivity and versatility

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12009990B1Hardware-based fault injection service
Publication Date: 2024.06.11 AMAZON TECH INC
  • US12009990B1 patent drawing
  • US12009990B1 patent drawing
  • US12009990B1 patent drawing

AI summary

Disclosed are various embodiments for simulating hardware-based faults in a cloud provider network. A first computing device can send a command to an offload card installed on a second computing device to introduce a simulated hardware fault into the second computing device. Then the first computing device can determine whether the second computing device has successfully recovered from the simulated hardware fault. Alternatively, an entry in an access control list (ACL) of a network switch can be modified to block network traffic to a first network interface of a host machine that is connected to the network switch. Then, a command can be sent to a second network interface of the host machine to instruct the host machine to perform a hard reset. Then it can be determined whether the host machine has successfully booted subsequent to the hard reset.