Hardware Fault Injection Service for Cloud Resilience Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Chaos engineering in virtualized cloud environments is limited, as cloud providers restrict tests on shared-tenant hosts and only allows simulation of virtualized computing resource availability, not physical resource changes, hindering comprehensive resilience testing.
Innovation Solution
A test service simulates hardware-based faults by disabling resources or introducing errors on host machines and network devices, monitoring recovery processes, and recording results for analysis, enabling more robust resilience testing across virtual and physical layers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If chaos engineering tests are performed in a virtualized environment, then virtualized computing resource availability can be tested, but physical resource changes cannot be simulated
Solution Approach 1:
The patent introduces a test service as an intermediary component that bridges the gap between virtualized environment constraints and physical resource testing requirements. The test service communicates with both the virtual machine manager and hardware components, enabling indirect access to physical resources through controlled interfaces while maintaining virtualization layer abstraction
Solution Approach 2:
The testing capability is segmented into multiple independent components: virtual resource testing, physical resource testing, and integrated testing. This allows the system to selectively enable physical resource fault injection only when appropriate, while maintaining virtualized environment safety constraints, thus expanding versatility without compromising reliability
2Productivity
If tests are performed on shared-tenant hosts, then resource utilization can be improved, but cloud providers limit the types of tests that can be performed
Solution Approach 1:
The patent applies local quality by enabling different types of tests on different portions of the host system. Physical resource fault injection can be targeted to specific hardware components or isolated regions, allowing comprehensive testing on shared hosts while maintaining security and stability for critical shared resources
Solution Approach 2:
The testing system dynamically adjusts the type and scope of tests based on host configuration, tenant requirements, and system state. This allows the system to maximize test variety on shared-tenant hosts while adapting to cloud provider constraints in real-time, thus maintaining both productivity and versatility
Data Source
AI summary
Disclosed are various embodiments for simulating hardware-based faults in a cloud provider network. A first computing device can send a command to an offload card installed on a second computing device to introduce a simulated hardware fault into the second computing device. Then the first computing device can determine whether the second computing device has successfully recovered from the simulated hardware fault. Alternatively, an entry in an access control list (ACL) of a network switch can be modified to block network traffic to a first network interface of a host machine that is connected to the network switch. Then, a command can be sent to a second network interface of the host machine to instruct the host machine to perform a hard reset. Then it can be determined whether the host machine has successfully booted subsequent to the hard reset.


