Failure Casting Hierarchy for Computer System Recovery
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex computer systems face challenges in managing failures due to their intricate nature, where traditional recovery methods are often unreliable and inefficient, particularly in handling unforeseen or unknown failures, leading to prolonged downtime and data loss.
Innovation Solution
The implementation of a failure casting system that uses a hierarchy to cast various failures into reboot-curable failures, allowing for systematic recovery through rebooting, thereby simplifying the recovery process and improving system reliability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional recovery methods are used to handle failures in complex computer systems, then the system can maintain operational complexity, but the recovery reliability and efficiency deteriorate due to unpredictable failure types and prolonged downtime
Solution Approach 1:
The patent applies universality by creating a unified failure casting mechanism that handles multiple types of failures (hardware, software, network, unknown) through a single standardized process. The failure casting logic universally translates all failure types into a common representation that triggers standardized recovery actions, making the recovery system multi-functional and applicable to diverse failure scenarios without requiring separate specialized handlers for each failure type.
Solution Approach 2:
The patent segments the complex failure management process into distinct modular components: failure detection module, failure casting logic, hierarchy navigation module, and recovery execution module. Each component handles a specific aspect of failure management independently, allowing the system to maintain operational complexity while improving recovery reliability through organized, manageable segments rather than a monolithic recovery system.
2Adaptability or versatility
If comprehensive failure handling for all failure types is implemented, then the system can handle diverse failures, but the recovery process complexity increases making it harder to manage
Solution Approach 1:
The patent introduces failure casting logic as an intermediary layer between failure detection and recovery execution. This mediator translates diverse failure types into a standardized failure representation, enabling the system to handle various failure types adaptively while keeping the recovery process simple and manageable. The intermediary abstracts the complexity of diverse failures behind a uniform interface.
Solution Approach 2:
The patent changes the parameter representation of failures by transforming specific failure details (hardware, software, network types) into a standardized failure parameter set that the recovery system can uniformly process. This parameter transformation allows the system to maintain high adaptability in handling different failure types while reducing recovery process complexity through consistent parameter structures.
3Reliability
If manual intervention is required for each failure type, then the system can provide precise recovery control, but the recovery time increases leading to prolonged downtime
Solution Approach 1:
The patent implements preliminary action by pre-configuring the failure hierarchy and casting rules during system setup. The failure casting logic and recovery procedures are prepared in advance, allowing the system to automatically translate and handle failures without requiring manual intervention for each failure type. This preliminary preparation maintains recovery precision while dramatically reducing downtime through automated rapid response.
Solution Approach 2:
The patent enables self-service by implementing an automated failure casting and recovery system that operates without human intervention. The system automatically detects failures, casts them into standardized representations, navigates the failure hierarchy, and executes appropriate recovery actions autonomously. This self-service capability provides precise recovery control through automated decision-making while minimizing downtime by eliminating manual response delays.
4Reliability
If extensive testing and design review are performed to eliminate bugs, then the software reliability improves, but the development time and cost increase significantly
Solution Approach 1:
The patent applies beforehand cushioning by implementing a safety net mechanism (failure casting and recovery system) that prepares for potential failures in advance rather than attempting to eliminate all bugs during development. This prior cushioning approach accepts that software will have bugs but provides automated mechanisms to quickly recover from them, thereby improving operational reliability without extending development time for exhaustive bug elimination.
Data Source
AI summary
A system and method for using failure casting to manage failures in a computer system. In accordance with an embodiment, the system uses a failure casting hierarchy to cast failures of one type into failures of another type. In doing this, the system allows incidents, problems, or failures to be cast into a (typically smaller) set of failures, which the system knows how to handle. In accordance with a particular embodiment, failures can be cast into a category that is considered reboot-curable. If a failure is reboot-curable then rebooting the system will likely cure the problem. Examples include hardware failures, and reboot-specific methods that can be applied to disk failures and to failures within clusters of databases. The system can even be used to handle failures that were hitherto unforeseen—failures can be cast into known failures based on the failure symptoms, rather than any underlying cause.


