Failure Recovery Apparatus with Dynamic Rule Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing failure recovery systems are ineffective for handling failures not covered by pre-defined countermeasure rules, requiring administrator expertise and being costly and time-consuming to obtain effective countermeasures for all types of failures, especially for low versatility application programs.
Innovation Solution
A failure recovery apparatus that includes a rule storage for countermeasure commands and a knowledge storage for selecting commands based on failure types, automatically creating new rules for unmatched failures and storing them for future occurrences, allowing for quick recovery of similar failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If pre-defined countermeasure rules are used for failure recovery, then recovery effectiveness for known failures is improved, but the system becomes ineffective for handling failures not covered by pre-defined rules
Solution Approach 1:
The system automatically generates new countermeasure rules by analyzing failure logs and executing trial countermeasures without administrator intervention. When a failure occurs, the system autonomously creates a rule based on the failure state and selected countermeasure, then stores it for future automatic handling of similar failures.
Solution Approach 2:
The system changes the state of the rule storage by dynamically adding new countermeasure rules based on actual failure occurrences. The rule storage evolves from a static pre-defined set to a dynamic collection that grows with experience, adapting to new failure types automatically.
2Adaptability or versatility
If administrator manually adds countermeasure rules for all failure types, then comprehensive failure coverage is improved, but administrator burden and time cost increase significantly
Solution Approach 1:
The system performs preliminary actions by automatically generating countermeasure rules in advance based on detected failure patterns. Instead of waiting for administrators to manually create rules after failures occur, the system proactively creates rules during normal operation or immediately after trial executions, preparing recovery strategies before they are needed.
Solution Approach 2:
The system serves itself by automatically analyzing failures, selecting appropriate countermeasures, and generating rules without requiring administrator expertise or manual intervention. This self-service capability eliminates the time loss associated with administrator involvement while maintaining comprehensive failure coverage.
3Reliability
If comprehensive countermeasure rules are obtained through experimentation, then recovery effectiveness for all failure types is improved, but cost and time requirements become prohibitive
Solution Approach 1:
The system conducts its own experimentation by automatically executing trial countermeasures when failures occur and evaluating their effectiveness. This self-conducted experimentation replaces expensive and time-consuming manual experimentation, allowing the system to learn and improve autonomously without external resource investment.
Solution Approach 2:
The system implements feedback loops where the results of executed countermeasures are analyzed and used to refine and update countermeasure rules. This continuous feedback mechanism allows the system to improve recovery effectiveness over time without requiring extensive upfront experimentation, as learning occurs continuously through actual operation.
4Ease of operation
If manual countermeasure execution is used, then flexibility in handling complex failures is improved, but recovery speed and automation level decrease
Solution Approach 1:
The system autonomously executes countermeasures by automatically selecting appropriate rules from its stored knowledge base and implementing them without administrator intervention. This self-service execution maintains flexibility through its learned rule set while achieving full automation, eliminating the trade-off between manual flexibility and automated speed.
Solution Approach 2:
The system accelerates the recovery process by using powerful automated rule matching and execution mechanisms that quickly identify and implement appropriate countermeasures. This acceleration maintains operational flexibility through intelligent rule selection while dramatically reducing recovery time compared to manual processes.
Data Source
AI summary
The failure knowledge storage 5 stores failure knowledge information describing countermeasure command selection information, of which a diversion to a failure other than failures described in a failure countermeasure rule in the rule storage 2 can be surmised. When the failure occurring in the service executor 11 matches the failure described in the failure countermeasure rule, the countermeasure retriever 3 executes the countermeasure command in the failure countermeasure rule on the service executor 11. It is decided whether or not a failure other than the failures described in the failure countermeasure rule matches a failure described in the failure knowledge information in the failure knowledge storage 5. When there is a match, the countermeasure command is read out of the rule storage 2 based on the selection information of the matched failure knowledge information and is executed on the service executor 11.


