Autonomous Instance Recovery in Distributed Processing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional public cloud information processing systems incur increased operating costs due to the necessity of standby systems for failure recovery, which are not efficiently managed.
Innovation Solution
An information processing system with two parallel distributed processing systems and a processing status storage unit that allows instances to monitor and autonomously recover from failures without the need for a standby system, by storing and retrieving processing status information to initiate restarts and report unsuccessful recoveries.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a standby information processing system is used for failure recovery, then service continuity is improved, but operating costs increase
Solution Approach 1:
The failed instance autonomously performs recovery by acquiring its own status information from the processing status storage unit and executing restart operations without requiring external standby systems or manual intervention. This self-service mechanism eliminates the need for dedicated standby infrastructure while maintaining service continuity.
Solution Approach 2:
The system pre-stores processing status information in the processing status storage unit before failures occur. This preliminary action enables instances to quickly acquire their status and execute recovery operations immediately upon failure detection, eliminating service interruption without requiring standby systems.
2Reliability
If conventional failure recovery mechanisms are used, then service continuity is maintained, but system complexity increases due to standby requirements
Solution Approach 1:
The invention extracts and stores only the essential processing status information in the processing status storage unit, separating this critical data from the complex standby system infrastructure. This extraction enables simple, direct access to status information by instances during recovery operations.
Solution Approach 2:
Instances autonomously manage their own recovery by independently acquiring status information and executing restart operations, eliminating the need for complex standby system coordination and external management mechanisms.
3Loss of energy
If instances autonomously manage recovery, then operating costs are reduced, but coordination and monitoring become more difficult
Solution Approach 1:
The processing status storage unit provides continuous feedback to instances about their recovery status and processing state. This feedback mechanism enables autonomous decision-making while maintaining transparent monitoring, as instances can query their status and the system can track recovery progress without complex coordination overhead.
Data Source
AI summary
A remote management system includes a device management service having multiple instances that performs parallel distributed processing through the instances and a worker service having multiple instances that performs parallel distributed processing through the instances. Each instance of the device management service acquires the status of processing related to the recovery of the relevant instance from counter information, and stores the status of the processing related to the recovery of the relevant instance in the counter information when having performed the processing. Each instance of the worker service acquires the status of processing related to the recovery of an instance of the device management service from the counter information, and stores the status of the processing related to the recovery of the instance of the device management service in the counter information when having performed the processing.


