Service Restoration Engine Phased Dependency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional methods for restoring application services in enterprise networks after a mass outage are time-consuming and resource-intensive due to complex interdependencies between services, often disrupting unaffected users and requiring the restart of entire application stacks.
Innovation Solution
The Faster Service Restoration (FSR) engine identifies dependencies and generates a dynamic run list to restore services in phases, maintaining minimum availability by staggering the restoration across clusters and executing healing scripts on target servers, allowing parallel restoration of services with no dependencies.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional methods are used to restore application services after a mass outage, then all services are restored eventually, but the restoration process is time-consuming and disrupts unaffected users
Solution Approach 1:
The patent segments the service restoration process into multiple phases based on service dependencies. Services are restored in successive phases where each phase restores a subset of services that have their dependencies satisfied, rather than restoring all services simultaneously or in a single sequential order. This segmentation allows parallel restoration of independent services while maintaining dependency constraints.
Solution Approach 2:
The patent implements dynamic service restoration by continuously monitoring service dependency satisfaction and automatically adjusting which services can be restored in each phase. The system dynamically determines the restoration sequence based on real-time dependency analysis, allowing flexible adaptation to different outage scenarios and service interdependencies without manual intervention.
2Reliability
If system engineers manually identify and restore services following an outage, then service restoration can be controlled, but the process requires multiple attempts and is resource-intensive
Solution Approach 1:
The patent implements self-service automation where the system automatically identifies affected services, analyzes their dependencies, determines the restoration sequence, and executes the restoration process without requiring system engineers to manually perform these tasks. The automated system manages the entire restoration workflow, reducing human intervention to monitoring and exception handling only.
Solution Approach 2:
The patent incorporates feedback mechanisms that continuously monitor service status, dependency satisfaction, and restoration progress. The system uses this feedback to dynamically adjust the restoration plan, verify successful restoration, and ensure that services are restored in the correct sequence while maintaining system stability throughout the process.
3Reliability
If entire application stacks are restarted to restore services, then all services are restored, but unaffected users experience disruption
Solution Approach 1:
The patent extracts and restores only the specific services that are affected by the outage and whose dependencies are satisfied, rather than restarting entire application stacks. This selective restoration approach isolates the restoration process to only the necessary services, leaving unaffected services running continuously without interruption to their users.
Solution Approach 2:
The patent applies local quality by treating different services with different restoration priorities and timing based on their specific dependency requirements. Each service is restored individually or in small groups when its specific dependencies are met, rather than applying a uniform restoration approach to all services. This allows services in different locations of the system to be restored at different times based on their local dependency conditions.
Data Source
AI summary
Techniques are disclosed for restoring application services in a computer network. One example method generally includes identifying a set of servers hosting an application and determining a plurality of successive phases for restoring the application. The method further includes identifying a first instance of a first service of the application executing on a first server of the set of servers hosting the application and determining other instances of the first service are unavailable on other servers of the set of servers hosting the application. The method further includes delaying restoration of additional services on the first server until at least a second instance of the first service is available on one or more servers of the set of servers hosting the application other than the first server.


