Network Service Fault Recovery via Automatic Resource Isolation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Networked computer services face performance degradation and service interruptions due to interdependent components, requiring human intervention for fault management, which can lead to suboptimal response times and errors during outages.
Innovation Solution
A system with network monitors collecting performance data and a rules-based control engine that automatically detects faults, deactivates or reroutes affected resources, and adjusts services to maintain responsiveness without human intervention, using a control engine to monitor and manage network health and dynamically adjust features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If human operators manually monitor and manage network services, then service reliability can be maintained through human judgment, but response time becomes unacceptably slow and errors increase during urgent outages
Solution Approach 1:
The system enables self-service by implementing an automated control engine that monitors network services, detects faults, and executes recovery actions without human intervention. The control engine continuously receives status data from network monitors, applies rules-based logic to determine fault conditions, and automatically deactivates or reroutes affected services, allowing the system to manage itself during outages.
Solution Approach 2:
The system implements feedback by establishing a closed-loop control mechanism where network monitors continuously collect status data from services, feed this information to the control engine, and receive automated control commands in response. This feedback loop enables real-time detection and response to faults, with the control engine adjusting service configurations based on current system state and predefined recovery rules.
2Adaptability or versatility
If human operators manually adjust network services during outages, then complex judgment calls can be made, but the limited number of connections that can be monitored reduces system scalability
Solution Approach 1:
The system replaces the mechanical system of human operators physically monitoring and adjusting services with an automated electronic control engine. This substitution eliminates the human operator bottleneck, allowing the system to monitor and manage a significantly larger number of network services and connections simultaneously without proportional increases in operational complexity.
Solution Approach 2:
The control engine provides universal functionality by implementing a rules-based logic framework that can handle multiple types of faults across diverse network services through a single unified system. The engine applies generalized recovery rules to various service types (search, email, messaging, etc.), making the system scalable and adaptable without requiring service-specific manual intervention procedures.
3Loss of time
If automated control commands are issued without human intervention, then response time improves during outages, but the complexity of the control system increases
Solution Approach 1:
The control system is segmented into distinct functional modules: network monitors that collect status data, a control engine that processes information and applies rules-based logic, and execution mechanisms that implement control commands. This segmentation allows each component to perform its specific function independently, reducing overall system complexity while enabling automated rapid response to faults.
Solution Approach 2:
The system performs preliminary action by pre-configuring rules-based logic and control commands before faults occur. Recovery procedures, deactivation rules, and rerouting instructions are established in advance, allowing the control engine to execute predetermined actions immediately upon detecting fault conditions, achieving rapid response without complex real-time decision-making.
Data Source
AI summary
A system and related techniques monitor the operation of networked computer services, such as Internet-based search or other services which may link to a set of server resources, supporting databases and other components. In current technology if a large-scale or other search service encounters a problem in one supporting platform, the overall response may be interrupted or hang. In embodiments for example a search engine which attempts to access a travel, shopping, news or other database or service in response to a user search request may suspend the delivery of the search results, or blank the page with a 404 or other error message, due to the. unavailability of one or two databases or other components. The service support platform of the invention on the other hand may monitor the computer services network and detect the occurrence of a failure or performance lag, and temporarily detach the faulted server or other resource from the larger network service. In embodiments a user's Web page may be adapted to omit a panel of information related to that suspended service or component, so that for example an icon for a travel map or hotel reservation tool may be grayed out or removed while still presenting the remainder of the services or options. Users may therefore still access the majority of the Web service they are attempting to use, without interruption or noticeable lag.


