Distributed System Hang Recovery via Centralized Watchdog
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Distributed computer systems face challenges in detecting and recovering from system hangs, which can lead to data loss and corruption, and introducing watchdog timers in a client-server model complicates the system, potentially causing a chain reaction of failures across machines.
Innovation Solution
Implementing a network thread in a distributed environment that manages requests and responses between client and server computers, with a timer set on the client to detect hang conditions, allowing for automatic reset requests to be sent to the server without the need for individual watchdog timers on each machine, enabling configurable time periods for response expectations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If watchdog timers are introduced to each client and server in a distributed system to detect system hangs, then system hang detection capability is improved, but system complexity increases significantly and may trigger chain reactions causing the whole distributed system to hang
Solution Approach 1:
The watchdog timer functionality is extracted from individual machines and centralized on a single designated machine. Instead of each machine running its own watchdog timer that could cause chain reactions, the detection capability is taken out and consolidated into a centralized service that monitors selected processes across the distributed system without the complexity and risk of distributed watchdog implementation.
Solution Approach 2:
A centralized watchdog service acts as an intermediary between the distributed system components. This mediator monitors processes across multiple machines from a single point, eliminating the need for each machine to independently manage watchdog timers. The intermediary approach simplifies the system architecture while maintaining comprehensive monitoring capability.
2Reliability
If clients and servers periodically wake up to reset watchdog timers to prevent system hangs, then system hang prevention is improved, but the blocking/waiting behavior of client-server model conflicts with the periodic wake-up requirement
Solution Approach 1:
The centralized watchdog service on the designated machine handles timer management autonomously without requiring client or server processes to wake up and reset timers. The service independently monitors process health and manages timeout logic, allowing client-server applications to maintain their blocking/waiting behavior while still achieving reliable hang prevention.
Solution Approach 2:
The centralized watchdog service acts as an intermediary that assumes responsibility for timer management. Instead of client and server processes needing to periodically wake up (which complicates their operation), the intermediary service handles all timer-related operations centrally, simplifying the operation of individual components while maintaining prevention capability.
3Reliability
If a system hang occurs on one machine in a distributed system, then local processing stops, but without proper detection mechanisms the hang triggers chain reactions causing the whole distributed system to hang
Solution Approach 1:
The detection capability is extracted from individual machines and centralized on a single designated machine. Instead of each machine running its own watchdog timer that could cause chain reactions, the detection capability is taken out and consolidated into a centralized service that monitors selected processes across the distributed system without the complexity and risk of distributed watchdog implementation.
Solution Approach 2:
The patent converts the potential harm of a single machine hang into a benefit by implementing a centralized watchdog that can detect and respond to hangs before they propagate. The centralized service monitors process health and can trigger recovery actions, transforming the harmful chain reaction effect into a controlled recovery scenario where the system self-heals without widespread failure.
Data Source
AI summary
Methods and apparatus, including computer program products, implementing and using techniques for resolving a hang condition in a distributed environment including a client computer and server computer(s). A network thread is provided, which passes requests from the client computer to the one or more server computers and responses from the server computers to the client computer. Worker thread(s) are provided on the server computers. The worker threads receive requests from the network thread, execute the requests, and pass responses to the requests back to the network thread. A request is sent from the client computer to a server computer through the network thread. A timer associated with the request is started on the client computer. The timer specifies a pre-defined time period for receiving a response to the request. When no response has been received within the pre-defined time period, a reset request is sent to the server computer.


