Distributed System Hang Recovery via Centralized Watchdog

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Distributed computer systems face challenges in detecting and recovering from system hangs, which can lead to data loss and corruption, and introducing watchdog timers in a client-server model complicates the system, potentially causing a chain reaction of failures across machines.

Innovation Solution

Implementing a network thread in a distributed environment that manages requests and responses between client and server computers, with a timer set on the client to detect hang conditions, allowing for automatic reset requests to be sent to the server without the need for individual watchdog timers on each machine, enabling configurable time periods for response expectations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If watchdog timers are introduced to each client and server in a distributed system to detect system hangs, then system hang detection capability is improved, but system complexity increases significantly and may trigger chain reactions causing the whole distributed system to hang

Engineering Contradiction:
Improvesystem hang detection capabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The watchdog timer functionality is extracted from individual machines and centralized on a single designated machine. Instead of each machine running its own watchdog timer that could cause chain reactions, the detection capability is taken out and consolidated into a centralized service that monitors selected processes across the distributed system without the complexity and risk of distributed watchdog implementation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

A centralized watchdog service acts as an intermediary between the distributed system components. This mediator monitors processes across multiple machines from a single point, eliminating the need for each machine to independently manage watchdog timers. The intermediary approach simplifies the system architecture while maintaining comprehensive monitoring capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If clients and servers periodically wake up to reset watchdog timers to prevent system hangs, then system hang prevention is improved, but the blocking/waiting behavior of client-server model conflicts with the periodic wake-up requirement

Engineering Contradiction:
Improvesystem hang preventionVSAvoidoperation simplicity
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The centralized watchdog service on the designated machine handles timer management autonomously without requiring client or server processes to wake up and reset timers. The service independently monitors process health and manages timeout logic, allowing client-server applications to maintain their blocking/waiting behavior while still achieving reliable hang prevention.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The centralized watchdog service acts as an intermediary that assumes responsibility for timer management. Instead of client and server processes needing to periodically wake up (which complicates their operation), the intermediary service handles all timer-related operations centrally, simplifying the operation of individual components while maintaining prevention capability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If a system hang occurs on one machine in a distributed system, then local processing stops, but without proper detection mechanisms the hang triggers chain reactions causing the whole distributed system to hang

Engineering Contradiction:
Improvedistributed system stabilityVSAvoidchain reaction failures
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The detection capability is extracted from individual machines and centralized on a single designated machine. Instead of each machine running its own watchdog timer that could cause chain reactions, the detection capability is taken out and consolidated into a centralized service that monitors selected processes across the distributed system without the complexity and risk of distributed watchdog implementation.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent converts the potential harm of a single machine hang into a benefit by implementing a centralized watchdog that can detect and respond to hangs before they propagate. The centralized service monitors process health and can trigger recovery actions, transforming the harmful chain reaction effect into a controlled recovery scenario where the system self-heals without widespread failure.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Data Source

PatentUS8453013B1System-hang recovery mechanisms for distributed systems
Publication Date: 2013.05.28 GOOGLE LLC
  • US8453013B1 patent drawing
  • US8453013B1 patent drawing
  • US8453013B1 patent drawing

AI summary

Methods and apparatus, including computer program products, implementing and using techniques for resolving a hang condition in a distributed environment including a client computer and server computer(s). A network thread is provided, which passes requests from the client computer to the one or more server computers and responses from the server computers to the client computer. Worker thread(s) are provided on the server computers. The worker threads receive requests from the network thread, execute the requests, and pass responses to the requests back to the network thread. A request is sent from the client computer to a server computer through the network thread. A timer associated with the request is started on the client computer. The timer specifies a pre-defined time period for receiving a response to the request. When no response has been received within the pre-defined time period, a reset request is sent to the server computer.