Clustered Health Checks for Distributed Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In modern distributed systems, checking the health status of numerous servers and data stores across hundreds or thousands of nodes is challenging, making it difficult to ensure application functionality and reduce service interruptions, as conventional methods are time-consuming and prone to human error.

Innovation Solution

A system and method for providing clustered health checks using a user interface to group and manage identifiers of computing resources, allowing for efficient health checks through a browser application, where identifiers are stored and associated with health check commands, enabling a single action to initiate checks on multiple resources, reducing human intervention and error.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional methods are used to check the status of servers and data stores individually, then the health status of each resource can be monitored, but the process becomes time-consuming and prone to human error when scaled to hundreds or thousands of nodes

Engineering Contradiction:
Improvehealth check accuracyVSAvoidhealth check duration
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the large-scale health check task by organizing computing resources into hierarchical groups (clusters, regions, availability zones). Instead of checking thousands of individual resources sequentially, the system divides them into manageable groups that can be checked in parallel, significantly reducing the overall time while maintaining comprehensive coverage through structured organization

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-configuring health check commands and grouping resources before actual health checks are needed. Resource groupings and check configurations are established in advance, allowing the system to execute health checks across hundreds or thousands of nodes efficiently without manual intervention during the actual check process

Inventive Principle:
Principle #10Preliminary action

2Reliability

If manual selection of each resource for health checks is performed, then individual resource status can be verified, but human error increases and efficiency decreases when dealing with large numbers of nodes

Engineering Contradiction:
Improveresource status monitoringVSAvoidhealth check efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system implements self-service by automatically selecting and checking resources based on pre-configured groupings and criteria. Once resources are organized into groups and health check parameters are set, the system autonomously performs health checks without requiring manual selection of each resource, eliminating human error and significantly improving efficiency while maintaining reliable monitoring of all designated resources

Inventive Principle:
Principle #25Self-service

3Reliability

If health checks are performed on hundreds or thousands of nodes, then comprehensive system health monitoring is achieved, but the complexity of managing and tracking individual resource status increases significantly

Engineering Contradiction:
Improvesystem health monitoring coverageVSAvoidhealth check management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by organizing thousands of computing resources into hierarchical groups (clusters, regions, availability zones). This structure breaks down the complexity of managing individual resource status into manageable hierarchical levels, where health check results can be aggregated and displayed at appropriate levels of the hierarchy, making comprehensive monitoring tractable despite the large number of nodes

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system merges individual resource health check results into grouped summaries. By combining status information from multiple resources within defined groups, the system presents consolidated health views that reduce complexity while maintaining comprehensive coverage. Users can see both individual resource status and aggregated group status, managing complexity through information merging at appropriate granularities

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20200358688A1Systems and methods for providing clustered health checks
Publication Date: 2020.11.12 T MOBILE US INC
  • US20200358688A1 patent drawing
  • US20200358688A1 patent drawing
  • US20200358688A1 patent drawing

AI summary

Systems and methods for providing clustered health checks are described herein. The systems and methods enable a user to select one or more selectable links to group together. The selectable links are associated with computing resources used by an application. Once selected, a parallel executor server receives the links and issues a health check command code to the computing resources associated with the group. The computing resources can be grouped together based on various reasons. The selectable links can be rendered in a graphical user interface of a health check application or an Internet browser.