Software Defined Failure Detection for Data Center Nodes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional data center monitoring systems are inflexible and struggle to adapt to varying data center topologies and conditions, leading to inefficiencies in failure detection and scalability, particularly in complex network environments.

Innovation Solution

A software-defined failure detection system with a scalable central controller that dynamically computes and updates monitoring topologies among failure detection agents, allowing for flexible and accurate monitoring of nodes across varying data center environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional monitoring systems are used with fixed topology coding, then implementation is simple, but the system cannot adapt to varying data center topologies and conditions

Engineering Contradiction:
Improveadaptability to varying data center topologiesVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The monitoring topology is transformed from a static, pre-coded configuration to a dynamic, software-defined structure. The central controller computes and distributes topology information to monitoring agents at runtime, allowing the system to adapt to changing data center topologies without requiring system redesign or recoding.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

A central controller is introduced as an intermediary between the data center infrastructure and the monitoring agents. This controller computes the monitoring topology based on current system state and distributes appropriate topology information to agents, enabling adaptive monitoring without direct complex interactions between all system components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If monitoring topology is fixed and coded into implementation, then deployment is straightforward, but the topology cannot be easily changed

Engineering Contradiction:
Improveflexibility to change topologyVSAvoidease of deployment
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The monitoring topology transitions from a static compiled configuration to a dynamic runtime-determined structure. Topology information is computed by the central controller based on current data center conditions and distributed to agents during operation, enabling flexible adaptation while maintaining simple initial deployment.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the parameter of topology representation from fixed code values to dynamic data structures that can be modified at runtime. This allows the same software implementation to adapt to different topologies by receiving updated topology information from the controller without requiring recompilation or reconfiguration.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If conventional solutions target flat network settings, then monitoring is simple, but the system fails in complex multi-network environments

Engineering Contradiction:
Improveaccuracy of failure detectionVSAvoidcomplexity of monitoring system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system assigns different monitoring responsibilities and topology information to different agents based on their local network context. Each agent receives topology information appropriate to its specific network environment, allowing accurate failure detection in complex multi-network settings while keeping each agent's processing requirements manageable.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The monitoring topology is dynamically computed to reflect the actual complex network structure, with the central controller analyzing the multi-network environment and distributing appropriate monitoring relationships to agents. This enables accurate failure detection in complex environments without requiring each agent to independently handle all complexity.

Inventive Principle:
Principle #15Dynamics

4Quantity of substance

If the system monitors more nodes, then coverage increases, but scalability and detection speed may be compromised

Engineering Contradiction:
Improvenumber of monitored nodesVSAvoidspeed of failure detection
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The monitoring system is segmented into multiple independent agents distributed across the data center, with a central controller that manages topology computation. This segmentation allows the system to scale to many nodes while maintaining detection speed, as each agent operates independently on its assigned monitoring tasks rather than a centralized process monitoring all nodes sequentially.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The monitoring topology is dynamically computed and optimized by the central controller to balance coverage and detection speed. The controller can adjust monitoring relationships based on current system conditions, node criticality, and network state, enabling the system to efficiently monitor large numbers of nodes while maintaining rapid failure detection capability.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10547499B2Software defined failure detection of many nodes
Publication Date: 2020.01.28 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US10547499B2 patent drawing
  • US10547499B2 patent drawing
  • US10547499B2 patent drawing

AI summary

Embodiments of the present systems and methods may provide the capability to monitor and detect failure of nodes in a data center environment by using a software defined failure detector that can be adjusted to varying conditions and data center topology. In an embodiment, a computer-implemented method for monitoring and detecting failure of electronic systems may comprise, in a system comprising a plurality of networked computer systems, defining at least one failure detection agent to monitor operation of other failure detection agents running on at least some of the electronic systems, and defining, at the controller, and transmitting, from the controller, topology information defining a topology of the failure detection agents to the failure detection agents, wherein the topology information includes information defining which failure detection agents each failure detection agent is to monitor.