Network Switch Failure Detection via Independent BMC Monitoring

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current network switching products face challenges in rapidly detecting and responding to failures, particularly in peer network switches, which can lead to significant network traffic loss and ripple-like effects due to delayed failure detection.

Innovation Solution

A network switching unit equipped with a baseboard management controller (BMC) that monitors the host CPU and network processing unit, detects failures, and notifies peer switches via a separate management link, enabling rapid failure detection and notification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If network switching products use traditional failure detection methods through network traffic monitoring, then the system structure remains simple, but the failure detection time is delayed and network traffic loss increases

Engineering Contradiction:
Improvefailure detection speedVSAvoidfailure detection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the network switching system into distinct functional components: network processing units for traffic forwarding, host CPUs for control functions, and separate BMCs for failure detection. Each component has dedicated monitoring capabilities, allowing parallel failure detection across multiple nodes without interfering with network traffic processing. This segmentation enables rapid failure detection while maintaining system reliability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces BMCs as intermediary components that specialize in failure detection and notification. These BMCs act as mediators between the network switching units and peer switches, providing dedicated failure monitoring capabilities independent of the main network processing functions. This intermediary approach enables rapid failure detection without burdening the network processing units.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If network switching products use a unified control structure for both network processing and failure monitoring, then the device complexity remains low, but the failure notification speed is delayed

Engineering Contradiction:
Improvefailure notification speedVSAvoidsystem architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent divides the system into separate functional modules: network processing units handling traffic, host CPUs managing control functions, and independent BMCs dedicated to failure detection and notification. This segmentation allows each component to operate independently with optimized functions, achieving rapid failure notification while maintaining manageable system complexity through clear functional separation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The BMC serves as an intermediary component that bridges the gap between network switching units and peer switches for failure notification. By introducing this dedicated intermediary, the system achieves rapid failure communication without requiring complex integration between network processing and monitoring functions, thus managing complexity while improving notification speed.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If network switching products rely on peer switches to detect each other's failures through network traffic, then the system structure remains simple, but the failure detection precision is insufficient and detection is delayed

Engineering Contradiction:
Improvefailure detection precisionVSAvoidfailure detection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service failure detection where each network switching unit's BMC continuously monitors the status of its own host CPU and network processing unit. This self-monitoring capability allows immediate detection of failures without waiting for external detection through network traffic anomalies, significantly improving both detection precision and speed.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent establishes feedback mechanisms where BMCs continuously monitor system status and immediately notify peer switches of detected failures. This real-time feedback loop ensures precise and timely failure detection, allowing the network to respond quickly to failures without waiting for indirect detection through traffic monitoring.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS9148337B2System and method for rapid peer node failure detection
Publication Date: 2015.09.29 DELL PROD LP
  • US9148337B2 patent drawing
  • US9148337B2 patent drawing
  • US9148337B2 patent drawing

AI summary

A system and method for rapid peer node failure detection including a network switching unit that includes a network processing unit configured to receive and forward network traffic using one or more ports, a host CPU coupled to the network processing unit and configured to manage the network processing unit, a link controller coupled to the host CPU and configured to couple the network switching unit to a peer network switching unit using a management link, and a baseboard management controller (BMC) coupled to the host CPU and the link controller. The link controller is separate and independent from the network processing unit. The BMC is configured to monitor the host CPU and the network switching unit, detect a failure in the network switching unit, and notify the peer network switching unit of the detected failure using the management link.