Network Switch Failure Detection via Independent BMC Monitoring
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current network switching products face challenges in rapidly detecting and responding to failures, particularly in peer network switches, which can lead to significant network traffic loss and ripple-like effects due to delayed failure detection.
Innovation Solution
A network switching unit equipped with a baseboard management controller (BMC) that monitors the host CPU and network processing unit, detects failures, and notifies peer switches via a separate management link, enabling rapid failure detection and notification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If network switching products use traditional failure detection methods through network traffic monitoring, then the system structure remains simple, but the failure detection time is delayed and network traffic loss increases
Solution Approach 1:
The patent segments the network switching system into distinct functional components: network processing units for traffic forwarding, host CPUs for control functions, and separate BMCs for failure detection. Each component has dedicated monitoring capabilities, allowing parallel failure detection across multiple nodes without interfering with network traffic processing. This segmentation enables rapid failure detection while maintaining system reliability.
Solution Approach 2:
The patent introduces BMCs as intermediary components that specialize in failure detection and notification. These BMCs act as mediators between the network switching units and peer switches, providing dedicated failure monitoring capabilities independent of the main network processing functions. This intermediary approach enables rapid failure detection without burdening the network processing units.
2Reliability
If network switching products use a unified control structure for both network processing and failure monitoring, then the device complexity remains low, but the failure notification speed is delayed
Solution Approach 1:
The patent divides the system into separate functional modules: network processing units handling traffic, host CPUs managing control functions, and independent BMCs dedicated to failure detection and notification. This segmentation allows each component to operate independently with optimized functions, achieving rapid failure notification while maintaining manageable system complexity through clear functional separation.
Solution Approach 2:
The BMC serves as an intermediary component that bridges the gap between network switching units and peer switches for failure notification. By introducing this dedicated intermediary, the system achieves rapid failure communication without requiring complex integration between network processing and monitoring functions, thus managing complexity while improving notification speed.
3Measurement precision
If network switching products rely on peer switches to detect each other's failures through network traffic, then the system structure remains simple, but the failure detection precision is insufficient and detection is delayed
Solution Approach 1:
The patent implements self-service failure detection where each network switching unit's BMC continuously monitors the status of its own host CPU and network processing unit. This self-monitoring capability allows immediate detection of failures without waiting for external detection through network traffic anomalies, significantly improving both detection precision and speed.
Solution Approach 2:
The patent establishes feedback mechanisms where BMCs continuously monitor system status and immediately notify peer switches of detected failures. This real-time feedback loop ensures precise and timely failure detection, allowing the network to respond quickly to failures without waiting for indirect detection through traffic monitoring.
Data Source
AI summary
A system and method for rapid peer node failure detection including a network switching unit that includes a network processing unit configured to receive and forward network traffic using one or more ports, a host CPU coupled to the network processing unit and configured to manage the network processing unit, a link controller coupled to the host CPU and configured to couple the network switching unit to a peer network switching unit using a management link, and a baseboard management controller (BMC) coupled to the host CPU and the link controller. The link controller is separate and independent from the network processing unit. The BMC is configured to monitor the host CPU and the network switching unit, detect a failure in the network switching unit, and notify the peer network switching unit of the detected failure using the management link.


