Link Aggregation Split-Brain Detection via Status Resolution Server
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-chassis link aggregation systems fail to distinguish between node failures and communication failures, leading to 'split-brain' scenarios that cause network congestion and errors due to forwarding loops.
Innovation Solution
A network system with a status resolution server that determines the operational status of peers within a link aggregation group, allowing for detection and recovery from split-brain scenarios by querying peers and reconfiguring network elements via a heartbeat protocol or central database, and establishing mediation links to maintain communication.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If link aggregation is implemented without status resolution mechanism, then network throughput and redundancy are improved, but split-brain scenarios occur causing network errors and congestion
Solution Approach 1:
The patent implements a feedback mechanism where the status resolution server continuously monitors peer status through heartbeat messages and mediation links. When a peer detects apparent failure, it queries the status resolution server which responds with authoritative status information, enabling dynamic adjustment of operational state and preventing split-brain scenarios while maintaining high availability
Solution Approach 2:
The status resolution server acts as an intermediary between peers in the link aggregation group. It receives status queries from peers and provides authoritative status information, mediating the failure detection process to distinguish between actual node failures and temporary communication issues, thereby preventing incorrect failover decisions
2Loss of time
If peer failure detection is implemented without status resolution server, then response time to failures is improved, but false failure detection occurs leading to split-brain scenarios
Solution Approach 1:
The system performs preliminary status verification through the status resolution server before declaring a peer failure. When a peer detects apparent failure through lost communication, it immediately queries the status resolution server for authoritative status confirmation, preventing premature failover decisions based on incomplete information
Solution Approach 2:
The status resolution server serves as an intermediary that provides authoritative status information to resolve ambiguity in failure detection. It receives status queries from peers and returns precise status information, enabling accurate distinction between actual failures and temporary communication issues without delaying rapid response
3Reliability
If multi-chassis LAG configuration is implemented, then node-level redundancy is improved, but communication failures between nodes cause split-brain scenarios
Solution Approach 1:
The status resolution server acts as a central intermediary that simplifies the complexity of multi-chassis LAG management. It centralizes status information and coordination functions, allowing peers to determine operational status through simple queries rather than complex peer-to-peer communication protocols, thereby reducing system complexity while maintaining redundancy
Solution Approach 2:
The status resolution server provides multiple functions within a single component: it maintains peer status information, processes status queries from peers, coordinates failover decisions, and prevents split-brain scenarios. This multi-functionality reduces overall system complexity by consolidating what would otherwise require multiple separate mechanisms
Data Source
AI summary
Various embodiments are described herein that provide a network system comprising a set of peers within a link aggregation group (LAG), the first set of peers including a first network element and a second network element and a status resolution server to connect to the set of peers within the link aggregation group, wherein one or more peers within the LAG is to query the status resolution server to determine an operational status of a peer in the set of peers in response to detection of an apparent failure of the peer.


