Network Device PFC Deadlock Root Cause Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods, such as the watchdog mechanism, cannot effectively address the root cause of PFC deadlocks in network egress port queues, leading to recurring communication disruptions and network breakdowns.
Innovation Solution
A network device uses an access control list to identify abnormal data flows and reports anomaly information to a management device, allowing operation and maintenance personnel to determine the source and destination of the issue, and visually present the PFC deadlock loop to locate and resolve the root cause.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the watchdog mechanism is used to resolve PFC deadlock, then the deadlock can be broken for an egress port queue, but the root cause cannot be identified and the deadlock may occur again
Solution Approach 1:
The patent implements a feedback mechanism where the network device sends anomaly information (including deadlock queue identifier, abnormal data flow identifier, and candidate root cause network device identifiers) to the management device. The management device then provides feedback by determining the actual root cause network device based on this information, enabling continuous improvement and prevention of recurring deadlocks
Solution Approach 2:
The management device acts as an intermediary between the network device and the root cause analysis. It receives anomaly information from the network device, processes this information to determine the actual root cause network device, and provides comprehensive feedback, thereby bridging the gap between deadlock detection and root cause resolution
2Loss of information
If anomaly information is reported to management device, then root cause can be identified, but system complexity increases
Solution Approach 1:
The patent segments the deadlock resolution system into distinct functional components: the network device responsible for detection and initial analysis (generating anomaly information), and the management device responsible for final root cause determination. This segmentation allows each component to focus on specific tasks, reducing overall system complexity while maintaining comprehensive root cause identification capability
Solution Approach 2:
The management device serves as an intermediary that consolidates complex analysis logic. Instead of distributing complex root cause analysis across multiple network devices, the intermediary management device centralizes this function, simplifying the architecture while enabling thorough root cause identification through coordinated information exchange
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
This application discloses a method, an apparatus, and a system for locating a root cause of a network anomaly, and a computer storage medium, and belongs to the field of network technologies. When a PFC deadlock occurs in a first egress port queue in a network device, the network device determines an abnormal data flow in the first egress port queue based on an access control list. Both an egress port and an ingress port of the abnormal data flow are uplink ports of the network device. The first egress port queue is any egress port queue in the network device. The network device sends anomaly information to a network management device, where the anomaly information includes an identifier of the abnormal data flow. The network management device transmits the identifier of the abnormal data flow to a display device for display by the display device.