LDP down fault positioning method and device
By using automated alarm and protocol status analysis, combined with testing, the complexity of LDP Down fault location has been solved, enabling rapid and accurate fault location and reducing the workload of maintenance personnel.
Patent Information
- Application Number
- CN202411299816.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-18
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-09-18
AI Technical Summary
LDP down fault location is complex and requires professional knowledge and experience, resulting in time-consuming and labor-intensive fault handling, inaccurate location, and impact on network operation.
By collecting and analyzing alarms and protocol status, and initiating tests on faulty network elements, the cause of LDP Down faults can be automatically located, including checking interface status, timer timeout, board status, and CPU utilization, reducing human error and omissions.
It achieves standardized and automated location of LDP Down faults, reducing the workload of maintenance personnel, improving the efficiency and accuracy of fault location, and reducing human error.
Smart Images

Figure CN119383063B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of LDP Down fault location, and in particular to an LDP Down fault location method and apparatus. Background Technology
[0002] LDP (Label Distribution Protocol) is a protocol used to distribute labels in Multiprotocol Label Switching (MPLS) networks. An LDP down failure refers to the loss or inability to establish an LDP session, leading to an interruption of the Label Switched Path (LSP) and consequently affecting normal network operation. Because it is a protocol-level failure, locating the problem is complex. Maintenance personnel need to spend a significant amount of time manually determining the fault location, causing prolonged delays and negatively impacting customers.
[0003] Currently, maintenance personnel need to use various protocols and commands based on their professional knowledge to analyze faults, which requires a significant amount of time to manually determine the cause of LDP down faults. On the one hand, this requires a high level of experience, demanding in-depth professional knowledge and extensive experience. On the other hand, even with experience, complex issues like routing fault location may still lead to oversights in practice, resulting in time-consuming, labor-intensive processes, low timely fault handling rates, and significant pressure on maintenance personnel. Therefore, it is necessary to standardize, streamline, and automate this process. Summary of the Invention
[0004] To address the above issues, this invention provides an LDP Down fault location method and apparatus. By collecting and analyzing alarms and protocol status of network elements experiencing LDP Down faults, and initiating tests on the faulty network elements, the cause of the LDP Down fault can be located, reducing the workload of network operation and maintenance personnel and improving operation and maintenance efficiency.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] In one embodiment of the present invention, an LDP Down fault location method is proposed, the method comprising:
[0007] Load network alarm data. When a device network element experiences a port down alarm and an LDP down alarm occurs continuously, determine that the LDP protocol down is caused by the device port down.
[0008] Check if the interface that established the LDP session has been shut down. If the interface has been shut down, undo shutdown on the interface and check if the LDP down has been recovered.
[0009] If the interface is not shut down, check if a command to cancel MPLS-related configurations has been executed. If so, execute the corresponding configuration command to restore the canceled configurations and check if LDP Down has been restored.
[0010] If the command to cancel MPLS-related configurations is not executed, check if the route to the LDP session peer exists. If the route does not exist, handle the IGP routing problem and check if LDP Down has been restored.
[0011] If the route exists, check if the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, check the board status, sub-board status, and the current CPU utilization of the device. If any abnormality is found, handle the abnormality and check if LDP Down has recovered.
[0012] If the LDP Hello-hold timer has not expired, check if the LDP Keepalive-hold timer has expired. If the LDP Keepalive-hold timer has expired, perform a Ping test on the network elements of the devices at both ends of the LDP session. If the Ping test fails, handle the Ping failure and check if the LDP Down has recovered. If the LDP Keepalive-hold timer has not expired, proceed to manual handling.
[0013] Further, check if the LDP Hello-hold timer has timed out, including:
[0014] View the count of Hello messages sent and received at both ends of the LDP session;
[0015] If the send or receive count remains unchanged, it indicates an error in Hello message transmission and reception, and the LDP Hello-hold timer has timed out.
[0016] Furthermore, if the LDP Hello-hold timer times out, check the board status. If a board in the status column is marked as abnormal, and the slot number in the SLOT column is the same as the slot number of the directly connected board, then an anomaly is identified. Check the sub-board status. If the logic_down column or the Init_result column is not marked as successful, then the sub-board status is abnormal. Check the current CPU utilization of the device. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization. Based on the query of the CPU utilization of the application modules within the board, find the modules with high CPU utilization. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization.
[0017] Further, check if the LDP Keepalive-hold timer has timed out, including:
[0018] View the count of Keepalive messages sent and received at both ends of the LDP session;
[0019] If the send or receive count remains unchanged, it indicates an error in Keepalive message transmission and reception, and the LDPKeepalive-hold timer has timed out.
[0020] In one embodiment of the present invention, an LDP Down fault location device is also provided, the device comprising:
[0021] The alarm collection and analysis module is used to load network alarm data. When a device network element has a port down alarm and LDP down alarms occur continuously, it is determined that the LDP protocol down is caused by the device port down.
[0022] The protocol status acquisition and analysis module is used to check whether the interface establishing the LDP session has been shut down. If the interface is shut down, it performs an undo shutdown on the interface and checks whether the LDP down has recovered. If the interface is not shut down, it checks whether a command to cancel MPLS-related configurations has been executed. If so, it executes the corresponding configuration command to restore the canceled configuration and checks whether the LDP down has recovered. If no command to cancel MPLS-related configurations has been executed, it checks whether a route to the LDP session peer exists. If the route does not exist, it handles IGP routing issues and checks whether the LDP down has recovered. If the route exists, it checks whether the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, it checks the board status, sub-board status, and the current CPU utilization of the device. If any abnormalities are found, they are handled, and the LDP down has recovered. If the LDP Hello-hold timer has not expired, it checks whether the LDP Keepalive-hold timer has expired. If the Keepalive-hold timer expires, a Ping test is performed on the network elements of the devices at both ends of the LDP session. If the Ping test fails, the Ping issue is handled and the LDP Down is checked to see if it has recovered. If the LDP Keepalive-hold timer does not expire, manual handling is performed.
[0023] Further, check if the LDP Hello-hold timer has timed out, including:
[0024] View the count of Hello messages sent and received at both ends of the LDP session;
[0025] If the send or receive count remains unchanged, it indicates an error in Hello message transmission and reception, and the LDP Hello-hold timer has timed out.
[0026] Furthermore, if the LDP Hello-hold timer times out, check the board status. If a board in the status column is marked as abnormal, and the slot number in the SLOT column is the same as the slot number of the directly connected board, then an anomaly is identified. Check the sub-board status. If the logic_down column or the Init_result column is not marked as successful, then the sub-board status is abnormal. Check the current CPU utilization of the device. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization. Based on the query of the CPU utilization of the application modules within the board, find the modules with high CPU utilization. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization.
[0027] Further, check if the LDP Keepalive-hold timer has timed out, including:
[0028] View the count of Keepalive messages sent and received at both ends of the LDP session;
[0029] If the send or receive count remains unchanged, it indicates an error in Keepalive message transmission and reception, and the LDPKeepalive-hold timer has timed out.
[0030] In one embodiment of the present invention, a computer device is also proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the aforementioned LDP Down fault location method.
[0031] In one embodiment of the present invention, a computer-readable storage medium is also provided, which stores a computer program for performing the LDP Down fault location method.
[0032] Beneficial effects:
[0033] 1. This invention incorporates expert capabilities, combines alarm information, network protocols, and test commands to automatically diagnose LDP Down faults, and performs real-time analysis and processing of inspection result data to ensure its timeliness.
[0034] 2. This invention provides a process-oriented analysis of LDP Down faults, helping network operations and maintenance personnel quickly locate service faults and reduce their workload.
[0035] 3. The fault location process of this invention is standardized and automated. Even if maintenance personnel are not familiar with the network or routing, they can still complete the fault location, reducing the possibility of human error and omission, improving the accuracy of fault diagnosis, and increasing the efficiency of fault location. Attached Figure Description
[0036] Figure 1 This is a schematic diagram of the LDP Down fault location method of the present invention;
[0037] Figure 2 This is a schematic diagram of the LDP Down fault location device of the present invention;
[0038] Figure 3 This is a schematic diagram of the computer device structure of the present invention. Detailed Implementation
[0039] The principles and spirit of the present invention will now be described with reference to several exemplary embodiments. It should be understood that these embodiments are provided merely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of this disclosure to those skilled in the art.
[0040] Those skilled in the art will recognize that embodiments of the present invention can be implemented as a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0041] According to an embodiment of the present invention, an LDP Down fault location method is proposed. Based on alarm collection and analysis, protocol status collection and analysis, and testing of the faulty network element, the fault location of LDP Down is achieved. The method mainly consists of six steps:
[0042] (1) Load network alarm data: Focus on device port down alarms and LDP down alarms.
[0043] (2) Check whether the interface used to establish the LDP session has been shut down.
[0044] (3) Check whether the command to cancel MPLS-related configurations has been executed.
[0045] (4) Check if the route exists.
[0046] (5) Check if the LDP Hello-hold timer has timed out.
[0047] (6) Check if the LDP Keepalive-hold timer has timed out.
[0048] The principles and spirit of the present invention will be explained in detail below with reference to several representative embodiments.
[0049] Figure 1 This is a schematic flowchart of the LDP Down fault location method of the present invention. Figure 1 As shown, the method includes:
[0050] 1. Load network alarm data. When a device network element experiences a port down alarm and an LDP down alarm occurs continuously, determine that the LDP protocol down is caused by the device port down.
[0051] 2. Check if the interface that established the LDP session has been shut down. If the interface has been shut down, undo shutdown on the interface and check if the LDP down has been restored.
[0052] 3. If the interface is not shut down, check whether the command to cancel MPLS-related configurations has been executed. If the command to cancel MPLS-related configurations has been executed, execute the corresponding configuration command to restore the canceled configurations and check whether LDP Down has been restored.
[0053] 4. If the command to cancel MPLS-related configurations is not executed, check if the route to the LDP session peer exists. If the route does not exist, handle the IGP routing problem and check if LDP Down has been restored.
[0054] 5. If the route exists, check if the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, check the board status, sub-board status, and the current CPU utilization of the device. If any abnormality is found, handle the abnormality and check if LDP Down has recovered.
[0055] 6. If the LDP Hello-hold timer has not expired, check if the LDP Keepalive-hold timer has expired. If the LDP Keepalive-hold timer has expired, perform a Ping test on the network elements of the devices at both ends of the LDP session. If the Ping test fails, handle the Ping failure and check if the LDP Down has recovered. If the LDP Keepalive-hold timer has not expired, proceed to manual handling.
[0056] It should be noted that although the operation of the method of the present invention has been described in a specific order in the above embodiments and figures, this does not require or imply that the operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0057] To provide a clearer explanation of the LDP Down fault location method described above, a specific embodiment will be used for illustration below. However, it is worth noting that this embodiment is only for better illustrating the present invention and does not constitute an improper limitation of the present invention.
[0058] Example:
[0059] S01: Load network alarm data: Focus on device port down alarms and LDP down alarms.
[0060] When a device network element experiences a port down alarm, followed by consecutive (adjustable by built-in settings, typically set to 3 times) LDP down alarms, it can be determined that the LDP protocol down is caused by the device port down.
[0061] Enter S02.
[0062] S02: Check if the interface used to establish the LDP session has been shut down.
[0063] Execute the command "display this" in the interface view of the LDP session. If the word "shutdown" appears, it means that the interface has been shut down.
[0064] If the interface is not shut down, proceed to S03.
[0065] If the interface is shut down, execute the command "undo shutdown" on this interface to start the interface. After execution, check if the LDP down has recovered. If it has recovered, the process ends; if it is still abnormal, proceed to S03.
[0066] How to determine if LDP Down has been restored:
[0067] Execute: "display mpls ldp interface${session interface}" to check if the corresponding Entity Status is Active. If it is Active, it indicates recovery; otherwise, it indicates an error. Entity Status: This displays the status of each interface, module, or service of the device network element, allowing you to understand the operational status of the device network element.
[0068] <xxx>display mpls 1dp interface GigabitEthernet6 / 0 / 0
[0069] LDP Interface Information in Pub1ic Network
[0070] ---------------------------------------------------------------------
[0071] Interface Name:GigabitEthernet6 / 0 / 0
[0072] LDP ID:61.141.29.117:0Transport Address:61.141.29.117
[0073] Entity Status:Active Effective MTU:3000
[0074] Configured He11o Ho1d Timer:15Sec
[0075] Negotiated He11o Hold Timer:15Sec
[0076] Configured He11o Send Timer:---
[0077] Configured Keepalive Ho1d Timer:45Sec
[0078] Configured Keepalive Send Timer:---
[0079] Configured Delay Timer:10Sec
[0080] Label Advertisement Mode:Downstream Unsolicited
[0081] He11o Message Sent / Rcvd:941654 / 942139(Message Count)
[0082] Autoconfiguration Source:---
[0083] mLDP P2MP Capability: Disabled
[0084] mLDP MP2MP Capability: Disabled
[0085] ---------------------------------------------------------------------
[0086] S03: Check if the command to cancel MPLS-related configurations has been executed.
[0087] Execute the command "display current-configuration".
[0088] (1) If the displayed information does not contain "mpls", it means that the MPLS-related configuration has been cancelled. This indicates that the "undo mpls" command was executed.
[0089] (2) If the displayed information does not contain "mpls ldp", it means that the MPLS LDP configuration has been cancelled. This indicates that the "undo mpls ldp" command was executed.
[0090] (3) If the displayed information does not contain "mpls ldp remote-peer", it means that the LDP remote session configuration has been deleted. This indicates that the command "undo mpls ldp remote peer" was executed.
[0091] If the command to cancel MPLS-related configurations is not executed, proceed to S04.
[0092] If a command to cancel MPLS-related configurations is executed, execute the corresponding configuration command to restore the canceled configurations.
[0093] (1)mpls
[0094] (2)mpls ldp
[0095] (3)mpls ldp remote-peer XXX
[0096] After execution, check if LDP Down has recovered. If it has recovered, the process ends; if it is still abnormal, proceed to S04.
[0097] S04: Check if the route exists.
[0098] Execute the command "display ip routing-table",
[0099] Check the Destination / Mask field to see if there is a route to the peer of the session.
[0100] If the route exists, proceed to S05.
[0101] If the route does not exist, rule out IGP routing issues. After ruling out IGP routing issues, check if LDP Down has recovered. If it has recovered, the process ends; if it is still abnormal, proceed to S05.
[0102] S05: Check if the LDP Hello-hold timer has timed out.
[0103] Execute the command "display mpls ldp interface".
[0104] Check if Hello messages are being sent normally at both ends of the session. Execute the command `display mplsldp interface` every 3 seconds to check the count of sent and received Hello messages. If the send or receive count remains unchanged after several consecutive executions of the command, it indicates an abnormality in Hello message transmission and reception, and the Hello-hold timer has timed out. Hello-hold: The "Hello" hold time. "Hello-hold" is a mechanism used to maintain stable neighbor relationships between network devices. In dynamic routing protocols, "Hello" messages are used to discover and maintain neighbor relationships. These protocols periodically send "Hello" messages to confirm the existence and reachability of neighbors. If no "Hello" message is received from a neighbor within a certain period, the neighbor is considered unreachable. "Hello-hold" is the length of time a device will wait if no "Hello" message is received; this time is usually longer than the "Hello" interval. This mechanism can reduce unnecessary routing fluctuations caused by brief network instability or latency.
[0105] If the LDP Hello-hold timer does not time out, proceed to S06.
[0106] If the LDP Hello-hold timer times out:
[0107] S05-1: Check the board status. First, execute `display dev` to check if the board's Status is normal. If it is not normal, determine the cause: board failure.
[0108] Use the display device tool to check the status of individual boards. If a board is marked as abnormal in the status column and the slot number in the SLOT column is the same as the slot number of the directly connected board, then it is considered abnormal.
[0109] S05-2: To check the sub-board status, execute `display device pic-status` and check the `logic_down` or `Init_result` column. If it is not "success", it indicates that the sub-board status is abnormal. `logic_down`: This refers to "logically offline," meaning that an interface or service has been logically and manually shut down. `Init_result`: This refers to the initialization result, used to describe the initialization status after device startup or the application of a configuration item. This status indicates whether the device has successfully completed the startup process or whether the configuration item has been correctly applied.
[0110] S05-3: Check the current CPU utilization of the device by executing `dis cpu-usage`. If the utilization exceeds the system threshold, determine the cause: abnormal CPU utilization.
[0111] S05-4: Locate modules with high CPU utilization. Execute `display cpu-usage slot x` to find modules with high CPU utilization based on the CPU utilization of application modules within the board. If the utilization exceeds the system threshold, determine the cause: abnormal CPU utilization.
[0112] <xxx>display cpu-usage slot 1
[0113] Cpu utilization statistics at 2023-02-27 14:48:49 350msSystem epu userate is:6%
[0114] Cpu utilization for five seconds:6%; one minute:6%; five minutes:6%.MaxCPU Usage:40%
[0115] Max CPU Usage stat.Time:2018-09-2100:00:37018ms
[0116]
[0117] CPU Usage Details are shown in Table 1 below:
[0118] Table 1
[0119]
[0120]
[0121] If any abnormalities are found during the above checks, handle the abnormalities. After handling, check if LDP Down has recovered. If it has recovered, the process ends; if it is still abnormal, proceed to S06.
[0122] S06: Check if the LDP Keepalive-hold timer has timed out.
[0123] Execute the command "display mpls ldp session".
[0124] Check if Keepalive messages are being sent normally at both ends of the session. Execute the command `displaympls ldp session` every 5 seconds to check the count of sent and received Keepalive messages. If the send or receive count remains unchanged after several consecutive executions of the command, it indicates an abnormality in Keepalive message transmission and reception, and the Keepalive-hold timer has timed out. "Keepalive-hold": This is a mechanism used to maintain the active state of a session or connection. "Keepalive" messages are used to detect and maintain whether an established connection is still valid. These messages are sent periodically to ensure that both ends of the connection can still communicate, even without data transmission. If the receiving end does not receive a "Keepalive" message within a predetermined time, it may assume the connection has been broken. "Keepalive-hold" specifies how long the system will wait before considering the connection invalid if no "Keepalive" message is received. This time is usually longer than the "Keepalive" interval to ensure that the connection is not mistakenly closed due to brief network delays or fluctuations.
[0125] If the Keepalive-hold timer does not time out, proceed to manual processing.
[0126] If the Keepalive-hold timer times out, perform a Ping test on the network elements of the devices at both ends of the session.
[0127] If the ping test is normal, the conclusion is: the other end is reachable by ping, and manual processing is required.
[0128] If the ping test is abnormal, the conclusion is: the ping test on the other end is unsuccessful.
[0129] If the ping test fails, address the ping failure issue. After addressing the issue, check if the LDP down condition has been resolved. If it has, the process ends; otherwise, proceed to manual intervention.
[0130] Based on the same inventive concept, this invention also proposes an LDP Down fault location device. The implementation of this device can refer to the implementation of the method described above, and repeated details will not be repeated. The term "module" used below can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0131] Figure 2 This is a schematic diagram of the LDP Down fault location device of the present invention. Figure 2 As shown, the device includes:
[0132] The alarm collection and analysis module 101 is used to load network alarm data. When a device network element experiences a port down alarm and an LDP down alarm occurs continuously, it is determined that the LDP protocol down is caused by the device port down.
[0133] The protocol status acquisition and analysis module 102 is used to check whether the interface establishing the LDP session has been shut down. If the interface is shut down, it performs an undo shutdown on the interface and checks whether the LDP down has recovered. If the interface is not shut down, it checks whether a command to cancel MPLS-related configurations has been executed. If such a command is executed, it executes the corresponding configuration command to restore the canceled configuration and checks whether the LDP down has recovered. If no command to cancel MPLS-related configurations has been executed, it checks whether a route to the LDP session peer exists. If the route does not exist, it handles IGP routing issues and checks whether the LDP down has recovered. If the route exists, it checks whether the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, it checks the board status, sub-board status, and the current CPU utilization of the device. If any abnormalities are found, they are handled, and the LDP down has recovered. If the LDP Hello-hold timer has not expired, it checks whether the LDP Keepalive-hold timer has expired. If the Keepalive-hold timer expires, a Ping test is performed on the network elements of the devices at both ends of the LDP session. If the Ping test fails, the Ping failure is handled, and the LDP Down is checked to see if it has recovered. If the LDP Keepalive-hold timer does not expire, manual handling is performed.
[0134] Check if the LDP Hello-hold timer has timed out, including:
[0135] View the count of Hello messages sent and received at both ends of the LDP session;
[0136] If the send or receive count remains unchanged, it indicates an error in Hello message transmission and reception, and the LDP Hello-hold timer has timed out.
[0137] If the LDP Hello-hold timer times out, check the board status. If a board in the status column is marked as abnormal, and the slot number in the SLOT column is the same as the slot number of the directly connected board, then an anomaly is identified. Check the sub-board status. If the logic_down or Init_result column is not marked as successful, then the sub-board status is abnormal. Check the current CPU utilization of the device. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization. Based on the query of the application module CPU utilization within the board, find the module with high CPU utilization. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization.
[0138] Check if the LDP Keepalive-hold timer has timed out, including:
[0139] View the count of Keepalive messages sent and received at both ends of the LDP session;
[0140] If the send or receive count remains unchanged, it indicates an error in Keepalive message transmission and reception, and the LDPKeepalive-hold timer has timed out.
[0141] It should be noted that although several modules of the LDP Down fault location device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.
[0142] Based on the aforementioned inventive concept, such as Figure 3 As shown, the present invention also proposes a computer device 200, including a memory 210, a processor 220, and a computer program 230 stored in the memory 210 and executable on the processor 220. When the processor 220 executes the computer program 230, it implements the aforementioned LDP Down fault location method.
[0143] Based on the aforementioned inventive concept, the present invention also proposes a computer-readable storage medium storing a computer program that executes the aforementioned LDP Down fault location method.
[0144] The LDP Down fault location method and apparatus proposed in this invention have the following advantages:
[0145] 1. This device incorporates expert capabilities and automatically performs fault diagnosis of LDP Down by combining alarm information, network protocols, and test commands. It also performs real-time analysis and processing of the inspection results data to ensure its timeliness.
[0146] 2. Conduct process-oriented analysis and handling of LDP Down faults to help network operations and maintenance personnel quickly locate the business faults and reduce their workload.
[0147] 3. The fault location process is standardized and automated, reducing the possibility of human error and omission, improving the accuracy of fault diagnosis, and increasing the efficiency of fault location.
[0148] While the spirit and principles of the invention have been described with reference to several specific embodiments, it should be understood that the invention is not limited to the disclosed specific embodiments, and the division of aspects does not imply that features in these aspects cannot be combined for benefit; such division is merely for ease of description. The invention is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
[0149] Regarding the limitation of the scope of protection of this invention, those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solution of this invention are still within the scope of protection of this invention.< / xxx> < / xxx>
Claims
1. A method for locating LDP Down faults, characterized in that, The method includes: Load network alarm data. When a device network element experiences a port down alarm and an LDP down alarm occurs continuously, determine that the LDP protocol down is caused by the device port down. Check if the interface that established the LDP session has been shut down. If the interface has been shut down, undo shutdown on the interface and check if the LDP down has been recovered. If the interface is not shut down, check if a command to cancel MPLS-related configurations has been executed. If so, execute the corresponding configuration command to restore the canceled configurations and check if LDP Down has been restored. If the command to cancel MPLS-related configurations is not executed, check if the route to the LDP session peer exists. If the route does not exist, handle the IGP routing problem and check if LDP Down has been restored. If the route exists, check if the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, check the board status, sub-board status, and the current CPU utilization of the device. If any abnormality is found, handle the abnormality and check if LDP Down has recovered. If the LDP Hello-hold timer has not expired, check if the LDP Keepalive-hold timer has expired. If the LDP Keepalive-hold timer has expired, perform a Ping test on the network elements of the devices at both ends of the LDP session. If the Ping test fails, handle the Ping failure and check if the LDP Down has recovered. If the LDP Keepalive-hold timer has not expired, proceed to manual handling.
2. The LDP Down fault location method according to claim 1, characterized in that, Check if the LDP Hello-hold timer has timed out, including: View the count of Hello messages sent and received at both ends of the LDP session; If the send or receive count remains unchanged, it indicates an error in Hello message transmission and reception, and the LDP Hello-hold timer has timed out.
3. The LDP Down fault location method according to claim 2, characterized in that, If the LDP Hello-hold timer times out, check the board status. If a board in the status column is marked as abnormal, and the slot number in the SLOT column is the same as the slot number of the directly connected board, then an anomaly is identified. Check the sub-board status. If the logic_down or Init_result column is not marked as successful, then the sub-board status is abnormal. Check the current CPU utilization of the device. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization. Based on the query of the application module CPU utilization within the board, find the module with high CPU utilization. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization.
4. The LDP Down fault location method according to claim 1, characterized in that, Check if the LDP Keepalive-hold timer has timed out, including: View the count of Keepalive messages sent and received at both ends of the LDP session; If the send or receive count remains unchanged, it indicates an error in Keepalive message transmission and reception, and the LDPKeepalive-hold timer has timed out.
5. An LDP Down fault location device, characterized in that, The device includes: The alarm collection and analysis module is used to load network alarm data. When a device network element has a port down alarm and LDP down alarms occur continuously, it is determined that the LDP protocol down is caused by the device port down. The protocol status acquisition and analysis module is used to check whether the interface establishing the LDP session has been shut down. If the interface is shut down, it performs an undo shutdown on the interface and checks whether the LDP down has recovered. If the interface is not shut down, it checks whether a command to cancel MPLS-related configurations has been executed. If so, it executes the corresponding configuration command to restore the canceled configuration and checks whether the LDP down has recovered. If no command to cancel MPLS-related configurations has been executed, it checks whether a route to the LDP session peer exists. If the route does not exist, it handles IGP routing issues and checks whether the LDP down has recovered. If the route exists, it checks whether the LDP Hello-hold timer has expired. If the LDP Hello-hold timer has expired, it checks the board status, sub-board status, and the current CPU utilization of the device. If any abnormalities are found, they are handled, and the LDP down has recovered. If the LDP Hello-hold timer has not expired, it checks whether the LDP Keepalive-hold timer has expired. If the Keepalive-hold timer expires, a Ping test is performed on the network elements of the devices at both ends of the LDP session. If the Ping test fails, the Ping issue is handled and the LDP Down is checked to see if it has recovered. If the LDP Keepalive-hold timer does not expire, manual handling is performed.
6. The LDP Down fault location device according to claim 5, characterized in that, Check if the LDP Hello-hold timer has timed out, including: View the count of Hello messages sent and received at both ends of the LDP session; If the send or receive count remains unchanged, it indicates an error in Hello message transmission and reception, and the LDP Hello-hold timer has timed out.
7. The LDP Down fault location device according to claim 6, characterized in that, If the LDP Hello-hold timer times out, check the board status. If a board in the status column is marked as abnormal, and the slot number in the SLOT column is the same as the slot number of the directly connected board, then an anomaly is identified. Check the sub-board status. If the logic_down or Init_result column is not marked as successful, then the sub-board status is abnormal. Check the current CPU utilization of the device. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization. Based on the query of the application module CPU utilization within the board, find the module with high CPU utilization. If the utilization exceeds the system threshold, then the cause is identified as abnormal CPU utilization.
8. The LDP Down fault location device according to claim 5, characterized in that, Check if the LDP Keepalive-hold timer has timed out, including: View the count of Keepalive messages sent and received at both ends of the LDP session; If the send or receive count remains unchanged, it indicates an error in Keepalive message transmission and reception, and the LDPKeepalive-hold timer has timed out.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1-4.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that performs the method according to any one of claims 1-4.
Citation Information
Patent Citations
IPRAN network fault positioning method and device
CN110650041A
IPRAN cloud private line fault positioning method and device
CN112468335A