Fault analysis method and device and related equipment
By determining the source and destination of network failures, combined with the online status and path analysis of nodes, the problem of inaccurate fault location in the existing technology is solved, and fast and accurate failure root cause analysis and visual display are achieved.
Patent Information
- Application Number
- CN202510399237.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
When the existing technology faces network failures in complex network environments, it is difficult for operation and maintenance personnel to quickly locate the root cause of the failure. The existing alarm methods lack intuitiveness and targeting, and require manual analysis.
By determining the source and destination of the path on-off fault, as the current node, it determines its online status, if it is not online, it will be recorded as a target fault, and it will traverse the next hop device until the final network device is found, and the root cause of the fault is inferred based on the path and alarm data.
It realizes that operation and maintenance personnel quickly lock network failure points, reduces manual analysis processes, improves the efficiency and accuracy of fault location, and provides a visual display of network on and off.
Smart Images

Figure CN120263620A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network communication technologies, and particularly to a fault analysis method, apparatus, and related devices. Background Art
[0002] With the development and popularization of networks, networking solutions have become increasingly complex and the networking scale has become increasingly large. For example, in campus networking solutions, wired devices, wireless devices, and even all-optical devices are generally involved. In scenarios such as campuses, the device scale is often in the tens of thousands, and the terminal scale is even larger. Facing various faults in the network, existing operation and maintenance systems usually prompt operation and maintenance personnel in the form of alarms, but basically single-point fault alarms are presented in a list form. When a connection or disconnection fault occurs in a certain device or link in the network, it usually triggers a series of fault alarms for related devices. It is very difficult for operation and maintenance personnel to determine the root cause of the problem in the first place when facing a large number of alarms, and they need to manually analyze and confirm these fault alarms to locate the root cause. Summary of the Invention
[0003] This application provides a fault analysis method, apparatus, and related devices.
[0004] In a first aspect, this application provides a fault analysis method, and the method includes:
[0005] Determine the source end and the destination end of a path connection or disconnection fault, where the source end is the first network device in the network managed by the controller that accesses a target terminal unable to access a target service, and the destination end is the second network device in the network managed by the controller that accesses a server carrying the target service;
[0006] Use the first network device as the current node, and determine whether the current node is online;
[0007] If it is determined that the current node is not online, record the current node as a target fault, and determine whether there is an accessible path between the current node and the destination end;
[0008] If it is determined that there is no accessible path between the current node and the destination end, use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online until the latest current node is the second network device.
[0009] Optionally, the method further includes:
[0010] If it is determined that the current node is not online, determine whether the link between the current node and the next-hop network device of the current node in the path is faulty;
[0011] If it is determined that there is a link failure between the current node and the next-hop network device of the current node in the path, record the link as the target failure, determine the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online.
[0012] Optionally, the method further includes:
[0013] If it is determined that there is a reachable path between the current node and the destination end, determine the recorded target failure as the root cause of the path on / off failure.
[0014] Optionally, the steps of determining the source end and the destination end of the path on / off failure include:
[0015] Based on the failure information that the target terminal cannot access the target service input by the administrator, the first network device accessing the target terminal in the network architecture managed by the controller is the source end, and the second network device accessing the server carrying the target service in the network architecture managed by the controller is the destination end.
[0016] In a second aspect, the present application provides a fault analysis device, and the device includes:
[0017] A determination unit, configured to determine the source end and the destination end of the path on / off failure, where the source end is the first network device accessing the target terminal that cannot access the target service in the network architecture managed by the controller, and the destination end is the second network device accessing the server carrying the target service in the network architecture managed by the controller;
[0018] A judgment unit, configured to use the first network device as the current node and judge whether the current node is online;
[0019] A recording unit, if the judgment unit determines that the current node is not online, the recording unit is configured to record the current node as the target failure; the judgment unit is further configured to judge whether there is a reachable path between the current node and the destination end;
[0020] If the judgment unit determines that there is no reachable path between the current node and the destination end, the judgment unit is further configured to use the next-hop network device of the current node in the path as the latest current node, and execute the step of judging whether the current node is online until the latest current node is the second network device.
[0021] Optionally, if the judgment determines that the current node is not online, the judgment unit is further configured to judge whether the link between the current node and the next-hop network device of the current node in the path is faulty;
[0022] If the determination unit determines that there is a link failure between the current node and the next-hop network device of the current node in the path, the recording unit is further configured to record this link as the target failure;
[0023] The determination unit is further configured to use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online.
[0024] Optionally, if the determination unit determines that there is a reachable path between the current node and the destination end, the determination unit is further configured to determine the recorded target failure as the root cause of the path on / off failure.
[0025] Optionally, when determining the source end and the destination end of the path on / off failure, the determination unit is specifically configured to:
[0026] Based on the failure information that the target terminal cannot access the target service input by the administrator, the first network device accessing the target terminal in the network managed by the controller is used as the source end, and the second network device accessing the server carrying the target service in the network managed by the controller is used as the destination end.
[0027] In a third aspect, an embodiment of the present application provides a fault analysis device, and the fault analysis device includes:
[0028] A memory for storing program instructions;
[0029] A processor for calling the program instructions stored in the memory and executing the steps of the method according to any one of the above first aspects according to the obtained program instructions.
[0030] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores computer-executable instructions for causing the computer to execute the steps of the method according to any one of the above first aspects.
[0031] In summary, the fault analysis method provided by the embodiment of the present application determines the source end and the destination end of the path on-off fault. Among them, the source end is the first network device that accesses the target terminal that cannot access the target service in the network managed by the controller, and the destination end is the second network device that accesses the server carrying the target service in the network managed by the controller; taking the first network device as the current node, and determining whether the current node is online; if it is determined that the current node is not online, recording the current node as the target fault, and determining whether there is a reachable path between the current node and the destination end; if it is determined that there is no reachable path between the current node and the destination end, taking the next-hop network device of the current node in the path as the latest current node, and executing the step of determining whether the current node is online until the latest current node is the second network device.
[0032] By using the fault analysis method provided by the embodiment of the present application, combined with specific fault scenarios, through end-to-end on-off troubleshooting, combined with the path and alarms, the ultimate root cause of the source node network disconnection can be deduced, which is convenient for the operation and maintenance personnel to directly lock the fault point and save the process of manual analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained according to these drawings of the embodiments of the present application.
[0034] Figure 1 It is a detailed flowchart of a fault analysis method provided by an embodiment of the present application;
[0035] Figure 2 It is a schematic diagram of a fault analysis process provided by an embodiment of the present application;
[0036] Figure 3 It is a schematic structural diagram of a fault analysis device provided by an embodiment of the present application;
[0037] Figure 4 It is a schematic diagram of the hardware structure of a fault analysis device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and do not limit the present application. The singular forms "a", "said", and "the" used in the present application and the claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to any or all possible combinations of one or more of the associated listed items.
[0039] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, in addition, the word "if" used may be interpreted as "when" or "while" or "in response to determining".
[0040] In the related art, as Figure 1 shown, if there is a backbone fiber failure of the OLT device in the network, a series of on-off fault alarms will be triggered under this fiber link, such as AP disconnection, ONU disconnection, etc. Generally, a single OLT backbone fiber can access dozens or hundreds of AP devices. If these alarms are only displayed in the form of a list in front of the operation and maintenance personnel, it is very difficult for the operation and maintenance personnel to determine the location of the fault point in the first time.
[0041] Currently, in addition to the alarm display in the form of a list, the network topology can also be visually displayed through the way of network fault visualization, and combined with the alarm data, the alarm data is displayed in the tip information of the device, which is convenient for the operation and maintenance personnel to intuitively analyze the network fault.
[0042] However, the device alarms are diverse and generally divided into many levels, such as emergency, important, minor, warning, etc. If the device alarms are simply displayed in the form of a visual topology, the operation and maintenance personnel still need to click on each device to view and analyze the alarm data of each device to draw a final conclusion. That is, the disadvantage of this technology is still not intuitive enough and lacks pertinence. It cannot visually display the network on-off fault and still requires manual analysis. For example, whether the fault point occurs on the entire device or a certain port of the device. Among a series of fault alarms, which one is the root cause of the network on-off. All these require manual analysis.
[0043] Exemplarily, referring to Figure 1 shown, it is a detailed flowchart of a fault analysis method provided by the embodiments of the present application, and the method includes the following steps:
[0044] Step 100: Determine the source end and the destination end of the path on-off fault.
[0045] In the embodiments of the present application, the source end is the first network device that accesses the target terminal that cannot access the target service in the network managed by the controller, and the destination end is the second network device that accesses the server carrying the target service in the network managed by the controller.
[0046] Specifically, when determining the source end and the destination end of the path connection fault, a preferred implementation is as follows:
[0047] Based on the fault information that the target terminal cannot access the target service input by the administrator, the first network device that accesses the target terminal in the network managed by the controller is used as the source end, and the second network device that accesses the server carrying the target service in the network managed by the controller is used as the destination end.
[0048] In practical applications, in an actual network, there may be multiple local faults. Therefore, when troubleshooting network connection and disconnection, it must be scenario-based, that is, troubleshooting the connection and disconnection for the source end and the destination end where the fault occurs.
[0049] For example, end-user Zhang San cannot access the network in the office area at noon one day. In this way, troubleshooting of the connection and disconnection can be carried out for the local topology from Zhang San (user terminal) to the egress switch.
[0050] For another example, end-user Zhang San cannot access the database service one day. At this time, troubleshooting of the connection and disconnection can be carried out for the network path from the network device (source end) accessed by the user terminal used by Zhang San in the network to the network device (destination end) accessing the database server in the network.
[0051] Furthermore, based on the relevant information of the fault path that the administrator can input, the source end and the destination end of the fault path in the network can be determined. Other methods can also be used to monitor the fault path to determine the source end and the destination end of the fault path. In the embodiments of the present application, no specific limitation is made here.
[0052] Step 110: Use the first network device as the current node and determine whether the current node is online.
[0053] Specifically, the controller maintains the device status of each network device in the network it manages. After determining the source end and the destination end of the fault path, starting from the source end, use the source end as the current node and determine the node status of the current node, that is, determine whether the current node is online, so as to determine whether the current node is the root cause of the path fault.
[0054] Step 120: If it is determined that the current node is not online, record the current node as the target fault and determine whether there is an accessible path between the current node and the destination end.
[0055] Specifically, when it is determined that the current node is not online (i.e., disconnected), record the fault that the current node is not online, and continue to determine whether there is a reachable path between the current node and the destination end.
[0056] If it is determined that there is a reachable path between the current node and the destination end, then determine the recorded target fault as the root cause of the path on / off fault.
[0057] That is to say, there is a reachable path between the current node and the destination end, but there is a fault in the path. The root cause of the fault is that the current node is not online.
[0058] Determine the fault that the current node is not online as the root cause of the path fault.
[0059] Step 130: If it is determined that there is no reachable path between the current node and the destination end, then use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online until the latest current node is the second network device.
[0060] Specifically, if it is determined that there is no reachable path between the current node and the destination end, it means that there are other reasons for the path fault later. At this time, continue to execute the fault analysis method:
[0061] Specifically, use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online. In the embodiments of the present application, this will not be elaborated here.
[0062] Furthermore, in the embodiments of the present application, the above fault analysis method further includes the following steps:
[0063] If it is determined that the current node is not online, then determine whether the link between the current node and the next-hop network device of the current node in the path is faulty.
[0064] That is to say, if the current node is not online, it is necessary to further determine whether the link between this node and the next-hop network device is normal. At this time, the following operations can be continued:
[0065] Determine whether the link between the current node and the next-hop network device of the current node in the path is faulty.
[0066] Specifically, if it is determined that the link between the current node and the next-hop network device of the current node in the path is faulty, record this link as the target fault, and use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online.
[0067] That is, if it is determined that the link between the current node and the next-hop network device of the current node in the path is abnormal, the link failure is determined as the root cause of the path failure.
[0068] That is, in an embodiment of the present application, the path from the source to the target and all on / off alarms are traversed to find the last (or multiple) device or link that causes the path to be blocked, mark the associated alarm object as a fault point, and display the root cause of the fault of the alarm.
[0069] By classifying the alarms of the operation and maintenance system and associating them with the network topology, the network connectivity can be accurately visualized. The operation and maintenance personnel can query the network topology within any range and see where the network is disconnected very intuitively, so as to achieve network connectivity visualization.
[0070] The following describes in detail the fault analysis process provided by the embodiment of the present application in combination with specific application scenarios. Figure 2 FIG. 1 is a schematic diagram of a fault analysis process provided in an embodiment of the present application. The fault analysis process includes the following steps:
[0071] Get fault path information;
[0072] Start traversing from the source end and determine whether the node currently traversed has a complete machine failure;
[0073] If it is determined that there is a whole machine failure, the failure point is recorded and it is determined whether there is a reachable path to the destination.
[0074] If it is determined that there is a reachable path to the destination, the whole machine failure is determined as the root cause of the path failure, and the process ends;
[0075] If it is determined that there is no reachable path to the destination, the next hop network device is used as the current node, and the step of determining whether the currently traversed node has a complete machine failure is performed;
[0076] Furthermore, it is also possible to determine whether the current node has a link failure of the front line;
[0077] If it is determined that the current node has a link failure of the front conductor, the fault point is recorded, and the link failure of the front conductor is determined as the root cause of the path failure;
[0078] If it is determined that there is no link failure of the front line at the current node, the process ends.
[0079] In this way, through the path analysis from the source to the destination, combined with the local topology path, and the alarms of the nodes and lines on the path, the root cause of the network disconnection of the source device (or terminal) can be found.
[0080] For example, see Figure 3As shown in the figure, it is a schematic structural diagram of a fault analysis device provided by an embodiment of the present application. The device includes:
[0081] A determination unit, configured to determine the source end and the destination end of the path on-off fault, where the source end is the first network device that accesses the target terminal that cannot access the target service in the network managed by the controller, and the destination end is the second network device that accesses the server carrying the target service in the network managed by the controller;
[0082] A judgment unit, configured to use the first network device as the current node and judge whether the current node is online;
[0083] A recording unit, if the judgment unit determines that the current node is not online, then the recording unit is configured to record the current node as the target fault; the judgment unit is further configured to judge whether there is an accessible path between the current node and the destination end;
[0084] If the judgment unit determines that there is no accessible path between the current node and the destination end, then the judgment unit is further configured to use the next-hop network device of the current node in the path as the latest current node and execute the step of judging whether the current node is online until the latest current node is the second network device.
[0085] Optionally, if the judgment determines that the current node is not online, then the judgment unit is further configured to judge whether the link between the current node and the next-hop network device of the current node in the path is faulty;
[0086] If the judgment unit determines that the link between the current node and the next-hop network device of the current node in the path is faulty, then the recording unit is further configured to record the link as the target fault;
[0087] The judgment unit is further configured to use the next-hop network device of the current node in the path as the latest current node and execute the step of judging whether the current node is online.
[0088] Optionally, if the judgment unit determines that there is an accessible path between the current node and the destination end, then the determination unit is further configured to determine the recorded target fault as the root cause of the path on-off fault.
[0089] Optionally, when determining the source end and the destination end of the path on-off fault, the determination unit specifically is configured to:
[0090] Based on the fault information that the target terminal cannot access the target service input by the administrator, use the first network device that accesses the target terminal in the network managed by the controller as the source end, and use the second network device that accesses the server carrying the target service in the network managed by the controller as the destination end.
[0091] The above units may be one or more integrated circuits configured to implement the above methods. For example, one or more Application Specific Integrated Circuits (ASICs), or one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain unit above is implemented in the form of a processing element scheduling program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these units may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0092] Furthermore, for the fault analysis device provided in the embodiments of the present application, from a hardware perspective, a schematic diagram of the hardware architecture of the fault analysis device can be seen in Figure 4 as shown. The fault analysis device may include: a memory 40 and a processor 41,
[0093] The memory 40 is used to store program instructions; the processor 41 calls the program instructions stored in the memory 40 and executes the above method embodiments according to the obtained program instructions. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0094] Optionally, the present application further provides a controller, including at least one processing element (or chip) for executing the above method embodiments.
[0095] Optionally, the present application further provides a program product, such as a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to cause the computer to execute the above method embodiments.
[0096] Here, the machine-readable storage medium may be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium may be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or a combination thereof.
[0097] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.
[0098] For the convenience of description, the above devices are described by dividing them into various units according to their functions. Of course, when implementing this application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0099] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0100] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0101] Moreover, these computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0102] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 in one block or multiple blocks.
[0103] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A fault analysis method, characterized in that, The method includes: Determine the source end and the destination end of the path on-off fault, where the source end is the first network device in the network managed by the controller that accesses the target terminal unable to access the target service, and the destination end is the second network device in the network managed by the controller that accesses the server carrying the target service; Take the first network device as the current node and determine whether the current node is online; If it is determined that the current node is not online, record the current node as the target fault and determine whether there is a reachable path between the current node and the destination end; If it is determined that there is no reachable path between the current node and the destination end, take the next-hop network device of the current node in the path as the latest current node and execute the step of determining whether the current node is online until the latest current node is the second network device.
2. The method according to claim 1, wherein The method further includes: If it is determined that the current node is not online, determine whether the link between the current node and the next-hop network device of the current node in the path is faulty; If it is determined that the link between the current node and the next-hop network device of the current node in the path is faulty, record the link as the target fault, determine the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online.
3. The method according to claim 1 or 2, characterized in that, The method further includes: If it is determined that there is a reachable path between the current node and the destination end, determine the recorded target fault as the root cause of the path on-off fault.
4. The method according to claim 1, characterized in that, The step of determining the source end and the destination end of the path on-off fault includes: Based on the fault information that the target terminal cannot access the target service input by the administrator, take the first network device in the network managed by the controller that accesses the target terminal as the source end, and take the second network device in the network managed by the controller that accesses the server carrying the target service as the destination end.
5. A fault analysis device, characterized in that, The device includes: A determination unit, configured to determine the source end and the destination end of the path on-off fault, where the source end is the first network device in the network managed by the controller that accesses the target terminal unable to access the target service, and the destination end is the second network device in the network managed by the controller that accesses the server carrying the target service; A judgment unit, configured to take the first network device as the current node and determine whether the current node is online; A recording unit, if the judgment unit determines that the current node is not online, the recording unit is configured to record the current node as the target fault; the judgment unit is further configured to determine whether there is a reachable path between the current node and the destination end; If the judgment unit determines that there is no reachable path between the current node and the destination end, the judgment unit is further configured to take the next-hop network device of the current node in the path as the latest current node and execute the step of determining whether the current node is online until the latest current node is the second network device.
6. The device according to claim 5, wherein If the determination determines that the current node is not online, the determination unit is further configured to determine whether the link between the current node and the next-hop network device of the current node in the path is faulty; If the determination unit determines that the link between the current node and the next-hop network device of the current node in the path is faulty, the recording unit is further configured to record the link as the target fault; The determination unit is further configured to use the next-hop network device of the current node in the path as the latest current node, and execute the step of determining whether the current node is online.
7. The apparatus according to claim 5 or 6, wherein If the determination unit determines that there is an accessible path between the current node and the destination end, the determination unit is further configured to determine the recorded target fault as the root cause of the path on / off fault.
8. The device according to claim 5, characterized in that, When determining the source end and the destination end of the path on / off fault, the determination unit is specifically configured to: Based on the fault information that the target terminal cannot access the target service input by the administrator, the first network device accessing the target terminal in the network configured by the controller is used as the source end, and the second network device accessing the server carrying the target service in the network configured by the controller is used as the destination end.
9. A fault analysis device, characterized in that, The fault analysis apparatus includes: a memory for storing program instructions; a processor for calling the program instructions stored in the memory and executing the steps of the method according to any one of claims 1-4 according to the obtained program instructions.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing the computer to execute the steps of the method according to any one of claims 1-4.