A device off-pipe fault analysis method and system, electronic device, and storage medium
By constructing a network topology and analyzing the number and location of abnormal devices, the system can quickly locate device disconnection faults, solving the problem of difficulty in identifying device disconnection faults in existing technologies and improving the efficiency and reliability of network operation and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- FIBERHOME TELECOMMUNICATION TECHNOLOGIES CO LTD
- Filing Date
- 2024-11-27
- Publication Date
- 2026-05-15
AI Technical Summary
Existing network management systems and technologies are insufficient to quickly identify and resolve equipment disconnection faults, leading to fault propagation and impacting network services and customer experience.
By scanning the device routing information within the management area, a network topology is constructed to identify abnormal devices. Based on the number and location information of abnormal devices, the location of the device out of management fault can be quickly located. The number of abnormal devices can be used to determine the out-of-management status of the device, thus realizing intelligent diagnosis and rapid troubleshooting.
It enables rapid identification of device disconnection faults, improves troubleshooting efficiency, reduces fault propagation, and enhances the efficiency and reliability of network operation and maintenance.
Smart Images

Figure CN119697001B_ABST
Abstract
Description
Technical Field
[0001] This disclosure belongs to the field of network operation and maintenance technology, and in particular relates to a method, system, electronic device and storage medium for analyzing equipment disconnection faults. Background Technology
[0002] As operator network topologies and services become more complex and diversified, the complexity of technology is growing exponentially, and network operation and maintenance are becoming increasingly complicated and difficult. Operation and maintenance personnel must deal with a large number of alarm messages generated by various devices, analyze the possible causes of failures, and manually troubleshoot, which is time-consuming, labor-intensive, and inefficient.
[0003] Existing network management systems and technologies are no longer sufficient to provide adequate support for operations and maintenance personnel, resulting in many faults not being detected and resolved in a timely manner. Faults continue to spread and escalate, eventually affecting business and customer experience, especially equipment detachment faults, which can lead to regional network failures in severe cases. Summary of the Invention
[0004] To address the aforementioned issues, this disclosure provides a method, system, electronic device, and storage medium for analyzing device disconnection faults. This solution locates the device disconnection fault by identifying the abnormal device's position within the network topology, enabling rapid troubleshooting.
[0005] To address the aforementioned technical problems, the first aspect of this invention proposes a method for analyzing equipment disconnection faults, the method comprising:
[0006] Scan the routing information of devices within the management area and construct the network topology based on the routing information;
[0007] Based on the routing status of each of the aforementioned devices, identify the abnormal devices that have experienced anomalies;
[0008] The device disconnection status is determined based on the number of abnormal devices, and the location of the device disconnection fault is determined based on the location information of the abnormal device in the network topology.
[0009] According to a preferred embodiment of the present invention, the step of determining the device disconnection status based on the number of abnormal devices, and determining the device disconnection fault location based on the device disconnection status using the location information of the abnormal devices in the network topology, includes:
[0010] When the number of abnormal devices exceeds a first preset number, the device disconnection status is a large-scale device disconnection status. The number of abnormal devices belonging to the first station device in the network topology is determined according to the location information. The first station device includes: devices directly connected to the data center network.
[0011] If the number of faulty devices belonging to the first station is greater than the second preset number, the fault location is determined to be the data center network;
[0012] If the number of abnormal devices belonging to the first station is not greater than the second preset number, the location of the device disconnection fault is determined based on the routing information of each abnormal device.
[0013] According to a preferred embodiment of the present invention, if the number of abnormal devices belonging to the first station equipment is not greater than a second preset number, then determining the location of the equipment disconnection fault based on the routing information of each abnormal device includes:
[0014] If the number of abnormal devices belonging to the first station device is not greater than the second preset number and is not equal to 0, determine whether each abnormal device has passed through the abnormal device belonging to the first station device based on the routing information of each abnormal device; if so, determine that the abnormal device belonging to the first station device has a disconnection fault.
[0015] If the number of abnormal devices belonging to the first station device is equal to 0, obtain the first routing information of the abnormal device before the abnormality and the second routing information after the abnormality.
[0016] Based on the first routing information and the second routing information, the fault location of the abnormal device is determined in the network topology.
[0017] According to a preferred embodiment of the present invention, the step of determining the device disconnection status based on the number of abnormal devices, and determining the device disconnection fault location based on the device disconnection status using the location information of the abnormal devices in the network topology, includes:
[0018] When the number of abnormal devices is not greater than a first preset number, obtain the current routing information of each abnormal device;
[0019] Determine the node IP of the last hop device that is alive in the routing path based on the current routing information;
[0020] When the node IP belongs to a device within the management area, log in directly to the last-hop device and determine the device that has experienced a disconnection fault based on the routing status of the neighbor routes of the last-hop device.
[0021] If the node IP does not belong to a device within the management area, the fault location is determined to be the data center network.
[0022] According to a preferred embodiment of the present invention, before determining the device disconnection status based on the number of abnormal devices, and determining the device disconnection fault location based on the device disconnection status using the location information of the abnormal devices in the network topology, the analysis method further includes:
[0023] Obtain alarm data from each of the abnormal devices, analyze the alarm data, and determine whether the alarm data is a momentary interruption alarm;
[0024] When the alarm data is a momentary interruption alarm, the corresponding abnormal device has not experienced a disconnection fault.
[0025] According to a preferred embodiment of the present invention, the analysis method further includes:
[0026] Real-time acquisition of the device status and connectivity of each of the aforementioned devices;
[0027] When the device status or connectivity of any of the devices changes, the device whose device status or connectivity has changed is designated as the first updating device, and the device associated with the first updating device is designated as the second updating device.
[0028] The routing information of the first and second updating devices is obtained by scanning, and the network topology is updated accordingly.
[0029] According to a preferred embodiment of the present invention, the step of using the device associated with the first updating device as the second updating device includes:
[0030] The device that communicates with the first updating device in the network topology is selected as the second updating device.
[0031] To address the aforementioned technical problems, a second aspect of the present invention provides a system for analyzing equipment disconnection faults, the system comprising:
[0032] The network topology construction module is used to scan the routing information of devices within the management area and construct the network topology based on the routing information.
[0033] The abnormal device analysis module is used to determine the abnormal devices that have anomalies based on the routing status of each device.
[0034] The device disconnection analysis module is used to determine the device disconnection status based on the number of abnormal devices, and based on the device disconnection status, determine the location of the device disconnection fault by using the location information of the abnormal device in the network topology.
[0035] To address the aforementioned technical problems, a third aspect of the present invention provides an electronic device, comprising:
[0036] Processor; and
[0037] A memory storing computer-executable instructions, which, when executed, cause the processor to perform the method described in any of the above embodiments.
[0038] To address the aforementioned technical problems, a fourth aspect of the present invention provides a computer storage medium, wherein the computer storage medium stores one or more programs, which, when executed by a processor, implement the method described in any of the above embodiments.
[0039] Compared with the prior art, this disclosure has the following advantages: This disclosure obtains the routing information of each device in the management area to construct the corresponding topology network, and then determines the abnormal devices that have anomalies based on the routing status of the devices. The number of abnormal devices determines the device disconnection status. Based on the device disconnection status, the location of the device disconnection fault is determined by the location information of the abnormal device in the network topology. This can quickly identify the device disconnection fault and perform intelligent fault diagnosis. The device disconnection status is determined by the number of abnormal devices, and then the location of the device disconnection fault is achieved by the location of the abnormal device in the network topology based on the device disconnection status, thus achieving the purpose of rapid troubleshooting.
[0040] Other features and advantages of this disclosure will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the disclosure. The objects and other advantages of this disclosure may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 A schematic flowchart of a method for analyzing equipment disconnection faults according to an embodiment of the present disclosure is shown;
[0043] Figure 2 A schematic diagram of an operator network topology according to an embodiment of the present disclosure is shown;
[0044] Figure 3 A schematic diagram of the operator network topology after removing device B according to an embodiment of the present disclosure is shown;
[0045] Figure 4 A second schematic flowchart of a device disconnection fault analysis method according to an embodiment of the present disclosure is shown.
[0046] Figure 5 A schematic diagram of a DCN-decentralized operator network topology according to an embodiment of the present disclosure is shown;
[0047] Figure 6 A schematic diagram of the operator network topology is shown when a primary station device malfunctions according to an embodiment of the present disclosure.
[0048] Figure 7 A schematic diagram illustrating the process of determining the location of a device disconnection fault based on the routing information of each abnormal device according to an embodiment of this disclosure is shown.
[0049] Figure 8 A schematic flowchart of a device disconnection fault analysis method according to an embodiment of the present disclosure is shown in part three.
[0050] Figure 9 A schematic diagram of the operator network topology is shown when devices E and F malfunction according to an embodiment of this disclosure;
[0051] Figure 10 A block diagram of a device disconnection fault analysis system according to an embodiment of the present disclosure is shown;
[0052] Figure 11 A schematic diagram of an electronic device structure according to an embodiment of the present disclosure is shown. Detailed Implementation
[0053] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0054] The same reference numerals in the accompanying drawings denote the same or similar elements, components, or parts, and therefore, repeated descriptions of the same or similar elements, components, or parts may be omitted below. It should also be understood that although terms such as first, second, third, etc., indicating numbers may be used herein to describe various devices, elements, components, or parts, these devices, elements, components, or parts should not be limited by these terms. That is, these terms are only used to distinguish one from another. For example, a first device may also be referred to as a second device, without departing from the essential technical solution of the invention. Furthermore, the terms "and / or" and "and / or" refer to all combinations including any one or more of the listed items.
[0055] Please see Figure 1 , Figure 1 This is a schematic diagram of one of the process flow diagrams for analyzing equipment disconnection faults provided by the present invention, such as... Figure 1 As shown, the method includes:
[0056] S11. Scan the routing information of devices within the management area and construct the network topology based on the routing information.
[0057] In this embodiment, as Figure 2 As shown, the operator's network topology includes the network management system, the DCN network, and various devices within the management area. The DCN network (Data Center Network) refers to the network architecture used to connect and manage various devices within the data center. Its main components include core switches, access switches, edge routers, firewalls and security devices, load balancers, etc. In this invention, it broadly refers to the network connecting the network management system and devices within the management area.
[0058] In this embodiment, routing information refers to information stored in routers or network devices that guides the path and direction of data packets transmission within the network. Network topology refers to the physical layout of various devices interconnected by transmission media. It refers to the specific physical (real) or logical (virtual) arrangement of the members constituting the network.
[0059] In this embodiment, the tracert command can be used to scan the routing information of devices throughout the management area. After obtaining the routing information of each management area, the network topology of each device can be constructed based on the communication relationship of the devices in the routing information.
[0060] In this embodiment, adding or removing a device or changing a fiber connection in the network topology will cause changes to the routing of its associated devices, such as... Figure 3 As shown, removing device B will change the routes for devices G and H. While device routing information is not static, it also doesn't change frequently. Therefore, it's necessary to update device routing information periodically. Specifically, this involves subscribing to changes in devices and fiber connections within the network management system. When a device or fiber connection changes, the associated devices are identified, and a new route scan is performed on these devices to obtain the new routing information.
[0061] Specifically, the system acquires the device status and connectivity of each device in real time. When the device status or connectivity of any device changes, the device with the changed status or connectivity is designated as the first update device, and the devices associated with the first update device are designated as the second update devices. The routing information of the first and second update devices is scanned to update the network topology. In this embodiment, device status can be usage, disabling, removal, addition, or change of purpose, and connectivity can be via I / O lines, bus connections, communication connections, or other connection methods. This solution uses the changed device and its connected devices as update devices when the device status or connectivity changes, and vertex scans the routing information of these update devices to update the network topology. This results in faster and more accurate network topology updates, avoiding resource waste caused by global scanning.
[0062] S12. Based on the routing status of each device, identify the abnormal device that is experiencing an anomaly.
[0063] In this embodiment, the routing status involves multiple aspects, including the physical status of the routing, the network connection status, and the software running status. If there are problems with the physical status or connection status of the device routing, it will be unable to communicate directly with other devices. At this time, it can be determined that the device is abnormal. If there are problems with the software running status of the device routing, it will generally lead to abnormal data transmission or slow data transmission speed. At this time, it can also be determined that the device is abnormal.
[0064] In this embodiment, the network management system can periodically check whether the routes of devices within the management area are reachable. Once a communication anomaly is detected in one or more devices, an alarm for the device communication anomaly is reported.
[0065] S13. Determine the device disconnection status based on the number of abnormal devices, and determine the location of the device disconnection fault based on the device disconnection status and the location information of the abnormal device in the network topology.
[0066] In this embodiment, most devices are generally in normal working condition. Once a large number of abnormal devices appear, it can be determined that the abnormal situation is quite serious. At the same time, the abnormal devices are at a high level or grade, such as core switches, access switches, edge routers, firewalls and security devices or load balancers in the data center network. In this case, it is necessary to quickly locate the anomaly in order to resolve the network problem.
[0067] In this embodiment, before determining that an abnormal device has experienced a disconnection fault, alarm data from each abnormal device is acquired and analyzed to determine whether the alarm data is a transient alarm. When the alarm data is a transient alarm, the corresponding abnormal device has not experienced a disconnection fault. A transient alarm is defined as an alarm whose generation and clearing time are less than a certain value (called the transient threshold). In simpler terms, transient alarms are those with a very short duration. These alarms are usually caused by network link instability, device interface failure, or other transient problems. Since a transient alarm indicates that the device has regained connectivity, further analysis of whether the device has experienced a disconnection fault is not conducted.
[0068] In this embodiment, the device disconnection status is first determined by the number of abnormal devices. An abnormality number standard can be set, and the abnormality number standard and the number of abnormal devices can be used to determine whether a large number of devices have disconnected. This allows for the selection of a method to determine the location of the disconnection and achieve rapid fault location. For example, the minimum number of abnormal devices in historical data when a large number of devices disconnected can be used as the abnormality number standard to determine the device disconnection status. Alternatively, the average number of abnormal devices in historical data when a large number of devices disconnected can be used as the abnormality number standard. Or, the number of nodes can be determined from the longest routing information from the DCN network to each device in the management area, and this number of nodes can be used as the abnormality number standard.
[0069] In this embodiment, after determining the device's disconnected state, a fault location method corresponding to the device's disconnected state is selected based on the device's disconnected state. The location of the device's disconnected fault is determined by the location information of each abnormal device in the network topology. For example, when a large number of devices are disconnected, the probability of each device's routing status having problems is very small. Therefore, at this time, it should be that one of the devices that is connected to multiple devices at the same time has become abnormal. The fault location can be quickly determined by the corresponding fault location method.
[0070] In this embodiment, the solution acquires routing information of each device within the management area to construct a corresponding topology network. Then, based on the routing status of the devices, it identifies abnormal devices that have malfunctioned. The number of abnormal devices determines the device's disconnection status. Based on the device's disconnection status, the location of the device's disconnection fault is determined by its position information in the network topology. This enables rapid identification of device disconnection faults and intelligent fault diagnosis. The solution determines the device's disconnection status by the number of abnormal devices and then locates the device disconnection fault by its position in the network topology, achieving rapid troubleshooting.
[0071] Please see Figure 4 , Figure 4This is a second schematic diagram of a method for analyzing equipment disconnection faults provided by the present invention, such as... Figure 4 As shown, step S13 includes the following methods:
[0072] S21. When the number of abnormal devices exceeds the first preset number, the device disconnection status is a large-scale device disconnection status. The number of abnormal devices belonging to the first station device in the network topology is determined according to the location information. The first station device includes: devices directly connected to the data center network.
[0073] In this embodiment, the first preset quantity is the abnormal quantity standard in the above embodiments, which will not be repeated here.
[0074] In this embodiment, by setting a first preset quantity, a predetermined standard is used to determine whether a large number of devices are out of control. In this way, different analysis methods can be selected to achieve rapid fault location and improve the efficiency of fault handling.
[0075] S22. If the number of faulty devices belonging to the first station is greater than the second preset number, the fault location is determined to be the data center network.
[0076] S23. If the number of abnormal devices belonging to the first station is not greater than the second preset number, the location of the device disconnection fault shall be determined according to the routing information of each abnormal device.
[0077] In this embodiment, when there are many instances of anomalies in the first-station equipment, and the probability of all first-station equipment malfunctioning simultaneously is low, this indirectly indicates that the data center network is experiencing anomalies. Prioritizing the inspection of the data center network connected to the first-station equipment can greatly improve the efficiency of fault handling. Specifically, for example, if the second preset number is 2, check the number of first-station equipment among the disconnected equipment. If the number of disconnected first-station equipment is greater than 2 (based on front-line troubleshooting experience, this is usually caused by DCN network anomalies), then it is considered to be caused by DCN disconnection. Figure 5 As shown.
[0078] If only one primary device has become unmanaged, check the historical routing information to see if all routes of the unmanaged devices pass through this primary device. If so, it indicates that the unmanaged status of this primary device caused a large number of devices to become unmanaged. Figure 6 As shown. Specifically, as... Figure 7 As shown, step S23 includes the following steps:
[0079] S31. If the number of abnormal devices belonging to the first station device is not greater than the second preset number and is not equal to 0, determine whether each abnormal device has passed through the abnormal device belonging to the first station device based on the routing information of each abnormal device; if so, determine that the abnormal device belonging to the first station device has a disconnection fault.
[0080] S32. If the number of abnormal devices belonging to the first station device is equal to 0, obtain the first routing information of the abnormal device before the abnormality and the second routing information after the abnormality.
[0081] S33. Based on the first routing information and the second routing information, determine the fault location of the abnormal device in the network topology.
[0082] In this embodiment, when a non-first-station device malfunctions, any faulty device will cause devices connected to the faulty device and subsequent connected devices to become disconnected. In this case, two faulty devices are randomly selected for route scanning. The scan results are compared with the historical routing information of these two devices to determine the specific location of the fault and present the diagnostic results to the user. For example, the routing information before the device malfunction is [ip1,ip2,ip3,ip4...], and the route scan result after the malfunction is: [ip1,ip2]. Please check the network status between ip2 and ip3.
[0083] Please see Figure 8 , Figure 8 This is a schematic diagram of the third step in the process of analyzing equipment disconnection faults provided by the present invention. Figure 8 As shown, step S13 includes the following methods:
[0084] S41. When the number of abnormal devices is not greater than the first preset number, obtain the current routing information of each abnormal device.
[0085] In this embodiment, the first preset quantity is the abnormal quantity standard in the above embodiments, which will not be repeated here.
[0086] S42. Determine the node IP of the last hop device that is alive in the routing path based on the current routing information.
[0087] In this embodiment, the surviving last-hop device is the last device in the routing path that can still communicate normally. The node IP refers to the IP address of the node server. In a distributed system, the node IP is the IP address of the node server used to connect to the network, provide data transmission and processing services, and is the unique identifier of the node server.
[0088] S43. When the node IP belongs to a device within the management area, log in directly to the last-hop device and determine the device that has experienced a disconnection fault based on the routing status of the neighbor routes of the last-hop device.
[0089] S44. When the node IP does not belong to a device within the management area, the fault location is determined to be the data center network.
[0090] In this embodiment, the unique identifier of the device is used to determine whether the device with the detachment fault belongs to the managed area. On the one hand, it can locate the location of the device, and on the other hand, it can also help staff determine the location of the device that is still working normally.
[0091] In this embodiment, as Figure 9 As shown, by scanning the routes of the disconnected devices E and F, the last surviving node IP is found in the scan results. Then, it is determined whether the IP belongs to a device within the managed area or a DCN network device. If it belongs to a device within the managed area, the last-hop device is directly logged into to check if its neighbor routes are normal, identifying the abnormal device and presenting it to the user as a diagnostic result. If the IP does not belong to a device within the managed area, it indicates a DCN network anomaly. The DCN network fault is then reported, and the IP is presented to the user as a diagnostic result. Figure 9 In the network topology shown, the last hop device is device D, and the neighbor node of device D is device C. If both device D and device C are normal, it means that devices E and F are both disconnected from management. If device D is abnormal while device C is normal, it means that device D is disconnected from management.
[0092] Please see Figure 10 , Figure 10 This is a block diagram of a device disconnection fault analysis system provided by the present invention, such as... Figure 10 As shown, the analysis system includes: a network topology construction module 11, an abnormal device analysis module 12, and a device disconnection analysis module 13.
[0093] In this embodiment, the network topology construction module 11 is used to scan the routing information of devices within the management area and construct the network topology based on the routing information.
[0094] In this embodiment, the abnormal device analysis module 12 is used to determine the abnormal devices that have anomalies based on the routing status of each device.
[0095] In this embodiment, the device disconnection analysis module 13 is used to determine the device disconnection status based on the number of abnormal devices, and based on the device disconnection status, determine the location of the device disconnection fault by using the location information of the abnormal device in the network topology.
[0096] In this embodiment, the device disconnection analysis module 13 is specifically used to determine the number of abnormal devices belonging to the first station device in the network topology based on the location information when the number of abnormal devices is greater than a first preset number, which is a large-scale device disconnection state; the first station device includes devices directly connected to the data center network; if the number of abnormal devices belonging to the first station device is greater than a second preset number, the fault location is determined to be the data center network; if the number of abnormal devices belonging to the first station device is not greater than the second preset number, the device disconnection fault location is determined based on the routing information of each abnormal device.
[0097] In this embodiment, the device disconnection analysis module 13 is specifically used to determine whether each abnormal device has passed through an abnormal device belonging to the first station device if the number of abnormal devices belonging to the first station device is not greater than a second preset number and is not equal to 0; if so, it is determined that the abnormal device belonging to the first station device has a disconnection fault; if the number of abnormal devices belonging to the first station device is equal to 0, it obtains the first routing information of the abnormal device before the abnormality and the second routing information after the abnormality; and determines the fault location of the abnormal device in the network topology based on the first routing information and the second routing information.
[0098] In this embodiment, the device disconnection analysis module 13 is specifically used to obtain the current routing information of each abnormal device when the number of abnormal devices is not greater than a first preset number, which is a non-mass device disconnection state; determine the node IP of the last hop device that is alive in the routing path based on the current routing information; when the node IP belongs to a device within the management area, directly log in to the last hop device, and determine the device that has disconnected from management based on the routing status of the neighbor routes of the last hop device; when the node IP does not belong to a device within the management area, determine the fault location as the data center network.
[0099] In this embodiment, the device disconnection analysis module 13 is also used to acquire alarm data of each abnormal device, analyze the alarm data, and determine whether the alarm data is a momentary interruption alarm; when the alarm data is a momentary interruption alarm, the corresponding abnormal device has not experienced a disconnection fault.
[0100] In this embodiment, the network topology construction module 11 is also used to obtain the device status and connectivity of each device in real time; when the device status or connectivity of any device changes, the device whose device status or connectivity has changed is designated as the first update device, and the device associated with the first update device is designated as the second update device; the routing information of the first update device and the second update device is obtained by scanning, and the network topology is updated.
[0101] In this embodiment, the network topology construction module 11 is specifically used to obtain the device that communicates with the first update device in the network topology, and use it as the second update device.
[0102] like Figure 11 As shown, this embodiment of the invention provides an electronic device, including a processor 1110, a communication interface 1120, a memory 1130, and a communication bus 1140, wherein the processor 1110, the communication interface 1120, and the memory 1130 communicate with each other through the communication bus 1140.
[0103] Memory 1130 is used to store computer programs;
[0104] The processor 1110, when executing the program stored in the memory 1130, implements any of the above analysis methods.
[0105] The electronic device provided in this embodiment of the invention includes a processor 1110 that executes a program stored in a memory 1130 to scan the routing information of devices within a management area and constructs a network topology based on the routing information; identifies abnormal devices based on the routing status of each device; determines the device disconnection status based on the number of abnormal devices; and determines the device disconnection fault location based on the device disconnection status and the location information of the abnormal device in the network topology.
[0106] The communication bus 1140 mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 1140 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, it is shown in the figure with only one thick line, but this does not indicate that there is only one bus or one type of bus.
[0107] The communication interface 1120 is used for communication between the above-mentioned electronic device and other devices.
[0108] The memory 1130 may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory 1130 may also be at least one storage device located remotely from the aforementioned processor 1110.
[0109] The processor 1110 mentioned above can be a general-purpose processor 1110, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0110] This invention provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors 1110 to implement the analysis method of any of the above embodiments.
[0111] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of the present invention is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0112] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.
Claims
1. A method for analyzing equipment detachment faults, characterized in that, The analytical method includes: Scan the routing information of devices within the management area and construct the network topology based on the routing information; Based on the routing status of each of the aforementioned devices, identify the abnormal devices that have experienced anomalies; The device disconnection status is determined based on the number of abnormal devices, and the location of the device disconnection fault is determined based on the location information of the abnormal devices in the network topology. The step of determining the device disconnection status based on the number of abnormal devices, and determining the device disconnection fault location based on the device disconnection status and the location information of the abnormal devices in the network topology, includes: When the number of abnormal devices exceeds a first preset number, the device disconnection status is a large-scale device disconnection status. The number of abnormal devices belonging to the first station device in the network topology is determined according to the location information. The first station device includes: devices directly connected to the data center network. If the number of faulty devices belonging to the first station is greater than the second preset number, the fault location is determined to be the data center network; If the number of abnormal devices belonging to the first station is not greater than the second preset number, the location of the device disconnection fault is determined based on the routing information of each abnormal device. When the number of abnormal devices is not greater than the first preset number, it is a non-mass device disconnection state, and the current routing information of each abnormal device is obtained. Determine the node IP of the last hop device that is alive in the routing path based on the current routing information; When the node IP belongs to a device within the management area, log in directly to the last-hop device and determine the device that has experienced a disconnection fault based on the routing status of the neighbor routes of the last-hop device. If the node IP does not belong to a device within the management area, the fault location is determined to be the data center network.
2. The analytical method according to claim 1, characterized in that, If the number of abnormal devices belonging to the first station equipment is not greater than the second preset number, then the location of the device disconnection fault is determined based on the routing information of each abnormal device, including: If the number of abnormal devices belonging to the first station device is not greater than the second preset number and is not equal to 0, determine whether each abnormal device has passed through the abnormal device belonging to the first station device based on the routing information of each abnormal device; if so, determine that the abnormal device belonging to the first station device has a disconnection fault. If the number of abnormal devices belonging to the first station device is equal to 0, obtain the first routing information of the abnormal device before the abnormality and the second routing information after the abnormality. Based on the first routing information and the second routing information, the fault location of the abnormal device is determined in the network topology.
3. The analytical method according to claim 1, characterized in that, Before determining the device disconnection status based on the number of abnormal devices, and determining the device disconnection fault location based on the device disconnection status and the location information of the abnormal devices in the network topology, the analysis method further includes: Obtain alarm data from each of the abnormal devices, analyze the alarm data, and determine whether the alarm data is a momentary interruption alarm; When the alarm data is a momentary interruption alarm, the corresponding abnormal device has not experienced a disconnection fault.
4. The analytical method according to any one of claims 1 to 3, characterized in that, The analytical method further includes: Real-time acquisition of the device status and connectivity of each of the aforementioned devices; When the device status or connectivity of any of the devices changes, the device whose device status or connectivity has changed is designated as the first updating device, and the device associated with the first updating device is designated as the second updating device. The routing information of the first and second updating devices is obtained by scanning, and the network topology is updated accordingly.
5. The analytical method according to claim 4, characterized in that, The step of designating the device associated with the first updating device as the second updating device includes: The device that communicates with the first updating device in the network topology is selected as the second updating device.
6. A system for analyzing equipment detachment faults, characterized in that, The analysis system includes: The network topology construction module is used to scan the routing information of devices within the management area and construct the network topology based on the routing information. The abnormal device analysis module is used to determine the abnormal devices that have anomalies based on the routing status of each device. The device disconnection analysis module is used to determine the device disconnection status based on the number of abnormal devices, and based on the device disconnection status, determine the location of the device disconnection fault through the location information of the abnormal device in the network topology; The device disconnection analysis module is specifically used to determine the number of abnormal devices belonging to the first-station device in the network topology based on the location information when the number of abnormal devices exceeds a first preset number, indicating a large-scale device disconnection state. The first-station device includes devices directly connected to the data center network. If the number of abnormal devices belonging to the first-station device exceeds a second preset number, the fault location is determined to be the data center network. If the number of abnormal devices belonging to the first-station device does not exceed the second preset number, the device disconnection fault location is determined based on the routing information of each abnormal device. The device disconnection analysis module is specifically used to obtain the current routing information of each abnormal device when the number of abnormal devices is no greater than a first preset number, indicating a non-mass device disconnection state; determine the node IP of the last hop device in the routing path based on the current routing information; when the node IP belongs to a device within the management area, directly log in to the last hop device and determine the device with the disconnection fault based on the routing status of the last hop device's neighbor routes; when the node IP does not belong to a device within the management area, the fault location is determined to be the data center network.
7. An electronic device, characterized in that, include: processor; as well as A memory storing computer-executable instructions, which, when executed, cause the processor to perform the method according to any one of claims 1-5.
8. A computer storage medium, characterized in that, in, The computer storage medium stores one or more programs, which, when executed by a processor, implement the method of any one of claims 1-5.