Network fault location methods, devices and storage media
By establishing a network topology and injecting fault events, and verifying and correcting node configuration information, the problem that existing technologies cannot be applied to real network scenarios is solved, and efficient network fault detection and repair are achieved.
Patent Information
- Application Number
- CN202411960582.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2044-12-27
AI Technical Summary
Existing network detection methods are not applicable to real-world network scenarios, have poor practicality, and cannot effectively locate faulty nodes in the network topology.
Establish the network topology of the target network, record the configuration information of each node in the non-fault state, generate fault events and inject them into the network topology, locate the faulty node by verifying the configuration information of the node after the fault event, and perform corrective processing.
It improves the practicality of network fault detection, enabling accurate location and repair of faulty nodes in the network topology, thus enhancing the accuracy and efficiency of fault detection.
Smart Images

Figure CN119766637B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a network fault location method, apparatus and storage medium. Background Technology
[0002] With the rapid development of cloud computing technology, open stack platforms have been widely used in enterprises and data centers. However, in open stack application scenarios, the network topology is relatively complex, involving various virtualization technologies, network protocols, and service deployment methods, which makes the prediction and localization of network faults extremely difficult.
[0003] In current network detection schemes, the virtual machine console determines connectivity between network nodes using network detection protocols. However, this approach is not suitable for real-world network scenarios and has poor practicality. Summary of the Invention
[0004] This application provides a network fault location method, apparatus, and storage medium, which solves the problem that the proposed network detection methods are not applicable to real network scenarios and have poor practicality. It can locate faulty nodes in the network topology corresponding to real network scenarios, greatly improving the practicality of fault detection.
[0005] To achieve the above objectives, this application adopts the following technical solution:
[0006] In a first aspect, this application provides a network fault location method, which includes: establishing a network topology corresponding to a target network and recording the configuration information of each node in the network topology when it is in a non-fault state; generating a fault event and injecting the fault event into the network topology, and recording the configuration information of each node in the network topology after the fault event occurs; and verifying the configuration information of each node after the fault event occurs based on the configuration information of each node when it is in a non-fault state, so as to locate the faulty node among the nodes.
[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the configuration information includes multiple types of configuration items; the method further includes:
[0008] In conjunction with the first aspect above, in one possible implementation, the method further includes: for each node among the nodes, based on multiple types of configuration items when the node is in a non-faulty state, verifying each configuration item of the node after the fault event occurs, obtaining verification results, and locating the faulty node in the network topology based on the verification results.
[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the configuration items include at least one of the following: node topology configuration, node running status configuration, and node system configuration.
[0010] In conjunction with the first aspect above, in one possible implementation, the method further includes: verifying whether the topology configuration of the node after the fault event occurs is the same as the topology configuration in the non-fault state, verifying whether the running state configuration after the fault event occurs is the same as the running state configuration in the non-fault state, and verifying whether the system configuration after the fault event occurs is the same as the system configuration in the non-fault state.
[0011] If the topology configuration, the running status configuration, and / or the system configuration are different, the node is marked as a fault node, and at least one first node other than the fault node in the network topology is selected; for each first node, it is verified whether the network connectivity status of the first node after the fault event occurs is the same as the network connectivity status when it is in a non-fault state.
[0012] If the network connectivity status is different, the first node is marked as a faulty node, and at least one second node other than the faulty node in the first node is selected; for each second node, it is verified whether the system configuration of the second node after the fault event occurs is the same as the system configuration of the node when it is in a non-faulty state; if the system configuration is different, the second node is marked as a faulty node.
[0013] In conjunction with the first aspect above, in one possible implementation, the method further includes: when the node is a faulty node, performing correction processing on the configuration information of the node after the fault event occurs based on the configuration information of the node when it is in a non-faulty state, so that the node changes from a faulty state to a non-faulty state; recording the correction scheme of the node; the correction scheme includes the correction process of the configuration information when the node changes from a faulty state to a non-faulty state.
[0014] In conjunction with the first aspect mentioned above, in one possible implementation, after corrective processing has been performed on each faulty node in the network topology, the method further includes: repeatedly performing the first operation until the availability of each node in the network topology has been detected; summarizing the availability of each node in the network topology to obtain the detection result of the network topology; the first operation includes: for each node in the network topology, detecting whether the network connectivity status of the node is connected; if it is connected, marking the availability of the node as available; otherwise, marking the availability of the node as unavailable.
[0015] In conjunction with the first aspect above, in one possible implementation, the fault event includes at least one of the following: network logical fault event, physical fault event, and protocol fault event; wherein, network logical fault events include simulated configuration error events and router port shutdown events; physical fault events include power outage events and network outage events; and protocol fault events include router fault events and firewall parameter setting error events.
[0016] Secondly, this application provides a network fault location device, which includes: a processing unit; the processing unit is configured to establish a network topology corresponding to a target network and record the configuration information of each node in the network topology when it is in a non-fault state; the processing unit is further configured to generate a fault event and inject the fault event into the network topology, and record the configuration information of each node in the network topology after the fault event occurs; the processing unit is further configured to verify the configuration information of each node after the fault event occurs based on the configuration information of each node when it is in a non-fault state, so as to locate the faulty node among the nodes.
[0017] In conjunction with the second aspect above, in one possible implementation, the configuration information includes multiple types of configuration items; the processing unit is specifically used to: for each node in each node, based on the multiple types of configuration items when the node is in a non-faulty state, verify each configuration item of the node after the fault event occurs, obtain the verification result, and locate the faulty node in the network topology based on the verification result.
[0018] In conjunction with the second aspect above, in one possible implementation, the configuration items include at least one of the following: node topology configuration, node running status configuration, and node system configuration.
[0019] In conjunction with the second aspect above, in one possible implementation, the processing unit is specifically configured to: verify whether the topology configuration of a node after the occurrence of the fault event is the same as its topology configuration in a non-faulty state, and verify whether the operating state configuration after the occurrence of the fault event is the same as its operating state configuration in a non-faulty state, and verify whether the system configuration after the occurrence of the fault event is the same as the system configuration in a non-faulty state; if the topology configuration, operating state configuration, and / or system configuration are different, then the node is marked as a faulty node, and at least one first node other than the faulty node in the network topology is selected; for each first node, verify whether the network connectivity status of the first node after the occurrence of the fault event is the same as its network connectivity status in a non-faulty state; if the network connectivity status is different, then the first node is marked as a faulty node, and at least one second node other than the faulty node among the at least one first node is selected; for each second node, verify whether the system configuration of the second node after the occurrence of the fault event is the same as its system configuration in a non-faulty state; if the system configuration is different, then the second node is marked as a faulty node.
[0020] In conjunction with the second aspect above, in one possible implementation, the processing unit is further configured to: when the node is a faulty node, perform corrective processing on the configuration information of the node after the fault event occurs, based on the configuration information of the node when it is in a non-faulty state, so that the node changes from a faulty state to a non-faulty state; record the correction scheme of the node; the correction scheme includes the correction process of the configuration information when the node changes from a faulty state to a non-faulty state.
[0021] In conjunction with the second aspect above, in one possible implementation, after all faulty nodes in the network topology have been corrected, the processing unit is specifically used to: repeatedly execute the first operation until the availability of all nodes in the network topology has been detected; and summarize the availability of all nodes in the network topology to obtain the detection result of the network topology.
[0022] The first operation includes: for each node in the network topology, checking whether the node's network connectivity status is connected; if it is connected, marking the node's availability as available; otherwise, marking the node's availability as unavailable.
[0023] In conjunction with the second aspect above, in one possible implementation, the fault event includes at least one of the following: network logical fault event, physical fault event, and protocol fault event; wherein, network logical fault events include simulated configuration error events and router port shutdown events; physical fault events include power outage events and network outage events; and protocol fault events include router fault events and firewall parameter setting error events.
[0024] Thirdly, this application provides a communication device comprising: a processor and a communication interface; the communication interface and the processor are coupled, the processor being configured to run computer programs or instructions to implement the network fault location method as described in the first aspect and any possible implementation thereof.
[0025] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a terminal, cause the terminal to perform the network fault location method as described in the first aspect and any possible implementation thereof.
[0026] Fifthly, this application provides a computer program product containing instructions that, when run on a computer, cause the computer to perform the network fault location method as described in the first aspect and any possible implementation thereof.
[0027] In a sixth aspect, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the network fault location method as described in the first aspect and any possible implementation thereof.
[0028] Specifically, the chip provided in this application also includes a memory for storing computer programs or instructions.
[0029] It should be noted that the aforementioned computer instructions may be stored, in whole or in part, on a computer-readable storage medium. This computer-readable storage medium may be packaged together with the processor of the device, or it may be packaged separately from the processor of the device; this application does not impose any limitation on this.
[0030] In a seventh aspect, this application provides a network fault location system, comprising: a network topology and a server, wherein the server is used to execute the network fault location method as described in the first aspect and any possible implementation thereof.
[0031] The descriptions of aspects two through seven in this application can be referenced to the detailed description of aspect one; and the beneficial effects of the descriptions of aspects two through seven can be referenced to the analysis of the beneficial effects of aspect one, which will not be repeated here.
[0032] In this application, the name of the aforementioned network fault location device does not limit the device or functional module itself. In actual implementation, these devices or functional modules may appear under other names. As long as the function of each device or functional module is similar to that of this application, it falls within the scope of the claims of this application and its equivalents.
[0033] These or other aspects of this application will become more readily apparent in the following description.
[0034] The above solution offers at least the following advantages: Based on the above technical solution, the network fault location method provided in this application establishes a network topology corresponding to the target network and records the configuration information of each node in the network topology when it is in a non-faulty state; it generates a fault event and injects the fault event into the network topology, recording the configuration information of each node in the network topology after the fault event occurs. In other words, in this technical solution, network faults occur based on the network topology corresponding to the target network, and network fault detection is also implemented based on the network topology corresponding to the target network. Furthermore, by verifying the configuration information of each node after the fault event occurs through the configuration information of each node in a non-faulty state, this technical solution can locate the faulty node in the network topology corresponding to the target network. This solves the problem that current network detection solutions, where the virtual machine console determines whether network nodes are connected through network detection protocols, are not applicable to real network scenarios and have poor practicality, greatly improving the practicality of network fault detection. Attached Figure Description
[0035] Figure 1 This is a flowchart of a network detection method and system based on OpenStack;
[0036] Figure 2 This application provides a schematic diagram of the architecture of a network fault location system.
[0037] Figure 3 A schematic diagram illustrating a network fault location method provided in an embodiment of this application;
[0038] Figure 4 This is a schematic diagram of the hardware structure of a communication device provided in an embodiment of this application;
[0039] Figure 5 A flowchart illustrating a network fault location method provided in this application embodiment;
[0040] Figure 6 A flowchart illustrating another network fault location method provided in this application embodiment;
[0041] Figure 7 A flowchart illustrating another network fault location method provided in this application embodiment;
[0042] Figure 8 A flowchart illustrating another network fault location method provided in this application embodiment;
[0043] Figure 9 A flowchart illustrating another network fault location method provided in this application embodiment;
[0044] Figure 10 A flowchart illustrating another network fault location method provided in this application embodiment;
[0045] Figure 11 A flowchart illustrating another network fault location method provided in this application embodiment;
[0046] Figure 12 This is a schematic diagram of the structure of a network fault location device provided in an embodiment of this application. Detailed Implementation
[0047] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0048] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0049] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.
[0050] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.
[0051] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0052] In the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0053] With the rapid development of cloud computing technology, OpenStack, as an open-source cloud computing management platform, has been widely used in enterprises and data centers. However, in OpenStack application scenarios, the network topology is relatively complex, involving various virtualization technologies, network protocols, and service deployment methods. This makes the prediction and location of network faults extremely difficult, undoubtedly posing a formidable challenge to operations and maintenance personnel.
[0054] Operators play a crucial role in the construction and maintenance of the backbone network within the overall network architecture, serving as an indispensable cornerstone for the continuous, stable, and rapid development of technology. Therefore, a stable and reliable network is a fundamental prerequisite. Network failures can lead to numerous serious consequences, posing significant risks and potential impacts on many industries and sectors that rely on the network.
[0055] Traditional network fault detection methods rely on manual simulation testing, which is inefficient and prone to overlooking potential fault factors. With the increasing complexity of networks and the massive growth brought about by network development, network security and stability face unprecedented challenges. For example, in recent years, large-scale service outages have occurred in networks, and the excessively long fault recovery time and untimely processing required for manual simulation testing have had a significant impact on customer service.
[0056] Among the currently proposed network detection schemes is a network detection method and system based on OpenStack, such as... Figure 1 As shown, the method includes:
[0057] 1. Create a service module project based on OpenStack, and create external and internal networks for the service module project.
[0058] 2. Create a virtual machine and configure it on the internal network.
[0059] 3. Create a router and add internal and external network interfaces to it.
[0060] 4. Add the network detection protocol to the service module project.
[0061] 5. Assign a floating IP address to the service module project and bind the floating IP address to the virtual machine.
[0062] 6. Disable the firewall in the virtual machine console.
[0063] 7. The virtual machine console uses network detection protocols to determine network connectivity.
[0064] The system includes: a network creation module, a virtual machine creation module, a router creation module, a network detection protocol addition module, a floating IP allocation module, a firewall disabling module, and a network judgment module.
[0065] In the network detection scheme described above, the virtual machine console determines whether network nodes are connected using a network detection protocol. However, this scheme is not suitable for real-world network scenarios and has poor practicality.
[0066] Furthermore, there are many types of network faults, and this technology cannot analyze and simulate other types of network faults, nor can it obtain corresponding corrective solutions, resulting in poor reliability.
[0067] Therefore, the network fault location method provided in this application establishes a network topology corresponding to the target network and records the configuration information of each node in the network topology when it is in a non-faulty state; generates a fault event and injects the fault event into the network topology, recording the configuration information of each node in the network topology after the occurrence of the fault event. In other words, the network fault in this technical solution occurs based on the network topology corresponding to the target network, and network fault detection is also implemented based on the network topology corresponding to the target network. Furthermore, by verifying the configuration information of each node after the occurrence of the fault event through the configuration information of each node in a non-faulty state, this technical solution can locate the faulty node in the network topology corresponding to the target network. This solves the problem that current network detection solutions, where the virtual machine console determines whether network nodes are connected through network detection protocols, are not applicable to real network scenarios and have poor practicality, greatly improving the practicality of network fault detection.
[0068] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0069] Figure 2 This is a schematic diagram of the architecture of a network fault location system provided in an embodiment of this application. Figure 2 As shown, the network fault location system includes: network topology 201 and server 202.
[0070] In this application, network topology 201 and server 202 are connected via a communication link. This communication link can be a wired communication link or a wireless communication link, and this application does not limit it in this regard.
[0071] In some embodiments, network topology 201 is generated based on a target network (e.g., a real network in a certain area). Network topology 201 includes multiple nodes. These nodes can be terminals, access network devices, core network devices, relay devices, etc., and this application does not impose any limitations on them.
[0072] Optionally, each node in network topology 201 can be developed and generated based on the OpenStack platform and can be supported by the OpenStack platform.
[0073] Optionally, network topology 201 can be edited and modified, and network topology 201 has customizable features.
[0074] In some embodiments, such as Figure 3 As shown, server 202 can be used to record and save the initial configuration information of each node in network topology 201 when it is successfully created. Each node is in a non-faulty state when creation is successful. Server 202 is also used to generate various types of fault events and inject these events into network topology 201, recording the configuration information of each node in network topology 201 when it is in a faulty state.
[0075] Furthermore, such as Figure 3 As shown, server 202 is also used to perform verification and analysis based on the initial configuration information of each node when it is successfully created and the configuration information when it is in a fault state, in order to locate the faulty node in network topology 201.
[0076] Furthermore, such as Figure 3 As shown, server 202 is also used to provide a correction scheme for each node in a fault state, and to repair the faulty node in each node so that the node changes from a faulty state to a non-faulty state.
[0077] Furthermore, such as Figure 3 As shown, server 202 is also used to verify the network topology after a node changes from a faulty state to a non-faulty state, and to detect the availability of each node in the network topology, so as to facilitate targeted analysis by operation and maintenance personnel and improve the efficiency of fault operation and maintenance.
[0078] Optionally, server 202 can be an open stack platform, which is mainly used to build and manage cloud computing environments and provide cloud computing services, including the management and allocation of resources such as compute, storage, and network.
[0079] Optionally, server 202 can be a standalone physical server, a server cluster consisting of multiple physical servers, a distributed file system, or at least one of the following cloud servers providing basic cloud computing services: cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Of course, the above is merely an exemplary description of server 202; server 202 can also be a relational database management system (e.g., MySQL, Oracle Database, or Microsoft SQL Server) and / or a database based on distributed file storage (e.g., MongoDB, HBase, or Cassandra), and this application does not impose any limitations in this regard.
[0080] When implemented in hardware, the various modules in server 202 can be integrated into, for example... Figure 4 The communication device shown is implemented in hardware. Specifically, as... Figure 4 As shown, the basic hardware structure of the communication device is introduced.
[0081] Figure 4 This is a schematic diagram of the hardware structure of a communication device provided in an embodiment of this application. Figure 4 As shown, the communication device includes at least one processor 401, a communication line 402, and at least one communication interface 404, and may also include a memory 403. The processor 401, memory 403, and communication interface 404 can be connected via the communication line 402.
[0082] The processor 401 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).
[0083] Communication line 402 may include a path for transmitting information between the aforementioned components.
[0084] Communication interface 404 is used to communicate with other devices or communication networks. It can use any transceiver-like device, such as Ethernet, radio access network (RAN), wireless local area network (WLAN), etc.
[0085] The memory 403 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of including or storing desired program code having the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0086] In one possible design, the memory 403 can exist independently of the processor 401, meaning the memory 403 can be an external memory of the processor 401. In this case, the memory 403 can be connected to the processor 401 via a communication line 402 to store execution instructions or application code, and its execution is controlled by the processor 401 to implement the network fault location method provided in the following embodiments of this application. In another possible design, the memory 403 can also be integrated with the processor 401, meaning the memory 403 can be an internal memory of the processor 401. For example, the memory 403 can be a cache, which can be used to temporarily store some data and instruction information.
[0087] As one possible implementation, processor 401 may include one or more CPUs, for example Figure 4 CPU0 and CPU1 in the example. As another possible implementation, the communication device may include multiple processors, such as... Figure 4 The processors 401 and 407 are included. As another possible implementation, the communication device may also include an output device 405 and an input device 406.
[0088] It should be noted that the various embodiments of this application can be referenced or learned from each other. For example, the same or similar steps, method embodiments, system embodiments and device embodiments can be referenced from each other without limitation.
[0089] Figure 5 A flowchart illustrating a network fault location method provided in this application embodiment, which can be applied to, for example... Figure 4 In the communication device shown. For example... Figure 5 As shown, the method includes the following S501-S503.
[0090] S501. Establish the network topology corresponding to the target network and record the configuration information of each node in the network topology when it is in a non-faulty state.
[0091] Optionally, the configuration information includes multiple types of configuration items. These configuration items include at least one of the following: node topology configuration, node running status configuration, and node system configuration.
[0092] It should be noted that the node's topology configuration reflects other nodes connected to that node in the network topology and their connections. The node's running status configuration reflects whether the node is powered on or powered off. The node's system configuration reflects the node's internal configuration information.
[0093] In one possible implementation, a network topology corresponding to the target network is established based on the OpenStack platform, and the configuration information of the network topology is recorded and saved (e.g., saved in a relational database) for subsequent analysis. For example, the configuration information of the nodes in the network topology after the failure event is verified based on the information to determine whether the node is a faulty node.
[0094] S502. Generate a fault event and inject the fault event into the network topology, recording the configuration information of each node in the network topology after the fault event occurs.
[0095] Optionally, the failure event includes at least one of the following: network logical failure event, physical failure event, and protocol failure event. Of course, the above is only an exemplary description of failure events, and failure events may also include topology failures (e.g., reconfiguration of nodes or addition of new nodes (network devices) to the main structure of the network, thereby changing the network topology).
[0096] Among them, network logical failure events include simulated configuration error events and router port shutdown events; physical failure events include power outage events and network outage events; and protocol failure events include router failure events and firewall parameter setting error events.
[0097] Understandably, a variety of different types of failure events can ensure the authenticity and comprehensiveness of network failure testing, and can truly reflect the problems that may be encountered in the actual network operating environment.
[0098] In one possible implementation, such as Figure 6 As shown, the implementation process of S502 includes:
[0099] 1. Read and save the configuration information of each node when it is in a non-faulty state (for example, read the configuration information of each node when it is in a non-faulty state from the database) for data comparison in subsequent verification processes.
[0100] 2. Select the type of fault event and prepare to inject it into the network topology.
[0101] 3. Configure the parameters for the fault event. The parameters for the fault event include, but are not limited to, at least one of the following: fault occurrence event, fault duration, and fault impact range. The fault impact range reflects the extent to which the fault event affects nodes within the network topology.
[0102] 4. Generate fault events and inject the fault event simulation into the network topology.
[0103] 5. Record the configuration information of each node in the network topology after a failure event. This configuration information includes: the topological relationship configuration of the nodes, the node's operational status configuration, and the node's system configuration.
[0104] The specific descriptions of the node topology configuration, node running status configuration, and node system configuration are as shown in the embodiment shown in S501, and will not be repeated here.
[0105] Among them, such as Figure 6 As shown, the process of generating a fault event includes steps 2-4 above.
[0106] S503. Based on the configuration information of each node in a non-faulty state, verify the configuration information of each node after a fault event occurs in order to locate the faulty node among the nodes.
[0107] In one possible implementation, corresponding to the configuration information including multiple types of configuration items, the implementation process of S503 above may include: for each node in each node, based on the multiple types of configuration items when the node is in a non-faulty state, verifying each configuration item of the node after the fault event occurs, obtaining the verification result, and locating the faulty node in the network topology based on the verification result.
[0108] Specifically, considering the various configuration items for nodes in a non-faulty state, the specific implementation process of S503 above is as shown in the embodiments in S701-S706, and will not be repeated here.
[0109] In another possible implementation, network fault detection tools (e.g., ping, traceroute, network packet capture) are used to quickly locate faulty nodes in the network topology. Simultaneously, the configuration information of each node after a network topology failure event is verified to assist in locating the faulty node. Then, the faulty node is corrected and repaired, and a fault correction plan is generated (e.g., adjusting network configuration, restarting, restoring backups, etc.).
[0110] The verification of configuration information of each node after a network topology failure event includes at least one of the following: power-on verification, network start / stop verification, network interconnection verification, and protocol input / output verification.
[0111] Based on the above technical solution, the network fault location method provided in this application establishes a network topology corresponding to the target network and records the configuration information of each node in the network topology when it is in a non-faulty state; generates a fault event and injects the fault event into the network topology, recording the configuration information of each node in the network topology after the fault event occurs. In other words, in this technical solution, network faults occur based on the network topology corresponding to the target network, and network fault detection is also implemented based on the network topology corresponding to the target network. Furthermore, by verifying the configuration information of each node after the fault event occurs through the configuration information of each node in a non-faulty state, this technical solution can locate the faulty node in the network topology corresponding to the target network. This solves the problem that current network detection solutions, where the virtual machine console determines whether network nodes are connected through network detection protocols, are not applicable to real network scenarios and have poor practicality, greatly improving the practicality of network fault detection.
[0112] Furthermore, this application simulates diverse failure events to ensure the authenticity and comprehensiveness of network failure testing, which can realistically reflect the problems that may be encountered in the actual network operating environment.
[0113] As one possible embodiment of this application, combined with Figure 5 ,like Figure 7 As shown, in the above S503, the configuration information of each node after a fault event is verified in order to locate the faulty node. This can be achieved through the following S701-S706.
[0114] S701. Verify whether the topology configuration of the node after a failure event is the same as the topology configuration in a non-failure state, and verify whether the running status configuration after a failure event is the same as the running status configuration in a non-failure state, and verify whether the system configuration after a failure event is the same as the system configuration in a non-failure state.
[0115] It is understandable that S701 can be understood as verifying the configuration information of nodes (including the node's topology relationship configuration, operating status configuration, and system configuration) after a network topology failure event occurs, in order to determine whether the network failure is caused by abnormal configuration information of the node.
[0116] S702. If the topology configuration is different, the running status configuration is different, and / or the system configuration is different, then the node will be marked as a faulty node.
[0117] In one possible implementation, the S702 implementation process may also include: if the configuration information of the two is inconsistent, then the node is marked as a faulty node.
[0118] It is understandable that the inconsistency in the configuration information indicates that the node failure is a network anomaly caused by abnormal configuration information.
[0119] S703. Filter out at least one first node in the network topology other than the faulty node; for each first node, verify whether the network connectivity state of the first node after the fault event is the same as the network connectivity state when it is in a non-faulty state.
[0120] The network connectivity state of a node when it is in a non-faulty state is called the connectivity state. The network connectivity state reflects the network connectivity between the node and the nodes connected to it on both sides.
[0121] In other words, verifying whether the network connectivity status of the first node after a failure event is the same as the network connectivity status when it is in a non-failure state is essentially: verifying whether the network connectivity status of the first node after a failure event is connected.
[0122] In one possible implementation, the S703 implementation process may include: calling a network connectivity testing tool (such as ping, traceroute, or network packet capture) to detect the network connectivity between the first node and the interconnected nodes on both sides, in order to verify whether the network connectivity status of the first node is connected after a failure event.
[0123] S704. If the network connectivity status is different, mark the first node as the faulty node.
[0124] In one possible implementation, the S704 implementation process may include: marking the first node as a faulty node when the network connectivity status of the first node is disconnected after a fault event occurs.
[0125] S705. Filter out at least one second node other than the faulty node in at least one first node; for each second node, verify whether the system configuration of the second node after the fault event is the same as the system configuration of the node when it is in a non-faulty state.
[0126] It should be noted that the system configuration of a node may change due to the influence of other failed nodes. Therefore, the S705 described above can be used to further screen for failed nodes to ensure node availability.
[0127] S706. If the system configurations are different, the second node will be marked as a faulty node.
[0128] Based on the above technical solution, the configuration information, network connectivity status and system configuration of the nodes were detected and verified to accurately locate each faulty node. Different types of faulty nodes were also classified to clarify the type of fault of each node, which will facilitate subsequent maintenance.
[0129] As one possible embodiment of this application, combined with Figure 5 ,like Figure 8 As shown, after S503, the network fault detection method also includes the following S801-S802.
[0130] S801. When a node is a faulty node, based on the configuration information of the node in a non-faulty state, the configuration information of the node after the fault event is corrected so that the node changes from a faulty state to a non-faulty state.
[0131] In one possible implementation, the S801 implementation process includes: corresponding to different fault conditions of the faulty node, correcting the configuration information of the node after the fault event occurs through the corresponding solution, so that the node changes from a faulty state to a non-faulty state.
[0132] The following sections describe three different fault scenarios corresponding to the faulty node, from scenario 1 to scenario 3, and detail the process of correcting the configuration information of the node after a fault event occurs.
[0133] Scenario 1: When a network anomaly is caused by an abnormal configuration information of a node, the configuration information of the node after the failure event is replaced with the configuration information of the node when it is in a non-faulty state, so as to correct the configuration information of the node when it is in a non-faulty state.
[0134] Scenario 2: If the network failure is caused by an abnormal network connectivity status, log in to the node console and reconfigure the network of the faulty node to ensure the network connectivity of the node.
[0135] Scenario 3: If the network anomaly at a node is caused by an abnormal system configuration, import the system configuration of the second node when it is in a non-faulty state to complete the correction.
[0136] S802, Correction scheme for recording nodes.
[0137] The correction scheme includes the process of correcting configuration information when a node changes from a faulty state to a non-faulty state.
[0138] In one example, corresponding to situation 1 above, the correction scheme lists the configuration information in the abnormal state (i.e., configuration information different from the configuration information in the non-faulty state); and lists and explains the corrected configuration information. For example, configuration information A is corrected to configuration information B.
[0139] In another example, corresponding to case 2 above, the correction scheme lists the nodes whose network configurations have been corrected and their node network configuration information.
[0140] In another example, corresponding to case 3 above, the system configuration before and after correction is listed in the correction plan.
[0141] Based on the above technical solution, when a node is a faulty node, the configuration information of the node after the fault event is corrected based on the configuration information of the node when it is in a non-faulty state, so that the node changes from a faulty state to a non-faulty state, and the correction scheme of the node is recorded so that the network topology can be maintained based on the correction scheme when the network topology fails in the future.
[0142] As one possible embodiment of this application, such as Figure 9 As shown, the implementation scheme of network fault detection method for locating, classifying, resolving faults and generating correction schemes includes the following steps 1-14.
[0143] 1. Read and save the configuration information of each node when it is in a non-faulty state and the configuration information of each node after a fault event occurs in the network topology.
[0144] 2. Verify that the configuration information of the two is consistent.
[0145] 3. If the configuration information of the two is inconsistent, the node will be marked as a faulty node.
[0146] 4. Correct inconsistent configuration information (e.g., correct the configuration information of a node after a failure event to the configuration information of the node when it is in a non-failure state).
[0147] 5. Generate a correction plan.
[0148] 6. If the configuration information of the two is consistent, then network tests are performed on the two ends of the interconnected network in at least one first node to determine the network connectivity status.
[0149] 7. Filter out at least one faulty node from the first node and mark the faulty node.
[0150] 8. Log in to the node control panel and reconfigure the network for the faulty node.
[0151] 9. Generate the correction scheme corresponding to this node.
[0152] 10. Export the system configuration of the second node after the failure event.
[0153] 11. Determine whether the system configuration of the second node after a failure event is the same as the system configuration of the node when it is in a non-failure state.
[0154] 12. If the system configuration of the second node is different, then the node is identified as a faulty node, and the system configuration of the second node in a non-faulty state is imported to complete the correction.
[0155] 13. Generate a correction plan.
[0156] 14. Summarize the correction schemes generated in steps 5, 9 and 13 to obtain the overall correction scheme, so that the network topology can be maintained based on the summarized scheme when a network topology failure occurs in the future.
[0157] As one possible embodiment of this application, combined with Figure 8 ,like Figure 10 As shown, after S802, the network fault detection method also includes a step of detecting the availability of the network topology, specifically including the following S1001.
[0158] S1001. Repeat the first operation until the availability of each node in the network topology has been detected. Summarize the availability of each node in the network topology to obtain the detection result of the network topology.
[0159] In some embodiments, the first operation includes: for each node in the network topology, detecting whether the network connectivity status of the node is connected; if it is connected, marking the availability of the node as available; otherwise, marking the availability of the node as unavailable.
[0160] In one possible implementation, the first operation includes: invoking a network connectivity testing tool (e.g., ping, traceroute, or network packet capture) to send detection data to the node to test its network connectivity. If the network connectivity testing tool receives a response from the node, it determines that the node is available, and based on the event and status code of the received response, it can further determine the node's availability level. For example, if the duration of the received response exceeds a preset duration, it determines that the node's network performance is poor and its availability level is low. If the network connectivity testing tool does not receive a response from the node, it determines that the node is unavailable.
[0161] Based on the above technical solution, the configuration information of each node in the network topology is obtained, and the network connections between each node in the network topology are traversed. The first operation is repeated until the availability of each node in the network topology has been detected. The availability of each node in the network topology is then summarized to obtain the network topology detection result. This technical solution can verify the availability of each node in the corrected network topology to determine whether the correction scheme is effective.
[0162] As one possible embodiment of this application, such as Figure 11 As shown, the specific implementation process of verifying each node in the corrected network topology can also be achieved through the following steps 1-8.
[0163] 1. Obtain the configuration information of each node in the network topology.
[0164] 2. Traverse the network connections between each node in the network topology.
[0165] 3. Use network connectivity testing tools (such as ping, traceroute, or network packet capture) to send test data to the node to test the node's network connectivity.
[0166] 4. If the network connectivity testing tool receives a response from a node, it will mark the node's availability as available.
[0167] In this context, when the network connectivity testing tool receives a response from a node, it indicates that the node's network connectivity is normal. Furthermore, based on the event and status code of the received response, the availability of the node can be further determined. For example, if the duration of the received response exceeds a preset duration, it indicates that the node's network performance is poor and its availability is low.
[0168] 5. If the network connectivity testing tool does not receive a response from the node, then the availability of that node is determined to be unavailable.
[0169] 6. Continue traversing the network topology to the next node.
[0170] 7. Determine if all nodes in the network topology have been traversed; if yes, proceed to step 8; otherwise, proceed to step 2 to continue traversing the next node in the network topology.
[0171] 8. Summarize the availability of each node in the network topology to obtain the network topology detection results.
[0172] Understandably, this technical solution uses network connectivity testing tools to quickly identify the cause and type of fault and automatically generate solutions, which shortens the time for troubleshooting and improves the accuracy and effectiveness of the solutions.
[0173] This application embodiment can divide the network fault location device into functional modules or functional units according to the above method example. For example, each function can be divided into a separate functional module or functional unit, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or in software functional modules or functional units. The module or unit division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0174] like Figure 12 The diagram shown is a structural schematic of a network fault location device 120 provided in an embodiment of this application. The network fault location device 120 includes a processing unit 1201.
[0175] Processing unit 1201 is used to establish the network topology corresponding to the target network and record the configuration information of each node in the network topology when it is in a non-fault state; processing unit 1201 is also used to generate a fault event and inject the fault event into the network topology, and record the configuration information of each node in the network topology after the fault event occurs; processing unit 1201 is also used to verify the configuration information of each node after the fault event occurs based on the configuration information of each node when it is in a non-fault state, so as to locate the faulty node among the nodes.
[0176] In one possible implementation, the configuration information includes multiple types of configuration items; the processing unit 1201 is specifically used to: for each node in each node, based on the multiple types of configuration items when the node is in a non-faulty state, verify each configuration item of the node after the fault event occurs, obtain the verification result, and locate the faulty node in the network topology based on the verification result.
[0177] In one possible implementation, the configuration items include at least one of the following: node topology configuration, node running status configuration, and node system configuration.
[0178] In one possible implementation, the processing unit 1201 is specifically configured to: verify whether the topology configuration of a node after a failure event is the same as its topology configuration in a non-failure state, and verify whether the operating state configuration after the failure event is the same as its operating state configuration in a non-failure state, and verify whether the system configuration after the failure event is the same as its system configuration in a non-failure state; if the topology configuration, operating state configuration, and / or system configuration are different, then the node is marked as a failure node, and at least one first node other than the failure node in the network topology is selected; for each first node, verify whether the network connectivity state of the first node after the failure event is the same as its network connectivity state in a non-failure state; if the network connectivity state is different, then the first node is marked as a failure node, and at least one second node other than the failure node among the at least one first node is selected; for each second node, verify whether the system configuration of the second node after the failure event is the same as its system configuration in a non-failure state; if the system configuration is different, then the second node is marked as a failure node.
[0179] In one possible implementation, the processing unit 1201 is further configured to: when the node is a faulty node, perform corrective processing on the configuration information of the node after the fault event occurs based on the configuration information of the node when it is in a non-faulty state, so that the node changes from a faulty state to a non-faulty state; record the correction scheme of the node; the correction scheme includes the correction process of the configuration information when the node changes from a faulty state to a non-faulty state.
[0180] In one possible implementation, after all faulty nodes in the network topology have been corrected, the processing unit 1201 is specifically used to: repeatedly execute the first operation until the availability of all nodes in the network topology has been detected; and summarize the availability of all nodes in the network topology to obtain the detection result of the network topology.
[0181] The first operation includes: for each node in the network topology, checking whether the node's network connectivity status is connected; if it is connected, marking the node's availability as available; otherwise, marking the node's availability as unavailable.
[0182] In one possible implementation, the fault event includes at least one of the following: network logical fault event, physical fault event, and protocol fault event; wherein, network logical fault events include simulated configuration error events and router port shutdown events; physical fault events include power outage events and network outage events; and protocol fault events include router fault events and firewall parameter setting error events.
[0183] In one possible implementation, the network fault location device 120 may further include a storage unit 1202. Figure 12 (shown in dashed box) The storage unit 1202 stores a program or instruction. When the processing unit 1201 executes the program or instruction, the network fault location device 120 can perform the network fault location method described in the above method embodiment.
[0184] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0185] This application provides a computer program product containing instructions that, when run on a computer, cause the computer to execute the network fault location method described in the above method embodiments.
[0186] This application also provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the network fault location method in the method flow shown in the above method embodiments.
[0187] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires; a portable computer disk drive; a hard disk drive; a random access memory (RAM); a read-only memory (ROM); an erasable programmable read-only memory (EPROM); a register; a hard disk drive; an optical fiber; a portable compact disc read-only memory (CD-ROM); an optical storage device; a magnetic storage device; or any suitable combination thereof; or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0188] Since the network fault location device, computer-readable storage medium, and computer program product in the embodiments of this application can be applied to the above method, the technical effects that can be obtained can also be referred to the above method embodiments. The embodiments of this application will not be repeated here.
[0189] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0190] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0191] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
[0192] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0193] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0194] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for locating network faults, characterized in that, The method includes: Establish the network topology corresponding to the target network, and record the configuration information of each node in the network topology when it is in a non-faulty state; A fault event is generated and injected into the network topology, and the configuration information of each node in the network topology is recorded after the fault event occurs; Based on the configuration information of each node in a non-faulty state, the configuration information of each node after the fault event occurs is verified in order to locate the faulty node among the nodes. The configuration information includes multiple types of configuration items; the configuration items include: the topology configuration of the node, the running status configuration of the node, and the system configuration of the node; The step of verifying the configuration information of each node after the failure event occurs, in order to locate the faulty node among the nodes, includes: For each of the nodes, it is verified whether the topology configuration of the node after the failure event is the same as the topology configuration in the non-failure state, and whether the running state configuration after the failure event is the same as the running state configuration in the non-failure state, and whether the system configuration after the failure event is the same as the system configuration in the non-failure state. If the topology configuration is different, the running state configuration is different, and / or the system configuration is different, then the node is marked as the fault node, and at least one first node other than the fault node in the network topology is selected; for each first node, it is verified whether the network connectivity status of the first node after the fault event occurs is the same as the network connectivity status when it is in a non-fault state. If the network connectivity states are not the same, the first node is marked as the faulty node, and at least one second node other than the faulty node among the at least one first node is selected; for each second node, it is verified whether the system configuration of the second node after the fault event occurs is the same as the system configuration of the node when it is in a non-faulty state; If the system configurations are different, the second node will be marked as the faulty node.
2. The method according to claim 1, characterized in that, The step of verifying the configuration information of each node after the failure event occurs, in order to locate the faulty node among the nodes, includes: For each of the nodes, based on the multiple types of configuration items when the node is in a non-faulty state, each configuration item of the node is verified one by one after the fault event occurs, and the verification result is obtained. Based on the verification result, the faulty node in the network topology is located.
3. The method according to claim 1, characterized in that, The method further includes: In the case that the node is a faulty node, based on the configuration information of the node when it is in a non-faulty state, the configuration information of the node after the fault event occurs is corrected so that the node changes from a faulty state to a non-faulty state. Record the correction plan for the node; the correction plan includes the correction process of the configuration information when the node changes from a fault state to a non-fault state.
4. The method according to claim 3, characterized in that, After corrective action has been taken at each faulty node in the network topology, the method further includes: Repeat the first operation until the availability of each node in the network topology has been detected. Summarize the availability of each node in the network topology to obtain the detection result of the network topology. The first operation includes: For each node in the network topology, detect whether the network connectivity status of the node is connected; If the node is connected, mark its availability as available; otherwise, mark its availability as unavailable.
5. The method according to claim 1, characterized in that, The failure event includes at least one of the following: network logical failure event, physical failure event, and protocol failure event; The network logical fault events include simulated configuration error events and router port shutdown events; the physical fault events include power outage events and network outage events; and the protocol fault events include router fault events and firewall parameter setting error events.
6. A network fault location device, characterized in that, The device includes a processing unit; The processing unit is used to establish the network topology corresponding to the target network and record the configuration information of each node in the network topology when it is in a non-faulty state. The processing unit is also configured to generate a fault event, inject the fault event into the network topology, and record the configuration information of each node in the network topology after the fault event occurs; The processing unit is further configured to verify the configuration information of each node after the fault event occurs, based on the configuration information of each node when it is in a non-fault state, so as to locate the faulty node among the nodes. The configuration information includes multiple types of configuration items; the configuration items include: the topology configuration of the node, the running status configuration of the node, and the system configuration of the node; The processing unit is specifically used for: For each of the nodes, it is verified whether the topology configuration of the node after the failure event is the same as the topology configuration in the non-failure state, and whether the running state configuration after the failure event is the same as the running state configuration in the non-failure state, and whether the system configuration after the failure event is the same as the system configuration in the non-failure state. If the topology configuration is different, the running state configuration is different, and / or the system configuration is different, then the node is marked as the fault node, and at least one first node other than the fault node in the network topology is selected; for each first node, it is verified whether the network connectivity status of the first node after the fault event occurs is the same as the network connectivity status when it is in a non-fault state. If the network connectivity states are not the same, the first node is marked as the faulty node, and at least one second node other than the faulty node among the at least one first node is selected; for each second node, it is verified whether the system configuration of the second node after the fault event occurs is the same as the system configuration of the node when it is in a non-faulty state; If the system configurations are different, the second node will be marked as the faulty node.
7. A communication device, characterized in that, include: A processor and a communication interface; the communication interface is coupled to the processor, the processor being used to run computer programs or instructions to implement the network fault location method as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed by a computer, perform the network fault location method as described in any one of claims 1-5.
9. A computer program product, characterized in that, The computer program product includes computer instructions that, when executed on a computer, cause the computer to perform the network fault location method as described in any one of claims 1-5.
Citation Information
Patent Citations
Method and equipment for verifying network equipment configuration in cloud network environment, and medium
CN113938378A
Method for acquiring network topology and indoor distribution system
US20240106732A1