Fault positioning method and device based on network architecture, equipment and storage medium

By generating a detection path in the switching network architecture and using the IPv4 source path option for packet loss detection, the problem of difficult fault location caused by switch forwarding anomalies is solved, achieving fast and accurate fault location and ensuring the stability of the data center.

CN120692148APending Publication Date: 2025-09-23MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510897233.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In a switching network architecture, when a core layer switch or access layer switch experiences forwarding anomalies, the fault cannot be automatically detected immediately, resulting in the inability to accurately locate the faulty board, affecting the stable operation of the data center.

Method used

By querying the connection relationship between the access layer and the core layer, multiple detection paths are generated, and the IPv4 source path option is used to control packet forwarding, perform packet loss detection, and determine the faulty network node based on the detection results.

Benefits of technology

It can quickly and accurately locate faulty equipment in the network architecture, improve the efficiency and accuracy of fault location, reduce the duration of faults, and ensure the stable operation of the data center.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692148A_ABST
    Figure CN120692148A_ABST
Patent Text Reader

Abstract

The invention provides a fault positioning method and device based on a network architecture, equipment and a storage medium. The method comprises the steps that the network architecture comprises a core layer and an access layer, the core layer comprises a plurality of first switches, a second switch connected with the first switches in the access layer is inquired, and a server connected with the second switch is inquired; a plurality of detection paths are generated, starting points and end points of the detection paths are two different servers, and the detection paths pass through the first switch and the second switch; performing packet loss detection on each detection path to obtain a detection result of each detection path; and based on the detection result of each detection path, a network node with a fault in the network architecture is determined, and the network node comprises at least one of the first switch, the second switch and the server. According to the invention, the efficiency and accuracy of fault positioning in a network architecture can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a fault location method, apparatus, device, and storage medium based on a network architecture. Background Art

[0002] With the development of cloud computing and large-scale data centers, network topologies based on switching architectures are becoming mainstream. Switching architectures divide complex networks into several layers, each focusing on specific functions. The access layer uses dual or multiple uplinks to perform load-balanced forwarding with the core layer using equal-cost multi-path routing. This load-balancing approach achieves load balancing across different paths. However, when forwarding anomalies occur on core or access layer switches, there are often no clear alarms, preventing automatic fault detection.

[0003] Related technologies require manual location of a large number of devices and boards one by one to troubleshoot hidden faults in network topology. When packet loss occurs or network quality degrades, the detection path is uncontrollable and the faulty board cannot be accurately locked in the shortest time, affecting the stable operation of the data center. Summary of the Invention

[0004] The embodiments of the present application provide a network architecture-based fault location method, apparatus, device, and storage medium, which can improve the efficiency and accuracy of fault location in the network architecture.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] An embodiment of the present application provides a fault location method based on a network architecture, wherein the network architecture includes a core layer and an access layer, and the core layer includes a plurality of first switches; the method includes:

[0007] Querying a second switch in the access layer connected to the first switch, and querying a server connected to the second switch;

[0008] generating a plurality of detection paths, wherein a starting point and an end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch;

[0009] Performing packet loss detection on each of the detection paths to obtain a detection result for each of the detection paths;

[0010] Based on the detection result of each detection path, a faulty network node in the network architecture is determined, wherein the network node includes at least one of the first switch, the second switch, and the server.

[0011] The present invention provides a network-based fault location device, including:

[0012] a path generation module, configured to query a second switch connected to the first switch in the access layer, and query a server connected to the second switch; and generate multiple detection paths, wherein the starting point and the end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch;

[0013] A path detection module, configured to perform packet loss detection on each detection path and obtain a detection result for each detection path;

[0014] A fault location module is used to determine a faulty network node in the network architecture based on the detection results of each detection path, wherein the network node includes at least one of the first switch, the second switch, and the server.

[0015] An embodiment of the present application provides an electronic device, comprising:

[0016] a memory for storing computer-executable instructions or computer programs;

[0017] The processor is configured to implement the network architecture-based fault location method provided in the embodiment of the present application when executing the computer-executable instructions or computer program stored in the memory.

[0018] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing a network architecture-based fault location method provided in an embodiment of the present application when executed by a processor.

[0019] An embodiment of the present application provides a computer program product, including a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by a processor, the network architecture-based fault location method provided in the embodiment of the present application is implemented.

[0020] The embodiments of the present application have the following beneficial effects:

[0021] The system queries the second switches connected to the first switch in the access layer and the servers connected to these second switches, analyzes the connection relationships between the core layer and the access layer, and between the access layer and the servers, and establishes associations between key devices and terminals in the network architecture. This lays the foundation for subsequent detection path generation and ensures that the detection covers relevant network nodes. Multiple detection paths are generated, starting and ending at different servers and passing through the first switch in the core layer and the second switch in the access layer. This allows detection signals to be effectively transmitted between key layers and nodes in the network architecture, ensuring that subsequent packet loss detection is performed along known and defined data transmission paths, facilitating the precise identification of potential fault points. Performing packet loss detection on each detection path directly reflects the data transmission quality along each detection path. Based on the detection results of each detection path, faulty network nodes in the network architecture are identified. By analyzing packet loss, a precise mapping from detection results to faulty nodes is achieved, enabling the rapid location of the specific network device experiencing the fault, improving the efficiency and accuracy of fault location within the network architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of an application mode of a network architecture-based fault location method provided in an embodiment of the present application;

[0023] Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application;

[0024] Figure 3A This is a schematic diagram of a first flow chart of a network architecture-based fault location method provided in an embodiment of the present application;

[0025] Figure 3B This is a second flow chart of the network architecture-based fault location method provided in an embodiment of the present application;

[0026] Figure 4A This is a third flow chart of the network architecture-based fault location method provided in an embodiment of the present application;

[0027] Figure 4B This is a fourth flow chart of a network architecture-based fault location method provided in an embodiment of the present application;

[0028] Figure 5 This is a schematic diagram of the network architecture provided by the embodiment of the present application;

[0029] Figure 6 This is a schematic diagram of the detection path structure provided by an embodiment of the present application;

[0030] Figure 7 This is a schematic diagram of the detection result analysis provided in the embodiment of the present application.

[0031] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of distinction between the advantages and disadvantages of the solutions or the priority in the implementation process. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0033] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0034] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0035] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0036] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0037] The collection and processing of relevant data (e.g., network architecture information) in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations when applied in instances, and the informed consent or separate consent of the personal information subject should be obtained. Subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.

[0038] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0039] 1) Internet Protocol v4 (IPv4): This is a connectionless protocol that operates on a packet-switched link layer (such as Ethernet). IPv4 uniquely identifies and locates devices through 32-bit IPv4 addresses, enabling data packets to be correctly forwarded from one device to another, providing basic IP network connectivity and communication capabilities.

[0040] 2) IPv4 Source Path Options: This is a special optional field in the IPv4 protocol that allows the source host to specify the routing path for a data packet from source to destination, rather than leaving it entirely to intermediate routers based on routing tables. It is used to embed a set of preset hop IP addresses in the data packet, controlling the hop-by-hop forwarding of the packet along the specified path within the network. When a network device receives such a packet, it modifies the destination address according to the order specified in the path options and forwards the packet to the next node in the path until it reaches the final destination.

[0041] 3) Internet Control Message Protocol (ICMP): It is an important sub-protocol in the TCP / IP protocol suite. Located at the network layer, it is mainly used to transmit control messages between IP hosts and routers and report various problems in network communications. Control messages refer to network messages such as whether the network is connected, whether the host is reachable, and whether the route is available.

[0042] 4) Clos Network Architecture (Clos): The network architecture is designed using a hierarchical model. The network architecture design includes: a core layer (interconnecting all access layer switches in a full mesh topology) and an access layer (aggregating traffic from servers and connecting directly to the core layer).

[0043] 5) Equal-Cost Multi-Path Routing (EMCP): This refers to the existence of multiple paths with equal costs to the same destination address. When the device supports equal-cost routing, forwarding traffic sent to the destination IP or destination network segment can be shared across different paths, achieving load balancing of network links and enabling rapid switching when a link fails.

[0044] 6) Detection flow: The detection flow is an end-to-end Internet packet explorer path of the Internet Control Message Protocol. In the embodiment of the present application, the path of the detection flow passes through the source server A → access layer switch A → core layer switch A inlet → core layer switch A outlet → access layer switch B → target server B, and the IP of each hop is explicitly specified through the IPv4 source path option.

[0045] 7) Packet Internet Groper (Ping): A connectivity check mechanism based on the Internet Control Message Protocol (ICMP), widely used to test the reachability between two network nodes. The basic principle of an IPG is that a source host sends an ICP message to a destination host. If the destination host is online and reachable, an ICP message is returned. By measuring the response, the PG can be used to quickly determine network connectivity.

[0046] With the development of cloud computing and large-scale data centers, network topologies using switching network architectures have gradually become mainstream. Switching network architectures divide complex networks into several layers. The access layer uses dual or multiple uplinks to perform load balancing forwarding with equal-cost multi-path routing with the core layer. By sharing traffic across different paths, load balancing of network links is achieved. When forwarding anomalies occur on the core layer switch or access layer switch board, there is often a lack of obvious alarms, resulting in the inability to automatically detect faults in the first place.

[0047] Related art solutions for probing network topology connectivity typically use Internet Packet Explorer (IPX) messages constructed using the Internet Control Message Protocol (ICMP). These messages are transmitted between two servers based on a default forwarding table. In a data center network with a switched network architecture, multiple equal-cost paths exist between core switches or access switches. The uplink path from the access switch to the core switch is the hashed result of multiple links. Each probe of the same target may pass through different core switch nodes or interfaces. If a switch card or interface fails, the IPX message in related art will pass through the faulty path during the probe, resulting in connectivity failure. However, equal-cost paths are uncontrollable and easily affected by the load mechanism of equal-cost paths. This makes it impossible to accurately determine the abnormal probe path in the shortest possible time. The forwarding path of the data packet performing connectivity probing in the probe path is dynamic, making it difficult to trace back the faulty path, determine the intermediate nodes, and accurately locate the faulty node in the faulty path. This results in a prolonged fault, impacting the stable operation of the data center.

[0048] The embodiments of the present application provide a network architecture-based fault location method, a network architecture-based fault location device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the efficiency and accuracy of fault location in the network architecture.

[0049] The following describes exemplary applications of the electronic devices provided in the embodiments of the present application. The devices provided in the embodiments of the present application can be implemented as various types of terminals, such as laptops, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, and in-vehicle terminals. They can also be implemented as servers. The following describes exemplary applications when the devices are implemented as terminals or servers.

[0050] See also Figure 1 , Figure 1 This is a schematic diagram of an application mode of a network architecture-based fault location method provided in an embodiment of the present application. In order to support a network architecture-based fault location application, an example is provided. Figure 1 The server 200, network 300, terminal device 400 and database 500 are involved. The terminal device 400 is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0051] In some embodiments, the embodiments of the present application can be implemented collaboratively by a server and a terminal device. For example, terminal device 400 sends a packet loss detection request to server 200 via network 300. Server 200 receives the packet loss detection request, and database 500 stores detection path information. The network architecture-based fault location method provided in the embodiments of the present application can locate the faulty network node based on the result of executing the packet loss detection request.

[0052] See also Figure 2 , Figure 2 is a structural diagram of an electronic device provided in an embodiment of the present application, Figure 2 The server 200 shown includes: at least one processor 410, a memory 450 and at least one network interface 420. The various components in the terminal device 400 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, Figure 2 Various buses are labeled as bus system 440 .

[0053] The processor 410 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0054] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.

[0055] The memory 450 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0056] In some embodiments, the memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0057] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and process hardware-based tasks;

[0058] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420 . Exemplary network interfaces 420 include Bluetooth, Wireless LAN (WiFi), and Universal Serial Bus (USB).

[0059] In some embodiments, the apparatus provided in the embodiments of the present application may be implemented in software. Figure 2 A network architecture-based fault location device 455 stored in memory 450 is shown. This device can be software in the form of a program or plug-in, and includes the following software modules: a path generation module 4551, a path detection module 4552, and a fault location module 4553. These modules are logical and can be arbitrarily combined or further separated based on the functions they implement. The functions of each module will be described below.

[0060] In some embodiments, the terminal or server can implement the network architecture-based fault location method provided in the embodiment of the present application by running various computer executable instructions or computer programs. For example, computer executable instructions can be microprogram-level commands, machine instructions or software instructions. The computer program can be a native program or software module in the operating system; it can be a local (Native) application (APPlication, APP); it can also be a small program that can be embedded in any APP, that is, a program that can be run only by downloading it to a browser environment. In short, the above-mentioned computer executable instructions can be instructions in any form, and the above-mentioned computer program can be an application, module or plug-in in any form.

[0061] The network architecture-based fault location method provided in the embodiment of the present application will be explained in combination with the exemplary application and implementation of the terminal provided in the embodiment of the present application.

[0062] The following describes the network architecture-based fault location method provided by an embodiment of the present application. As previously mentioned, the electronic device that implements the network architecture-based fault location method of the embodiment of the present application can be a terminal, a server, or a combination of the two. Therefore, the execution entity of each step will not be repeated below.

[0063] See also Figure 3A , Figure 3A This is a first flow chart of the fault location method based on network architecture provided by the embodiment of the present application, which will be combined with Figure 3A The steps shown are explained, Figure 3A The executive body is Figure 1 Server 200 in.

[0064] In step 301, a second switch connected to a first switch in an access layer is queried, and a server connected to the second switch is queried.

[0065] As an example, the network architecture is a switched network architecture (Clos Network Architecture, Clos). The network architecture adopts a hierarchical model design, dividing the complex network design into several layers, each focusing on certain specific functions. The network architecture design includes: a core layer and an access layer. The core layer includes multiple first switches, and the access layer includes multiple second switches. Each first switch in the core layer is connected to at least one second switch in the access layer. The second switch connected to the first switch in the access layer is queried to determine the connection relationship between the first switch and the second switch. Each second switch is connected to at least one server. The server connected to each second switch is queried to determine the connection relationship between the second switch and the server.

[0066] As an example, step 301 is actually to traverse the entire network architecture in a top-down order before generating the detection path. Specifically, first determine the first switch in the network architecture whose interface is in the "up" (activated, connected) state. The first switch whose interface is in the "up" state can connect to the next level of network equipment. Traverse to the next layer along the interface of the first switch in the "up" state, query the second switch connected to the interface of the first switch, and the second switch further distributes the network signal of the first switch to each terminal area. Then continue to traverse the second switch whose interface is in the "up" state, and query the server connected to the interface of the second switch. The server is the data processing and storage core in the network, and is usually connected to a specific port of the second switch. Through the above traversal, a complete network structure reflecting the connection relationship between the first switch, the second switch and the server is obtained, which provides a basis for the subsequent construction of the detection path.

[0067] In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the network architecture provided in an embodiment of the present application. The network architecture 501 includes a core layer 502 and an access layer 503. The core layer 502 includes multiple first switches, for example, a first switch A and a first switch B. The access layer 503 includes multiple second switches, for example, a second switch A to a second switch n. The multiple second switches are connected to at least one server, for example, server A to server n.

[0068] In step 302, a plurality of detection paths are generated.

[0069] As an example, the starting point and the end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch.

[0070] In some embodiments, step 302 can be implemented by the following method: randomly selecting two servers from the network architecture, and querying the first switch and the second switch connecting the two servers; based on the queried Internet Protocol addresses of the first switch and the second switch, and the Internet Protocol addresses of the two servers, generating a detection path.

[0071] As an example, the number of detection paths is a preset number that can cover all services. The starting point and end point of the detection path are two randomly selected servers in the network architecture. These two servers will serve as the source server and the target server respectively. Starting from the source server, the first switch and the second switch connected to the two servers are queried. The detection path passes through the first switch and the second switch to reach another target server. For the source server, query the connected second switch (access layer switch), and further query the first switch (core layer switch) connected to the second switch. For the target server, query the connected second switch, and then query the first switch connected to the second switch. Obtain the Internet Protocol addresses of the first switch and the second switch passed by the query, as well as the Internet Protocol addresses of the two servers in the network architecture, combine them according to the logic and order of network data transmission, repeat the above steps, and generate a preset number of multiple detection paths. The multiple paths can be generated serially or in parallel, and this application does not impose any restrictions on this.

[0072] Through the embodiments of the present application, by randomly selecting different server pairs and querying the first switch and the second switch connected to these servers, multiple detection paths are generated, which can cover different server and switch combinations, monitor the connectivity of different links and nodes in the network, and promptly discover potential network problems.

[0073] In some embodiments, querying the first switch and the second switch connected to two servers can be achieved by the following method: querying the first second switch connected to the first server; querying each first switch connected to the first second switch, and querying the second second switch connected to the first switch and not connected to the first server.

[0074] As an example, run a command line to query all first and second switches connected to the first server that are in the "up" state and record the IP addresses of the corresponding second switch interfaces. "Up" indicates that the board is powered on or operational; only second switches in the "up" state can function normally and participate in data transmission. Then, query each first switch connected to the first second switch and run a command line to query all second second switches connected to the first switch but not to the first server, recording the corresponding Internet Protocol addresses of the interfaces. This gradually builds a complete list of switch paths connecting the two servers to generate a detection path.

[0075] Through the embodiment of the present application, by querying the first second switch connected to the first server and recording the IP address of its interface, the starting point of the network path is ensured to be accurate, and each first switch connected to the first second switch is queried. The second second switches connected to these first switches and not connected to the first server are further queried, and the network path is gradually expanded to build a complete switch path list. Pay attention to the second switch in the "up" state to ensure that the selected network device can work normally and participate in data transmission. Avoid invalid path construction caused by selecting unavailable devices and improve the effectiveness of the detection path.

[0076] In step 303, packet loss detection is performed on each detection path to obtain a detection result for each detection path.

[0077] In some embodiments, any two different servers include a first server and a second server, and step 303 can be implemented by the following method: based on the order of the network nodes passed by the detection path representation, the Internet Protocol address of the first server, the Internet Protocol address of the first switch, the Internet Protocol address of the second switch, and the Internet Protocol address of the second server are written into the header of the data packet; based on the header of the data packet, the data packet is forwarded and processed to obtain the detection result of the detection path.

[0078] As an example, for each detection path, based on the order of the network nodes passed through as represented by the detection path, the Internet Protocol address of the first server, the Internet Protocol address of the first switch, the Internet Protocol address of the second switch, and the Internet Protocol address of the second server are written into the header of the data packet according to the source path option of the fourth version of the Internet Protocol (IPv4). The source path option of the fourth version of the Internet Protocol (IPv4) specifies the order of the network nodes passed through. The IPv4 source path option is a special optional field in the IPv4 protocol that specifies the IP address list of the intermediate nodes in the detection path. It allows the source host to specify the routing path of the data packet from the source to the destination, rather than having the intermediate routers independently determine it based on the routing table. It is used to embed a set of preset hop IP addresses in the data packet to control the hop-by-hop forwarding of the packet along the specified path in the network.

[0079] The format of the IPv4 source path option includes: Type (1 byte): 131 (LSRR) or 137 (SS RR), Length (1 byte): total length of the option (including type, length, pointer and IP list), Pointer (1 byte): indicates the location of the IP address that needs to be processed, IP List (4 bytes × N): IP address of the intermediate node.

[0080] LSRR is a loose source routing record, and packets are forwarded according to the series of router addresses specified in the LSRR. The sending host lists the IP addresses of multiple routers in the packet. When a packet is forwarded from one router to the next, it checks whether the IP address of the current router matches the next router address specified in the LSRR. If so, the packet is forwarded to the next designated router. If not, the router forwards the packet according to its own routing table until it reaches the next router specified in the LSRR. SSRR is a strict source routing record, and packets must be forwarded strictly according to the router addresses specified in the SSRR. Each forwarding step must pass through a designated router. If a packet encounters a router whose IP address does not match the next router address specified in the SSRR during forwarding, the packet is discarded.

[0081] The forwarding process is performed on the data packets to be forwarded, based on the order of the network nodes along the probe path represented by the source path option. This process is achieved by sending an Internet Control Message Protocol (ICMP) probe packet via a command line tool. The ICP is a key subprotocol in the TCP / IP protocol suite, located at the network layer and primarily used to transmit control messages between IP hosts and routers. The ICP sends a command to the target server, sending a series of request messages (i.e., multiple data packets). The probe results returned by the probe packets are used to determine whether there are any anomalies in the probe path. The ICP (Packet Internet Groper) is a connectivity detection mechanism based on the ICP. It is widely used to test the reachability between two network nodes and returns an ICP message as the detection result of the probe path.

[0082] For example, see Figure 6, set the IP address of server 6031 to 5.5.5.5, the IP address of server 6032 to 6.6.6.6, and specify in the IPv4 source path option: 1st hop: 35.1.1.3 (the interface connecting the second switch 6021 and server 6031), 2nd hop: 13.1.1.1 (H1 / 0 / 1 port of card A of the first switch 6011), 3rd hop: 14.1.1.4 (the interface connecting the second switch 6022 and the first switch 6011), 4th hop: 4 6.1.1.6 (the interface connecting server 6032 and second switch 6022). At this time, a detection path with a fixed forwarding order is constructed: server 6031 → second switch 6021 → H1 / 0 / 1 interface of board A of first switch 6011 → H1 / 0 / 2 interface of board A of first switch 6011 → second switch 6022 → server 6032 (i.e., the red detection path 6041). The order of the data packet forwarding nodes corresponding to the above detection path is written into the data packet header.

[0083] Starting from the first server (source server), the data packet sends a request message in the order represented by the detection path. The second server (target server) will reply with a response message. The detection result shows whether each data packet forwarded according to the specified IP sequence has received an echo response, including the timeout period. If there is no reply message, a timeout message will be displayed. At this time, there is an abnormality in the detection path, and the abnormal path information is recorded. The detection path has packet loss. When an abnormality is found in the detection path (such as packet loss), the abnormal detection path information is recorded, including the IP address and interface information of each network node along the detection path, as well as the specific location where the abnormality occurred during the detection process, so as to facilitate subsequent troubleshooting and analysis.

[0084] Through the embodiments of the present application, the IP addresses of the source server, switch, and target server are written into the packet header in sequence, and the forwarding path of the packet is controlled by the source path option of IPv4, thereby ensuring the accuracy and controllability of the detection. Based on the ICMP protocol, a detection packet is sent and the return results are analyzed to promptly detect packet loss in the network and effectively identify fault points or performance bottlenecks in the network. When packet loss is found, detailed path information and abnormal location are recorded to provide a basis for subsequent fault repair, help quickly locate and solve the problem, and reduce the impact of network failures on the business.

[0085] In step 304, based on the detection results of each detection path, a faulty network node in the network architecture is determined.

[0086] In some embodiments, the network node includes at least one of a first switch, a second switch, and a server. Figure 3B , Figure 3B1 is a second flow chart of a fault location method based on network architecture provided in an embodiment of the present application. Figure 3A Step 304 in the above example can be accomplished by Figure 3B Steps 3041 to 3042 in the embodiment are implemented as described below.

[0087] In step 3041, when the detection result of the detection path is abnormal, the first switch, the second switch, and the server included in the detection path are used as network nodes to be processed.

[0088] As an example, when the detection result of the detection path is abnormal, it indicates that there is packet loss in the data packet forwarding in the detection path, and no echo information returned by the target server is received. The network nodes passed in the backtracking detection path are the nodes to be processed, which are nodes in the detection path that have not been checked in detail before or may have faults, which can be the first switch or the second switch.

[0089] In step 3042, fault location processing is performed on the network node to be processed to obtain the network node with faults in the network architecture.

[0090] In some embodiments, step 3042 can be implemented by the following method: when the network node to be processed is the first switch, and the detection results of all detection paths passing through the first switch are abnormal, the first switch is determined to be a faulty network node; when the network node to be processed is the second switch, and only the detection results of the detection paths passing through the second switch are abnormal, the second switch is determined to be a faulty network node; when the network node to be processed is the server, and only the detection results of all detection paths passing through the server are abnormal, the server is determined to be a faulty network node.

[0091] As an example, fault location processing is performed on the network node to be processed, and the detection results of multiple equal-cost routing paths passing through the same first switch are analyzed. Equal-cost multi-path routing means that there are multiple paths with equal costs to the same destination address. When the second switch in the detection path can forward data packets normally, it is determined that the network node to be processed is the first switch. When the detection results of all detection paths passing through the first switch are abnormal, it indicates that the first switch cannot forward the data packet normally, resulting in the detection path failing to reach the second server on time, and the second server does not return message information. It is determined that the first switch is a faulty network node, and the status information of each interface of the first switch is checked, including the physical status of the interface (such as whether it is in the "up" state), the link layer status (such as whether there is an error count), and the sending and receiving statistics of the data packet for subsequent fault repair.

[0092] For example, see Figure 7 The two equal-cost path routes of the detection path 7041 and the detection path 7042 pass through the first switch. The analysis of the detection results of the detection path passing through the first switch 7011 is shown in Table 1 below, which is described in detail below.

[0093] Table 1

[0094]

[0095] Among them, × means that the detection result is not connected and there is packet loss. The detection result indicates connectivity, no packet loss, and that both detection path 7041 and detection path 7042 pass through second switch 7021 and second switch 7022. The source server and target server are the same, indicating that second switch 7021 and second switch 7022 are both functioning normally. At this point, the network node to be processed is determined to be first switch 7011. The detection result through first switch 7011 indicates packet loss, while the detection result through first switch 7012 is normal. Therefore, the faulty node can be determined to be first switch 7011. Similarly, if the detection result of detection path 7041 indicates no packet loss, while the detection result of detection path 7042 indicates packet loss, the faulty network node can also be determined to be first switch 7012.

[0096] As an example, if after analyzing the detection results of multiple equal-cost routing paths passing through the same second switch, it is determined that the network node to be processed is the second switch, it indicates that the first switch passed through in the detection path is normal and the second switch may have a fault. Check the status information of each interface of the second switch. If only the detection result of the detection path passing through the second switch is abnormal, then the second switch is determined to be a faulty network node.

[0097] For example, detection path 7043 and detection path 7044 both pass through the first switch 7011 and the second switch 7023. The source server and the target server are the same, namely server 7031 and server 7033. The analysis of the detection results is shown in Table 2 below, which is described in detail below.

[0098] Table 2

[0099]

[0100] Among them, × means that the detection result is not connected and there is packet loss. It indicates that the detection result is connected and there is no packet loss. The detection result of detection path 7043 shows packet loss, and the detection result of detection path 7044 does not show packet loss. Detection path 7043 passes through the second switch 7021, and detection path 7044 passes through the second switch 7023. The detection results of the first switch 7011 and the second switch 7023 are both normal. At this time, the network node to be processed is the second switch 8021. The detection result of the detection path 7043 passing through the second switch 7021 is abnormal. Therefore, it can be determined that the faulty network node is the second switch 7021.

[0101] When the detection results of the first switch and the second switch through which the detection path passes are normal, it is determined that the network node to be processed is the server. The detection results of the detection path that only passes through the server are analyzed. If all detection results are abnormal, the server is determined to be a faulty network node.

[0102] Through the embodiments of the present application, by classifying and analyzing different types of switches, the detection results of equivalent routing paths passing through the nodes to be processed are compared and analyzed, network nodes that can normally forward data packets are excluded, the nodes to be processed are determined, and the faulty network nodes are accurately located based on the detection results of the detection paths passing through the nodes to be processed, thereby improving the accuracy of fault location and reducing the possibility of misjudgment.

[0103] Through the embodiments of the present application, by analyzing the detection results of each detection path, the first switch, the second switch and the server in the detection path with the abnormality can be accurately identified as the network nodes to be processed, the fault scope can be quickly narrowed down, and the focus can be placed on the switch that may have the fault, rather than blindly searching the entire network, thereby improving the efficiency of fault location and reducing the time and workload of troubleshooting.

[0104] In some embodiments, after step 304 , the fault may be eliminated by performing at least one of the following: isolating the faulty network node; or switching the detection flow to another detection path.

[0105] As an example, after determining the faulty network node, the corresponding interface of the faulty switch is isolated or switched to other equivalent detection paths to eliminate the anomaly of the detection path. The determined faulty network node is isolated from the network to prevent it from continuing to affect the normal operation of the network. This is achieved by disconnecting it from adjacent nodes. For example, if the first switch fails, its interface with the upstream and downstream devices can be temporarily closed, or physically disconnected (such as unplugging the network cable, etc.). An equal-cost path refers to the existence of multiple paths with equal costs to the same destination address. When the device supports equal-cost routing, the forwarding traffic sent to the destination IP or destination network segment can be shared through different paths to achieve load balancing of the network link.

[0106] Through the embodiments of the present application, when an abnormality occurs in the detection path, by switching the traffic to other equivalent detection paths and isolating the faulty network node, the faulty node can be prevented from continuing to affect the normal operation of the network, network connectivity can be quickly restored, and the continuity of network services can be ensured.

[0107] In some embodiments, the specific problem of the faulty node is repaired. If it is a software configuration problem, log in to the management interface of the switch, recheck and correct the relevant configuration to ensure that the configuration meets the requirements for normal network operation. If it is a hardware problem, replace the faulty interface board or the entire switch device. After the repair is completed, it is necessary to re-perform the detection test to verify whether the repair operation has successfully eliminated the abnormal detection path, send the detection packet again, and check whether a normal echo response can be received and whether there are no abnormal situations such as packet loss and timeout. If the detection result shows that the network has returned to normal, it means that the fault has been successfully eliminated, and the repaired node can be reconnected to the network, indicating that the faulty network node is correctly located.

[0108] The network architecture-based fault location method provided by the embodiments of the present application has the following beneficial effects:

[0109] By querying each primary switch in the core layer for its connected secondary switches in the access layer and the servers connected to each secondary switch, multiple detection paths are generated, covering different combinations of servers and switches. This allows connectivity monitoring of different links and nodes in the network architecture, enabling timely identification of potential issues. Using the IPv4 source path option, the IP addresses of the source server, switch, and destination server are sequentially written into the packet header to ensure the accuracy of the detection path, ensuring packet forwarding along the designated path and improving detection controllability. Based on the detection results of the detection path, the primary and secondary switches in the abnormal path are identified as potential nodes for investigation. By analyzing the detection results of the potential nodes, the faulty network node is precisely identified, effectively narrowing the scope of troubleshooting and reducing the time required to locate the faulty network node, achieving accurate and rapid fault location. After identifying the faulty node, measures are taken to isolate the faulty network node or switch the detection flow to another equally costly detection path to prevent the faulty node from further impacting network operations. Equal-cost routing is used to achieve traffic balancing, quickly restoring network connectivity, and ensuring network service continuity.

[0110] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0111] With the development of cloud computing and large-scale data centers, network topologies using switching architectures have become increasingly mainstream. Switching architectures divide complex networks into several layers, sharing traffic across different paths to achieve load balancing on network links. Related technologies typically use Internet Packet Explorer messages constructed using the Internet Control Message Protocol to detect network connectivity. These messages are transmitted between two servers based on a default forwarding table. In data center networks using switching architectures, multiple equal-cost paths exist between core or access switches. The uplink path from an access switch to a core switch is the hash of multiple links, and each probe to the same target may pass through different core switch nodes or interfaces. If a core or access switch experiences forwarding anomalies on a card, the unpredictable detection path often results in a lack of clear alarms. This makes it difficult to accurately identify the abnormal detection path and faulty node immediately, resulting in prolonged failures and impacting the stable operation of the data center.

[0112] In an embodiment of the present application, multiple probe flows covering all service flow paths are constructed between the core layer and the access layer to perform connectivity detection based on the information of the uplink and downlink ports connecting the core layer and the access layer. The routing path of a specified Internet Control Message Protocol probe packet is displayed through the source path option of Internet Protocol version 4, so that the probe flow strictly forwards data packets according to the specified path node sequence. The probe flow supports Internet Control Message Protocol detection in the network architecture. When packet loss occurs in the probe flow, the switches in the core layer and access layer in the probe path where the packet loss occurred are traced back, and the packet loss situation of multiple equal-cost paths is combined to accurately locate the faulty node, significantly shortening the fault location time.

[0113] The following is a description with reference to the accompanying drawings. Figure 4A , Figure 4A This is a third flow chart of the network architecture-based fault location method provided in an embodiment of the present application, which will be described in detail in conjunction with the steps shown in FIG4 .

[0114] In step 401, it is determined whether a detection path is to be constructed.

[0115] In some embodiments, it is determined whether a detection path needs to be constructed in the current network environment, so as to detect connectivity of the current network environment based on the detection path.

[0116] When the judgment result of step 401 is no, the current process ends.

[0117] When the determination result of step 401 is yes, step 402 is executed. In step 402, a plurality of detection paths are constructed based on the Internet Protocol addresses of the switches in the network architecture.

[0118] In some embodiments, an Internet Protocol address of a switch is obtained in a network architecture. The network architecture is a switching network architecture designed in a hierarchical model, including a core layer and an access layer. The core layer includes multiple first switches, and the access layer includes multiple second switches. Each first switch in the core layer is connected to at least one second switch in the access layer, and each second switch is connected to at least one server. In some embodiments, see Figure 5 , Figure 5 This is a schematic diagram of the network architecture provided in an embodiment of the present application. The network architecture 501 includes a core layer 502 and an access layer 503. The core layer 502 includes multiple first switches, for example, a first switch A and a first switch B. The access layer 503 includes multiple second switches, for example, a second switch A to a second switch n. The second switch is connected to at least one server, for example, server A to server n.

[0119] For the first switch in the core layer, a command line is executed to query all boards in the first switch that are in the "up" state. "Up" indicates that the board is enabled or running. The first switch to which a board in the "up" state belongs is able to operate normally and participate in data transmission. The board integrates various electronic components and interfaces, etc., to implement the device's data processing and transmission functions. For each first switch to which a board in the "up" state belongs, the interface of the first switch is queried and the corresponding Internet Protocol address (IP address) of the interface is recorded. The IP address uniquely identifies the interface in the network architecture. The second switch in the access layer connected to the interface is queried and all boards in the second switch that are in the "up" state are queried, and the IP address of the corresponding interface of the second switch is recorded. The second switch connected to the first switch is queried by executing a command line to query all interfaces connected to servers that are in the "up" state. The IP address of the interface and the IP information of the server to which the interface is connected are recorded. Each second switch is connected to at least one server. Through this query and traversal process, a complete network structure is obtained, reflecting the connection relationship between the first switch, the second switch, and the server.

[0120] For example, for the first switch A in the core layer 502, query the second switch A connected to the first switch A in the access layer 503, and for the second switch A, query the server A connected to the second switch A, and obtain one of the connection relationships: first switch A→second switch A→server A.

[0121] Select any two servers as the source server and the target server (corresponding to the first server and the second server mentioned above), obtain the Internet Protocol address of the first switch passing through the source server and the target server, the Internet Protocol address of the second switch, and the Internet Protocol addresses of the source server and the target server (corresponding to the starting point and the end point of the above detection path), and combine them according to the logic and sequence of network data transmission. For each board of each first switch and each second switch connected to the uplink port, select at least one server and the target server to build a detection flow. The detection flow is an end-to-end Internet Control Message Protocol Internet Packet Explorer path, and the network nodes passed by the Internet Packet Explorer path corresponding to the detection flow are used as the detection path. According to the order of network nodes specified in the source path option of the fourth version of the Internet Protocol (IPv4), the IPv4 source path option is a special optional field in the IPv4 protocol that specifies the IP address list of intermediate nodes in the detection path. It allows the source host to specify the routing path of the data packet from the source to the destination, rather than having the intermediate routers decide it independently according to the routing table. It is used to embed a set of preset hop IP addresses in the data message to control the hop-by-hop forwarding of the message along the specified path in the network.

[0122] The format of the IPv4 Source Path Option includes: Type (1 byte): 131 (LSRR) or 137 (SSRR); Length (1 byte): The total length of the option (including type, length, pointer, and IP list); Pointer (1 byte): Indicates the location of the IP address to be processed; IP list (4 bytes × N): The IP addresses of intermediate nodes. An LSRR is a loose source routing record, meaning packets are forwarded according to the series of router addresses specified in the LSRR. An SSRR is a strict source routing record, meaning packets must be forwarded strictly according to the router addresses specified in the SSRR, passing through the specified routers at each step. Multiple detection paths are determined based on the node forwarding order specified in the Source Path Option to cover all service requirements.

[0123] In some embodiments, see Figure 6 , Figure 6 This is a schematic diagram of the detection path construction provided by an embodiment of the present application; the first switch 6011 includes a board A, and board A has two interfaces in "up" state: H1 / 0 / 1 (IP address 13.1.1.1) and H1 / 0 / 2 (IP address 24.1.1.1), wherein interface H1 / 0 / 1 is connected to the second switch 6021, and the IP address of the server interface T1 / 0 / 3 on the second switch 6021 is 35.1.1.3, and the corresponding IP address of the server 6031 is 35.1.1.5, and the interface H1 / 0 / 2 is connected to the second switch 6022, and the IP address of the interface T1 / 0 / 3 of the server of the second switch 6022 is 46.1.1.4, and the corresponding IP address of the server 6032 is 46.1.1.6. In the process of constructing the detection path, server 6031 is used as the source server and server 6032 is used as the target server. The Internet Packet Explorer operation is performed on server 6032 from server 6031. The Internet Packet Explorer is a connectivity detection mechanism based on the Internet Control Message Protocol. The source host sends an Internet Control Message Protocol message to the target host. If the target host is online and reachable, the Internet Control Message Protocol message is returned.

[0124] For example, set the IP address of server 6031 to 5.5.5.5 and the IP address of server 6032 to 6.6.6.6. In the IPv4 source path option, specify: 1st hop: 35.1.1.3 (the interface connecting the second switch 6021 and server 6031), 2nd hop: 13.1.1.1 (H1 / 0 / 1 port of card A of the first switch 6011), 3rd hop: 14.1.1.4 (the interface connecting the second switch 6022 and the first switch 6031). 011), hop 4: 46.1.1.6 (the interface connecting server 6032 and second switch 6022). At this point, a detection path with a fixed forwarding order is constructed: server 6031 → second switch 6021 → H1 / 0 / 1 interface on card A of first switch 6011 → H1 / 0 / 2 interface on card A of first switch 6011 → second switch 6022 → server 6032 (i.e., red detection path 6041). Repeating the above steps can also generate the following detection path: server 6031 → second switch 6021 → first switch 6012 → second switch 6022 → server 6032 (i.e., blue detection path 6042).

[0125] After executing step 402, the judgment of step 403 is executed. In step 403, a detection packet of the Internet Control Message Protocol is sent to determine whether there is any abnormality in the detection result.

[0126] In some embodiments, a probe packet of the Internet Control Message Protocol is sent through a command line tool. The Internet Control Message Protocol is an important subprotocol in the TCP / IP protocol family. This protocol is located in the network layer and is mainly used to transmit control messages between IP hosts and routers. By sending an Internet Packet Explorer command, a series of request messages, i.e., multiple data packets, are sent to the target server. The detection results returned by the probe packet are used to determine whether there is an abnormality in the detection path. The Internet Packet Explorer is a connectivity detection mechanism based on the Internet Control Message Protocol and is widely used to test whether two network nodes are reachable.

[0127] When the judgment result of step 403 is yes, step 404 is executed. In step 404, a faulty network node in the abnormal detection path is determined.

[0128] In some embodiments, see Figure 4B , Figure 4B This is a fourth flow chart of a network architecture-based fault location method provided in an embodiment of the present application. Figure 4A Step 404 in the above example can be performed by executing Figure 4B Steps 4041 to 4046 in are implemented as described below.

[0129] In step 4041, the detection path information with abnormalities is recorded.

[0130] In some embodiments, after sending a request message in the detection path, the target server will reply with a response message. The detection result shows whether each data packet forwarded according to the specified IP sequence has received an echo response, including a timeout period. If no response message is received, a timeout message is displayed. At this time, there is an abnormality in the detection path, and the abnormal path information is recorded. The detection path has packet loss. When an abnormality is found in the detection path (such as packet loss), the abnormal detection path information is recorded, including the IP address and interface information of each network node (such as a switch, router, etc.) along the detection path, as well as the specific location where the abnormality occurred during the detection process (such as between which node and the next node the packet loss occurred), the packet loss rate, the timeout period and other detailed information for subsequent troubleshooting and analysis.

[0131] After step 4041, step 4042 is executed. In step 4042, the network node to be processed in the detection path is traced back to determine whether there is any abnormality in all detection flows passing through the first switch.

[0132] In some embodiments, the system traces back to network nodes in the detection path that have not yet been processed and may have problems. The pending network nodes are nodes in the detection path that have not been thoroughly inspected or may have problems, and may be the first switch or the second switch. For all detection paths that forward data through the first switch, the system queries the connectivity detection results of each detection path, such as the forwarding record of each data packet, whether there is packet loss, whether the timeout period is normal, and other information, to determine whether there are any anomalies in all detection flows passing through the first switch.

[0133] When the judgment result of step 4042 is yes, step 4043 is executed. In step 4043, it is determined that the faulty network node with the abnormality is the first switch.

[0134] In some embodiments, if all detection flows passing through the first switch are abnormal, the faulty network node with the abnormality is determined to be the first switch, and the status information of each interface of the first switch is checked, including the physical status of the interface (such as whether it is in the "up" state), the link layer status (such as whether there is an error count), and the sending and receiving statistics of the data packet, so as to facilitate subsequent fault repair. When the echo result shows that there is packet loss in the detection path, the detection results of the equivalent routing path passing through the first switch are analyzed and compared to determine whether the first switch is a faulty network node. Equal-cost multi-path routing refers to the existence of multiple paths with equal costs to the same destination address. When the device supports equal-cost routing, the forwarded data packets sent to the target server can be shared by different detection paths to achieve load balancing of the network link.

[0135] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the detection result analysis provided by an embodiment of the present application; the core layer includes a first switch 7011 and a first switch 7012, the access layer includes a second switch 7021, a second switch 7022, a second switch 7023 and a second switch 7024, the access layer connects multiple servers, including server 7031, server 7032, server 7033 and server 7034, and constructs four detection paths: detection path 7041 (black line), detection path 7042 (blue line), detection path 7043 (red line) and detection path 7044 (green line). The message of each detection path is initiated by the source server, passes through the second switch, the first switch, and the second switch, and reaches the target server, and the target server returns an echo message, and the echo message includes echo success, delay and packet loss.

[0136] For example, the two equal-cost path routes of detection path 7041 and detection path 7042 are analyzed as shown in Table 1 below, which is described in detail below.

[0137] Table 1

[0138]

[0139] Among them, × means that the detection result is not connected and there is packet loss. The detection results indicate connectivity and no packet loss. Both detection paths 7041 and 7042 pass through second switch 7021 and 7022, and the source server and target server are the same. Therefore, it can be concluded that both second switches are normal. The detection results through first switch 7011 show packet loss, while the detection results through first switch 7012 are normal. Therefore, the faulty node can be determined to be first switch 7011. Furthermore, if all detection paths through first switch 7011 show packet loss, first switch 7011 can also be determined to be the faulty network node. Similarly, if the detection results for detection path 7041 show no packet loss, while the detection results for detection path 7042 show packet loss, the faulty network node can also be determined to be first switch 7012.

[0140] When the judgment result of step 4042 is no, step 4044 is executed. In step 4044, it is determined that the faulty network node with the abnormality is the second switch.

[0141] In some embodiments, when the judgment result of step 4042 is no, it indicates that the first switch passed through in the detection path is normal, and the second switch may have a fault. The status information of each interface of the second switch is checked, including the physical status of the interface (such as whether it is in the "up" state), the link layer status (such as whether there is an error count), and the sending and receiving statistics of the data packet. A comparative analysis is performed based on the detection results of the equivalent routing path passing through the second switch to determine the second switch with a fault for subsequent fault repair.

[0142] For example, see Figure 7 , detection path 7043 and detection path 7044 both pass through the first switch 7011 and the second switch 7023. The source server and the target server are the same, both server 7031 and server 7033. The analysis of the detection results is shown in Table 2 below, which is described in detail below.

[0143] Table 2

[0144]

[0145] Among them, × means that the detection result is not connected and there is packet loss. It indicates that the detection result is connected and there is no packet loss. The detection result of detection path 7043 shows packet loss, and the detection result of detection path 7044 does not show packet loss. Detection path 7043 passes through the second switch 7021, and detection path 7044 passes through the second switch 7023. Therefore, it is determined that the faulty network node is the second switch 7021.

[0146] After the faulty network node is determined, step 4045 is executed. In step 4045, an operation of eliminating the abnormal detection path is performed.

[0147] In some embodiments, after determining the faulty network node, the faulty switch corresponding interface can be isolated or switched to another redundant path to eliminate the abnormality of the detection path. The determined faulty network node is isolated from the network to prevent it from continuing to affect the normal operation of the network. This can be achieved by disconnecting it from adjacent nodes. For example, if the first switch fails, its interfaces with the upstream and downstream devices can be temporarily closed, or the connection can be physically disconnected (such as unplugging the network cable).

[0148] In step 4046, if the detection result of the detection path returns to normal, it is determined that the faulty network node is successfully located.

[0149] In some embodiments, repairs are performed on specific problems of the faulty node. If it is a software configuration problem, such as an incorrect access control list configuration that causes abnormal packet discarding, it is necessary to log in to the switch's management interface, recheck and correct the relevant configuration, and ensure that the configuration meets the requirements for normal network operation. If it is a hardware problem, such as physical damage or performance degradation of certain interfaces of the switch, it is necessary to replace the faulty interface board or the entire switch device. After the repair is completed, it is necessary to re-perform the detection test to verify whether the repair operation has successfully eliminated the abnormal detection path, send the detection packet again, and check whether a normal echo response can be received, and whether there are no abnormal situations such as packet loss and timeout. If the detection results show that the network has returned to normal, it means that the fault has been successfully eliminated, and the repaired node can be reconnected to the network, indicating that the faulty network node has been correctly located.

[0150] Continue to see Figure 4A When the judgment result of step 403 is no, step 405 is executed. In step 405, the judgment of whether the detection result is abnormal is continued.

[0151] In some embodiments, connectivity tests are continued on other detection paths to determine whether there are other abnormal detection results.

[0152] The network architecture-based fault location method provided in the embodiments of the present application has the following beneficial effects:

[0153] By constructing multiple detection paths, covering all service flow paths, comprehensive network connectivity detection is achieved, ensuring the detection of potential anomalies within the network. Leveraging the source path option of Internet Protocol version 4 (IPv4), the routing path of specified Internet Control Message Protocol (ICMP) probe packets is precisely displayed, ensuring that the probe flow of the detection path strictly forwards packets according to the specified path node sequence, enhancing detection accuracy and controllability. In the event of packet loss on the network, the detection path can be traced back to the switches in the core and access layers where packet loss occurred. By combining the packet loss information of multiple equal-cost paths, the faulty node can be precisely located, effectively improving fault location efficiency and significantly reducing fault location time. Furthermore, by isolating the corresponding interfaces of the faulty switches or switching to other redundant paths, abnormal detection paths are promptly eliminated, ensuring normal network operation. Specific issues at the faulty nodes are repaired, and the repair results are verified through re-detection testing, further improving network reliability and stability.

[0154] The following continues to describe the exemplary structure of the network architecture-based fault location device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, such as Figure 2As shown, the software modules in the network architecture-based fault location device 455 stored in the memory 450 may include: a path generation module 4551, used to query the second switch connected to the first switch in the access layer, and query the server connected to the second switch; generate multiple detection paths, wherein the starting point and end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch; a path detection module 4552, used to perform packet loss detection on each detection path to obtain a detection result for each detection path; a fault location module 4553, used to determine a network node with a fault in the network architecture based on the detection result of each detection path, wherein the network node includes at least one of the first switch, the second switch, and the server.

[0155] In some embodiments, the path generation module 4551 is also used to randomly select two servers from the network architecture and query the first switch and the second switch connecting the two servers; and generate a detection path based on the queried Internet Protocol addresses of the first switch and the second switch, and the Internet Protocol addresses of the two servers.

[0156] In some embodiments, the path generation module 4551 is further configured to query the first second switch connected to the first server; query each first switch connected to the first second switch; and query the second second switch connected to the first switch but not connected to the first server.

[0157] In some embodiments, the two different servers include a first server and a second server, and the path detection module 4552 is further used to write the Internet Protocol address of the first server, the Internet Protocol address of the first switch, the Internet Protocol address of the second switch, and the Internet Protocol address of the second server into the header of the data packet based on the order of the network nodes passed by the detection path; forward the data packet based on the header of the data packet to obtain the detection result of the detection path.

[0158] In some embodiments, the fault location module 4553 is also used to, when the detection result of the detection path is abnormal, treat the first switch, the second switch and the server included in the detection path as the network nodes to be processed; perform fault location processing on the network nodes to be processed, and obtain the network nodes with faults in the network architecture.

[0159] In some embodiments, the fault location module 4553 is further used to determine that the first switch is a faulty network node when the network node to be processed is a first switch and the detection results of all detection paths passing through the first switch are abnormal; when the network node to be processed is a second switch and only the detection results of the detection paths passing through the second switch are abnormal, then the second switch is determined to be a faulty network node; when the network node to be processed is a server and only the detection results of all detection paths passing through the server are abnormal, then the server is determined to be a faulty network node.

[0160] In some embodiments, after determining the faulty network node in the network architecture based on the detection results of each detection path, the fault location module 4553 is also used to perform at least one of the following: isolating the faulty network node; switching the detection flow to other detection paths.

[0161] An embodiment of the present application provides a computer program product, comprising a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the electronic device to perform the network architecture-based fault location method described in the embodiment of the present application.

[0162] The embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions or computer programs. When the computer-executable instructions or computer programs are executed by a processor, the processor will execute the network architecture-based fault location method provided by the embodiment of the present application, for example, Figure 3A The fault location method based on the network architecture is shown.

[0163] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.

[0164] In some embodiments, computer-executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0165] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0166] By way of example, computer-executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected by a communication network.

[0167] In summary, through the embodiments of the present application, multiple detection paths are generated to cover different combinations of servers and switches, potential problems are discovered in a timely manner, and the IP addresses of the source server, switch, and target server are written into the data packet header in sequence based on the IPv4 source path option to ensure the controllability of the detection path. By analyzing the detection results of the nodes to be processed, the network nodes with faults are accurately determined, the scope of fault investigation is effectively narrowed, the time for locating the faulty network nodes is reduced, and the faulty nodes are accurately and quickly located. After determining the faulty node, measures are taken to isolate the faulty network node or switch the detection flow to other equivalent detection paths to prevent the faulty node from continuing to affect the network operation, quickly restore network connectivity, and ensure network service continuity.

[0168] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A fault location method based on network architecture, characterized in that: The network architecture includes a core layer and an access layer, the core layer includes a plurality of first switches; and the method includes: querying a second switch in the access layer connected to the first switch, and querying a server connected to the second switch; generating a plurality of detection paths, wherein a starting point and an end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch; Performing packet loss detection on each of the detection paths to obtain a detection result for each of the detection paths; Based on the detection result of each detection path, a faulty network node in the network architecture is determined, wherein the network node includes at least one of the first switch, the second switch, and the server.

2. The method according to claim 1, characterized in that The generating of multiple detection paths includes: Randomly selecting two servers from the network architecture, and querying a first switch and a second switch connected to the two servers; The detection path is generated based on the queried Internet Protocol addresses of the first switch and the second switch, and the Internet Protocol addresses of the two servers.

3. The method according to claim 2, characterized in that The two servers include a first server and a second server, and the querying of a first switch and a second switch connecting the two servers includes: querying a first second switch connected to the first server; Each first switch connected to the first second switch is queried, and a second second switch connected to the first switch and not connected to the first server is queried.

4. The method according to claim 1, wherein The two different servers include a first server and a second server; The performing packet loss detection on each of the detection paths to obtain a detection result of the detection path includes: writing the Internet Protocol address of the first server, the Internet Protocol address of the first switch, the Internet Protocol address of the second switch, and the Internet Protocol address of the second server into a header of a data packet based on the order of the network nodes represented by the detection path; The data packet is forwarded based on the header of the data packet to obtain a detection result of the detection path.

5. The method according to claim 1, wherein The determining of a faulty network node in the network architecture based on a detection result of each detection path includes: If the detection result of the detection path is abnormal, the first switch, the second switch and the server included in the detection path are used as network nodes to be processed; Perform fault location processing on the network node to be processed to obtain the network node with fault in the network architecture.

6. The method according to claim 5, characterized in that The performing fault location processing on the network node to be processed to obtain a network node with a fault in the network architecture includes: If the network node to be processed is a first switch and detection results of all detection paths passing through the first switch are abnormal, determining that the first switch is a faulty network node; If the network node to be processed is a second switch and only the detection result of the detection path passing through the second switch is abnormal, then the second switch is determined to be a faulty network node; If the network node to be processed is a server, and only when the detection results of all detection paths passing through the server are abnormal, the server is determined to be a faulty network node.

7. A fault location device based on network architecture, characterized in that: The device comprises: a path generation module, configured to query a second switch connected to the first switch in the access layer, and query a server connected to the second switch; and generate multiple detection paths, wherein the starting point and the end point of the detection path are two different servers, and the detection path passes through the first switch and the second switch; A path detection module, configured to perform packet loss detection on each detection path and obtain a detection result for each detection path; A fault location module is used to determine a faulty network node in the network architecture based on the detection results of each detection path, wherein the network node includes at least one of the first switch, the second switch, and the server.

8. An electronic device, characterized in that: The electronic device comprises: a memory for storing computer-executable instructions or computer programs; The processor is configured to implement the network architecture-based fault location method according to any one of claims 1 to 6 when executing the computer-executable instructions or computer program stored in the memory.

9. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the network architecture-based fault location method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the network architecture-based fault location method according to any one of claims 1 to 6 is implemented.