Gray fault detection and positioning method and system based on hybrid in-band network telemetry

By combining active and passive in-band network telemetry methods, a highly efficient gray fault detection and location system was designed, which solves the problems of large bandwidth consumption and untimely detection in the existing technology, and realizes fast and accurate fault location and traffic rerouting.

CN116436770BActive Publication Date: 2026-05-19SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANDONG COMP SCI CENTNAT SUPERCOMP CENT IN JINAN
Filing Date
2023-04-18
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing active in-band network telemetry methods consume a large amount of bandwidth in gray fault detection, resulting in untimely detection and high resource consumption. On the other hand, passive in-band network telemetry is affected by the tidal distribution of service traffic and cannot fully perceive the network status, resulting in low efficiency in gray fault detection and location.

Method used

Active in-band network telemetry and passive in-band network telemetry are effectively integrated. The server collects hop-by-hop telemetry information carried by passive INT probe packets for a first detection. If a fault is found, active INT probe packets are sent for a second detection. Combined with a distributed server and a virtual SDN network controller, path priorities are set for fault location, and traffic is adjusted in a timely manner through a rerouting mechanism.

Benefits of technology

It achieves fast and accurate gray fault detection and location, reduces bandwidth consumption, improves detection efficiency and resource utilization, and can complete fault location and traffic rerouting within seconds, reducing the waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116436770B_ABST
    Figure CN116436770B_ABST
Patent Text Reader

Abstract

The application provides a grey fault detection and positioning method and system based on hybrid in-band network telemetry, and relates to the field of fault detection.The method comprises the following steps: a server collects hop-by-hop telemetry information of a passive INT detection packet, performs a first detection on whether a fault exists, and sends a second detection instruction of a fault path to a controller of a virtual SDN network; the controller sends an active INT detection packet to the server, and performs a second detection on the path with the fault in the first detection; a source server re-routes data flow of the path information with the real fault; the controller sets priorities for all the path information with the real fault, compares the paths according to the priorities, and obtains a fault position; the controller feeds back the fault position to the server, and the server finds all the paths related to the fault position and ages in advance.The application integrates the active in-band network telemetry and the passive in-band network telemetry, makes up for the deficiency of a single telemetry method, and improves the efficiency and reliability of the network telemetry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of network fault detection technology, and particularly relates to a gray fault detection and location method and system based on hybrid in-band network telemetry. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Data centers (DCs) are crucial carriers of next-generation information and communication technologies such as 5G, artificial intelligence, and cloud computing. They are applied in many aspects of production and daily life and have significant research value. Through continuous integration and transformation, traditional data centers are gradually evolving into hyperscale data centers. A data center network (DCN) is a specially designed network used to interconnect a large number of computing and storage nodes within a data center. Data center networks support various services carried by the data center, such as web services, distribution, high-performance computing, data analytics, and data storage, requiring scalability, efficiency, and reliability. However, network failures caused by hardware, software, and human error are unavoidable, necessitating continuous monitoring and rapid fault detection, location, and recovery.

[0004] Network failures typically refer to a state where a network cannot provide normal service or has reduced service quality due to hardware problems, software vulnerabilities, virus intrusion, etc. Generally, network failures can be divided into two categories. The first category of network failures is explicit failures, which are caused by the equipment that builds the network, mainly including network cards, network cables, routers, switches, modems, etc. Explicit failures are usually accompanied by obvious symptoms, such as hardware damage or abnormal link disconnection. Using some simple methods, such as the PING command and the Tracert command, professionals can easily detect and handle such failures before the damage escalates. Explicit failures are highly destructive, but they are short-lived, easy to handle, and the damage they can cause is very limited.

[0005] However, another type of fault, known as gray faults, is more complex and more harmful. Gray faults are defined as a form of differential observability. More precisely, a system is defined as experiencing a gray fault when at least one application observes that the system is unhealthy, but an observer observes the system as healthy. Gray faults are generally difficult to detect and may persist for a long time. Furthermore, manual detection and location of fault points are difficult and time-consuming, and can cause significant damage to the data center network during fault handling. Therefore, establishing a rapid and reliable gray fault detection and location mechanism is crucial to minimizing the adverse effects of gray faults.

[0006] Network measurement is a crucial technology for achieving network awareness and control. Comprehensive, systematic, and efficient network measurement profoundly impacts future network operational efficiency. Traditional network measurement can be categorized into active measurement, passive measurement, and hybrid measurement based on the measurement method. Active measurement proactively sends probe packets to the network under test according to specific measurement needs. Due to the influence of internal network factors, the probe packets undergo a series of characteristic changes. By analyzing these changes, network status information and performance parameters can be obtained. Passive measurement acquires, records, and analyzes data packets at key devices and nodes in the network to obtain network status and performance parameters. Compared to active measurement, passive measurement has less impact on the network and yields more accurate results because it does not inject additional probe packets. However, since the measurement is only deployed at key devices and nodes, passive measurement can only obtain local network status information and cannot perceive the entire network. In addition, the practical application effect is limited by network device performance and network bandwidth, which may cause a certain degree of measurement accuracy loss.

[0007] Hybrid measurement scientifically integrates active and passive measurement, leveraging the advantages of both for more efficient and accurate network measurement. Traditional network measurement methods, due to their simple deployment, have been widely used in network management. However, with the increasing scale of networks and surging traffic, traditional network measurement technologies exhibit various problems, such as low accuracy of measurement algorithms, poor universality of measurement languages, and low intelligence in measurement task configuration, making them unsuitable for future network requirements. The emergence and development of Software-Defined Networking (SDN) has made fine-grained network measurement and refined network management possible. As an emerging network architecture, SDN decouples control and forwarding functions, enabling efficient and unified management of network behavior through a controller. It makes the underlying network logic transparent, simplifies the complexity of network measurement logic, and allows switches to collect network measurement data, thus achieving efficient and reliable measurement. However, deploying additional measurement mechanisms may consume limited network resources, and centralized control planes also have performance bottlenecks.

[0008] Compared to traditional measurement solutions and software-defined network measurement solutions, network telemetry is considered an ideal and effective measurement alternative, offering better accuracy, scalability, and performance. In-band network telemetry (INT), as a typical application of network telemetry, has received widespread attention from academia and industry. INT is an emerging network telemetry framework driven by a programmable data plane (PDP). INT combines packet forwarding with network measurement. Data packets contain telemetry instructions, which are processed and executed by programmable network elements. Therefore, network elements not only forward data packets but also participate in network measurement tasks. When a data packet carrying telemetry instructions passes through a device, the telemetry instructions instruct the INT device what network information to collect and insert into the data packet. Therefore, INT is an effective way to obtain network status information, providing accurate real-time data for network operations, management, and maintenance (OAM).

[0009] The inventors discovered that in-band telemetry (INT) can currently be divided into two main categories: active and passive. Active in-band network telemetry carries hop-by-hop telemetry data by constructing INT probe packets. Therefore, its focus is on designing efficient path planning algorithms. Passive in-band network telemetry relies on traffic flows to carry hop-by-hop telemetry information. Therefore, its focus is usually on designing efficient task orchestration algorithms. Active in-band network telemetry is characterized by flexible probe path construction but high bandwidth overhead. Passive in-band network telemetry has low bandwidth overhead but is affected by the tidal distribution of traffic flow.

[0010] In-band telemetry (INT) is well-suited for applications such as fault detection due to its flexible programmability, real-time monitoring, high signal-to-noise ratio, and flow-by-flow network awareness. However, only a very small number of studies have explored the application of in-band network telemetry in gray fault detection and localization. Since most of these studies employ active in-band network telemetry methods, they result in significant bandwidth consumption and suffer from problems such as system complexity, high resource consumption, and insufficient detection timeliness. Summary of the Invention

[0011] To overcome the shortcomings of the prior art, this invention provides a gray fault detection and localization method and system based on hybrid in-band network telemetry. It effectively integrates active in-band network telemetry and passive in-band network telemetry and applies them to the detection and localization of gray faults. It designs an efficient and complete gray fault detection and localization method based on hybrid in-band network telemetry, which makes up for the shortcomings of single telemetry methods, further improves the efficiency and reliability of network telemetry, and can quickly detect equipment and link faults and respond accordingly.

[0012] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0013] The first aspect of this invention provides a gray fault detection and localization method based on hybrid in-band network telemetry.

[0014] A gray fault detection and localization method based on hybrid in-band network telemetry includes the following steps:

[0015] Step 1: The server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a first detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a second detection command for the faulty path to the controller of the virtual SDN network.

[0016] Step 2: The controller receives the secondary detection command and sends an active INT probe packet to the server. The server forwards the active INT probe packet and performs a secondary detection on the path that was faulty in the first detection to confirm whether there is a real fault. The server then sends the path information that is actually faulty to the controller.

[0017] Step 3: The controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information.

[0018] Step 4: All servers in the distributed server perform steps 1 to 3 above, uploading information on all truly faulty paths in the network to the controller;

[0019] Step 5: The controller sets priorities for all paths with actual faults, compares the paths based on the priorities, and obtains the fault location.

[0020] Step 6: The controller feeds back the fault location to the server, which then searches for all paths related to the fault location and ages them out in advance.

[0021] Preferably, the server collects hop-by-hop telemetry information carried in passive INT probe packets, obtains all feasible path information between the source and the target, and performs a fault detection on the path information, specifically:

[0022] Set up a local path information table on the server, and record the aging time and secondary detection time of each path entry in the path information table;

[0023] After receiving telemetry information, the server will add the path information extracted from the telemetry information to the local path information table, or update the aging time of path table entries with the same path.

[0024] When the aging time of a path entry is 0, the path entry is deleted from the path information table.

[0025] A fault is identified when the secondary detection time of a path entry in the path information table is 0.

[0026] Preferably, the telemetry information includes the identifiers of the switches through which the passive INT probe packets and active INT probe packets pass, the ingress port ID of the switch, and the egress port ID of the switch.

[0027] Preferably, the aging time and secondary testing time should adhere to the following constraints:

[0028] agetime ≥ stdtime + prtime.

[0029] Among them, prtime refers to the time required for the INT packet to be transmitted from the sender to the receiver during the secondary detection process; agetime is the aging time; and stdtime is the secondary detection time.

[0030] Preferably, the server forwards active INT probe packets to perform a secondary check on paths that showed a fault in the first check, confirming whether a fault actually exists. Specifically:

[0031] If source A sends an active INT probe packet to destination B on path P to perform a secondary detection on path P, then:

[0032] If destination B receives an active INT probe packet from source A before the aging time is 0, destination B will update the aging time of path entry P, indicating that path P is not faulty.

[0033] If target B does not receive an active INT probe packet from source A before the aging time of P reaches 0, it indicates that path P is indeed faulty.

[0034] Preferably, the controller prioritizes all paths with actual faults, compares the paths based on these priorities, and determines the fault location. Specifically:

[0035] Step 1: The controller retrieves the location of the source and destination in the data center network from each path table entry, and sets priority attributes for the source and destination respectively, defined as Source(Pod,Tor,Server,Priority) and Destination(Pod,Tor,Server,Priority);

[0036] Step 2: When the controller receives the first fault path information No.1, it compares the source / destination location of No.1 with the source / destination location of No.i, using No.1 as the reference.

[0037] Step 3: Set the appropriate priority for No.i according to the priority setting rules;

[0038] Step 4: Compare the path entries according to their priority, and compare the ones with higher priority first to find the fault location.

[0039] The preferred priority setting rule is as follows:

[0040] (1) If the Source of No.i has a different pod compared to the Source of No.1, then the priority of the Source of No.i is 1;

[0041] Otherwise, if Source No.i has the same pod as Source No.1, the priority of Source No.i is set based on the following:

[0042] 1) If Source No.i has the same tor and the same server as Source No.1, then Source No.i has a priority of 4.

[0043] 2) If the Source of No.i has the same tor as the Source of No.1, but a different server, then the Source of No.i has a source priority of 3.

[0044] 3) If the Source of No.i has a different tor than the Source of No.1, then the priority of the Source of No.i is 2;

[0045] (2) If the destination of No.i has a different pod compared to the destination of No.1, then the priority of the destination of No.i is 1;

[0046] Otherwise, if Destination No.i has the same pod as Destination No.1, the priority of Destination No.i is set based on the following:

[0047] 1) If Destination No.i has the same tor and the same server as Destination No.1, then Destination No.i has a priority of 4.

[0048] 2) If the destination of No.i has the same tor as the destination of No.1, but a different server, then the source priority of the destination of No.i is 3.

[0049] 3) If the destination of No.i has a different tor than the destination of No.1, then the priority of the destination of No.i is 2;

[0050] Finally, the priority of No.i = the priority of No.i's Source + the priority of No.i's Destination.

[0051] The second aspect of the present invention provides a gray fault detection and location system based on hybrid in-band network telemetry.

[0052] A gray fault detection and location system based on hybrid in-band network telemetry includes:

[0053] The primary detection module is configured as follows: the server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a primary detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a secondary detection command for the faulty path to the controller of the virtual SDN network.

[0054] The secondary detection module is configured as follows: the controller receives the secondary detection instruction, sends an active INT probe packet to the server, the server forwards the active INT probe packet, performs a secondary detection on the path that has a fault in the primary detection, confirms whether there is a real fault, and sends the path information that has a real fault to the controller.

[0055] The rerouting module is configured such that the controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information.

[0056] The acquisition module is configured such that all servers in the distributed server execute the above-mentioned detection module to the rerouting module, and upload the information of all paths in the network that actually have faults to the controller;

[0057] The fault location module is configured such that the controller sets priorities for all paths with actual faults, compares the paths based on the priorities, and obtains the fault location.

[0058] The feedback module is configured such that the controller feeds back the fault location to the server, and the server searches for all paths related to the fault location and ages them out in advance.

[0059] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of the gray fault detection and location method based on hybrid in-band network telemetry as described in the first aspect of the present invention.

[0060] The fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the gray fault detection and localization method based on hybrid in-band network telemetry as described in the first aspect of the present invention.

[0061] The above one or more technical solutions have the following beneficial effects:

[0062] To overcome the shortcomings of single telemetry methods and further improve the efficiency and reliability of network telemetry, this invention effectively integrates active in-band network telemetry and passive in-band network telemetry and applies them to the detection and localization of gray faults. This improves the problem of active in-band network telemetry consuming a large amount of bandwidth in gray fault detection and increases detection efficiency. At the same time, an efficient fault localization method is designed to avoid the waste of a large amount of computing resources and improve localization efficiency.

[0063] This invention provides an efficient and complete gray fault detection and location framework that can monitor device and link status in real time, quickly detect device and link faults and respond accordingly, and achieve gray fault detection and location within seconds.

[0064] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0065] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0066] Figure 1 This is a flowchart of the method in the first embodiment.

[0067] Figure 2 The first embodiment is a flowchart of the hybrid INT workflow.

[0068] Figure 3 This is a flowchart of the data traffic rerouting process based on source routing for the first embodiment.

[0069] Figure 4 This is a system structure diagram of the second embodiment. Detailed Implementation

[0070] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0071] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.

[0072] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0073] The overall concept proposed in this invention is as follows:

[0074] This invention designs a lightweight, hybrid INT-based method for rapid gray fault detection and localization, capable of accurately and quickly detecting and locating gray faults, and providing a complete gray fault detection and localization framework for fat-tree data center networks. Balancing resource and efficiency, it can detect network gray faults in real time and immediately reroute affected traffic, while completing fault localization within seconds.

[0075] The main contributions of this invention are summarized as follows:

[0076] A mechanism for rapid detection of gray network faults in DCN based on hybrid INT is proposed. Furthermore, this represents a valuable exploration of telemetry applications in hybrid in-band networks. Specifically, all feasible paths in the network are collected through passive INT, and network faults are detected based on telemetry information. Simultaneously, to improve detection accuracy, a secondary detection mechanism based on active INT is designed, which actively sends simplified probe packets to perform secondary detection on possible faulty paths.

[0077] A remote control was designed to implement a secondary detection mechanism and centralized fault location. The server should upload timed-out path entries to the controller, which will then decide whether to perform secondary detection or centralized fault location. A rapid location method was also designed, capable of quickly and accurately completing the location within a specified time, requiring only minimal computation.

[0078] A feedback mechanism for the remote centralized controller was introduced. Specifically, once the fault point is accurately located, the controller quickly sends the fault location back to the server. Then, the server marks all paths related to the fault location in the path information table and ages them out in advance.

[0079] Example 1

[0080] This embodiment discloses a gray fault detection and localization method based on hybrid in-band network telemetry.

[0081] A gray fault detection and localization method based on hybrid in-band network telemetry includes the following steps:

[0082] Step 1: The server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a first detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a second detection command for the faulty path to the controller of the virtual SDN network.

[0083] Step 2: The controller receives the secondary detection command and sends an active INT probe packet to the server. The server forwards the active INT probe packet and performs a secondary detection on the path that was faulty in the first detection to confirm whether there is a real fault. The server then sends the path information that is actually faulty to the controller.

[0084] Step 3: The controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information.

[0085] Step 4: All servers in the distributed server perform steps 1 to 3 above, uploading information on all truly faulty paths in the network to the controller;

[0086] Step 5: The controller sets priorities for all paths with actual faults, compares the paths based on the priorities, and obtains the fault location.

[0087] Step 6: The controller feeds back the fault location to the server, which then searches for all paths related to the fault location and ages them out in advance.

[0088] This invention presents a lightweight, fast gray fault detection and localization method based on hybrid INT, which infers gray faults from telemetry data. The workflow of the proposed method is as follows: Figure 1 As shown, it is divided into five stages: hybrid in-band network telemetry, network fault detection, traffic rerouting, network fault location, and network fault feedback.

[0089] First, the server continuously collects hop-by-hop telemetry information (blue arrows) carried by data packets. These packets are also called passive INT packets. Most flows in a data center network are frequently active, so all feasible paths in the network can be obtained in a short time. The collected telemetry information includes the identifiers of the switches they pass through, as well as the corresponding ingress and egress port IDs, which constitute the basic structure of the path. By extracting the telemetry information, the path information between the source and destination can be obtained.

[0090] Then, the server maintains a path information table, recording the agetime (agetime) and secondary detection time (sdtime) for each path entry. Upon receiving telemetry information, the server adds the extracted path information to its local path information table or updates the agetime of path entries with the same path. When a network failure occurs, data packets cannot pass through the affected path, causing the sdtime of the relevant entry P = {A, ..., B} in the path information table to equal 0. Following SDN design paradigms, this invention introduces an external controller for information notification and fault location. The controller then instructs source A to send an INT probe packet on path P for secondary detection (gray arrow). This INT probe packet is also called an active INT packet. If destination B receives the INT probe packet from source A before agetime equals 0, it updates the agetime of path entry P, indicating that the path is not faulty. Otherwise, if destination B has not received the INT probe packet before agetime equals 0 in P, the confirmation fails.

[0091] Once a failure is confirmed, subsequent traffic will be rerouted to other feasible paths (green arrows) to prevent packet loss. Simultaneously, the servers upload the faulty paths to the remote controller, which uses difference comparisons to pinpoint the fault location in all paths.

[0092] Finally, after locating the fault point, the controller quickly reports the fault point to the server. The server pre-sets the agetime and stdtime of all path entries related to the fault point in the local path information table to 0.

[0093] Specifically:

[0094] (I) Hybrid In-Band Network Telemetry

[0095] It is feasible to obtain all feasible paths in the network using passive INT and infer network failure using a timeout mechanism. However, this is not rigorous because, considering an extreme case, if no data packets pass through a path for an extended period, triggering the timeout mechanism, it might be mistakenly identified as a failed path. Observations show that although most traffic in data center networks is active, this situation cannot be completely avoided. Therefore, this invention designs an additional secondary detection mechanism based on active INT to address the problem of false fault identification. First, a secondary detection time (stdtime) is set for each path information entry. When stdtime = 0 for a path, an INT probe packet is sent to that path for secondary detection. Since the vast majority of links in the network are healthy and active, the triggering of the secondary detection mechanism is extremely rare. Therefore, the secondary detection mechanism only consumes a small amount of bandwidth. In summary, the hybrid INT formed by combining active and passive INT significantly reduces the error rate of fault detection using a single telemetry method and greatly reduces bandwidth consumption, saving network resources.

[0096] Furthermore, in the original INT model, devices with INT functionality were expected to expose sufficient internal device state, including switch ID, ingress / egress port ID, queue depth, and queuing latency. However, in this invention, since only whether a link is up or down is of concern, it is not necessary to obtain all the aforementioned internal states to make the system more lightweight. Therefore, this invention simplifies the format of INT packets, collecting only the following three types of internal device states. Note that when using the term INT packet, it does not refer to active INT packets or passive INT packets, but both. Switch_id (8 bits): The identifier of the switch. The controller assigns a unique ID to each switch. Ingress_port_id (8 bits): The ingress port ID of the INT packet entering the switch. Egress_port_id (8 bits): The egress port ID of the INT packet leaving the switch.

[0097] When service packets pass through the network, the switches along the path insert INT information after the IP header of the service packets. The information collected in the INT probe includes the IDs of the switches they pass through, as well as the corresponding ingress and egress port IDs. This INT workflow and probe format are combined as follows: Figure 2 As shown.

[0098] In this way, the server can continuously collect telemetry information to obtain all feasible paths between the source and destination. Secondly, each probe packet collector, which is also a server, stores these feasible paths in a path information table, with each path entry having an aging time. For each newly acquired path, if the path already exists in the path information table, its aging time is updated; otherwise, it is added to the table. This method of collecting path information does not require injecting a large number of probes into the network, thus having a smaller impact on the network.

[0099] Thus, this invention uses two types of packets: passive INT packets and active INT packets. In addition, there are SR packets used for traffic rerouting. Therefore, this invention contains three types of packets. Different IP protocol numbers were designed to distinguish them. The specific formats of these three packet types are shown in Table 1.

[0100] Table 1 Package Format Details

[0101]

[0102] When a probe packet travels through the network, the switches along the way process it accordingly based on the IP protocol number. For example, if a packet's IP protocol number is 0x700 or 0x702, it indicates that it is an INT packet. The switches along the way will then insert the INT information after the IP header of the INT packet. If a packet's IP protocol number is 0x701, it indicates that it is an SR packet. The switches along the way will then forward this packet strictly according to the SR forwarding rules.

[0103] (II) Gray Fault Detection

[0104] Upon receiving INT messages, the INT information carried in these messages is parsed and stored in the path information table on each server at the receiving end. Each path entry records all switches from the sender to the receiver, as well as the corresponding ingress and egress ports. For example, the path information table for H0 is shown in Table 2, which illustrates all servers that can reach H0 and all possible paths.

[0105] Table 2H0 Path Information Table

[0106]

[0107] In addition, in the path entry, there are two time values, namely stdtime and agetime. They are the core components of fault detection. Specifically, when the stdtime of a path entry is 0, the path should be re-detected. When the agetime of a path entry is 0, the path entry is considered invalid, and at the same time, the entry is deleted from the path information table. It is not difficult to understand from the above description that setting reasonable agetime and stdtime is very important. In fact, the values of agetime and stdtime are very flexible, but the following constraints should be followed:

[0108] agetime ≥ stdtime + prtime.

[0109] Where prtime refers to the time required for the INT packet to be transmitted from the sender to the receiver during the re-detection process.

[0110] The following introduces the fault detection process. As Figure 1 shown, a path path from H2 to H9

[0113] , = {H2, T1_3, T1_2, L1_4, L1_1, S2_1, S2_3, L5_1, L5_3, T4_2, T4_4, H9}. Due to the interruption of S2_3 and L5_1, some data packets carrying INT information are discarded and cannot reach the target server H9. Because sdtime < agetime, the sdtime of the entry P corresponding to path i in the path information table of server H9 will first tend to 0. At this time, the controller will notify server H2 to send an INT probe packet on path i for re-detection. If server H9 receives the INT probe packet sent by server H2 before agetime = 0, it updates the agetime of path entry P, indicating that path i has no fault. Otherwise, if server H9 still does not receive the INT probe packet before the agetime of path entry P is 0, it is confirmed that a fault has occurred.

[0111] (3) SR Rerouting Mechanism

[0112] When the destination server detects a fault, it will immediately upload the fault path information to the controller. Then, the controller notifies the corresponding source server of the fault path information. Finally, the source server will use source routing technology to re-route the affected data traffic in a timely manner according to the latest path information table. Figure 3 shows the data traffic rerouting process based on source routing.

[0113] In computer networks, source routing allows the sender of a data packet to specify the route the packet takes through the network, typically by marking the route in the packet header. Figure 3 In this process, UDP packets are used to carry the SR payload. Simultaneously, to inform the packet that it is an SR packet, the IP protocol number is set to "0x701". 512 bits are reserved between the IP and UDP headers for the SR label stack. Additionally, 8 bits are allocated to each SR label to represent the switch output port ID, meaning each switch can support a maximum of 256 output ports. Therefore, the SR stack includes all switch egress ports on the specified path. When an SR packet traverses the network, the switch parses the SR labels in the SR stack sequentially and forwards it from the designated port. In summary, path entries identified as faulty paths by fault detection are disabled, and subsequent traffic is rerouted to other feasible paths to prevent packet loss.

[0114] (iv) Location mechanism for gray faults

[0115] Simply detecting faults and rerouting potentially affected traffic is not a complete solution. A better approach is to accurately pinpoint the fault and implement targeted solutions.

[0116] In data center networks, even a single point of failure can affect multiple paths; this observation can be used for precise network fault localization. However, because distributed servers do not share a global network view, a single faulty path entry on a single server is insufficient to pinpoint the exact location of the network fault. Therefore, all faulty path entries from the distributed servers should be uploaded to the controller and stored in a faulty path information table. The controller then gradually narrows down the network fault to a single link between two devices by identifying commonalities among all faulty path entries in the table. For example, all paths affected by a link failure between L1 and S2 are shown in Table 2, with the entries ordered according to the order received by the controller.

[0117] Table 2 Fault Path Information in the Controller

[0118]

[0119]

[0120] First, in the first round of comparison, the controller compares No.1 and No.2, identifying their similarities and obtaining the following result: H2→T1_3→T1_2→L1_4→L1_1→S2_1→S2_3→L5_1→L5_3→T4_2. Then, the result of the first round is compared with No.3, revealing their shared path as S2_3→L5_1→L5_3→T4_2. After the second round of comparison, if the fault location is still not accurately pinpointed, iterative operation should continue. The result of the second round is compared with No.4, yielding the precise fault location as S2_3→L5_1. Theoretically, this method can accurately locate network faults, which is a commonly used method in fault localization.

[0121] However, data center networks are typically large-scale, and a single point of failure can generate hundreds or even thousands of fault paths. This simple comparison method is not only computationally intensive but also inflexible. For example, in the previous example, the results of the first round of comparison were not very effective because it still included too many devices and links, making precise location still very difficult. However, when comparing No.1 with No.8, the precise fault location can be found directly: S2_3→L5_1. Compared to the previous method that required three rounds of comparison to obtain a result, this comparison is clearly more meaningful and efficient. Therefore, for No.1, to achieve higher efficiency, a path entry similar to No.8 should be selected and allowed to participate in the comparison first. Therefore, this invention improves upon the above method. Specifically, the controller can obtain the source and destination positions in the fat-tree data center network from each path entry and set their Priority attributes, defined as Source(Pod, Tor, Server, Priority) and Destination(Pod, Tor, Server, Priority), respectively. When the controller receives the first fault path information No.1, it should immediately begin the location process to minimize the time required from the occurrence of the fault to successful location. Therefore, based on No.1, the source / destination positions of No.1 and No.i should be compared respectively, and then the corresponding priority should be set for No.i.

[0122] The priority setting rules are as follows:

[0123] (1) If the Source of No.i has a different pod compared to the Source of No.1, then the priority of the Source of No.i is 1;

[0124] Otherwise, if Source No.i has the same pod as Source No.1, the priority of Source No.i is set based on the following:

[0125] 1) If Source No.i has the same tor and the same server as Source No.1, then Source No.i has a priority of 4.

[0126] 2) If the Source of No.i has the same tor as the Source of No.1, but a different server, then the Source of No.i has a source priority of 3.

[0127] 3) If the Source of No.i has a different tor than the Source of No.1, then the priority of the Source of No.i is 2;

[0128] (2) If the destination of No.i has a different pod compared to the destination of No.1, then the priority of the destination of No.i is 1;

[0129] Otherwise, if Destination No.i has the same pod as Destination No.1, the priority of Destination No.i is set based on the following:

[0130] 1) If Destination No.i has the same tor and the same server as Destination No.1, then Destination No.i has a priority of 4.

[0131] 2) If the destination of No.i has the same tor as the destination of No.1, but a different server, then the source priority of the destination of No.i is 3.

[0132] 3) If the destination of No.i has a different tor than the destination of No.1, then the priority of the destination of No.i is 2;

[0133] Finally, the priority of No.i = the priority of No.i's Source + the priority of No.i's Destination.

[0134] Finally, the path entries are compared based on priority to obtain the accurate fault location, requiring only 1-2 rounds. This method requires minimal computation, making it highly efficient and flexible.

[0135] (v) Feedback mechanism for gray faults

[0136] After fault location, accurate fault location should be fully utilized to improve system efficiency and reliability. Specifically, after fault location, the controller quickly feeds back the accurate fault location to the server. The server searches the path information table for all paths related to the fault location and ages them prematurely, and promptly reroutes the affected data traffic using source routing based on the latest path information table. Setting up a feedback mechanism has the following three advantages: (a) Data packets can be rerouted to non-faulty paths before reaching the faulty path, thus avoiding data packet loss. (b) It skips the secondary confirmation process, reducing bandwidth costs and improving system efficiency. (c) It avoids repeated fault detection and location, reduces detection delays caused by waiting for path aging time, saves system computing resources, and improves system timeliness.

[0137] Example 2

[0138] This embodiment discloses a gray fault detection and location system based on hybrid in-band network telemetry.

[0139] like Figure 4 As shown, a gray fault detection and location system based on hybrid in-band network telemetry includes:

[0140] The primary detection module is configured as follows: the server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a primary detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a secondary detection command for the faulty path to the controller of the virtual SDN network.

[0141] The secondary detection module is configured as follows: the controller receives the secondary detection instruction, sends an active INT probe packet to the server, the server forwards the active INT probe packet, performs a secondary detection on the path that has a fault in the primary detection, confirms whether there is a real fault, and sends the path information that has a real fault to the controller.

[0142] The rerouting module is configured such that the controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information.

[0143] The acquisition module is configured such that all servers in the distributed server execute the above-mentioned detection module to the rerouting module, and upload the information of all paths in the network that actually have faults to the controller;

[0144] The fault location module is configured such that the controller sets priorities for all paths with actual faults, compares the paths based on the priorities, and obtains the fault location.

[0145] The feedback module is configured such that the controller feeds back the fault location to the server, and the server searches for all paths related to the fault location and ages them out in advance.

[0146] Example 3

[0147] The purpose of this embodiment is to provide a computer-readable storage medium.

[0148] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the gray fault detection and location method based on hybrid in-band network telemetry as described in Embodiment 1 of this disclosure.

[0149] Example 4

[0150] The purpose of this embodiment is to provide an electronic device.

[0151] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the gray fault detection and location method based on hybrid in-band network telemetry as described in Embodiment 1 of this disclosure.

[0152] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0153] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.

[0154] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A gray fault detection and localization method based on hybrid in-band network telemetry, characterized in that, Includes the following steps: Step 1: The server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a first detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a second detection command for the faulty path to the controller of the virtual SDN network. Step 2: The controller receives the secondary detection command and sends an active INT probe packet to the server. The server forwards the active INT probe packet and performs a secondary detection on the path that was faulty in the first detection to confirm whether there is a real fault. The server then sends the path information that is actually faulty to the controller. Step 3: The controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information. Step 4: All servers in the distributed server perform steps 1 to 3 above, uploading information on all truly faulty paths in the network to the controller; Step 5: The controller assigns priorities to all paths with actual faults, compares the paths based on these priorities, and determines the fault location. Specifically, setting priorities and comparing paths based on priorities include: Step 5.1: The controller obtains the location of the source and destination in the data center network from each path table entry, and sets priority attributes for the source and destination respectively, defined as Source(Pod, Tor, Server, Priority) and Destination(Pod, Tor, Server, Priority); Step 5.2: When the controller receives the first fault path information No.1, it compares the source / destination location of No.i with the source / destination location of No.1, using No.1 as the reference. Step 5.3: Set the appropriate priority for No.i according to the following priority setting rules: (1) If the Source of No.i has a different pod compared to the Source of No.1, then the Source of No.i has a priority of 1; Otherwise, if Source No.i has the same pod as Source No.1, the priority of Source No.i is set based on the following: 1) If Source No.i has the same tor and the same server as Source No.1, then Source No.i has a priority of 4. 2) If the Source of No.i has the same tor as the Source of No.1, but a different server, then the Source of No.i has a source priority of 3. 3) If the Source of No.i has a different tor than the Source of No.1, then the priority of the Source of No.i is 2; (2) If the destination of No.i has a different pod compared to the destination of No.1, then the priority of the destination of No.i is 1; Otherwise, if Destination No.i has the same pod as Destination No.1, the priority of Destination No.i is set based on the following: 1) If Destination No.i has the same tor and the same server as Destination No.1, then Destination No.i has a priority of 4. 2) If the destination of No.i has the same tor as the destination of No.1, but a different server, then the source priority of the destination of No.i is 3. 3) If the destination of No.i has a different tor than the destination of No.1, then the priority of the destination of No.i is 2; Finally, the priority of No.i = the priority of No.i's Source + the priority of No.i's Destination; Step 5.4: The controller compares the path entries according to the priority. The path entry with higher priority is compared with the baseline path No.1 first. The fault location is obtained by identifying the commonalities among all fault path entries. Step 6: The controller feeds back the fault location to the server, which then searches for all paths related to the fault location and ages them out in advance.

2. The gray fault detection and localization method based on hybrid in-band network telemetry as described in claim 1, characterized in that, The server collects hop-by-hop telemetry information carried in passive INT probe packets, obtains all feasible path information between the source and the target, and performs a fault check on the path information, specifically: Set up a local path information table on the server, and record the aging time and secondary detection time of each path entry in the path information table; After receiving telemetry information, the server will add the path information extracted from the telemetry information to the local path information table, or update the aging time of path table entries with the same path. When the aging time of a path entry is 0, the path entry is deleted from the path information table. A fault is identified when the secondary detection time of a path entry in the path information table is 0.

3. The gray fault detection and localization method based on hybrid in-band network telemetry as described in claim 1, characterized in that, The telemetry information includes the identifiers of the switches through which the passive INT probe packets and active INT probe packets pass, the ingress port ID of the switch, and the egress port ID of the switch.

4. The gray fault detection and location method based on hybrid in-band network telemetry as described in claim 2, characterized in that, The aging time and secondary testing time should adhere to the following constraints: agetime ≥ stdtime + prtime. Among them, prtime refers to the time required for the INT packet to be transmitted from the sender to the receiver during the secondary detection process; agetime is the aging time; and stdtime is the secondary detection time.

5. The gray fault detection and location method based on hybrid in-band network telemetry as described in claim 1, characterized in that, The server forwards proactive INT probe packets and performs a secondary check on paths that showed a fault in the first check to confirm whether a fault actually exists. Specifically: If source A sends an active INT probe packet to destination B on path P to perform a secondary detection on path P, then: If destination B receives an active INT probe packet from source A before the aging time is 0, destination B will update the aging time of path entry P, indicating that path P is not faulty. If target B does not receive an active INT probe packet from source A before the aging time of P reaches 0, it indicates that path P is indeed faulty.

6. A gray fault detection and location system based on hybrid in-band network telemetry, employing the gray fault detection and location method based on hybrid in-band network telemetry as described in any one of claims 1-5, characterized in that: include: The primary detection module is configured as follows: the server collects hop-by-hop telemetry information carried by the passive INT probe packets, obtains all feasible path information between the source and the target, performs a primary detection on whether there is a fault in the path information, and if the detection result is that there is a fault, it sends a secondary detection command for the faulty path to the controller of the virtual SDN network. The secondary detection module is configured as follows: the controller receives the secondary detection instruction, sends an active INT probe packet to the server, the server forwards the active INT probe packet, performs a secondary detection on the path that has a fault in the primary detection, confirms whether there is a real fault, and sends the path information that has a real fault to the controller. The rerouting module is configured such that the controller sends the faulty path information to the source server of the faulty path information, and the source server reroutes the data traffic of the faulty path information. The acquisition module is configured such that all servers in the distributed server execute the above-mentioned detection module to the rerouting module, and upload the information of all paths in the network that actually have faults to the controller; The fault location module is configured such that the controller sets priorities for all paths with actual faults, compares the paths based on the priorities, and obtains the fault location. The feedback module is configured such that the controller feeds back the fault location to the server, and the server searches for all paths related to the fault location and ages them out in advance.

7. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the gray fault detection and localization method based on hybrid in-band network telemetry as described in any one of claims 1-5.

8. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the gray fault detection and location method based on hybrid in-band network telemetry as described in any one of claims 1-5.