Network detection method and device, computer equipment and storage medium

By acquiring and utilizing network topology information to generate a set of probe links and task configurations, the problem of insufficient automated management in existing RoCE network probing solutions is solved, achieving full-scenario coverage and low-redundancy fault location, which is suitable for large-scale clusters and multi-tenant scenarios.

CN121864644APending Publication Date: 2026-04-14BEIJING KINGSOFT CLOUD NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing RoCE network probing solutions rely on manual configuration or fixed templates, which cannot automatically perceive the network topology, resulting in missing or misconfigured probe links. They are difficult to adapt to the automated management of large-scale networks and cannot adaptively adjust according to actual topology changes. They also suffer from insufficient probe coverage and excessive performance overhead, making it difficult to meet the needs of multi-tenant and ultra-large-scale clusters.

Method used

By acquiring the topology information of the target network, a set of probe links is generated, and probe task configuration is generated based on the link set. Probe links covering the network between switches and between the graphics processor and the switch are automatically generated, achieving automated, full-scenario coverage and low-redundancy rapid fault location.

Benefits of technology

It achieves automated management of RoCE network probing, ensuring complete coverage and low redundancy of probing links, improving the efficiency and accuracy of fault location, and is suitable for large-scale clusters and multi-tenant scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121864644A_ABST
    Figure CN121864644A_ABST
Patent Text Reader

Abstract

The invention relates to a network detection method and apparatus, a computer device and a storage medium. The method comprises the steps of obtaining network topology information in a target network; a detection link set used for active detection is generated based on the network topology information, each detection link represents the corresponding relation between a detection starting point and a detection terminal point in one-time active detection, the detection starting point is a switch, and the detection terminal point is a switch or a graphics processor; generating detection task configuration according to the detection link set; and sending the detection task configuration to a corresponding switch to execute active detection through the switch. Therefore, a detection link set covering the interchangers and between the graphics processor and the interchangers can be automatically generated, configuration is generated and issued for execution, and the problems that traditional manual or fixed template configuration is prone to mismatching, cannot adapt to large-scale clusters, and is high in detection redundancy and low in fault positioning efficiency are solved; automation, full-scene coverage, low redundancy and rapid fault positioning of RoCE network detection are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of network operation and maintenance technology, and in particular to a network detection method, device, computer equipment and storage medium. Background Technology

[0002] With the deployment of large-scale artificial intelligence computing and high-performance computing clusters, network architectures based on Remote Direct Memory Access over Converged Ethernet (RoCE) are widely used for high-speed communication between graphics processing unit (GPU) clusters. In large RoCE networks, to achieve proactive discovery and rapid location of network faults, it is usually necessary to configure Network Quality Analyzer (NQA) probes on switches. By detecting indicators such as network link connectivity and latency, the stability and reliability of communication between GPUs are ensured.

[0003] Existing RoCE network active probing solutions mainly rely on manual methods or configuration based on fixed templates. In the manual method, maintenance personnel typically select the source node, destination node, and corresponding probing path level based on experience, and then issue NQA probing tasks one by one on the switch. In the template method, several fixed probing path combinations are predefined, and probing tasks are generated in batches according to the template and issued to the switch for execution.

[0004] However, manual configuration lacks the ability to automatically perceive network topology and the connection relationship between graphics processors and switches, making it prone to missed or misconfigured probe links and unsuitable for automated management of large-scale networks. Template-based methods, with fixed probe paths, cannot adaptively adjust to changes in actual network topology and the dynamic changes in the number of graphics processors. Furthermore, existing technologies struggle with flexible probe link orchestration across different probe levels from Tier 0 to Tier 3 and lack a sampling-based probe link selection mechanism for large-scale RoCE networks, resulting in insufficient probe coverage, excessive performance overhead, and an inability to meet the needs for proactive probe and rapid localization in multi-tenant and ultra-large-scale cluster scenarios. Summary of the Invention

[0005] In view of this, in order to solve the above-mentioned technical problems or some of the technical problems, the embodiments of the present invention provide a network detection method, apparatus, computer equipment and storage medium.

[0006] In a first aspect, embodiments of the present invention provide a network detection method, comprising: Obtain network topology information in the target network, the network topology information including the connection relationship between each switch in the target network, and the access port information between the graphics processor in the target network and the connected switch; Based on the network topology information, a set of detection links for active detection is generated. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. Generate a detection task configuration based on the set of detection links; The detection task configuration is sent to the corresponding switch to perform active detection through the switch.

[0007] In one possible implementation, obtaining network topology information in the target network includes: Obtain the connection relationships between switches in the target network as the first topology information; Obtain information about the switches and access ports to which each graphics processor in the target network is connected; A second topology information is generated based on the correspondence between each graphics processor and the access switch, as well as the access port information; Based on the first topology information and the second topology information, the network topology information is constructed, and the network topology information is used to determine the correspondence between the detection start point and the detection end point in the detection link set.

[0008] In one possible implementation, generating a set of probe links for active probing based on the network topology information includes: Based on the network topology information, the connection paths between switches at different detection levels are determined, as well as the connection path between the switches and the graphics processor. Based on the connection path, each detection start point and corresponding detection end point are determined, so as to generate multiple detection links based on the detection start point and detection end point, thus obtaining the detection link set.

[0009] In one possible implementation, generating the probe task configuration based on the probe link set includes: For each detection link in the detection link set, obtain the detection start point, detection end point, and detection level of the detection link; Obtain the set of detection parameters corresponding to the detection level, wherein the set of detection parameters includes at least one of the following: transmission interval parameter, timeout parameter, hop count parameter, and message parameter; Based on the detection start point, the detection end point, and the detection parameter set, a detection task configuration corresponding to the detection link is generated. The detection task configuration is used to instruct the switch to perform active detection of the corresponding detection link.

[0010] In one possible implementation, the step of configuring and sending the probe task to the corresponding switch to perform active probe through the switch includes: Determine the target switch corresponding to the detection task configuration, and send the detection task configuration to the switch; If the current detection level is level zero, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor connected to the target switch as the detection ending point. The virtual interface is a gateway used to access the graphics processor. If the current detection level is the first level, then the loopback interface of the target switch is used as the detection start point and the loopback interfaces of other switches are used as the detection end point to perform active detection. The loopback interface represents the detection endpoint when active detection is performed between switches. If the current detection level is the second level, then the virtual interface of the target switch is used as the detection starting point and the virtual interfaces of other switches are used as the detection ending point to perform active detection; If the current detection level is the third level, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor deployed across the switch as the detection endpoint.

[0011] In one possible implementation, after generating the set of probe links for active probing based on the network topology information, the method further includes: Obtain the detection level and number of links corresponding to each detection link in the detection link set; The sampling ratio corresponding to each detection level is determined based on the number of links; Based on the sampling ratio, a portion of the detection links are selected from the set of detection links as sampling detection links; Based on the sampling and detection link, a corresponding sampling and detection task configuration is generated, and the sampling and detection task configuration is sent to the corresponding switch so that sampling and detection can be performed through the switch; Based on the execution results of the sampling and detection, update the detection status of the detection links in the detection link set that did not participate in the sampling and detection.

[0012] In one possible implementation, if the current detection level is the third level, the method further includes: The switches in the target network are divided into multiple switch clusters, and the switches in each switch cluster are numbered according to the same numbering rule to obtain switch numbers. For each switch, obtain the set of connected graphics processors, and obtain the first switch number in the current switch cluster for the first switch currently connected to each graphics processor; Each graphics processor is associated with a second switch in another switch cluster, and the second switch has the same number as the first switch. Active probing is performed using the virtual interface of the second switch as the starting point and the associated graphics processor as the ending point. Alternatively, each graphics processor can be associated with a third switch in the current switch cluster, with the third switch number being different from the first switch number; Active probing is performed using the virtual interface of the third switch as the starting point and the associated graphics processor as the ending point.

[0013] In a second aspect, embodiments of the present invention provide a network detection device, comprising: The acquisition module is used to acquire network topology information in the target network. The network topology information includes the connection relationship between each switch in the target network, as well as the access port information between the graphics processor in the target network and the connected switch. The first generation module is used to generate a set of detection links for active detection based on the network topology information. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. The second generation module is used to generate a detection task configuration based on the detection link set; The sending module is used to send the detection task configuration to the corresponding switch so that active detection can be performed through the switch.

[0014] Thirdly, embodiments of the present invention provide a computer device, including: a processor and a memory, wherein the processor is configured to execute a network detection program stored in the memory to implement the network detection method described in any one of the first aspects above.

[0015] Fourthly, embodiments of the present invention provide a storage medium storing one or more programs, which can be executed by one or more processors to implement the network detection method described in any one of the first aspects.

[0016] The network probing scheme provided in this invention obtains network topology information of the target network, including the connection relationships between switches in the target network and the access port information between the graphics processor and the connected switches. Based on the network topology information, a set of probe links for active probing is generated. Each probe link represents the correspondence between the probe start point and the probe end point in an active probe, where the probe start point is a switch and the probe end point is a switch or a graphics processor. A probe task configuration is generated according to the probe link set. The probe task configuration is sent to the corresponding switch to execute active probing. Thus, a set of probe links covering between switches and between graphics processors and switches can be automatically generated and the configuration can be generated and sent for execution. This solves the problems of traditional manual or fixed template configurations being prone to omissions and misconfigurations, unable to adapt to large-scale clusters, high probe redundancy, and inefficient fault location. It achieves automated, full-scenario coverage, low redundancy, and rapid fault location for RoCE network probing. Attached Figure Description

[0017] Figure 1 A flowchart illustrating a network detection method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the four-layer detection link hierarchy in the network detection method provided in this embodiment of the invention; Figure 3 A visual schematic diagram of the third-level sampling group detection strategy in the network detection method provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of a network detection system provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the structure of a network detection device provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] To facilitate understanding of the embodiments of the present invention, further explanations and descriptions will be provided below with reference to the accompanying drawings and specific embodiments. These embodiments do not constitute a limitation on the embodiments of the present invention.

[0020] Figure 1 This is a flowchart illustrating a network detection method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method specifically includes: S11. Obtain network topology information in the target network. The network topology information includes the connection relationship between each switch in the target network, as well as the access port information between the graphics processor in the target network and the connected switch.

[0021] The network probing method provided in this invention is applicable to RoCE network environments in large-scale data centers, and can be applied to computing device clusters consisting of multiple graphics processing unit (GPU) servers and multi-level switches. The network management system or centralized control node acts as the execution entity, automatically generating and scheduling active probing tasks based on the access relationships and network topology of the GPUs and switches in the target network, and then distributing these tasks to each switch for execution. By configuring and executing network quality analysis probing tasks on the switch side, continuous monitoring of the GPU communication links and network path status is achieved. This method is suitable for multi-tenant, large-scale clusters and other application scenarios with high requirements for network reliability and fault location efficiency.

[0022] In this embodiment, the target network can be a RoCE network, composed of RoCE series switches (e.g., RoCE_TOR switch: an access switch directly connected to the graphics processing unit (GPU), equipped with a Vlanif interface (GPU gateway) and a Loopback interface; RoCE_AGG switch: an aggregation switch used to connect multiple RoCE_TOR switches and RoCE_CORE switches, undertaking data forwarding and aggregation functions; RoCE_CORE switch: a core switch, the backbone of the entire RoCE network, responsible for high-speed data forwarding across aggregation areas.) and graphics processing units (GPUs). Network topology information is used to describe the connection relationships between switches and graphics processing units in the target network (e.g., the connection relationships between switches (network backbone layer topology) and the GPU and switch access port information (terminal access layer topology)), and is the basic data for automatically generating probe links. The graphics processing unit (GPU) is a network terminal computing device that connects to the RoCE_TOR switch through a physical port.

[0023] Specifically, a list of switches participating in RoCE communication in the target network is obtained. Port configuration information and port adjacency information are retrieved from each switch. Port adjacency information indicates whether a physical or logical connection has been established between that port and another switch port. Based on the port adjacency information, interconnected switches are associated to form connection relationship data. This connection relationship includes at least the identifiers of the interconnected switches, the corresponding port identifiers, and the connection direction. Through this method, the topology of the switches in the target network can be obtained, representing the basic forwarding path between different switches in the network.

[0024] When obtaining graphics processor (GPU) access relationships, the focus is not on the physical location of the GPU itself, but rather on the mapping relationship between its network access point and switch ports. Specifically, this includes: obtaining a list of GPUs participating in RoCE communication in the target network from the compute cluster management system or resource management system; for each GPU, obtaining its network address information and its network subnet information to determine the range of access switches corresponding to that GPU; and obtaining the port configuration and virtual interface configuration corresponding to the network subnet from the top-of-rack access switch (RoCE_TOR switch) in the target network. The virtual interface acts as the network gateway for the GPU, that is, the RoCE_TOR Vlanif (virtual interface) acts as the GPU gateway.

[0025] Based on the network address of the graphics processor and the port or virtual interface configuration of the corresponding access switch, each graphics processor is associated with its corresponding access switch and access port, forming access port information between the graphics processor and the access switch, which represents the switch to which each graphics processor is connected.

[0026] After acquiring the above two types of information, they are organized into network topology information. The network topology information includes at least: the access port information of the switches and the graphics processors they are connected to, and the connection relationships between the switches. This network topology information serves as the basic data for generating the probe link set, dividing the probe layers, and generating the probe task configuration, and is continuously called and reused.

[0027] In one possible implementation, the connection relationships between switches in the target network are obtained as first topology information; the switches and access port information connected to each graphics processor in the target network are obtained; second topology information is generated based on the correspondence between each graphics processor and the connected switches and the access port information; and network topology information is constructed based on the first and second topology information, which is used to determine the correspondence between the detection start point and the detection end point in the detection link set.

[0028] In this embodiment, switches in the target network are uniformly identified and managed to determine the set of switches participating in target network communication. Port configurations and port adjacency information are obtained from each switch, where port adjacency information indicates whether a physical or logical connection has been established between different switch ports. Based on the port adjacency information, switches with existing connections are paired and associated to generate connection relationships between switches. A connection relationship includes at least a switch identifier, a corresponding port identifier, and a connection direction. The connection relationships between switches are stored as first topology information.

[0029] When acquiring graphics processor (GPU) access information, the set of GPUs participating in communication within the target network is determined. For each GPU, its network address information and its network subnet information are obtained to locate the range of access switches corresponding to that GPU. The port configuration and virtual interface configuration corresponding to the network subnet are obtained from the access switches. The virtual interface acts as the network gateway for the GPU, carrying its network communication traffic. Based on the matching relationship between the GPU's network address and its port and virtual interface configurations, the switch to which each GPU is connected and its corresponding access port information are determined.

[0030] The graphics processor identifier, access switch identifier, and access port identifier are associated to generate the access relationship between the graphics processor and the switch. The access relationship is used as the second topology information to characterize the access location of the graphics processor at the switch level and its network entry characteristics in the target network.

[0031] The switch connections in the first topology information are integrated with the graphics processor (GPU) access connections in the second topology information to form a unified network topology. This network topology allows us to determine the reachable paths from any GPU, through its access switches and intermediate switches, to another GPU or switch. The constructed network topology information is used to determine the correspondence between each probe start point and probe end point in the probe link set.

[0032] S12. Generate a set of probe links for active probing based on network topology information. Each probe link represents the correspondence between the probe start point and the probe end point in an active probing. The probe start point is a switch, and the probe end point is a switch or a graphics processor.

[0033] In this embodiment, the probe link set is a structured link set covering the entire Tier0-Tier3 level, generated based on network topology information. Each probe link contains key information such as a unique identifier, probe start point (switch and corresponding interface), probe end point (switch interface or GPU IP), hierarchical attributes, and communication method, which is the direct basis for generating probe task configurations in the future.

[0034] The probe originates from the switch initiating the probe and its designated interface. The interface type can be a Vlanif interface (GPU gateway) or a Loopback interface (logical interface), and must match the functionality of the corresponding probe level. The probe destination is the target object receiving the probe packets, which falls into two categories: the switch's Loopback interface (used for monitoring links between Tier 1 level switches) and the GPU's IP address (used for Tier 0 and Tier 3 level GPU activity detection). This destination must also be compatible with the probe origin's tier attributes and communication method. Tier attributes characterize the Tier 0-Tier 3 level identifier for each probe link, directly determining the link's probe logic, interface selection, and communication method.

[0035] Specifically, initial probe links are generated based on network topology information according to Tier 0-Tier 3 rules. Tier 0 establishes a Layer 2 communication link from the RoCE_TOR switch VLAN interface (GPU gateway) to the local GPU IP; Tier 1 generates bidirectional backbone links between Loopback interfaces of different tiers of switches; Tier 2 establishes cross-access layer links between different RoCE_TOR switch VLAN interfaces; and Tier 3 forms a Layer 3 communication link from the RoCE_TOR switch VLAN interface across groups or within the same group to the remote GPU IP, supporting arbitrary combinations of links at each level. The probe links generated at each level are then used to create a probe link set.

[0036] To reduce network load and switch CPU usage in large-scale cluster scenarios, a sampling and probing mechanism is introduced to optimize the link set. For Tier 3, a grouping strategy is implemented based on tenant and port dimensions. In cross-group scenarios, the GPUs connected to the RoCE_TOR are grouped according to preset rules, and gateways of other groups with the same RoCE_TOR number probe the corresponding group of GPUs. In scenarios within the same group, they are grouped by consecutive intervals, and gateways of other RoCE_TORs within the group probe them sequentially. At the same time, the legality and redundancy of all links at all levels are checked, and invalid and duplicate links are eliminated. Finally, a standardized probing link set containing key information such as the probe start point, end point, communication method, and hierarchical attributes is formed.

[0037] In one possible implementation, the connection paths between switches at different detection levels are determined based on network topology information, as well as the connection path between the switches and the graphics processor; based on the connection paths, each detection start point and its corresponding detection end point are determined, so as to generate multiple detection links based on the detection start point and detection end point, thus obtaining a set of detection links.

[0038] In this embodiment, based on the connection relationships between switches recorded in the network topology information and the access relationships between graphics processors and access switches, a set of candidate probe endpoints reachable when a switch is used as a probe starting point is determined. The candidate probe endpoint set includes other switches directly or indirectly connected to the probe starting point switch, as well as graphics processors reachable through the probe starting point switch, thus providing a basic correspondence for generating probe links. According to the candidate probe endpoint set, each probe starting point switch is sequentially combined with its corresponding probe endpoint to generate multiple probe links. Each probe link represents the correspondence between the probe starting point and the probe endpoint in an active probe. The generated probe link set is used for the generation and distribution of subsequent probe task configurations to control the switches to perform active probes according to the probe link set.

[0039] S13. Generate the detection task configuration based on the detection link set.

[0040] In this embodiment, the core attributes of each probe link are defined based on the probe link set, including probe level (Tier0-Tier3), probe start point and corresponding interface (Vlanif / Loopback), probe end point (GPU or switch) and specific identifier or interface (GPU IP or switch Loopback interface IP), communication method (Layer 2, Layer 3) and other key information.

[0041] According to the preset configuration rules, standardized NQA probe parameters are matched for each probe link: the probe packet type is uniformly specified as ICMP-echo, and differentiated parameters such as packet sending frequency, timeout, TTL (Time to Live), and Payload are assigned based on link hierarchy attributes and network scenario requirements (e.g., the packet sending frequency of Tier 1 backbone links is higher than that of Tier 3 cross-domain links to ensure real-time monitoring of backbone stability). Management attributes for the probe tasks are then added, including: task name, unique identifier, tenant affiliation, and execution priority. Finally, the key information, probe parameters, and management attributes obtained from each probe link are integrated and structured into a single probe task configuration according to the format required by the Netconf protocol. All configurations are then aggregated to form a complete probe task configuration set, supporting batch deployment and transactional execution.

[0042] In one possible implementation, for each probe link in the probe link set, the probe start point, probe end point, and probe level of the probe link are obtained; the probe parameter set corresponding to the probe level is obtained, and the probe parameter set includes at least one of the following: transmission interval parameter, timeout parameter, hop count parameter, and packet parameter; based on the probe start point, probe end point, and probe parameter set, a probe task configuration corresponding to the probe link is generated, and the probe task configuration is used to instruct the switch to perform active probes on the corresponding probe link.

[0043] In this embodiment, for each probe link in the probe link set, the corresponding probe start point, probe end point, and probe level are parsed from the probe link. The probe start point represents the switch performing the active probe, the probe end point represents the target object of this active probe, and the probe level indicates the probe type to which the probe link belongs. Subsequently, based on the probe level, a set of probe parameters corresponding to that probe level is obtained from a pre-established hierarchical parameter mapping relationship. The probe parameter set includes at least one or more of the following: transmission interval parameter, timeout parameter, hop count parameter, and packet parameters. These parameters are used to limit the execution frequency, waiting time, reachable path range, and characteristics of the probe packets. Specific probe parameters can be set according to the needs of the probe and are not specifically limited here.

[0044] After obtaining the set of detection parameters, the detection task is structured and configured based on the detection start point, detection end point, and the set of detection parameters, generating a detection task configuration that corresponds one-to-one with the detection link. The detection task configuration includes at least interface information to identify the detection start point, address information to identify the detection end point, and detection parameter information corresponding to the detection layer, which instructs the detection start point switch to perform active detection of the corresponding detection link to the detection end point according to the detection parameters.

[0045] S14. Send the probe task configuration to the corresponding switch to perform active probe through the switch.

[0046] In this embodiment, the target switch corresponding to the detection start point is determined based on the detection start point information recorded in the detection task configuration. Subsequently, the detection task configuration is sent to the target switch, and the configuration is loaded onto the target switch, enabling it to recognize the detection task content relevant to itself. After the detection task configuration is loaded, the target switch is controlled to periodically or as needed perform active detection operations based on the detection start point, detection end point, and detection parameter information contained in the configuration. The detection results generated during the detection process are used for subsequent status analysis and fault location.

[0047] In one possible implementation, the target switch corresponding to the probe task configuration is determined, and the probe task configuration is sent to the switch. If the current probe level is level zero, active probe is performed with the virtual interface of the target switch as the probe start point and the graphics processor connected to the target switch as the probe end point. The virtual interface is the gateway used to access the graphics processor. If the current probe level is level one, active probe is performed with the loopback interface of the target switch as the probe start point and the loopback interfaces of other switches as the probe end points. The loopback interface represents the probe endpoint when active probe is performed between switches. If the current probe level is level two, active probe is performed with the virtual interface of the target switch as the probe start point and the virtual interfaces of other switches as the probe end points. If the current probe level is level three, active probe is performed with the virtual interface of the target switch as the probe start point and the graphics processor deployed across switches as the probe end point.

[0048] In this embodiment, based on the detection start point identifier recorded in the detection task configuration, the target switch corresponding to the detection task configuration is determined, and the detection task configuration is sent to the target switch, enabling the target switch to load and recognize the detection task configuration. The detection task configuration includes detection level information, which instructs the target switch to perform active detection using the detection start point interface type and detection end point type corresponding to the current detection level.

[0049] When the current detection level is level zero, the target switch selects the virtual interface (i.e., RoCE_TOR Vlanif (GPU gateway)) used to access the graphics processor from its own interface as the detection starting point. The virtual interface acts as the network gateway for the graphics processor and is used to carry the communication traffic of the graphics processor. At the same time, based on the access relationship information recorded in the detection task configuration, the target switch determines the graphics processor connected to the virtual interface as the detection endpoint and controls the target switch to send probe packets to the graphics processor (GPU IP) through the virtual interface to perform active detection of the graphics processor's access status.

[0050] When the current detection level is the first level, the target switch selects a loopback interface (RoCE_TOR Loopback, RoCE_AGG Loopback, RoCE_CORE Loopback) from its own interfaces as the detection start point. The loopback interface is a logical interface configured on the switch to identify itself and serve as the detection endpoint when the switches perform active detection. At the same time, based on the switch connection relationship recorded in the detection task configuration, the loopback interfaces of other switches are determined as the detection endpoints, and the target switch is controlled to send detection packets to the detection endpoints to perform active detection of the communication path between switches.

[0051] When the current detection level is the second level, the target switch selects its own virtual interface used to access the graphics processor as the detection starting point, and determines the corresponding virtual interfaces on other switches as the detection endpoints based on the switch association information recorded in the detection task configuration. This controls the target switch to perform active detection between the virtual interfaces of the switches to detect the network connectivity status between different switch access layers.

[0052] When the current detection level is the third level, the target switch also uses its own virtual interface as the detection starting point, and determines the graphics processor deployed under other switches as the detection endpoint based on the cross-switch access relationship recorded in the detection task configuration. Then, it controls the target switch to send detection messages to the graphics processor through the virtual interface to perform active detection of the graphics processor communication path in the cross-switch scenario.

[0053] In this way, the target switch can automatically select the corresponding detection start interface type and detection end object according to different detection levels, and execute active detection matching the detection level according to the detection task configuration, thereby realizing the unified distribution and execution of multi-level detection tasks.

[0054] In practice, the device identifier (e.g., RoCE_TOR-1, RoCE_AGG-2) corresponding to the detection starting point of each configuration is extracted from the generated detection task configuration set, and the corresponding physical switch is matched as the target switch. Batch distribution is performed using the Netconf protocol. After receiving the detection task configuration, the target switch stores it in the runtime configuration file and synchronizes it to the local NQA module.

[0055] After receiving the configuration, the switch's native NQA module parses the probe level and performs active probing according to the corresponding rules. First, it identifies the probe level. For Tier 0 (GPU access layer probing), the NQA module parses the virtual interface (Vlanif) identifier of the target switch in the configuration. Combining this with the local interface configuration table, it locates the IP address of that virtual interface (i.e., the GPU gateway address). Once the interface status is confirmed to be Up, it is used as the probe starting point. If the interface status is abnormal, an alarm is triggered and the probe on that link is paused. Based on the list of GPU IPs connected to the target switch associated with the configuration (derived from the GPU-switch binding relationship), the probe endpoint is determined to be the IP address of the GPU directly connected to the local RoCE_TOR switch.

[0056] The system employs a Layer 2 communication method (no routing forwarding required, direct transmission based on the VLAN broadcast domain) to construct probe packets according to configured parameters (such as ICMP-echo messages, packet sending frequency, timeout time, etc.) and send them from the Vlanif interface (virtual interface) to the target GPU IP. The NQA module records the packet sending results in real time (such as whether a response is received, response delay, packet loss rate) and stores them in the local probe result cache.

[0057] For Tier 1 (network backbone link monitoring), the NQA module extracts the loopback interface identifier (e.g., Loopback0) of the target switch in the configuration, queries the local interface configuration to obtain the IP address of that interface, and uses it as the starting point for probing, ensuring that even fluctuations in physical ports do not affect the probing execution. Based on the loopback interface IPs of other switches in the configuration (derived from the network topology relationship between switches), the probe endpoint is locked to the loopback interface IP of the cross-tier switch (e.g., RoCE_TOR corresponds to RoCE_AGG, RoCE_AGG corresponds to RoCE_CORE), and bidirectional probing is supported (e.g., RoCE_TOR-1→RoCE_AGG-1, RoCE_AGG-1→RoCE_TOR-1).

[0058] ICMP-echo probe messages are constructed according to configuration parameters and sent from the starting Loopback interface to the ending Loopback interface, focusing on monitoring the connectivity and transmission stability of the backbone link; the NQA module records data such as round-trip delay and packet loss rate for each round of probes, providing a basis for backbone network fault diagnosis.

[0059] For Tier 2 (cross-access layer connectivity monitoring), the virtual interface (Vlanif) identifier of the target switch in the parsing configuration is used as the starting point for probing after confirming that the interface status is normal. This ensures consistency with the starting interface type of Tier 0 / Tier 3, simplifying the configuration logic. Based on the Vlanif interface IPs of other RoCE_TOR switches in the configuration (derived from the cross-access layer network topology), the probe endpoint is locked to the GPU gateway interface IP of another RoCE_TOR switch (i.e., the virtual interface (Vlanif) of another switch). This does not involve the GPU itself, focusing only on the connectivity between access layer gateways. ICMP-echo probe packets are sent according to the configuration parameters, supporting automatic adaptation of Layer 2 or Layer 3 communication methods based on network VLAN division, verifying whether the network path between different access layer switches is unobstructed. The probe results are cached locally in real time, waiting for periodic reporting by the Telemetry protocol.

[0060] For Tier 3 (cross-domain GPU liveness detection), similar to Tier 0 / Tier 2, the probe starts from the VLAN interface (GPU gateway) of the target switch (RoCE_TOR), ensuring the interface is in normal condition and routeable to the VLAN to which the target GPU belongs. Based on the GPU IPs deployed across switches in the configuration (derived from cross-group / same-group GPU grouping relationships), the probe endpoint is locked to a remote GPU IP accessed by another RoCE_TOR switch, and this GPU has been selected as a representative probe target through the sampling grouping policy (cross-group same-number grouping / G same-group continuous interval grouping). A Layer 3 communication method (relying on routing forwarding to achieve cross-domain reachability) is used, and probe packets are constructed according to the configuration parameters and sent from the VLAN interface to the remote GPU IP. The NQA module records the probe response results, while following the minimum link change mechanism for GPU changes. If a GPU is added during the probe process, the probe links of existing groups are supplemented first; if a GPU is reduced, only the probe configuration of the corresponding group is adjusted, without affecting the probe execution of other groups and tiers, ensuring probe stability and low redundancy.

[0061] As an example, the table below shows the detection start point, detection end point, and function corresponding to different detection levels.

[0062]

[0063] When specifically implementing the network detection method of this embodiment, the following steps are performed: 1. Collect the switch-level topology of the RoCE network and the mapping relationship between GPUs and access switches.

[0064] 2. Probe Link Generation: Based on the collected GPU-switch hierarchical information, Tier 0–Tier 3 probe links are generated. Links with any combination of tiers are supported. To reduce network load and switch CPU usage, a sampling probe concept is introduced, selecting some links according to a GPU grouping (by tenant and port) strategy, reducing the number of probe tasks without affecting overall coverage.

[0065] 3. Probe Task Configuration Generation: Based on the link generation results, generate the NQA probe task configuration, including: probe source device, destination device, probe packet type (ICMP-echo), packet transmission frequency, timeout, TTL, payload, and other parameters. 4. NQA Probe Task Distribution: The generated NQA configuration is distributed to the switch through the control layer using Netconf, ensuring transactional and batch distribution capabilities, and supporting dynamic modification, stopping, and starting of NQA probe tasks.

[0066] 5. Probe Task Management and Optimization: Based on cluster topology or GPU changes, receive MQ, automatically adjust NQA probe links, and support manual differential updates.

[0067] As an example, such as Figure 2 The diagram shows a four-layer probe link hierarchy in the network probing method provided in this embodiment of the invention. 1. Tier 0 (GPU access layer activation), located in the area between the RoCE_TOR on the left and the directly connected GPUs. Probing logic: Based on the direct connection relationship of "RoCE_TOR→GPU" in the diagram, the probe starts from the Vlanif interface (GPU gateway) of the RoCE_TOR and ends with the GPU IP connected to the RoCE_TOR (GPUs numbered "1-64" in the diagram). Layer 2 communication is used (no need for cross-switch forwarding) to verify whether the local GPU is properly connected, corresponding to the GPU activation (Layer 2 communication) function.

[0068] 2. Tier 1 (Network backbone link monitoring), located in the area between RoCE_TOR and RoCE_AGG in the diagram. Detection logic: Follow the steps in the diagram: "RoCE_TOR..." The connection relationship of "RoCE_AGG" starts from the Loopback interface of RoCE_TOR and ends at the Loopback interface of RoCE_AGG. It adopts three-layer communication and monitors the link connectivity and stability between backbone switches. The corresponding link monitoring function (RoCE_CORE is not marked in the figure, but the logic can be extended to the connection between RoCE_AGG and RoCE_CORE).

[0069] 3. Tier 2 (Cross-Access Layer Connectivity Monitoring), Location in the diagram: The area on the right where multiple RoCE_TORs are associated (RoCE_TORs appear repeatedly and are interconnected in the diagram). Detection logic: Follow the "RoCE_TOR" in the diagram. The "RoCE_TOR" cross-switch relationship starts with the Vlanif interface of one RoCE_TOR and ends with the Vlanif interface of another RoCE_TOR (both are GPU gateways). It does not need to go through the GPU. The core function is to verify whether the network path between different access layer switches is unobstructed, corresponding to the network connectivity monitoring function.

[0070] 4. Tier 3 (Cross-Domain GPU Liveness Detection): Location in the diagram: The area on the right between the RoCE_TOR and non-directly connected GPUs (across multiple RoCE_TORs). Detection Logic: Following the indirect connection relationship of "RoCE_TOR → Cross-Switch GPU" in the diagram, starting from the Vlanif interface of the local RoCE_TOR and ending at the GPU IP connected to other RoCE_TORs (GPU number "M" in the diagram), it uses Layer 3 communication (requiring cross-switch forwarding). The core verification is "cross-domain GPU connectivity," corresponding to the function of cross-RoCE_TOR GPU liveness detection (Layer 3 communication), and implicitly includes sampling grouping logic (in the diagram, "N" and "M" represent representative GPUs after grouping).

[0071] In one possible implementation, after generating a set of probe links for active probing based on network topology information, the method further includes: Obtain the detection level and number of links corresponding to each detection link in the detection link set; determine the sampling ratio corresponding to each detection level based on the number of links; select a portion of the detection links from the detection link set as sampling detection links according to the sampling ratio; generate the corresponding sampling detection task configuration based on the sampling detection links, and send the sampling detection task configuration to the corresponding switch to execute the sampling detection through the switch; update the detection status of the detection links in the detection link set that did not participate in the sampling detection according to the execution result of the sampling detection.

[0072] In this embodiment, the set of probe links is traversed and parsed to obtain the probe level corresponding to each probe link. The probe links are then categorized and counted according to the probe level to obtain the number of probe links under each probe level. The probe level is used to characterize the type and coverage of the probe links, and the number of links is used to reflect the scale of the probe links under that probe level.

[0073] After obtaining the number of links corresponding to each detection layer, a higher sampling ratio is allocated to detection layers with more links, and a lower sampling ratio is allocated to detection layers with fewer links, based on the differences in link scale across different detection layers. This controls the overall scale of active detection while ensuring detection coverage. The sampling ratio is used to limit the proportion of detection links participating in sampling detection at the corresponding detection layer. Based on the sampling ratio, a subset of detection links is selected from the detection link set as sampling detection links. The selection process is executed independently within each detection layer to ensure that sampling between different detection layers does not affect each other and that the selected sampling detection links can represent the overall link status at the corresponding detection layer. Detection links that are not selected do not participate in this active detection.

[0074] After determining the sampling and probing links, a corresponding sampling and probing task configuration is generated based on these links. The sampling and probing task configuration includes at least the probing start point, probing end point, and probing parameter information matching the probing level for each sampling and probing link. Subsequently, the sampling and probing task configuration is sent to the corresponding switch, causing the switch to load and execute the sampling and probing task, thereby performing active probing on the sampling and probing links through the switch. After the sampling and probing is completed, the probing results for each sampling and probing link are obtained. Based on the probing results of the sampling and probing links within the same probing level or the same path range, the status of the probing links that did not participate in the sampling and probing is inferred and updated to reflect the overall status of the network links under the corresponding probing level. This reduces probing overhead while continuously maintaining the target network's probing status (for example, if the sampling and probing results show that all probing links are normal, then the status of the sampled probing links is determined to be normal communication).

[0075] In one possible implementation, if the current detection level is the third level, the method further includes: The switches in the target network are divided into multiple switch clusters. For each switch cluster, the switches are numbered using the same numbering rule to obtain switch numbers. For each switch, the set of connected graphics processors (GPUs) is obtained, along with the first switch number within the current switch cluster that each GPU is currently connected to. Each GPU is associated with a second switch in another switch cluster, with the second switch number being the same as the first switch number. Active probing is performed, using the virtual interface of the second switch as the probe starting point and the associated GPU as the probe endpoint. Alternatively, each GPU is associated with a third switch in the current switch cluster, with the third switch number being different from the first switch number. Active probing is performed, using the virtual interface of the third switch as the probe starting point and the associated GPU as the probe endpoint.

[0076] In this embodiment, the switches in the target network are first divided into multiple switch clusters based on the topology of the target network and the hierarchical attributes of the switches. Each switch cluster represents a group of switches that are the same or similar in terms of network structure and functional role. For each switch cluster, the switches within the cluster are sequentially numbered according to a unified numbering rule to obtain switch numbers, thereby logically forming a corresponding switch association relationship between switches with the same number in different switch clusters (for example, cluster A contains switch numbers 1-8, and cluster B contains switch numbers 1-8).

[0077] After completing the switch cluster partitioning and numbering, for each switch, the set of graphics processors currently connected to it is obtained, and the first switch number of the first switch connected to each graphics processor in the current switch cluster is determined. The first switch number is used to uniquely identify the access position of the graphics processor in its own switch cluster, providing a basis for subsequently establishing switch associations across or within a cluster.

[0078] Based on the first switch number, two optional probe associations are established for each graphics processor (GPU). The first type is a cross-cluster association, where the GPU is associated with a second switch in another switch cluster, where the second switch has the same number as the first switch in its cluster. This allows switches with the same number in different switch clusters to act as probe starting points for each other, performing cross-cluster active probes on the associated GPU. In this case, the virtual interface of the second switch is used as the probe starting point, and the associated GPU is used as the probe endpoint, controlling the second switch to perform the active probe.

[0079] The second type is the intra-cluster association method, which associates the graphics processor (GPU) with a third switch in the current switch cluster. The third switch has a different identifier in the cluster than the first switch. This method establishes a probe relationship between switches with different identifiers within the same cluster, enabling collaborative probes of the GPU by multiple switches within the cluster. In this association method, the virtual interface of the third switch serves as the probe starting point, and the GPU as the probe endpoint, controlling the third switch to perform active probes.

[0080] By utilizing the above method and the switch cluster partitioning and unified numbering rules, the automatic establishment of the detection relationship between the graphics processor and different switches is realized. This enables active detection to cover the communication paths between different switch clusters as well as the multi-switch scenarios within the same switch cluster. This ensures detection coverage while avoiding random changes in the detection links, thereby improving the stability and maintainability of the detection task.

[0081] The network probing method provided in this invention automatically acquires and models the access relationships between switches and graphics processors in the target network. This allows for the automatic generation of active probing links covering different levels and switch clusters without manual intervention, avoiding omissions and misconfigurations caused by manual configuration and improving the accuracy and consistency of active probing in large-scale networks. Furthermore, by introducing layered and sampling probing mechanisms, the number of probing tasks required is significantly reduced while ensuring effective coverage of critical links and typical paths. This reduces the probing load on the switch side and additional network overhead, making the solution applicable to large-scale graphics processor clusters and multi-switch deployment scenarios, balancing probing coverage and system performance.

[0082] Furthermore, by establishing stable detection associations based on switch cluster numbering rules, and adjusting only the local detection configuration when devices are added or removed, the detection link has good stability and scalability. This facilitates continuous monitoring and rapid location of target network faults, thereby improving the overall network reliability and operational efficiency.

[0083] As an example, such as Figure 3 The diagram shown is a visualization of the third-level sampling group probing strategy in the network probing method provided in this embodiment of the invention. Specifically, it uses switch clusters (Group), RoCE_TOR switches, and GPU grouping examples to intuitively demonstrate the rules for cross-Group probing and same-Group probing.

[0084] Figure 3 In the diagram, G1, G2, and G16 represent switch clusters, RoCE_TOR-1 to 8 represent switches within the cluster, and the number range represents GPU groups. This is explained in two scenarios: Left Group 2: Same as Group Probe, indicating cross-RoCE_TOR probes within the same cluster. Core rule: Within the same Group, the GPUs connected to a RoCE_TOR are grouped according to consecutive intervals, and the gateways of other RoCE_TORs in the cluster probe each other sequentially. Taking Group 2 as an example, G2 contains RoCE_TOR-1~8. GPU grouping method: The GPUs connected to each RoCE_TOR are divided into 7 groups according to consecutive intervals (the image only shows a partial grouping). The GPUs of G2-ROCE_TOR-2 are divided into groups "1-10" (Group 1: GPUs 1~10); the GPUs of G2-ROCE_TOR-3 are divided into groups "11-19" (Group 2: GPUs 11~19); the GPUs of G2-ROCE_TOR-4 are divided into groups "20-28" (Group 3: GPUs 20~28); and so on, until the "56-64" group of G2-ROCE_TOR-8 (Group 7: GPUs 56~64).

[0085] The detection execution logic is as follows: A gateway for a RoCE_TOR within a group only probes GPUs in groups corresponding to other RoCE_TORs to avoid duplication. For example, the gateway for G2-ROCE_TOR-1 specifically probes GPUs in groups "1-10" of G2-ROCE_TOR-2 and GPUs in groups "11-19" of G2-ROCE_TOR-3; the gateway for G2-ROCE_TOR-2 specifically probes GPUs in groups "11-19" of G2-ROCE_TOR-3 and GPUs in groups "20-28" of G2-ROCE_TOR-4. Ultimately, this ensures that all GPU groups of all RoCE_TORs within the group are probed, but each switch only undertakes a portion of the probe task, reducing the load.

[0086] Group 1 on the right: Cross-Group Probes, indicating cross-RoCE_TOR probes between different clusters. Core rule: Within different groups, the GPUs connected to RoCE_TOR are grouped according to the same group number, and gateways with the same RoCE_TOR number in other groups can detect each other.

[0087] Figure 3 The system involves three groups: G1, G2, and G16. The GPU grouping method is as follows: all RoCE_TORs within a group are connected to GPUs and uniformly divided into 8 groups (each group has the same number of GPUs). For example, GPUs in G2-ROCE_TOR-1 are divided into groups "1-8" (Group 1: GPUs 1-8) and "9-16" (Group 2: GPUs 9-16). GPUs in G16-ROCE_TOR-1 are also divided into 8 groups such as "1-8" and "9-16" (consistent with the grouping rule of G2). All RoCE_TORs in all groups are divided into 8 groups according to this rule to ensure that the group numbers correspond. The detection execution logic is as follows: A RoCE_TOR numbered N in a certain group only detects GPUs in the same group as RoCE_TORs numbered N in other groups. For example, the gateway for G2-ROCE_TOR-1 specifically detects GPUs in groups "1-8" of G16-ROCE_TOR-1 and groups "1-8" of G1-ROCE_TOR-1; the gateway for G2-ROCE_TOR-2 specifically detects GPUs in groups "9-16" of G16-ROCE_TOR-2; and the gateway for G2-ROCE_TOR-8 specifically detects GPUs in groups "57-64" of G16-ROCE_TOR-8. This covers core cross-cluster scenarios without requiring all RoCE_TORs in Group A to detect all RoCE_TORs in Group B.

[0088] As an example, such as Figure 4 The diagram shows a network detection system according to an embodiment of the present invention, comprising three core components: Agent-NQA Module, Controller, and Analyzer. Specifically, the Agent-NQA Module (switch-side functional module) is a native detection unit built into the switch. The Agent refers to the proxy module, which is a functional carrier deployed on the switch; the NQA Module, or "Network Quality Analyzer Module," is the execution end of the detection task. It is responsible for constructing ICMP-echo and other detection packets as needed and sending them to the target endpoint (GPU or other switch), while simultaneously recording the detection results (such as response latency, packet loss rate, and connectivity status) in real time. It is the direct execution carrier for Tier 0 to Tier 3 layer 4 detection tasks.

[0089] The diagram shows the following connection logic: directly bound to the Switch, with "NQA - Tier 0", "NQA - Tier 1", "NQA - Tier 2", and "NQA - Tier 3" clearly marked below, indicating that all four layer probing tasks are executed by this module on the switch; at the same time, the probing results are reported to the Analyzer through the probing data interaction arrows.

[0090] The diagram labels "Int - Loop (Loopback Interface)" and "Int - Vlanif (Vlanif Interface)" correspond to the rule that Tier 1 uses the Loopback interface, while Tier 0 / Tier 2 / Tier 3 uses the Vlanif interface as the starting point for probing. RoCE Fabric, or "RoCE network fabric architecture," refers to the network backbone composed of RoCE_TOR, RoCE_AGG, and RoCE_CORE switches. It serves as the execution scenario for Tier 1 (backbone link monitoring) and Tier 2 (cross-access layer connectivity monitoring) probing tasks, reflecting the overall network topology.

[0091] XPath is a language used to locate nodes in an XML document. "XPath: NQA" means that Analyzer uses XPath syntax to extract NQA probe information from switch configuration or data, ensuring the accuracy of data collection. This is a technical detail reflecting Analyzer's monitoring and collection functions.

[0092] 2. The Controller (control layer module) is the system's decision-making and management center, responsible for coordinating the entire lifecycle management of probing tasks. Task distribution: The "IP Topology Collection" function collects the IP topology of the RoCE cluster (IP interconnection relationships between switches and between switches and GPUs) and GPU access information, stores it in the basic information database, and then determines the NQA probing combination (such as probing start point, end point, and level) based on this data. Finally, the probing configuration is distributed to the Agent - NQA Module of the switch via the Netconf protocol (Network Configuration Protocol).

[0093] It also supports operations such as modifying, stopping, and starting probe tasks, such as adjusting the packet sending frequency of Tier 1 and pausing the probe links of Tier 0 that have been repaired. Cluster-level NQA probe tasks can be generated according to preset rules (such as sampling grouping strategies and topology dynamic update rules) to ensure that the tasks are continuously executed and automatically updated with the network topology / GPU status (such as automatically supplementing Tier 3 probe links when a GPU is added).

[0094] Interact with the Agent - NQA Module via the NQA List (NQA Probe Task List).

[0095] 3. The Analyzer module is responsible for receiving and analyzing the probe data and outputting the operation and maintenance results.

[0096] The probe data sent by the switch Agent - NQA Module (such as the connectivity status of Tier 3 cross-domain GPUs and the packet loss rate of Tier 1 backbone links) is received through the Telemetry protocol and then cleaned and integrated (such as classified and statistically analyzed by level).

[0097] It supports linking white-box monitoring items (such as switch port traffic and GPU utilization) and displays the detection results in a hierarchical manner (such as GPU access status corresponding to Tier0 and backbone link status corresponding to Tier1), while generating "alarms / maintenance events" (such as triggering an alarm when the Tier2 cross-access layer link is interrupted).

[0098] Figure 5 This is a schematic diagram of the structure of a network detection device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the device specifically includes: The acquisition module 51 is used to acquire network topology information in the target network. The network topology information includes the connection relationship between each switch in the target network, and the access port information between the graphics processor in the target network and the connected switch. The first generation module 52 is used to generate a set of detection links for active detection based on the network topology information. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. The second generation module 53 is used to generate a detection task configuration based on the detection link set; The sending module 54 is used to send the detection task configuration to the corresponding switch so that active detection can be performed through the switch.

[0099] In one possible implementation, the acquisition module is specifically used to acquire the connection relationship between each switch in the target network as the first topology information; Obtain information about the switches and access ports to which each graphics processor in the target network is connected; A second topology information is generated based on the correspondence between each graphics processor and the access switch, as well as the access port information; Based on the first topology information and the second topology information, the network topology information is constructed, and the network topology information is used to determine the correspondence between the detection start point and the detection end point in the detection link set.

[0100] In one possible implementation, the first generation module is specifically used to determine the connection path between switches at different detection levels based on the network topology information, and to determine the connection path between the switch and the graphics processor. Based on the connection path, each detection start point and corresponding detection end point are determined, so as to generate multiple detection links based on the detection start point and detection end point, thus obtaining the detection link set.

[0101] In one possible implementation, the second generation module is specifically used to obtain the detection start point, detection end point and detection level of each detection link in the detection link set; Obtain the set of detection parameters corresponding to the detection level, wherein the set of detection parameters includes at least one of the following: transmission interval parameter, timeout parameter, hop count parameter, and message parameter; Based on the detection start point, the detection end point, and the detection parameter set, a detection task configuration corresponding to the detection link is generated. The detection task configuration is used to instruct the switch to perform active detection of the corresponding detection link.

[0102] In one possible implementation, the sending module is specifically used to determine the target switch corresponding to the probe task configuration and send the probe task configuration to the switch; If the current detection level is level zero, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor connected to the target switch as the detection ending point. The virtual interface is a gateway used to access the graphics processor. If the current detection level is the first level, then the loopback interface of the target switch is used as the detection start point and the loopback interfaces of other switches are used as the detection end point to perform active detection. The loopback interface represents the detection endpoint when active detection is performed between switches. If the current detection level is the second level, then the virtual interface of the target switch is used as the detection starting point and the virtual interfaces of other switches are used as the detection ending point to perform active detection; If the current detection level is the third level, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor deployed across the switch as the detection endpoint.

[0103] In one possible implementation, the acquisition module is further configured to acquire the detection level and number of links corresponding to each detection link in the detection link set; The second generation module is further configured to determine the sampling ratio corresponding to each detection level based on the number of links; Based on the sampling ratio, a portion of the detection links are selected from the set of detection links as sampling detection links; Based on the aforementioned sampling and detection link, a corresponding sampling and detection task configuration is generated; The sending module is further configured to send the sampling and detection task configuration to the corresponding switch so that the sampling and detection can be performed through the switch; Based on the execution results of the sampling and detection, update the detection status of the detection links in the detection link set that did not participate in the sampling and detection.

[0104] In one possible implementation, the processing module 55 is used to divide the switches in the target network into multiple switch clusters, and number the switches in each switch cluster according to the same numbering rule to obtain switch numbers. For each switch, obtain the set of connected graphics processors, and obtain the first switch number in the current switch cluster for the first switch currently connected to each graphics processor; Each graphics processor is associated with a second switch in another switch cluster, and the second switch has the same number as the first switch. Active probing is performed using the virtual interface of the second switch as the starting point and the associated graphics processor as the ending point. Alternatively, each graphics processor can be associated with a third switch in the current switch cluster, with the third switch number being different from the first switch number; Active probing is performed using the virtual interface of the third switch as the starting point and the associated graphics processor as the ending point.

[0105] The device provided in this embodiment may be as follows: Figure 5 The apparatus shown can perform, for example Figure 1 All steps of the network detection method are then implemented to achieve... Figure 1 For details on the technical effectiveness of the network detection method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0106] Figure 6 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. Figure 6 The computer device 600 shown includes at least one processor 601, a memory 602, at least one network interface 604, and other user interfaces 603. The various components in the computer device 600 are coupled together via a bus system 605. It is understood that the bus system 605 is used to implement communication between these components. In addition to a data bus, the bus system 605 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 6 The general designated all buses as Bus System 605.

[0107] The user interface 603 may include a display, keyboard, or clicking device (e.g., mouse, trackball, touchpad, or touchscreen).

[0108] It is understood that the memory 602 in this embodiment of the invention can be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate Synchronous DRAM (DDRSDRAM), Enhanced Synchronous DRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), and Direct Rambus RAM (DRRAM). The memory 602 described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0109] In some implementations, memory 602 stores elements, executable units or data structures, or subsets thereof, or extended sets thereof: operating system 6021 and application program 6022.

[0110] The operating system 6021 includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application program 6022 includes various applications, such as a media player and a browser, used to implement various application functions. The program implementing the method of this embodiment can be included in the application program 6022.

[0111] In this embodiment of the invention, by calling the program or instructions stored in the memory 602, specifically the program or instructions stored in the application program 6022, the processor 601 executes the method steps provided in each method embodiment, including, for example: Obtain network topology information in the target network, the network topology information including the connection relationship between each switch in the target network, and the access port information between the graphics processor in the target network and the connected switch; Based on the network topology information, a set of detection links for active detection is generated. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. Generate a detection task configuration based on the set of detection links; The detection task configuration is sent to the corresponding switch to perform active detection through the switch.

[0112] The methods disclosed in the above embodiments of the present invention can be applied to processor 601, or implemented by processor 601. Processor 601 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 601 or by instructions in the form of software. The processor 601 may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present invention can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software units in the decoding processor. The software units may be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory 602. Processor 601 reads the information in memory 602 and, in conjunction with its hardware, completes the steps of the above method.

[0113] It is understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described herein, or combinations thereof.

[0114] For software implementation, the techniques described herein can be implemented by units that perform the functions described herein. The software code can be stored in memory and executed by a processor. The memory can be implemented in the processor or external to the processor.

[0115] The computer device provided in this embodiment may be as follows: Figure 6 The device shown can perform, for example Figure 1 All steps of the network detection method are then implemented to achieve... Figure 1 For details on the technical effectiveness of the network detection method shown, please refer to [link / reference]. Figure 1 The relevant descriptions are presented concisely and will not be elaborated upon here.

[0116] This invention also provides a storage medium (computer-readable storage medium). This storage medium stores one or more programs. The storage medium may include volatile memory, such as random access memory; it may also include non-volatile memory, such as read-only memory, flash memory, hard disk, or solid-state drive; and it may also include combinations of the above types of memory.

[0117] One or more programs in the storage medium can be executed by one or more processors to implement the network probing method described above that is executed on the device side.

[0118] The processor is used to execute a network probing program stored in memory to implement the following steps of a network probing method executed on the device side: Obtain network topology information in the target network, the network topology information including the connection relationship between each switch in the target network, and the access port information between the graphics processor in the target network and the connected switch; Based on the network topology information, a set of detection links for active detection is generated. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. Generate a detection task configuration based on the set of detection links; The detection task configuration is sent to the corresponding switch to perform active detection through the switch.

[0119] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0120] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented in hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0121] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A network detection method, characterized in that, include: Obtain network topology information in the target network, the network topology information including the connection relationship between each switch in the target network, and the access port information between the graphics processor in the target network and the connected switch; Based on the network topology information, a set of detection links for active detection is generated. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. Generate a detection task configuration based on the set of detection links; The detection task configuration is sent to the corresponding switch to perform active detection through the switch.

2. The method according to claim 1, characterized in that, The acquisition of network topology information in the target network includes: Obtain the connection relationships between switches in the target network as the first topology information; Obtain information about the switches and access ports to which each graphics processor in the target network is connected; A second topology information is generated based on the correspondence between each graphics processor and the access switch, as well as the access port information; Based on the first topology information and the second topology information, the network topology information is constructed, and the network topology information is used to determine the correspondence between the detection start point and the detection end point in the detection link set.

3. The method according to claim 1, characterized in that, The generation of a set of probe links for active probing based on the network topology information includes: Based on the network topology information, the connection paths between switches at different detection levels are determined, as well as the connection path between the switches and the graphics processor. Based on the connection path, each detection start point and corresponding detection end point are determined, so as to generate multiple detection links based on the detection start point and detection end point, thus obtaining the detection link set.

4. The method according to claim 1, characterized in that, The step of generating the detection task configuration based on the detection link set includes: For each detection link in the detection link set, obtain the detection start point, detection end point, and detection level of the detection link; Obtain the set of detection parameters corresponding to the detection level, wherein the set of detection parameters includes at least one of the following: transmission interval parameter, timeout parameter, hop count parameter, and message parameter; Based on the detection start point, the detection end point, and the detection parameter set, a detection task configuration corresponding to the detection link is generated. The detection task configuration is used to instruct the switch to perform active detection of the corresponding detection link.

5. The method according to claim 1, characterized in that, The step of sending the detection task configuration to the corresponding switch to perform active detection through the switch includes: Determine the target switch corresponding to the detection task configuration, and send the detection task configuration to the switch; If the current detection level is level zero, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor connected to the target switch as the detection ending point. The virtual interface is a gateway used to access the graphics processor. If the current detection level is the first level, then the loopback interface of the target switch is used as the detection start point and the loopback interfaces of other switches are used as the detection end point to perform active detection. The loopback interface represents the detection endpoint when active detection is performed between switches. If the current detection level is the second level, then the virtual interface of the target switch is used as the detection starting point and the virtual interfaces of other switches are used as the detection ending point to perform active detection; If the current detection level is the third level, then active detection is performed with the virtual interface of the target switch as the detection starting point and the graphics processor deployed across the switch as the detection endpoint.

6. The method according to claim 1, characterized in that, After generating the set of probe links for active probing based on the network topology information, the method further includes: Obtain the detection level and number of links corresponding to each detection link in the detection link set; The sampling ratio corresponding to each detection level is determined based on the number of links; Based on the sampling ratio, a portion of the detection links are selected from the set of detection links as sampling detection links; Based on the sampling and detection link, a corresponding sampling and detection task configuration is generated, and the sampling and detection task configuration is sent to the corresponding switch so that sampling and detection can be performed through the switch; Based on the execution results of the sampling and detection, update the detection status of the detection links in the detection link set that did not participate in the sampling and detection.

7. The method according to claim 5, characterized in that, If the current detection level is the third level, the method further includes: The switches in the target network are divided into multiple switch clusters, and the switches in each switch cluster are numbered according to the same numbering rule to obtain switch numbers. For each switch, obtain the set of connected graphics processors, and obtain the first switch number in the current switch cluster for the first switch currently connected to each graphics processor; Each graphics processor is associated with a second switch in another switch cluster, and the second switch has the same number as the first switch. Active probing is performed using the virtual interface of the second switch as the starting point and the associated graphics processor as the ending point. Alternatively, each graphics processor can be associated with a third switch in the current switch cluster, with the third switch number being different from the first switch number; Active probing is performed using the virtual interface of the third switch as the starting point and the associated graphics processor as the ending point.

8. A network detection device, characterized in that, include: The acquisition module is used to acquire network topology information in the target network. The network topology information includes the connection relationship between each switch in the target network, as well as the access port information between the graphics processor in the target network and the connected switch. The first generation module is used to generate a set of detection links for active detection based on the network topology information. Each detection link represents the correspondence between the detection start point and the detection end point in an active detection. The detection start point is a switch, and the detection end point is a switch or a graphics processor. The second generation module is used to generate a detection task configuration based on the detection link set; The sending module is used to send the detection task configuration to the corresponding switch so that active detection can be performed through the switch.

9. A computer device, characterized in that, include: A processor and a memory, the processor being configured to execute a network probing program stored in the memory to implement the network probing method according to any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores one or more programs, which can be executed by one or more processors to implement the network detection method according to any one of claims 1 to 7.