A network anomaly perception method, device and related equipment

By acquiring the GPU card number information of Leaf devices and servers through the controller and using a dual-view method to determine network anomalies, the problem of GPU server VLAN access not conforming to the same track communication was solved, and rapid fault location and efficient operation and maintenance were achieved.

CN118694679BActive Publication Date: 2026-01-06NEW H3C TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410833471.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-01-06
Estimated Expiration
2044-06-25

AI Technical Summary

Technical Problem

In AI computing, the VLAN access of GPU servers is not networked according to the same-track communication requirements, which leads to an increase in cross-Leaf Layer 3 traffic, which cannot achieve maximum performance and may even cause AI computing tasks to take longer or be interrupted.

Method used

The controller obtains the GPU card number information of each Leaf device and server to determine whether the network requirements are met. A network anomaly detection method with both Leaf and Server perspectives is used to determine whether the network is abnormal.

Benefits of technology

Quickly locate fault points, improve operational efficiency, ensure GPU computing power reaches maximum performance, and avoid delays or interruptions in AI computing tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118694679B_ABST
    Figure CN118694679B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent computing centers, in particular to a network anomaly perception method and device and related equipment. The method comprises the following steps: acquiring GPU card number information corresponding to N server network cards accessed by each Leaf downward interface, wherein the downward interface of one Leaf belongs to the same VLAN, and each server network card directly connected with the downward interface of the Leaf belongs to the same network segment; acquiring network segment information to which M blocks of network cards of each server belong and GPU card number information corresponding to each network card; judging whether the GPU card number corresponding to the N server network cards accessed by each Leaf downward interface meets networking requirements, obtaining a judgment result in the Leaf dimension; judging whether the network segment to which the M blocks of network cards of each server belong and the GPU card number corresponding to each network card meet the networking requirements, obtaining a judgment result in the server dimension; and determining whether the networking is abnormal based on the judgment result in the Leaf dimension and the judgment result in the server dimension.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent computing center technology, and in particular to a method, apparatus and related equipment for network anomaly detection. Background Technology

[0002] In today's rapidly evolving AI (Artificial Intelligence) computing landscape, users are choosing to build intelligent, lossless data centers using GPU (Graphics Processing Unit) servers coupled with high-bandwidth, low-latency RoCE (RDMA over Converged Ethernet) networks. GPU servers typically have multiple network interface cards (NICs) and multiple GPU cards, connected via Leaf devices. When connecting GPU servers to VLANs (Virtual Local Area Networks), they must adhere to the same-track communication requirements for networking and communication. Failure to do so will result in unnecessary cross-Leaf Layer 3 traffic, preventing the GPU from reaching its maximum performance potential, leading to longer AI computing task times, or even task interruptions. Summary of the Invention

[0003] This application provides a method, apparatus, and related equipment for network anomaly detection.

[0004] In a first aspect, this application provides a network anomaly detection method applied to a controller, the method comprising:

[0005] Obtain the GPU card number information corresponding to the N server network cards connected to the downlink interface of each Leaf device in the network. Among them, the downlink port of a Leaf belongs to the same VLAN, and the server network cards directly connected to the downlink port of the Leaf belong to the same network segment.

[0006] Obtain the network segment information of each of the M network cards of each server included in the network, and the GPU card number information corresponding to each network card;

[0007] Determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first group of networks, and obtain the judgment result of the Leaf dimension;

[0008] Determine whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second group of networks, and obtain the judgment result at the server level;

[0009] Based on the judgment results of the Leaf dimension and the server dimension, it is determined whether the network topology is abnormal.

[0010] Optionally, the step of determining whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first network group, and obtaining the determination result at the Leaf dimension, includes:

[0011] For each Leaf device, determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device match the preset GPU numbering rules;

[0012] If a match is found, the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device are determined to meet the first networking requirements; otherwise, the first networking requirements are determined not to be met.

[0013] Optionally, the steps for determining whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second network group, and obtaining the determination result at the server level, include:

[0014] For each server, determine whether the network segment to which the server's M network cards belong matches the preset network segment allocation rules. If they match, determine that the network segment to which the server's M network cards belong meets the second network requirements; otherwise, determine that it does not meet the second network requirements.

[0015] Determine whether the GPU card number of the server matches the preset GPU numbering rule. If it matches, determine that the GPU card number of the server meets the second network requirement; otherwise, determine that it does not meet the second network requirement.

[0016] Optionally, the preset GPU numbering rules include: the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device are the same, and the GPU card numbers of each server are incremented.

[0017] The preset network segment allocation rules include: each network card belongs to a different network segment, and the network segment to which each network card belongs corresponds one-to-one with the network segments included in the DRMA group composed of each Leaf device.

[0018] Optionally, the step of determining whether the network topology is abnormal based on the judgment results at the Leaf dimension and the server dimension includes:

[0019] If it is determined that the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device all meet the preset GPU numbering rules, the network segments to which the M network cards of each server belong all meet the preset network segment allocation rules, and the GPU card numbers of each server all meet the preset GPU numbering rules, then the network configuration is determined to be normal; otherwise, the network configuration is determined to be abnormal.

[0020] Secondly, this application provides a network anomaly detection device applied to a controller, the device comprising:

[0021] The acquisition unit is used to acquire the GPU card number information corresponding to the N server network cards connected to the downlink interface of each Leaf device in the network. Among them, the downlink port of a Leaf belongs to the same VLAN, and the server network cards directly connected to the downlink port of the Leaf belong to the same network segment.

[0022] The acquisition unit is further configured to acquire, respectively, the network segment information to which the M network cards of each server in the network belong, and the GPU card number information corresponding to each network card;

[0023] The judgment unit is used to determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first group of networks, and to obtain the judgment result of the Leaf dimension;

[0024] The judgment unit is also used to determine whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second group of networks, and to obtain the judgment result at the server level.

[0025] The determining unit is used to determine whether the network is abnormal based on the judgment results of the Leaf dimension and the judgment results of the server dimension.

[0026] Optionally, when determining whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first network group, and obtaining the determination result at the Leaf dimension, the determination unit is specifically used for:

[0027] For each Leaf device, determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device match the preset GPU numbering rules;

[0028] If a match is found, the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device are determined to meet the first networking requirements; otherwise, the first networking requirements are determined not to be met.

[0029] Optionally, when determining whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second network group, and obtaining the determination result at the server level, the determination unit is specifically used for:

[0030] For each server, determine whether the network segment to which the server's M network cards belong matches the preset network segment allocation rules. If they match, determine that the network segment to which the server's M network cards belong meets the second network requirements; otherwise, determine that it does not meet the second network requirements.

[0031] Determine whether the GPU card number of the server matches the preset GPU numbering rule. If it matches, determine that the GPU card number of the server meets the second network requirement; otherwise, determine that it does not meet the second network requirement.

[0032] Optionally, the preset GPU numbering rules include: the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device are the same, and the GPU card numbers of each server are incremented.

[0033] The preset network segment allocation rules include: each network card belongs to a different network segment, and the network segment to which each network card belongs corresponds one-to-one with the network segments included in the DRMA group composed of each Leaf device.

[0034] Optionally, when determining whether the network is abnormal based on the judgment results at the Leaf dimension and the server dimension, the determining unit is specifically used for:

[0035] If it is determined that the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device all meet the preset GPU numbering rules, the network segments to which the M network cards of each server belong all meet the preset network segment allocation rules, and the GPU card numbers of each server all meet the preset GPU numbering rules, then the network configuration is determined to be normal; otherwise, the network configuration is determined to be abnormal.

[0036] Thirdly, embodiments of this application provide a network anomaly detection device, which includes:

[0037] Memory, used to store program instructions;

[0038] A processor is configured to invoke program instructions stored in the memory and execute the steps of the method as described in any one of the first aspects above, according to the obtained program instructions.

[0039] Fourthly, embodiments of this application also provide a computer-readable storage medium storing computer-executable instructions for causing a computer to perform the steps of the method as described in any of the first aspects above.

[0040] In summary, the network anomaly detection method provided in this application embodiment, applied to a controller, includes: acquiring GPU card number information corresponding to N server network cards connected to the downlink interfaces of each Leaf device in the network, wherein the downlink port of a Leaf belongs to the same VLAN, and the server network cards directly connected to the downlink port of the Leaf belong to the same network segment; acquiring the network segment information to which M network cards of each server belong, and the GPU card number information corresponding to each network card; determining whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the first network requirements, and obtaining a Leaf-level judgment result; determining whether the network segment to which the M network cards of each server belong and the GPU card numbers corresponding to each network card meet the second network requirements, and obtaining a server-level judgment result; and determining whether the network is abnormal based on the Leaf-level judgment result and the server-level judgment result.

[0041] Using the network anomaly detection method provided in this application, the controller comprehensively displays all possible fault points in the system through a dual perspective of leaf / server, plus the internal perspective of the server. At the same time, it helps users quickly locate fault points by displaying a topology summary. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments of this application or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings of the embodiments of this application.

[0043] Figure 1 A detailed flowchart of a network anomaly detection method provided in an embodiment of this application;

[0044] Figure 2 This is a schematic diagram of a network topology for a parallel track network.

[0045] Figure 3 This is a schematic diagram of a network anomaly detection method provided in an embodiment of this application;

[0046] Figure 4 This is a schematic diagram of the structure of a network anomaly detection device provided in an embodiment of this application;

[0047] Figure 5 This is a schematic diagram of the hardware architecture of a network anomaly detection device provided in an embodiment of this application. Detailed Implementation

[0048] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to limit the application. The singular forms “a,” “the,” and “the” as used in this application and claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.

[0049] It should be understood that although the terms first, second, third, etc., may be used to describe various information in embodiments of this application, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" may also be interpreted as "when," "when," or "in response to a determination."

[0050] Intelligent computing centers are often characterized by large scale and rapid expansion. Quickly locating the point of failure among thousands or even tens of thousands of network cards and GPU cards is one of the challenges in the industry. Moreover, the faults vary, and it is easy to miss some faults if the analysis is only from the perspective of network equipment or servers.

[0051] The network anomaly detection solution provided in this application displays the entire network information from multiple angles through a two-way perspective, from the Leaf and Server (server, host) viewpoints. It also checks the connection relationships of the GPUs inside the server, inferring all possible faults across the entire network. While refining the possible faults, the topology summary display helps users quickly locate the abnormal devices and servers, which can greatly improve operation and maintenance efficiency.

[0052] For example, see Figure 1 The diagram shown is a detailed flowchart of a network anomaly detection method provided in an embodiment of this application. This method is applied to a controller and includes the following steps:

[0053] Step 100: Obtain the GPU card number information corresponding to the N server network cards connected to the downlink interfaces of each Leaf device in the network.

[0054] Among them, the downlink port of one Leaf belongs to the same VLAN, and the network cards of each server directly connected to the downlink port of this Leaf belong to the same network segment.

[0055] In practical applications, refer to Figure 2The diagram shows a network topology diagram for a parallel network: From the server's perspective, assume the server has a total of 6 network cards, including 4 parameter network cards and 2 storage network cards (not shown in the diagram). The 4 network cards belong to 4 different network segments and are connected to 4 Leaf devices respectively. Each of the 4 network cards of the server occupies one interface of the 4 Leaf devices (occupying a total of 4 interfaces).

[0056] From Leaf's perspective, assuming the first Leaf device has four interfaces for connecting to the server's parameter network cards, and each interface connects to the first parameter network card of each of the four servers, then these four network cards belong to the same network segment. The first network cards of the four servers communicate with each other via Layer 2, which is defined as belonging to the same communication track. Similarly, the second network cards of the four servers are connected to the interfaces of the second Leaf device used to connect to the server's parameter network cards, and the second network cards of the four servers communicate with each other via Layer 2, which is defined as belonging to the same communication track; ..., then there are four communication tracks.

[0057] In this embodiment of the application, on the controller side, with each Leaf device as the execution granularity, the controller can obtain the GPU card number information corresponding to the N server network cards connected to the downlink interface of each Leaf device in real time or from the local database.

[0058] For example, obtain the GPU IDs corresponding to the network cards of servers 1-4 connected to the Leaf-1 downlink interface; obtain the GPU IDs corresponding to the network cards of servers 1-4 connected to the Leaf-2 downlink interface; obtain the GPU IDs corresponding to the network cards of servers 1-4 connected to the Leaf-3 downlink interface; obtain the GPU IDs corresponding to the network cards of servers 1-4 connected to the Leaf-4 downlink interface.

[0059] Step 110: Obtain the network segment information of each of the M network cards of each server included in the network, and the GPU card number information corresponding to each network card.

[0060] In this embodiment of the application, on the controller side, with each server as the execution granularity, the controller can obtain the network segment information to which each network card (parameter network card) of each server belongs and the GPU number information corresponding to each network card in real time / from the local database.

[0061] For example, obtain the network segment information to which the four network cards of server 1 belong, and the GPU card number information corresponding to each of the four network cards.

[0062] Step 120: Determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first group of networks, and obtain the judgment result of the Leaf dimension.

[0063] Specifically, in this embodiment of the application, when determining whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first network group, and obtaining the determination result at the Leaf dimension, a preferred implementation method is as follows:

[0064] For each Leaf device, determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device match the preset GPU numbering rules;

[0065] If a match is found, the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device are determined to meet the first networking requirements; otherwise, the first networking requirements are determined not to be met.

[0066] In other words, from the Leaf device perspective, it is determined whether the GPU number corresponding to each server network card connected to the downlink interface of each Leaf device meets the requirements of the same-track network. If a Leaf device meets the requirements, it is determined that the GPU cards corresponding to the N server network cards connected to the downlink interface of that Leaf device are in normal status; if a Leaf device does not meet the requirements, it is determined that there is an abnormal situation among the GPU cards corresponding to the N server network cards connected to the downlink interface of that Leaf device.

[0067] Specifically, in this embodiment, the preset GPU numbering rule includes: the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device are the same. In other words, it determines whether the GPU card numbers connected to the downlink interface of each Leaf device are the same.

[0068] Step 130: Determine whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second group of networks, and obtain the judgment result at the server level.

[0069] In this embodiment of the application, when determining whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the second network requirements, and obtaining the server-level determination result, a preferred implementation is as follows:

[0070] For each server, determine whether the network segment to which the server's M network cards belong matches the preset network segment allocation rules. If they match, determine that the network segment to which the server's M network cards belong meets the second network requirements; otherwise, determine that it does not meet the second network requirements.

[0071] Determine whether the GPU card number of the server matches the preset GPU numbering rule. If it matches, determine that the GPU card number of the server meets the second network requirement; otherwise, determine that it does not meet the second network requirement.

[0072] In other words, from the server perspective,

[0073] Specifically, in this embodiment of the application, the preset GPU numbering rules include: the GPU card numbers of each server are incremented; the preset network segment allocation rules include: each network card belongs to a different network segment, and the network segment to which each network card belongs corresponds one-to-one with the network segments included in the DRMA group composed of each Leaf device.

[0074] Step 140: Based on the judgment results of the Leaf dimension and the server dimension, determine whether the network topology is abnormal.

[0075] In this embodiment of the application, when determining whether the network is abnormal based on the judgment results of the Leaf dimension and the judgment results of the server dimension, a preferred implementation method is as follows:

[0076] If it is determined that the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device all meet the preset GPU numbering rules, the network segments to which the M network cards of each server belong all meet the preset network segment allocation rules, and the GPU card numbers of each server all meet the preset GPU numbering rules, then the network configuration is determined to be normal; otherwise, the network configuration is determined to be abnormal.

[0077] The network anomaly detection process provided in this application embodiment will be described in detail below with reference to specific application scenarios. For example, see [link to relevant documentation]. Figure 3 The diagram shown is a logical schematic of a network anomaly detection method provided in this application embodiment. At the Leaf level, information on N servers accessing the same VLAN (same network segment) is obtained, along with information on the N GPU cards corresponding to the N network cards of the N servers. It is determined whether all GPU card numbers accessed in the same VLAN are identical. If so, it is determined that all servers within the DRMA group are accessing normally at the Leaf level. Otherwise, if the GPU numbers corresponding to M network cards are identical, it is determined that (NM) servers have connected to the wrong device (e.g., connected to Leaf devices in other DRMA groups).

[0078] At the Server level, the X network cards configured on the server correspond to the X GPU cards. It is determined whether each GPU card has a specified number, whether the GPU card numbers are sequential, and whether the network cards of a single Server belong to different network segments. Furthermore, it is verified that the network segments of each network card correspond one-to-one with the network segments of all Leaf devices included in the DRMA group. If a GPU card is found to be without a specified number, or if the GPU card numbers are not sequential, then an unnumbered GPU card or a GPU card without a network card connection is identified, indicating a network anomaly. If two network cards of a server belong to the same network segment, then a network card of that server is connected to the wrong VLAN (interface). If the network segments of a server's network cards do not correspond one-to-one with the network segments of the Leaf devices in the DRMA group, then a server is connected to the wrong Leaf. It should be noted that if all the above determinations are true, then at the Server level, all servers within the DRMA group are considered to have normal access.

[0079] As can be seen from the above, the network is considered to be without abnormalities only when it is confirmed from both the Leaf and Server dimensions that all servers within the DRMAgroup are connected normally.

[0080] For example, see Figure 4 The diagram shown is a structural schematic of a network anomaly detection device provided in an embodiment of this application. This device is applied to a controller and includes:

[0081] The acquisition unit 40 is used to acquire the GPU card number information corresponding to the N server network cards connected to the downlink interface of each Leaf device in the network. Among them, the downlink port of a Leaf belongs to the same VLAN, and the server network cards directly connected to the downlink port of the Leaf belong to the same network segment.

[0082] The acquisition unit 40 is further configured to acquire, respectively, the network segment information to which the M network cards of each server in the network belong, and the GPU card number information corresponding to each network card;

[0083] The judgment unit 41 is used to determine whether the GPU card number corresponding to the N server network cards connected to the downlink interface of each Leaf device meets the requirements of the first group of networks, and to obtain the judgment result of the Leaf dimension.

[0084] The judgment unit 41 is further used to determine whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second group of networks, and to obtain the judgment result at the server level.

[0085] The determining unit 42 is used to determine whether the network is abnormal based on the judgment results of the Leaf dimension and the judgment results of the server dimension.

[0086] Optionally, when determining whether the GPU card numbers corresponding to the N server network cards connected to the downlink interfaces of each Leaf device meet the requirements of the first network group, and obtaining the determination result at the Leaf dimension, the determination unit 41 is specifically used for:

[0087] For each Leaf device, determine whether the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device match the preset GPU numbering rules;

[0088] If a match is found, the GPU card numbers corresponding to the N server network cards connected to the downlink interface of the Leaf device are determined to meet the first networking requirements; otherwise, the first networking requirements are determined not to be met.

[0089] Optionally, when determining whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the requirements of the second network group, and obtaining the server-level determination result, the determination unit 41 is specifically used for:

[0090] For each server, determine whether the network segment to which the server's M network cards belong matches the preset network segment allocation rules. If they match, determine that the network segment to which the server's M network cards belong meets the second network requirements; otherwise, determine that it does not meet the second network requirements.

[0091] Determine whether the GPU card number of the server matches the preset GPU numbering rule. If it matches, determine that the GPU card number of the server meets the second network requirement; otherwise, determine that it does not meet the second network requirement.

[0092] Optionally, the preset GPU numbering rules include: the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device are the same, and the GPU card numbers of each server are incremented.

[0093] The preset network segment allocation rules include: each network card belongs to a different network segment, and the network segment to which each network card belongs corresponds one-to-one with the network segments included in the DRMA group composed of each Leaf device.

[0094] Optionally, when determining whether the network is abnormal based on the judgment results at the Leaf dimension and the server dimension, the determining unit 42 is specifically used for:

[0095] If it is determined that the GPU card numbers corresponding to the N server network cards connected to the downlink interface of each Leaf device all meet the preset GPU numbering rules, the network segments to which the M network cards of each server belong all meet the preset network segment allocation rules, and the GPU card numbers of each server all meet the preset GPU numbering rules, then the network configuration is determined to be normal; otherwise, the network configuration is determined to be abnormal.

[0096] These units can be one or more integrated circuits configured to implement the above methods, such as one or more Application Specific Integrated Circuits (ASICs), one or more digital signal processors (DSPs), or one or more Field Programmable Gate Arrays (FPGAs). Alternatively, when one of these units is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a Central Processing Unit (CPU) or other processor capable of calling program code. Furthermore, these units can be integrated together to form a system-on-a-chip (SOC).

[0097] Furthermore, regarding the network anomaly detection device provided in this application embodiment, from a hardware perspective, the hardware architecture schematic diagram of the network anomaly detection device can be found in [reference needed]. Figure 5 As shown, the network anomaly detection device may include: a memory 50 and a processor 51.

[0098] The memory 50 is used to store program instructions; the processor 51 calls the program instructions stored in the memory 50 and executes the above method embodiment according to the obtained program instructions. The specific implementation method and technical effect are similar, and will not be described again here.

[0099] Optionally, this application also provides a controller, including at least one processing element (or chip) for performing the above method embodiments.

[0100] Optionally, this application also provides a program product, such as a computer-readable storage medium storing computer-executable instructions for causing the computer to perform the above-described method embodiments.

[0101] Here, a machine-readable storage medium can be any electronic, magnetic, optical, or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, a machine-readable storage medium can be: RAM (Random Access Memory), volatile memory, non-volatile memory, flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.

[0102] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, which can take the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.

[0103] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0104] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0106] Furthermore, these computer program instructions can also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0108] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A network anomaly perception method, characterized by, The method is applied to a controller and comprises the following steps: Respectively acquiring GPU card number information corresponding to N server network cards accessed by downlink interfaces of each Leaf device included in a network, wherein a downlink port of a Leaf device belongs to a same VLAN, and each server network card directly connected to the downlink port of the Leaf device belongs to a same network segment; Respectively acquiring network segment information to which M network cards of each server included in the network belong and GPU card number information corresponding to each network card; Judging whether the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device meets a first network requirement, to obtain a Leaf dimension judgment result; Judging whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet a second network requirement, to obtain a server dimension judgment result; Based on the Leaf dimension judgment result and the server dimension judgment result, determining whether the network is abnormal.

2. The method of claim 1, wherein, The step of judging whether the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device meets the first network requirement, to obtain the Leaf dimension judgment result, comprises the following steps: For each Leaf device, judging whether the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of the Leaf device matches a preset GPU number rule; If the match is true, it is determined that the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of the Leaf device meets the first network requirement; otherwise, it is determined that the first network requirement is not met.

3. The method of claim 1 or 2, wherein, The step of judging whether the network segment to which the M network cards of each server belong and the GPU card number corresponding to each network card meet the second network requirement, to obtain the server dimension judgment result, comprises the following steps: For each server, judging whether the network segment to which the M network cards of the server belong matches a preset network segment allocation rule, if the match is true, it is determined that the network segment to which the M network cards of the server belong meets the second network requirement; otherwise, it is determined that the second network requirement is not met; Judging whether the GPU card number of the server matches the preset GPU number rule, if the match is true, it is determined that the GPU card number of the server meets the second network requirement; otherwise, it is determined that the second network requirement is not met.

4. The method of claim 3, wherein, The preset GPU number rule comprises that the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device is the same, and each GPU card number of each server is incremental; The preset network segment allocation rule comprises that each network card belongs to different network segments, and the network segment to which each network card belongs is in one-to-one correspondence with each network segment included in a DRMA group composed of each Leaf device.

5. The method of claim 3, wherein, The step of determining whether the network is abnormal based on the Leaf dimension judgment result and the server dimension judgment result comprises the following steps: If it is determined that the GPU card number corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device meets the preset GPU number rule, the network segment to which the M network cards of each server belong meets the preset network segment allocation rule, and the GPU card number of each server meets the preset GPU number rule, it is determined that the network is normal; otherwise, it is determined that the network is abnormal.

6. A network anomaly perception apparatus characterized by comprising: The device is applied to a controller and comprises the following components: The acquisition unit is configured to acquire GPU card number information corresponding to N server network cards accessed by downlink interfaces of each Leaf device included in the networking, respectively, wherein the downlink interface of one Leaf belongs to the same VLAN, and each server network card directly connected to the downlink interface of the Leaf belongs to the same network segment; The acquisition unit is further configured to acquire network segment information to which M blocks of network cards of each server included in the networking belong and GPU card number information corresponding to each network card, respectively; The judgment unit is configured to judge whether the GPU card numbers corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device satisfy a first networking requirement, to obtain a judgment result in the Leaf dimension; The judgment unit is further configured to judge whether the network segments to which the M blocks of network cards of each server belong and the GPU card numbers corresponding to each network card satisfy a second networking requirement, to obtain a judgment result in the server dimension; The determination unit is configured to determine whether the networking is abnormal based on the judgment result in the Leaf dimension and the judgment result in the server dimension.

7. The apparatus of claim 6, wherein, When judging whether the GPU card numbers corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device satisfy the first networking requirement to obtain the judgment result in the Leaf dimension, the judgment unit is specifically configured to: For each Leaf device, judge whether the GPU card numbers corresponding to the N server network cards accessed by the downlink interface of the Leaf device match a preset GPU number rule; If the match is true, it is determined that the GPU card numbers corresponding to the N server network cards accessed by the downlink interface of the Leaf device satisfy the first networking requirement; otherwise, it is determined that the first networking requirement is not satisfied.

8. The apparatus of claim 6 or 7, wherein, When judging whether the network segments to which the M blocks of network cards of each server belong and the GPU card numbers corresponding to each network card satisfy the second networking requirement to obtain the judgment result in the server dimension, the judgment unit is specifically configured to: For each server, judge whether the network segment to which the M blocks of network cards of the server belong matches a preset network segment allocation rule, if the match is true, it is determined that the network segment to which the M blocks of network cards of the server belong satisfies the second networking requirement; otherwise, it is determined that the second networking requirement is not satisfied; Judge whether the GPU card number of the server matches the preset GPU number rule, if the match is true, it is determined that the GPU card number of the server satisfies the second networking requirement; otherwise, it is determined that the second networking requirement is not satisfied.

9. The apparatus of claim 8, wherein, The preset GPU number rule includes that the GPU card numbers corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device are the same, and the GPU card numbers of each server are incremental; The preset network segment allocation rule includes that each network card belongs to different network segments, and the network segments to which each network card belongs are one-to-one corresponding to each network segment included in a DRMA group composed of each Leaf device.

10. The apparatus of claim 8, wherein, When determining whether the networking is abnormal based on the judgment result in the Leaf dimension and the judgment result in the server dimension, the determination unit is specifically configured to: If it is determined that the GPU card numbers corresponding to the N server network cards accessed by the downlink interfaces of each Leaf device satisfy the preset GPU number rule, the network segments to which the M network cards of each server belong satisfy the preset network segment allocation rule, and the GPU card numbers of each server satisfy the preset GPU number rule, it is determined that the networking is normal; otherwise, it is determined that the networking is abnormal.

11. A network anomaly perception apparatus characterized by comprising: The network anomaly perception device comprises: a memory for storing program instructions; a processor for calling the program instructions stored in the memory and executing the steps of the method according to any one of claims 1-5.

12. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer executable instructions for causing the computer to execute the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Task execution method and device, storage medium and electronic equipment

    CN115509749A

  • Network topology operation and maintenance method and device and related equipment

    CN118055028A