Network topology management method, device and system and storage medium
By sending topology discovery request messages and analyzing response messages in the southbound network of the artificial intelligence device cluster, the complex problem of southbound network management is solved, and efficient management and maintenance of southbound networks is achieved.
Patent Information
- Application Number
- CN202510444558.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The southbound network management in the field of artificial intelligence device clusters is complex and difficult to maintain, especially as the cluster size continues to expand.
After the first device receives the start topology discovery instruction, it sends a topology discovery request message to the switch, broadcasts to the second device, and receives and parses the topology discovery response message, establishes a network connection to obtain network topology information.
It realizes efficient management and maintenance of the southbound network, improves the efficiency and accuracy of network topology management, and ensures the integrity of network topology information and network connection compliance.
Smart Images

Figure CN119996220A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of network management, and in particular to a network topology management method, device, system and storage medium. Background Art
[0002] In the field of artificial intelligence device clusters, major manufacturers are committed to developing their own southbound networks. Given the increasing requirements for data processing speed and efficiency, the design of southbound networks has become particularly critical. In order to achieve the goal of high-bandwidth, low-latency data transmission, various manufacturers have adopted a variety of technical routes when developing southbound networks, including standard Ethernet, RDMA (Remote Direct Memory Access) networks, and bus-level protocols. Although this diversity has enriched the implementation of southbound networks, it has also greatly increased the complexity of their management. As the scale of clusters continues to expand, how to efficiently monitor and maintain the southbound network has become a major challenge that needs to be solved urgently. Summary of the invention
[0003] The present invention provides a network topology management method, device, system and storage medium to solve the defects of complex and difficult maintenance of southbound network management in the field of artificial intelligence device clusters in related technologies.
[0004] The present invention provides a network topology management method, which is applied to a first device and includes: In the case of receiving a start topology discovery instruction, sending a topology discovery request message to the switch, so that the switch broadcasts the topology discovery request message to a second device, where the second device is a device other than the first device connected to the switch in a southbound network, the southbound network comprising a plurality of server nodes and a plurality of switches, each server node being provided with at least one device, and each device being connected to at least one switch; receiving a topology discovery response message forwarded by the switch, where the topology discovery response message is replied to the switch by the second device when the second device passes the check on the topology discovery request message; The topology discovery response message is parsed, and based on the identification information of the second device obtained through the analysis and the local identification information, a network connection is established with the second device to obtain the network topology information of the first device.
[0005] A network topology management method provided by the present invention also includes: Based on the network topology information, sending a topology keep-alive request message to the switch, so that the switch unicasts the topology keep-alive request message to a target device, where the target device is a device in the second device that establishes a network connection with the first device; receiving a topology keep-alive response message forwarded by the switch, and parsing the topology keep-alive response message to obtain a keep-alive status, wherein the topology keep-alive response message is replied to the switch by the target device based on a check result of the topology keep-alive request message; When the keep-alive state is failure, continue to send topology keep-alive request messages until the number of failures exceeds a threshold, and then resend the topology discovery request message.
[0006] According to a network topology management method provided by the present invention, the identification information includes a node identification, a device identification and a port identification.
[0007] The present invention provides a network topology management method, which is applied to a second device and includes: receiving a topology discovery request message broadcasted by a switch, wherein the topology discovery request message is sent by a first device to the switch when receiving an instruction to start topology discovery, wherein the first device is a device other than the second device connected to the switch in a southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, each server node is provided with at least one device, and each device is connected to at least one switch; The topology discovery request message is checked, and if the check passes, a topology discovery response message is replied to the switch, so that the switch forwards the topology discovery response message to the first device, and the first device is used to parse the topology discovery response message, and establish a network connection with the second device based on the identification information of the second device obtained by the analysis and the local identification information, so as to obtain the network topology information of the first device.
[0008] According to a network topology management method provided by the present invention, the checking of the topology discovery request message includes: Parsing the topology discovery request message to obtain a port identifier and a group identifier of a source device; Compare the port identifier of the source device with the local port identifier to obtain a first comparison result, and compare the group identifier of the source device with the local group identifier to obtain a second comparison result; When the first comparison result and the second comparison result are both consistent, it is determined that the check is passed, otherwise an alarm log is generated.
[0009] A network topology management method provided by the present invention also includes: receiving a topology keep-alive request message unicast by the switch, where the topology keep-alive request message is sent by the first device to the switch when a network connection is established between the first device and the second device; The topology keep-alive request message is checked, and based on the check result, a topology keep-alive response message is replied to the switch, so that the switch forwards the topology keep-alive response message to the first device, and the first device is used to parse the topology keep-alive response message to obtain the keep-alive status.
[0010] According to a network topology management method provided by the present invention, the checking of the topology keep-alive request message includes: Parsing the topology keep-alive request message to obtain identification information of a source device, identification information of a target device, and a group identification of the source device; Compare the node identifier in the identification information of the target device with the local node identifier to obtain a first result, compare the port identifier in the identification information of the source device with the local port identifier to obtain a second result, and compare the group identifier of the source device with the local group identifier to obtain a third result; When the first result, the second result, and the third result are all consistent, checking whether the connection information exists locally, and if so, determining that the keep-alive state is successful, otherwise determining that the keep-alive state is failed, the connection information being determined based on the identification information of the source device and the identification information of the target device; When any one of the first result, the second result and the third result is inconsistent, an alarm log is generated.
[0011] The present invention also provides a network topology management device, which is applied to a first device and includes: a sending unit, configured to send a topology discovery request message to the switch upon receiving a topology discovery start instruction, so that the switch broadcasts the topology discovery request message to a second device, wherein the second device is a device other than the first device connected to the switch in a southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, each server node is provided with at least one device, and each device is connected to at least one switch; a receiving unit, configured to receive a topology discovery response message forwarded by the switch, wherein the topology discovery response message is replied to the switch by the second device when the second device passes the check on the topology discovery request message; An establishing unit is used to parse the topology discovery response message, and based on the identification information of the second device obtained by the analysis and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0012] The present invention also provides a network topology management device, which is applied to a second device and includes: a message receiving unit, configured to receive a topology discovery request message broadcast by a switch, wherein the topology discovery request message is sent by a first device to the switch when receiving an instruction to start topology discovery, wherein the first device is a device other than the second device connected to the switch in a southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, each server node is provided with at least one device, and each device is connected to at least one switch; A message reply unit is used to check the topology discovery request message and, if the check passes, reply a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device. The first device is used to parse the topology discovery response message and, based on the identification information of the second device obtained by the analysis and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0013] The present invention also provides a device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, any of the network topology management methods described above is implemented.
[0014] The present invention also provides a network topology management system, comprising a management node, a plurality of server nodes and a plurality of switches, each server node is provided with at least one device as described above, each device comprises at least one port, and each port is connected to a switch; The server node is used to obtain network topology information of each local device and report the network topology information of each device to the management node; The management node is used to obtain information of each switch based on the switch management network, and manage the network topology information of each device and the information of each switch.
[0015] According to a network topology management system provided by the present invention, the management node is also used to assign a node identifier to each server node, and uniformly assign IP addresses to all ports of all devices on all server nodes; The server node is further configured to generate a global device identifier for each device based on the node identifier and the local device identifier of each device.
[0016] According to a network topology management system provided by the present invention, a plurality of management nodes are provided.
[0017] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the network topology management method described in any one of the above is implemented.
[0018] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the network topology management method described above is implemented.
[0019] The network topology management method, device, system and storage medium provided by the present invention can obtain the identification information of the second device by sending a topology discovery request message, receiving a topology discovery response message, parsing the topology discovery response message and other steps, and establish a network connection with the second device accordingly to obtain the network topology information of the first device. Since the southbound network includes multiple server nodes, each server node is provided with at least one device, and each device can be used as both a first device and a second device, according to the above steps, all network topology information of the entire southbound network can be collected, and the management and maintenance of the southbound network can be realized according to these network topology information. In addition, the first device automatically sends a topology discovery request message when receiving a start topology discovery instruction, without manual configuration or checking of the network connection, which helps to improve the efficiency and accuracy of network topology management. By broadcasting the topology discovery request message through the switch, each device connected to the switch in the network can be covered, ensuring the integrity of the network topology information. The topology discovery response message is replied by the second device when the request message is checked and passed, so that the compliance and validity of the network connection between the first device and the second device can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0021] Figure 1 It is a structural schematic diagram of the southbound network provided by the present invention; Figure 2 It is one of the structural schematic diagrams of the network topology management system provided by the present invention; Figure 3 This is the second structural diagram of the network topology management system provided by the present invention; Figure 4 It is one of the flow diagrams of the network topology management method provided by the present invention; Figure 5 This is the second flow chart of the network topology management method provided by the present invention; Figure 6 It is a schematic diagram of topology discovery provided by the present invention; Figure 7 It is a schematic diagram of topology keep-alive provided by the present invention; Figure 8 It is one of the structural schematic diagrams of the network topology management device provided by the present invention; Fig. 9 This is the second structural diagram of the network topology management device provided by the present invention; Fig.10 It is a structural schematic diagram of the device provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] In the field of artificial intelligence device clusters, the southbound network refers to the network formed by connecting the service side ports of the cluster to other devices. Here, the artificial intelligence device can be a GPU (Graphics Processing Unit), a GPGPU (General-purpose computing on Graphics Processing Units), a TPU (Tensor Processing Unit), etc., and the present invention does not specifically limit this. It should be understood that the service side port of the cluster refers to the port used by each device in the cluster for business data transmission. Through these ports, the cluster can form a unified whole to jointly process business requests and data transmission.
[0024] In order to solve the management problem of the southbound network, the present invention fully describes the process of topology discovery and topology maintenance of the southbound network, and formulates a management system to implement the management solution. In order to facilitate the understanding of the technical solution of the present invention, the technical solution of the present invention is mainly introduced below by taking the southbound network of the GPU cluster as an example.
[0025] A GPU cluster usually includes multiple server nodes, each of which is deployed with multiple GPU accelerator cards (hereinafter referred to as "GPU cards"). Each of these GPU cards has multiple ports, and each port is connected to a switch. Figure 1 is a schematic diagram of the structure of the southbound network provided by the present invention, such as Figure 1 As shown, it only takes server node 0 as an example. Server node 0 is deployed with 8 GPU cards, GPU0 to GPU7, and each GPU card is set with 8 ports, each port is connected to one of the 8 switches, switch0 to switch7. Such an architecture allows each GPU card to communicate efficiently with other GPU cards in the cluster through the switch, thereby building a complex network topology, namely the southbound network. It should be noted that Figure 1 The ports on each GPU card connected to the external switches (ie, switch0 to switch7) are the service-side ports of the cluster.
[0026] It should be noted that a server node refers to an independent computing device or processing unit that undertakes a specific task in the network. They are interconnected through the network to jointly complete data processing, storage and transmission tasks. For example, a server node can be a GPU server. Inside the server node, GPU cards can also be interconnected through PCIE (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) buses, switches, etc. ( Figure 1 Not shown in the figure), since the present invention is mainly to achieve the management of the southbound network, the connection network between the GPU cards inside the server node will not be described in detail here.
[0027] An embodiment of the present invention provides a network topology management system, which includes a management node, multiple server nodes and multiple switches. Each server node is provided with at least one device, each device includes at least one port, each port is connected to a switch, and here, these ports on each device constitute the service side port of the cluster. In the network topology management system, each server node is used to obtain the network topology information of each local device and report the network topology information of each device to the management node; the management node is used to obtain the information of each switch based on the switch management network, and manage the network topology information of each device and the information of each switch. It should be understood that topology refers to the physical connection relationship between entities within the guide network, and network topology management refers to the process of configuring, monitoring and maintaining these connection relationships.
[0028] It should be noted that the above-mentioned device refers to an artificial intelligence device, which can be a GPU, GPGPU, TPU, etc. The embodiment of the present invention mainly introduces the GPU as an example. When the device is a GPU, the device refers to a GPU card.
[0029] Figure 2 This is one of the structural diagrams of the network topology management system provided by the present invention, such as Figure 2 As shown, the management server 1 and the management server 2 in the figure are both management nodes, GPU server 0, GPU server 1, ..., GPU server 63 are all server nodes, and switch0, switch1, ..., switch7 are switches. Among them, the management node (i.e., the management server) can access the server management network, and allocate and manage resources for each server node (i.e., each GPU server) through the server management network. At the same time, the management node can also access the switch management network and collect switch information through the switch management network. It should be understood that the management node can be an ordinary GPU server or other server. Regardless of the type of server, the server needs to be able to access the server management network and the switch management network at the same time.
[0030] Specifically, in order to achieve the management of the southbound network, the present invention designs two management software, one as a server role, the software name is fabric-server, deployed on the management node, the management node supports centralized deployment and distributed deployment. When the reliability requirement is high, the distributed deployment method can be used to prevent the occurrence of fabric-server software abnormalities or server hardware abnormalities such as power failure on the server where it is located; the other as a client role, the software name is fabric-client, deployed on the GPU server (i.e., server node). It should be noted that when the management node adopts a distributed deployment method, one of the management servers can be used as the master (i.e., the main management server), and the other management servers can be slaves (i.e., backup management servers).
[0031] Specifically, when the GPU server and the switch cluster are connected through the transmission medium, that is, connected to the southbound network, the user can configure the management IP (Internet Protocol) address of the GPU server, the management IP address of the management server, and the management IP address of the switch. At the same time, the user can configure an IP address for each port of all devices (that is, all GPU cards) in the southbound network according to their own network planning requirements. If the user does not configure it, the management server can also uniformly assign IP addresses to the ports of all devices in the southbound network.
[0032] After completing the corresponding configuration, the fabric-client on each GPU server can send a topology discovery request message in the southbound network to broadcast the GPU information on the GPU server in the southbound network to other fabric-clients that are network-reachable. Here, the specific GPU information includes the GPU MAC (Media Access Control) address, GPU port ID, GPU IP address, etc. When the fabric-client on other GPU servers receives this information, it replies its own GPU information to the requester. After the requester receives it, the requester side can establish a network connection that is network-reachable. For example, GPU1.port5 and GPU9. port 5 can establish a network connection and communicate with each other. Here, GPU1.port 5 refers to port 5 of GPU1, and GPU9. port 5 refers to port 5 of GPU9, where both GPU1 and GPU9 are GPU cards.
[0033] According to the same logic, all ports of all GPU cards in the cluster will send topology discovery request messages, and each GPU card will collect all connection information in the southbound network that is reachable to its own network. At this point, the network topology relationship of each GPU card is established. The fabric-client on each GPU server can report the network topology information of all GPU cards on the server to the fabric-server, so that the fabric-server will have the network topology connection relationship of all fabric-clients, making the connection relationship on the fabric-server global and complete. In addition, the fabric-server will also collect switch information through the switch management network, so that the fabric-server has the entire information of the GPU server and the switch.
[0034] It should be noted that since there are many switch manufacturers and their management methods are different, they cannot provide communication interaction uniformly through business ports or in-band methods. Therefore, for the sake of solution compatibility, the switch management port and GPU server can be used uniformly for communication interaction. In addition, in order to facilitate the management node to access the switch management network and the server management network, the switch management network and the server management network can be directly deployed into a unified management network. However, the switch management network and the server management network are usually not in the same network. The purpose of this setting is to prevent the server and switch from interoperating with each other. Therefore, in this scenario, fabric-server can be deployed on a jump machine that can access both the switch management network and the server management network.
[0035] Based on the above embodiment, the management node is also used to assign a node identifier to each server node, and to uniformly assign IP addresses to all ports of all devices on all server nodes; the server node is also used to generate a global device identifier for each device based on the node identifier and the local device identifier of each device.
[0036] Specifically, the above-mentioned topology discovery request and other protocol messages are all sent in the southbound network and have nothing to do with the management node (i.e., the management server). The management node mainly provides auxiliary functions. For example, the management node can assign node identifiers to each server node (i.e., all GPU servers that assume the client role), uniformly assign IP addresses to all ports of all devices on all server nodes, obtain relevant information about switches, etc. Here, the node identifier can be expressed as a node id, which refers to a unique identifier assigned to each node in the network, used to distinguish and identify different server nodes, that is, the global number of each server node, which is unique.
[0037] After the node identifier is allocated, the server node can generate the global device identifier of each device based on the node identifier and the local device identifier of each device on the node. Here, the local device identifier can be expressed as the local gpu id, which refers to the unique identifier of each device (such as each GPU card) on the GPU server to which it belongs; the global device identifier (which can be expressed as the global gpu id) refers to the unique identifier of each device in the entire device cluster. The following will introduce in detail the deployment process of the management software fabric-server and fabric-client, as well as the process of generating the global device identifier of each device.
[0038] Figure 3 This is the second structural diagram of the network topology management system provided by the present invention. Figure 3 As shown, first, the user decides which servers assume the server role, that is, as management servers. A single server can be used for centralized management, or multiple servers can be used for distributed management. If a single server fails, it will cause abnormalities in the fabric-server network and affect the normal management of the network topology. Therefore, distributed deployment is recommended.
[0039] After determining which servers will assume the server role, users can install the fabric-server installation package on these management servers and start the fabric-server service. Subsequently, users can set the configuration file of the fabric-server service according to the networking requirements. The file specifically includes the following contents: ① The IPV4 address list of the management server (there can be multiple management servers that assume the server role); ② Set the TCP port of the transport layer, the default value is 65330; ③ The IPV4 address list of fabric-client members, that is, the IPV4 address list of all GPU servers managed by this management server.
[0040] After completing the configuration file settings, you need to restart the fabric-server service. During the service restart process, if there are multiple management servers (that is, there are multiple management nodes), that is, a distributed scenario, one is selected from the IPV4 address list of the management server as the fabric-server master node, and the others are fabric-server slave nodes. In other words, one can be selected from multiple management servers as the master (i.e., the main management server), and the others are slaves (i.e., backup management servers). Here, when electing the master, the election can be performed according to the PCB (Printed Circuit Board) serial number or UUID (Universally Unique Identifier) of each management server. After the master node is elected, the master node will subsequently assume tasks such as resource allocation, and submit information such as the results of resource allocation to other slave nodes for backup. If the master node fails, a master node can be re-elected from all slave nodes.
[0041] For example, Figure 3 As shown in the figure, the user has selected two servers as management servers, where management server 1 is the master and management server 2 is the slave. Fabric-server software is deployed on both management server 1 and management server 2. Figure 3 The linux user and linux kernel in the Linux system represent the user space and kernel space respectively. The user space refers to the memory space where ordinary user programs run, and the kernel space refers to the memory space where the operating system kernel runs. Figure 3As shown in the figure, the fabric-server software installation package is deployed in the user space, while the network card driver runs in the kernel space. The fabric-server service can interact with the network card driver.
[0042] After completing the deployment of fabric-server, users can start the fabric-client deployment process. First, users also decide which GPU servers will assume the client role, install the fabric-client installation package on each selected GPU server, and start the fabric-client service. Then, users set the configuration file of each fabric-client service according to the networking requirements. The file specifically includes the following contents: ① The IPV4 address list of all management servers that assume the server role; ② Set the TCP port of the transport layer. If not set, it will default to 65330; ③ Configure the group identifier (i.e. clique id) for each GPU card on this GPU server. If the user does not configure it, it will default to 0. If the user configures it, GPU cards with the same clique id can communicate in the southbound network; ④ The IPV4 address list of the southbound network is configured according to each port of each GPU card. If the user does not configure it, it can be uniformly allocated by the fabric-server master in the management node. Here, each GPU server includes one or more GPU cards, and each GPU card has one or more ports. Here, an IPV4 address needs to be configured for each port.
[0043] After completing the configuration file settings, you need to restart the fabric-client service. During the service restart process, if the user has not configured the IPV4 address of each port of each GPU card, apply to the fabric-servermaster for the IPV4 address of each port of all GPU cards. At the same time, apply to the fabric-server master for the node id (i.e., node identifier) of this GPU server. For each GPU server, after obtaining the node identifier, a global device identifier (i.e., global gpu id) can be generated for each GPU card based on the node identifier and the local device identifier of each GPU (i.e., local gpu id). For example, assuming there are 64 GPU servers, each GPU server has 8 GPU cards, and a total of 64×8=512 GPU cards. Based on the node id of the server where each GPU card is located and the local gpu id of the GPU card, a global unified number can be generated for each GPU card, that is, the global gpu id. Here, the local gpu id can be a unique number for each GPU card on the GPU server to which it belongs, such as Figure 1There are 8 GPU cards on server node 0, and their local gpu ids are 0 to 7. Through the global gpu id, the corresponding GPU card can be located in the entire GPU cluster.
[0044] For example, taking the case that each GPU server reserves a maximum of 8 GPU cards, the rules for generating the global gpu id are as follows: the lower 3 bits are the local gpu id, and the remaining high bits are the node id. For example, if the node id is 1 and the local gpu id is 2, the generated global gpu id is 0xa (hexadecimal representation), and its corresponding binary representation is 1010, and the decimal representation is 10. Taking the case that each GPU server reserves a maximum of 16 GPUs, the rules for the global gpu id are as follows: the lower 4 bits are the local gpu id, and the remaining high bits are the node id.
[0045] For each GPU card, after generating the global gpu id of the GPU card, a corresponding MAC address can be generated for each port on the GPU card. The MAC address generation rule is as follows: 00:00:00:node_id:gpu_id:port_id, where node_id refers to the node identifier, gpu_id refers to the global gpu id of the GPU card, and port_id is the port identifier, which refers to the unique number of each port on the GPU card to which it belongs.
[0046] After generating the global GPU ID and MAC address, you can configure this information into the register of the GPU card. Then, apply for the UUID of the GPU server from the fabric-server master, and the fabric-server master will return a UUID of the GPU server to the fabric-client.
[0047] During the restart of the fabric-client service, the clique id of each GPU card is configured to the manager thread of the GPU kernel driver (i.e. Figure 3The fabric-manager-thread module in the gpu driver shown in Figure 1 is used to carry the clique id in the subsequent topology discovery and topology keepalive. After completing various configurations, the fabric-client service can query the fabric-manager-thread module for the current status. If the current status of the fabric-manager-thread module is ready, the service status is migrated to the active status. If the fabric-manager-thread module is not ready or the service status migration fails, you can view it through the service log command. Finally, after the service status is successfully migrated to the active status, the fabric-manager-thread module can be notified to start the topology discovery process.
[0048] In addition, if Figure 3 As shown, a monitoring tool (i.e., monitor tool) is installed on each GPU server. This tool is user-oriented. Users can use this tool to view the network topology information of all GPU cards on the GPU server for operation and maintenance.
[0049] It should be noted that the network topology management provided by the present invention is mainly implemented through a topology discovery process and a topology preservation process. The specific processes of topology discovery and topology preservation are described in detail below.
[0050] Based on any of the above embodiments, Figure 4 It is one of the flow charts of the network topology management method provided by the present invention, such as Figure 4 As shown, the method is applied to a first device, and the method includes: Step 410, upon receiving an instruction to start topology discovery, sending a topology discovery request message to the switch, so that the switch broadcasts the topology discovery request message to a second device, wherein the second device is a device other than the first device connected to the switch in a southbound network, wherein the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch.
[0051] It should be noted that the execution subject of the method provided in the embodiment of the present invention is the first device, which refers to the device that initiates the topology discovery request in the southbound network. The second device refers to other devices in the southbound network that are connected to the same switch as the first device. For example, assume that the southbound network includes 64 GPU cards, namely GPU0~GPU63, and each GPU card has 8 ports, respectively represented as port0~port7. For each GPU card, its port port0 is connected to the switch switch0, port1 is connected to the switch switch1,..., port7 is connected to the switch switch7. In this case, when GPU0 is the first device, other GPU cards (ie, GPU1~GPU63) are the second devices; when GPU1 is the first device, other GPU cards (ie, GPU0, GPU2~GPU63) are the second devices, and so on. It should be understood that for each GPU card in the southbound network, it is both the first device and the second device.
[0052] Specifically, when the fabric-manager-thread module driven by the GPU kernel receives the start topology discovery command sent by the fabric-client service, the topology detection process starts. The following is an example using GPU0 as the first device.
[0053] First, create a raw socket in the kernel-driven fabric-manager-thread module to trigger each port of GPU0 to periodically call the socket programming interface to broadcast and send a topology discovery request message. Starting a topology discovery instruction refers to a signal or command that triggers the first device (such as GPU0) to start executing the topology discovery process. Here, broadcast refers to a communication method of sending data packets to all devices in the network. In an embodiment of the present invention, broadcast means that the topology discovery request message will be sent to all switches connected to GPU0, and each switch will forward the message to all devices connected to it (i.e., other GPU cards in the southbound network).
[0054] It can be understood that the topology discovery request message is a data packet in a specific format sent by the first device (such as GPU0) to obtain network topology information. Specifically, the message header field of the topology discovery request message may include ETH (DMAC, SMAC, PROTO), IPV4 header, UDP header, TOPO header, Payload, etc. The meaning of each field is shown in Table 1 below: Table 1 Header of topology discovery request message
[0055] Table 2 TOPO header of topology discovery request message As shown in Table 2, in the topology discovery request message, the TOPO message header includes a total of 32 bytes, and currently uses bytes 0-19, of which bytes 0-3 encapsulate the message type opcode, group identifier clique id and status field, and the value corresponding to the opcode can be 0x1, indicating that the message type is a discover message (i.e., a topology discovery request message); bytes 4-7 encapsulate the node identifiers of the source GPU server and the target GPU server; bytes 8-11 encapsulate the device identifiers of the source GPU card and the target GPU card (i.e., the global number of the GPU card, which is unique); bytes 12-15 encapsulate the port numbers on the source GPU card and the target GPU card (i.e., the port identifier, which is unique inside each GPU card); bytes 16-19 are reserved fields. It should be noted that when GPU0 is the first device, the message header TOPO header encapsulates the src node id, src gpu id, and src port id related to GPU0, and the dst related information is all 0.
[0056] For GPU0, each port will send a topology discovery request message, which will be broadcast on the switch connected to each port, and other GPU cards connected to the same switch will receive the topology discovery request message. For example, for port 0 on GPU0, after sending a topology discovery request message, it will be broadcast on switch switch0, and port 0 ports of GPU1~GPU63 connected to switch0 will receive the request message.
[0057] Step 420: Receive a topology discovery response message forwarded by the switch. The topology discovery response message is replied to the switch by the second device when the second device passes the check on the topology discovery request message.
[0058] Specifically, other GPU cards (i.e., second devices) can receive topology discovery request messages through corresponding socket programming interfaces. After receiving the topology discovery request messages, they will perform a series of checks to confirm the legitimacy and validity of the messages. For example, they can check whether the source MAC address carried in the message matches a known legitimate device address to prevent potential network attacks or misoperations.
[0059] For another example, when deploying a southbound network, it is usually hoped that the connections between ports are regular, so that once a problem occurs, it is easy to troubleshoot. For example, the port0 port of GPU0 (whose port id is 0) is connected to the port0 port of other GPU cards, and the port1 port is connected to the port1 port of other GPU cards, and so on. Therefore, after receiving the topology discovery request message, other GPU cards can check whether the src port id carried in the message is consistent with the port id of the local receiving port (that is, check whether the port id of the sending port and the receiving port of the request message are consistent). If they are inconsistent, an alarm log is generated to remind the user to check the network connection. The main purpose of doing this is to hope that all GPU ports with the same port id are connected to the same switching network, which is convenient for wiring and troubleshooting network problems.
[0060] For example, other GPU cards can check whether the clique ID carried in the message is consistent with the clique ID of the GPU card. If not, an alarm log is generated to remind the user to check the network connection and the network configuration of the switch, and the topology discovery request message is discarded.
[0061] If the check of the topology discovery request message passes, that is, all fields in the message are checked to be legal, then the topology discovery response message can be replied according to the protocol requirements. Here, the topology discovery response message is a data packet in a specific format sent to the switch according to the protocol requirements by the second device (i.e., other GPU cards) after receiving and successfully verifying the topology discovery request message. The message header of the topology discovery response message is the same as the message header of the topology discovery request message, but the opcode value encapsulated in the message header TOPO is 0x2, indicating that the message type is discover response (i.e., topology discovery response message), and at the same time encapsulates its own source MAC address, encapsulates its own clique id, writes its own node id, gpu id, port id into the dst node id, dst gpu id, dst port id fields, and calls the socket programming interface to send the topology discovery response message.
[0062] After receiving the topology discovery response message replied by the second device, the switch may forward the response message to the first device, such as GPU0.
[0063] Step 430: parse the topology discovery response message, and establish a network connection with the second device based on the identification information of the second device obtained through the analysis and the local identification information, to obtain the network topology information of the first device.
[0064] Specifically, after receiving the topology discovery response message, GPU0 can parse the message to obtain the identification information of the second device carried in the message (i.e., dst node id, dst gpu id, dst port id) and the status field status. If the status field is success, GPU0 can obtain a complete connection information, i.e., src [nodeid, gpu id, port id] - dst [node id, gpu id, port id]. Here, src [node id, gpu id, port id] is the identification information of the first device (i.e., local identification information), and dst [nodeid, gpu id, port id] is the identification information of the second device. Based on these identification information, the first device can establish a network connection with the second device.
[0065] Specifically, the first device can match the identification information of the second device obtained through parsing with the connection information stored locally. If the match is successful, it indicates that the connection information with the second device already exists, and the topology discovery response message can be directly discarded; if the match fails, the new connection information can be added to the local storage. Based on the matching or newly added connection information, the first device can try to establish a network connection with the second device, that is, use the appropriate network protocol and port number to communicate to ensure the correct transmission and reception of data.
[0066] After the first device stores the connection information locally, it can obtain the network topology information of the first device. Here, the network topology information refers to the information that records the connection relationship between the first device and other devices in the southbound network. For example, for GPU0, its network topology information includes information such as the node identifier, device identifier, and port identifier of each GPU card interconnected with GPU0 through a switch. This information exists in the form of connection information, such as "src [node id, gpu id, portid]- dst [node id, gpu id, port id]", which represents the connection relationship between the source device (ie, the first device) and the destination device (ie, the second device). By constructing such network topology information, it is possible to understand the connection status of the entire southbound network and provide a basis for subsequent network communication and resource management.
[0067] Following the same logic as GPU0, each GPU card in the entire southbound network will obtain its own complete network topology information, which records the information of each GPU card that can be interconnected with it through the switch. After completing the topology discovery process, the entire network topology relationship is established. At this time, the fabric-client deployed on each GPU server can obtain the network topology information of all GPU cards on the server and report this information to the fabric-server for unified management and maintenance.
[0068] According to the method provided by the embodiment of the present invention, the first device can obtain the identification information of the second device by sending a topology discovery request message, receiving a topology discovery response message, parsing the topology discovery response message, etc., and establish a network connection with the second device accordingly to obtain the network topology information of the first device. Since the southbound network includes multiple server nodes, each server node is provided with at least one device, and each device can be used as both a first device and a second device, according to the above steps, all network topology information of the entire southbound network can be collected, and the management and maintenance of the southbound network can be realized according to these network topology information. In addition, the first device automatically sends the topology discovery request message when receiving the instruction to start the topology discovery, without manual configuration or checking of the network connection, which helps to improve the efficiency and accuracy of network topology management. By broadcasting the topology discovery request message through the switch, each device connected to the switch in the network can be covered, ensuring the integrity of the network topology information. The topology discovery response message is replied by the second device when the request message is checked and passed, so that the compliance and validity of the network connection between the first device and the second device can be ensured.
[0069] Based on any of the above embodiments, the method further includes: Step 440: Based on the network topology information, send a topology keep-alive request message to the switch, so that the switch unicasts the topology keep-alive request message to a target device, where the target device is a device in the second device that establishes a network connection with the first device.
[0070] It should be noted that the topology keepalive request message is a message used to confirm the network connection status and maintain the network topology information. When a device in the network (such as a GPU card) establishes a network connection, in order to ensure the continuity and effectiveness of the connection, this message can be sent regularly to verify the survival status of the network connection. If no response is received within a period of time, or the response indicates that the connection is disconnected, appropriate measures (such as resending the topology discovery request message) can be taken to restore the network connection.
[0071] Specifically, when the first device establishes a complete network connection with the target device, the topology preservation process can be started for this connection. Here, the target device refers to a device that has established a network connection with the first device in the southbound network. For example, taking GPU0 as the first device, when the port 5 of GPU0 establishes a network connection with the port 5 of GPU9, the topology preservation process can be started for the network connection between the two based on the network topology information between the two. Here, GPU9 is the target device.
[0072] Specifically, after GPU0 and GPU9 establish a network connection, a TCP (Transmission Control Protocol) connection can be established between the server node where GPU0 is located and the server node where GPU9 is located, that is, a TCP connection can be established between the two GPU servers, which helps to reduce resource consumption. Through the established TCP connection, GPU0 can send a topology keep-alive request message to GPU9.
[0073] Here, if each port of GPU0 (such as port0~port7) has established a network connection with each port of GPU9, a topology keep-alive request message is sent based on each connection port of GPU0. Each port of GPU0 can periodically call the corresponding socket programming interface to send a message, which contains the node identifier, device identifier, port identifier and other identification information of the source device (i.e., the first device GPU0) and the target device (i.e., GPU9), so that the switch can accurately unicast the message to the target device.
[0074] It is understandable that the message header field of the topology keep-alive request message may include ETH (DMAC, SMAC, PROTO), IPV4 header, UDP header, TOPO header, Payload, etc. The meaning of each field is shown in Table 3 below: Table 3 Header of the topology keepalive request message
[0075] Table 4 TOPO header of topology keepalive request message As shown in Table 4, in the topology keep-alive request message, the TOPO message header includes a total of 32 bytes, of which bytes 0-19 are currently used, of which bytes 0-3 encapsulate the message type opcode, group identifier clique id and status field, and the value corresponding to the opcode can be 0x3, indicating that the message type is a keep-alive message (i.e., a topology keep-alive request message); bytes 4-7 encapsulate the node identifiers of the source GPU server and the target GPU server; bytes 8-11 encapsulate the device identifiers of the source GPU card and the target GPU card (i.e., the global number of the GPU card, which is unique); bytes 12-15 encapsulate the port numbers on the source GPU card and the target GPU card (i.e., the port identifier, which is unique inside each GPU card); bytes 16-19 are reserved fields. It should be noted that when GPU0 is used as the first device (i.e., the source device) and GPU9 is used as the target device, the message header TOPO header encapsulates the src node id, src gpu id, src port id related to GPU0, and the dst node id, dst gpu id, dst port id related to GPU9.
[0076] Step 450, receiving a topology keep-alive response message forwarded by the switch, and parsing the topology keep-alive response message to obtain a keep-alive status, wherein the topology keep-alive response message is replied to the switch by the target device based on the inspection result of the topology keep-alive request message.
[0077] Specifically, for GPU0, each port that has established a network connection will send a topology keep-alive request message, which will be unicast on the switch connected to the corresponding port, and the corresponding target device can receive the topology keep-alive request message through the corresponding socket programming interface. After receiving the topology keep-alive request message, the target device will perform a series of checks to confirm the legitimacy and validity of the message. For example, it can be checked whether the dst node id in the message is consistent with the node id of the server where the device (that is, the GPU card) is located, whether the src port id in the message is consistent with the port id of the local receiving port, and whether the clique id carried in the message is consistent with the clique id of the device. If not, an alarm log is generated to remind the user to check the network connection.
[0078] For another example, check the connection information in the message, that is, whether src [node id, gpu id, port id] - dst [node id, gpu id, port id] exists in the local machine. If not, fill the error information into the status field of the message.
[0079] If all fields in the topology keep-alive request message are checked to be legal, the target device can reply the topology keep-alive response message to the switch according to the protocol requirements. Here, the topology keep-alive response message is a data packet in a specific format sent to the switch according to the protocol requirements after the target device receives and checks the topology keep-alive request message. The message header of the topology keep-alive response message is the same as the message header of the topology keep-alive request message, but the opcode value encapsulated in the message header TOPO is 0x4, indicating that the message type is keep-alive response (i.e., topology keep-alive response message), and at the same time encapsulates its own clique id, and calls the socket programming interface to send the topology keep-alive response message.
[0080] After receiving the topology keep-alive response message from the target device, the switch can forward the response message to the first device, such as GPU0. GPU0 receives these response messages by monitoring and parses the status field in the message to determine whether the keep-alive status is successful or failed.
[0081] Step 460: When the keep-alive state is failed, continue to send topology keep-alive request messages until the number of failures exceeds a threshold, and then resend the topology discovery request message.
[0082] Specifically, if the status field obtained by parsing is success, the keep-alive status is determined to be successful, otherwise the keep-alive status is determined to be failed. If it fails, the first device (such as GPU0) can continue to send topology keep-alive request messages to the target device. If no successful response is received after sending the request message for multiple consecutive times (such as 3 times), it will eventually be determined that the keep-alive has failed and an alarm log will be generated. At this time, the first device can resend the topology discovery request message to try to re-establish the network connection with the target device.
[0083] Based on any of the above embodiments, the identification information includes a node identification, a device identification and a port identification.
[0084] Specifically, identification information refers to a set of data that can uniquely identify a device or port in the network, which may include node identification, device identification, and port identification. Among them, the node identification refers to the unique number of the server node where the device is located, the device identification refers to the global unique number of the device in the cluster, and the port identification refers to the local unique number of each port on the device. These identification information together constitute the unique identity of the device, so that each device in the network can be accurately identified and located.
[0085] Based on any of the above embodiments, Figure 5 This is the second flow chart of the network topology management method provided by the present invention, such as Figure 5 As shown, the method is applied to a second device, and the method includes: Step 510: receiving a topology discovery request message broadcasted by the switch, wherein the topology discovery request message is sent by the first device to the switch when receiving a topology discovery start instruction, wherein the first device is a device other than the second device connected to the switch in the southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, wherein each server node is provided with at least one device, and each device is connected to at least one switch; Step 520, check the topology discovery request message, and if the check passes, reply a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device, and the first device is used to parse the topology discovery response message, and establish a network connection with the second device based on the identification information of the second device obtained by the analysis and the local identification information, so as to obtain the network topology information of the first device.
[0086] It should be noted that the execution subject of the method provided in the embodiment of the present invention is the second device, which refers to a device that receives a topology discovery request in the southbound network and replies to a response message. The first device refers to a device that is connected to the same switch as the second device and initiates a topology discovery request. The specific operations performed by the first device and the second device in the topology discovery process in the embodiment of the present invention can refer to the relevant embodiments of the topology discovery process when the first device is the execution subject in the above embodiment, and the embodiment of the present invention will not be repeated here.
[0087] According to the method provided by the embodiment of the present invention, the first device can obtain the identification information of the second device by sending a topology discovery request message, receiving a topology discovery response message, parsing the topology discovery response message, etc., and establish a network connection with the second device accordingly to obtain the network topology information of the first device. Since the southbound network includes multiple server nodes, each server node is provided with at least one device, and each device can be used as both a first device and a second device, according to the above steps, all network topology information of the entire southbound network can be collected, and the management and maintenance of the southbound network can be realized according to these network topology information. In addition, the first device automatically sends the topology discovery request message when receiving the instruction to start the topology discovery, without manual configuration or checking of the network connection, which helps to improve the efficiency and accuracy of network topology management. By broadcasting the topology discovery request message through the switch, each device connected to the switch in the network can be covered, ensuring the integrity of the network topology information. The topology discovery response message is replied by the second device when the request message is checked and passed, so that the compliance and validity of the network connection between the first device and the second device can be ensured.
[0088] Based on any of the above embodiments, in step 520, the checking of the topology discovery request message includes: Step 521: parse the topology discovery request message to obtain the port identifier and group identifier of the source device.
[0089] Specifically, after receiving the topology discovery request message through the socket programming interface, the second device can parse the message to obtain various fields in the message, including src port id, clique id, etc. Here, src port id is the port identifier of the source device (i.e., the first device), and clique id is the group identifier of the source device.
[0090] Step 522: compare the port identifier of the source device with the local port identifier to obtain a first comparison result, and compare the group identifier of the source device with the local group identifier to obtain a second comparison result.
[0091] Step 523: If the first comparison result and the second comparison result are both consistent, determine that the check is passed; otherwise, generate an alarm log.
[0092] Specifically, after parsing the corresponding fields in the message, the port identifier of the source device can be compared with the local port identifier, that is, check whether the src port id is consistent with the port id of the local receiving port to obtain a first comparison result. If the first comparison result is inconsistent, an alarm log can be generated to remind the user to check the network connection. The main purpose of doing this is to hope that all ports with the same port id of all GPU cards are connected to the same switching network, which is convenient for wiring and troubleshooting network problems.
[0093] At the same time, the group ID of the source device can be compared with the local group ID, that is, check whether the clique ID in the message is consistent with the clique ID of the device, and obtain the second comparison result. If the second comparison result is inconsistent, an alarm log is generated to remind the user to check the network connection and the network configuration of the switch, and the topology discovery request message is discarded. This is because only GPU cards with the same clique ID can communicate in the southbound network.
[0094] If the first comparison result and the second comparison result are both consistent, it indicates that the check of each field in the topology discovery request message is passed, and the second device can reply with a topology discovery response message according to the protocol requirements.
[0095] Based on any of the above embodiments, the method further includes: Step 530: receiving a topology keep-alive request message unicast by the switch, where the topology keep-alive request message is sent by the first device to the switch when a network connection is established with the second device; Step 540, check the topology keep-alive request message, and based on the check result, reply a topology keep-alive response message to the switch, so that the switch forwards the topology keep-alive response message to the first device, and the first device is used to parse the topology keep-alive response message to obtain the keep-alive status.
[0096] It should be noted that, for the specific operations performed by the first device and the second device in the topology keep-alive process in the embodiment of the present invention, reference can be made to the relevant embodiments of the topology keep-alive process when the first device is the execution subject in the above embodiment, and the embodiments of the present invention will not be repeated here.
[0097] Based on any of the above embodiments, in step 540, the checking of the topology keepalive request message includes: Step 541: parse the topology keep-alive request message to obtain identification information of the source device, identification information of the target device, and the group identification of the source device.
[0098] Specifically, after the second device receives the topology keep-alive request message through the socket programming interface, it can parse the message to obtain the fields in the message, including src node id, src gpu id, src port id, cliqueid, dst node id, dst gpu id, dst port id, etc. Here, src node id, src gpu id, src portid is the identification information of the source device (i.e., the first device), dst node id, dst gpu id, dst port id is the identification information of the target device, and clique id is the group identifier of the source device.
[0099] Step 542, compare the node identifier in the identification information of the target device with the local node identifier to obtain a first result, compare the port identifier in the identification information of the source device with the local port identifier to obtain a second result, and compare the group identifier of the source device with the local group identifier to obtain a third result; Step 543: if the first result, the second result, and the third result are all consistent, check whether the connection information exists locally; if so, determine that the keep-alive state is successful; otherwise, determine that the keep-alive state is failed; the connection information is determined based on the identification information of the source device and the identification information of the target device; Step 544: if any one of the first result, the second result and the third result is inconsistent, generate an alarm log.
[0100] Specifically, after parsing the corresponding fields in the message, the node ID of the target device can be compared with the node ID of the server, that is, checking whether the dst node id is consistent with the node id of the server to obtain a first result. If the first result is inconsistent, an alarm log is generated.
[0101] Secondly, compare the port ID of the source device with the local port ID, that is, check whether the src port id is consistent with the port id of the local receiving port, and obtain the second result. If the first result is inconsistent, an alarm log can be generated to remind the user to check the network connection. The main purpose of doing this is to hope that all ports with the same port id of all GPU cards are connected to the same switching network, which is convenient for wiring and troubleshooting network problems.
[0102] At the same time, the group ID of the source device is compared with the local group ID, that is, the cliqueid in the message is checked to see if it is consistent with the clique id of the device, and the third result is obtained. If the third result is inconsistent, an alarm log is generated to remind the user to check the network connection and the network configuration of the switch, and the topology keep-alive request message is discarded. This is because only GPU cards with the same clique id can communicate in the southbound network.
[0103] In addition, it is also necessary to check whether the connection information in the message exists locally. The connection information here includes the identification information of the source device (src node id, src gpu id, src port id) and the identification information of the target device (dst nodeid, dst gpu id, dst port id). If it does not exist, fill the error information in the status field of the request message. If all the fields of the request message are checked to be legal, the topology keep-alive response message can be replied according to the protocol requirements.
[0104] Based on any of the above embodiments, Figure 6 is a schematic diagram of topology discovery provided by the present invention, such as Figure 6 As shown in the figure, when the kernel-driven fabric-manager-thread module receives the start topology discovery command issued by the fabric-client service, it starts topology detection. The following takes GPU0 as an example. The specific steps are as follows: Step A1: Create a raw socket in the fabric-manager-thread module of the kernel driver, send a topology discovery request message based on each port of GPU0, and listen at the same time (to receive the topology discovery response message).
[0105] Specifically, each port of GPU0 (i.e., port0~port7) can periodically (e.g., set to 1 second, or adjusted according to the performance requirements of topology convergence) call the socket programming interface sendto to send a broadcast message, i.e., send a discover request message (also called a topology discovery request message), which will be broadcast on the southbound switch connected to the corresponding port, and other GPU cards connected to the same switch will receive the topology discovery request message. For example, Figure 6 As shown in the figure, for port 0 of GPU0, after it sends a discover request message, the message will be broadcast on the southbound switch switch0, that is, the southbound switch switch0 transparently transmits the discover request, so that other GPU cards connected to switch0 (i.e. Figure 6 GPU1~GPU63) in the GPU card cluster shown in receives the request message.
[0106] Step A2: After receiving the discover request message through the socket programming interface recvfrom, other GPUs check whether each field in the message is legal.
[0107] Specifically, check whether the src port id in the message is consistent with the port id of the receiving port, and check whether the clique id in the message is consistent with the clique id of the GPU card. If they are inconsistent, an alarm log is generated to remind the user to check the network connection and the network configuration of the switch. If all fields are checked to be legal, reply to the discover response message (i.e., the topology discovery response message) according to the protocol requirements, encapsulate your own source MAC, encapsulate your own cliqueid, write your own node id, gpu id, port id into the dst node id, dst gpu id, dst port id fields, and call the socket programming interface to send the message.
[0108] In step A3, the switch transparently transmits all discover response messages to GPU0, that is, the switch forwards each topology discovery response message received to GPU0.
[0109] Step A4: GPU0 receives and parses the response message, and establishes a network connection based on the parsing result.
[0110] Specifically, after parsing, the dst information in the message (i.e., dst node id, dst gpu id, dstport id) can be obtained. If the status field in the message is success, a complete connection information is obtained: src [nodeid, gpu id, port id]- dst [node id, gpu id, port id]. According to the same logic, each GPU in the entire southbound network will obtain a complete network topology information, recording the information of each GPU neighbor that can be connected to it through the southbound switch. If the connection information in the topology discovery response message already exists locally, the message can be directly discarded.
[0111] Based on any of the above embodiments, Figure 7 is a schematic diagram of topology preservation provided by the present invention, such as Figure 7 As shown in the figure, once the fabric-manager-thread module establishes a complete connection information, it can start the keep-alive process for this connection. The following takes GPU0 as an example to introduce it. The specific steps are as follows: Step B1, create a TCP socket in the kernel-driven fabric-manager-thread module for listening (to receive response messages), establish a TCP connection with the neighbor (one TCP connection is established between every two server nodes, which helps to reduce resource consumption), and send a topology keep-alive request message based on each connection of GPU0.
[0112] Specifically, each port of GPU0 that has established a network connection can periodically (for example, set it to 3 seconds and adjust it according to the performance requirements of topology convergence) call the socket programming interface send to send a unicast message, that is, send a keep-alive request message (also called a topology keep-alive request message). The message encapsulates the src node id, src gpu id, srcport id of GPU0 and the dst node id, dst gpu id, and dst port id of the destination GPU. Figure 7 As shown, the topology keep-alive request message will be unicast on the southbound switch connected to the corresponding port, that is, the southbound switch transparently transmits the keep-alive request to the corresponding GPU card.
[0113] Step B2: After receiving the topology keep-alive request message through the socket programming interface read, other GPUs check whether each field and connection information in the message is legal.
[0114] Specifically, check whether the dst node id in the message is consistent with the node id of the local machine, check whether the srcport id in the message is consistent with the port id of the receiving port, and check whether the clique id in the message is consistent with the cliqueid of the GPU card. If they are inconsistent, an alarm log is generated to remind the user to check the network connection and the network configuration of the switch. In addition, it is necessary to check whether the connection information in the message exists in the local machine. If not, fill the error information into the status field in the topology keep-alive request message. If all the fields are checked to be legal, reply to the keep-alive response message (i.e., the topology keep-alive response message) according to the protocol requirements, encapsulate your own clique id, and call the socket programming interface to send the response message.
[0115] Step B3: The switch transparently transmits all keep-alive response messages, that is, the switch forwards each topology keep-alive response message received to GPU0.
[0116] Step B4: GPU0 receives the response message and parses it to obtain the status information in the message. If the status is success, the keep-alive is successful, otherwise the keep-alive fails. If it fails, continue to send keep-alive requests. If it fails three times in a row, the keep-alive is finally judged as a failure and an alarm log is generated. Subsequently, the topology discovery request message is used to try to re-establish the connection.
[0117] The present invention solves the topology management problem of the southbound network. In the entire network management mechanism, the unified allocation of global resources is firstly realized through the management node, such as the allocation of node ID, the IP address of the southbound network, etc., and the single point failure problem is solved through distributed deployment. Then, by sending a topology discovery request message in the southbound network, all the connection relationships of the entire southbound network can be collected, and the connection information can be maintained through the topology keep-alive request message. The whole solution has both functional considerations and maintainability guarantees.
[0118] The network topology management device provided by the present invention is described below. The network topology management device described below and the network topology management method described above can be referenced to each other.
[0119] Based on any of the above embodiments, Figure 8 This is one of the structural diagrams of the network topology management device provided by the present invention, such as Figure 8 As shown, the device is applied to a first device, and the device includes: The sending unit 810 is configured to send a topology discovery request message to the switch when receiving a topology discovery start instruction, so that the switch broadcasts the topology discovery request message to a second device, where the second device is a device other than the first device connected to the switch in a southbound network, where the southbound network includes a plurality of server nodes and a plurality of switches, each server node is provided with at least one device, and each device is connected to at least one switch; The receiving unit 820 is configured to receive a topology discovery response message forwarded by the switch, where the topology discovery response message is replied to the switch by the second device when the second device passes the topology discovery request message check; The establishing unit 830 is used to parse the topology discovery response message, and based on the identification information of the second device obtained by the analysis and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0120] In the apparatus provided by the embodiment of the present invention, the first device can obtain the identification information of the second device by sending a topology discovery request message, receiving a topology discovery response message, parsing the topology discovery response message, etc., and establish a network connection with the second device accordingly to obtain the network topology information of the first device. Since the southbound network includes multiple server nodes, each server node is provided with at least one device, and each device can be used as both a first device and a second device, according to the above steps, all network topology information of the entire southbound network can be collected, and the management and maintenance of the southbound network can be realized according to these network topology information. In addition, the first device automatically sends a topology discovery request message when receiving the instruction to start topology discovery, without manual configuration or checking of network connection, which helps to improve the efficiency and accuracy of network topology management. By broadcasting the topology discovery request message through the switch, each device connected to the switch in the network can be covered, ensuring the integrity of the network topology information. The topology discovery response message is replied by the second device when the request message is checked and passed, so that the compliance and validity of the network connection between the first device and the second device can be ensured.
[0121] Based on any of the above embodiments, the device further includes: a request sending unit, configured to send a topology keep-alive request message to the switch based on the network topology information, so that the switch unicasts the topology keep-alive request message to a target device, wherein the target device is a device in the second device that establishes a network connection with the first device; a response receiving unit, configured to receive a topology keep-alive response message forwarded by the switch, and parse the topology keep-alive response message to obtain a keep-alive status, wherein the topology keep-alive response message is replied to the switch by the target device based on a check result of the topology keep-alive request message; The state judgment unit is used to continue sending the topology keep-alive request message when the keep-alive state is failure, and resend the topology discovery request message after the number of failures exceeds a threshold.
[0122] Based on any of the above embodiments, the identification information includes a node identification, a device identification and a port identification.
[0123] Based on any of the above embodiments, Fig. 9 This is a second structural diagram of the network topology management device provided by the present invention, such as Fig. 9 As shown, the device is applied to a second device, and the device includes: A message receiving unit 910 is configured to receive a topology discovery request message broadcast by a switch, wherein the topology discovery request message is sent by a first device to the switch when receiving an instruction to start topology discovery, wherein the first device is a device other than the second device connected to the switch in a southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, wherein each server node is provided with at least one device, and each device is connected to at least one switch; The message reply unit 920 is used to check the topology discovery request message, and if the check passes, reply a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device. The first device is used to parse the topology discovery response message and establish a network connection with the second device based on the identification information of the second device obtained by the analysis and the local identification information, so as to obtain the network topology information of the first device.
[0124] In the apparatus provided by the embodiment of the present invention, the first device can obtain the identification information of the second device by sending a topology discovery request message, receiving a topology discovery response message, parsing the topology discovery response message, etc., and establish a network connection with the second device accordingly to obtain the network topology information of the first device. Since the southbound network includes multiple server nodes, each server node is provided with at least one device, and each device can be used as both a first device and a second device, according to the above steps, all network topology information of the entire southbound network can be collected, and the management and maintenance of the southbound network can be realized according to these network topology information. In addition, the first device automatically sends a topology discovery request message when receiving the instruction to start topology discovery, without manual configuration or checking of network connection, which helps to improve the efficiency and accuracy of network topology management. By broadcasting the topology discovery request message through the switch, each device connected to the switch in the network can be covered, ensuring the integrity of the network topology information. The topology discovery response message is replied by the second device when the request message is checked and passed, so that the compliance and validity of the network connection between the first device and the second device can be ensured.
[0125] Based on any of the above embodiments, the message reply unit 920 is specifically used for: Parsing the topology discovery request message to obtain a port identifier and a group identifier of a source device; Compare the port identifier of the source device with the local port identifier to obtain a first comparison result, and compare the group identifier of the source device with the local group identifier to obtain a second comparison result; When the first comparison result and the second comparison result are both consistent, it is determined that the check is passed, otherwise an alarm log is generated.
[0126] Based on any of the above embodiments, the device further includes: a request receiving unit, configured to receive a topology keep-alive request message unicast by the switch, wherein the topology keep-alive request message is sent by the first device to the switch when a network connection is established with the second device; A response reply unit is used to check the topology keep-alive request message and, based on the check result, reply a topology keep-alive response message to the switch, so that the switch forwards the topology keep-alive response message to the first device, and the first device is used to parse the topology keep-alive response message to obtain a keep-alive status.
[0127] Based on any of the above embodiments, the response reply unit is specifically used for: Parsing the topology keep-alive request message to obtain identification information of a source device, identification information of a target device, and a group identification of the source device; Compare the node identifier in the identification information of the target device with the local node identifier to obtain a first result, compare the port identifier in the identification information of the source device with the local port identifier to obtain a second result, and compare the group identifier of the source device with the local group identifier to obtain a third result; When the first result, the second result, and the third result are all consistent, checking whether the connection information exists locally, and if so, determining that the keep-alive state is successful, otherwise determining that the keep-alive state is failed, the connection information being determined based on the identification information of the source device and the identification information of the target device; When any one of the first result, the second result and the third result is inconsistent, an alarm log is generated.
[0128] Fig.10 The following is an example of a physical structure diagram of a device, such as Fig.10 As shown, the electronic device may include: a processor 1010 , a communication interface 1020 , a memory 1030 and a communication bus 1040 , wherein the processor 1010 , the communication interface 1020 , and the memory 1030 communicate with each other via the communication bus 1040 . The processor 1010 can call the logic instructions in the memory 1030 to execute a network topology management method, which is applied to a first device. The method includes: when receiving a start topology discovery instruction, sending a topology discovery request message to the switch so that the switch broadcasts the topology discovery request message to a second device, where the second device is a device other than the first device connected to the switch in a southbound network, and the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch; receiving a topology discovery response message forwarded by the switch, where the topology discovery response message is replied to the switch by the second device when the topology discovery request message is checked and passed; parsing the topology discovery response message, and based on the identification information of the second device obtained by the parsing and the local identification information, establishing a network connection with the second device to obtain the network topology information of the first device.
[0129] The processor 1010 can call the logic instructions in the memory 1030 to execute the network topology management method, which is applied to the second device, and the method includes: receiving a topology discovery request message broadcast by the switch, the topology discovery request message is sent by the first device to the switch when receiving the instruction to start topology discovery, the first device is other devices other than the second device connected to the switch in the southbound network, the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch; checking the topology discovery request message, and if the check passes, replying a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device, the first device is used to parse the topology discovery response message, and based on the identification information of the second device obtained by the analysis and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0130] In addition, the logic instructions in the above-mentioned memory 1030 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the relevant technology. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0131] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the network topology management method provided by the above methods. The method is applied to a first device, and the method includes: when receiving an instruction to start topology discovery, sending a topology discovery request message to a switch so that the switch broadcasts the topology discovery request message to a second device, wherein the second device is a device other than the first device connected to the switch in a southbound network, wherein the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch; receiving a topology discovery response message forwarded by the switch, wherein the topology discovery response message is replied to the switch by the second device when the topology discovery request message is checked and passed; parsing the topology discovery response message, and based on the identification information of the second device obtained by parsing and the local identification information, establishing a network connection with the second device to obtain the network topology information of the first device.
[0132] When the computer program is executed by the processor, the computer can execute the network topology management method provided by the above methods, which is applied to the second device, and includes: receiving a topology discovery request message broadcast by the switch, the topology discovery request message is sent by the first device to the switch when receiving an instruction to start topology discovery, the first device is other devices other than the second device connected to the switch in the southbound network, the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch; checking the topology discovery request message, and if the check passes, replying a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device, the first device is used to parse the topology discovery response message, and based on the identification information of the second device obtained by parsing and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0133] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the network topology management method provided by the above-mentioned methods, the method being applied to a first device, the method comprising: upon receiving an instruction to start topology discovery, sending a topology discovery request message to a switch so that the switch broadcasts the topology discovery request message to a second device, the second device being a device other than the first device connected to the switch in a southbound network, the southbound network comprising a plurality of server nodes and a plurality of switches, each server node being provided with at least one device, and each device being connected to at least one switch; receiving a topology discovery response message forwarded by the switch, the topology discovery response message being replied to the switch by the second device when the topology discovery request message is checked and passed; parsing the topology discovery response message, and establishing a network connection with the second device based on the identification information of the second device obtained by the parsing and the local identification information, to obtain the network topology information of the first device.
[0134] When the computer program is executed by the processor, it is implemented to execute the network topology management method provided by the above-mentioned methods, which is applied to the second device, and the method includes: receiving a topology discovery request message broadcast by the switch, the topology discovery request message is sent by the first device to the switch when receiving the instruction to start topology discovery, the first device is other devices other than the second device connected to the switch in the southbound network, the southbound network includes multiple server nodes and multiple switches, each server node is provided with at least one device, and each device is connected to at least one switch; checking the topology discovery request message, and if the check passes, replying a topology discovery response message to the switch, so that the switch forwards the topology discovery response message to the first device, the first device is used to parse the topology discovery response message, and based on the identification information of the second device obtained by the parsing and the local identification information, establish a network connection with the second device to obtain the network topology information of the first device.
[0135] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0136] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A network topology management method, characterized in that: The method is applied to a first device, and the method includes: In the case of receiving a start topology discovery instruction, sending a topology discovery request message to the switch, so that the switch broadcasts the topology discovery request message to a second device, where the second device is a device other than the first device connected to the switch in a southbound network, the southbound network comprising a plurality of server nodes and a plurality of switches, each server node being provided with at least one device, and each device being connected to at least one switch; receiving a topology discovery response message forwarded by the switch, where the topology discovery response message is replied to the switch by the second device when the second device passes the check on the topology discovery request message; The topology discovery response message is parsed, and based on the identification information of the second device obtained through the analysis and the local identification information, a network connection is established with the second device to obtain the network topology information of the first device.
2. The network topology management method according to claim 1, characterized in that: Also includes: Based on the network topology information, sending a topology keep-alive request message to the switch, so that the switch unicasts the topology keep-alive request message to a target device, where the target device is a device in the second device that establishes a network connection with the first device; receiving a topology keep-alive response message forwarded by the switch, and parsing the topology keep-alive response message to obtain a keep-alive status, wherein the topology keep-alive response message is replied to the switch by the target device based on a check result of the topology keep-alive request message; When the keep-alive state is failure, continue to send topology keep-alive request messages until the number of failures exceeds a threshold, and then resend the topology discovery request message.
3. The network topology management method according to claim 1 or 2, characterized in that: The identification information includes a node identification, a device identification and a port identification.
4. A network topology management method, characterized in that: The method is applied to a second device, and the method includes: receiving a topology discovery request message broadcasted by a switch, wherein the topology discovery request message is sent by a first device to the switch when receiving an instruction to start topology discovery, wherein the first device is a device other than the second device connected to the switch in a southbound network, wherein the southbound network includes a plurality of server nodes and a plurality of switches, each server node is provided with at least one device, and each device is connected to at least one switch; The topology discovery request message is checked, and if the check passes, a topology discovery response message is replied to the switch, so that the switch forwards the topology discovery response message to the first device, and the first device is used to parse the topology discovery response message, and establish a network connection with the second device based on the identification information of the second device obtained by the analysis and the local identification information, so as to obtain the network topology information of the first device.
5. The network topology management method according to claim 4, characterized in that: The checking of the topology discovery request message includes: Parsing the topology discovery request message to obtain a port identifier and a group identifier of a source device; Compare the port identifier of the source device with the local port identifier to obtain a first comparison result, and compare the group identifier of the source device with the local group identifier to obtain a second comparison result; When the first comparison result and the second comparison result are both consistent, it is determined that the check is passed, otherwise an alarm log is generated.
6. The network topology management method according to claim 4, characterized in that: Also includes: receiving a topology keep-alive request message unicast by the switch, where the topology keep-alive request message is sent by the first device to the switch when a network connection is established between the first device and the second device; The topology keep-alive request message is checked, and based on the check result, a topology keep-alive response message is replied to the switch, so that the switch forwards the topology keep-alive response message to the first device, and the first device is used to parse the topology keep-alive response message to obtain the keep-alive status.
7. The network topology management method according to claim 6, characterized in that: The checking of the topology keep-alive request message includes: Parsing the topology keep-alive request message to obtain identification information of a source device, identification information of a target device, and a group identification of the source device; Compare the node identifier in the identification information of the target device with the local node identifier to obtain a first result, compare the port identifier in the identification information of the source device with the local port identifier to obtain a second result, and compare the group identifier of the source device with the local group identifier to obtain a third result; When the first result, the second result, and the third result are all consistent, checking whether the connection information exists locally, and if so, determining that the keep-alive state is successful, otherwise determining that the keep-alive state is failed, the connection information being determined based on the identification information of the source device and the identification information of the target device; When any one of the first result, the second result and the third result is inconsistent, an alarm log is generated.
8. A device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the network topology management method according to any one of claims 1 to 7 is implemented.
9. A network topology management system, characterized in that: It comprises a management node, a plurality of server nodes and a plurality of switches, each server node is provided with at least one device as claimed in claim 8, each device comprises at least one port, and each port is connected to a switch; The server node is used to obtain network topology information of each local device and report the network topology information of each device to the management node; The management node is used to obtain information of each switch based on the switch management network, and manage the network topology information of each device and the information of each switch.
10. The network topology management system according to claim 9, characterized in that: The management node is also used to assign node identifiers to each server node, and uniformly assign IP addresses to all ports of all devices on all server nodes; The server node is further configured to generate a global device identifier for each device based on the node identifier and the local device identifier of each device.
11. The network topology management system according to claim 9, characterized in that: There are multiple management nodes.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the network topology management method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and device for realizing keep-alive mechanism
CN101635675A
Method and device for discovering network topology in software defined network SDN (Software Defined Network)
CN105721318A
A non-centralized service cluster system for position service and a fault detection method
CN109873713A
Network system, network message processing method and device and storage medium
CN116366455A
Model deployment method, system and equipment based on cluster topological structure and medium
CN117155791A
Cited By
Computing system, computing management method, and electronic device
CN121349379A
Automatic routing construction method based on topology information carried by registration message and controller
CN121887705A