Topology management method and system, electronic equipment and storage medium

Through the method of automatically reporting information on the client and automatically generating topology information on the server, the problem of cumbersome configuration and complex management in multi-device interconnect topology management is solved, automated configuration and centralized management are realized, and system efficiency and stability are improved.

CN120263658APending Publication Date: 2025-07-04广州壁仞智能科技有限公司 +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510466227.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, multi-device interconnect topology management has problems such as cumbersome configuration process, lack of centralized management and high restart costs, especially in multi-device parallel computing systems, which affects the performance and efficiency of the system.

Method used

Each client automatically reports device link information and switch port information, and the server automatically generates external interconnect topology information of the device and issues it to each client for analysis and configuration, realizing automated topology management.

Benefits of technology

It simplifies the configuration process, improves configuration efficiency, reduces management complexity, and realizes centralized management and efficient distributed training of all interconnected devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263658A_ABST
    Figure CN120263658A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence chips, and provides a topology management method and system, electronic equipment and a storage medium, and the method comprises the steps: receiving equipment link information and switch port information reported by each client; based on a model distributed training mode and the equipment link information and the switch port information of each client, generating equipment external interconnection topology information; and issuing the device external interconnection topology information to each client, so that each client analyzes the device external interconnection topology information and performs topology configuration on each local interconnection device based on an analysis result, and the interconnection devices are devices connected with a switch box. According to the invention, each client automatically reports the local equipment link information and the switch port information, and the server automatically generates the external interconnection topology information of the equipment and issues the external interconnection topology information to each client, so that automatic topology configuration is realized, the configuration process is greatly simplified, and the configuration efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence chips, and particularly to a topology management method, system, electronic device, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence and big data technologies, the complexity of deep learning models has been increasing day by day, and the amount of data to be processed has also shown an explosive growth trend. To address this challenge, the computing power of a single device has gradually become unable to meet the requirements. Therefore, multi-device parallel computing has become a key means to improve the training efficiency and performance of deep learning models. In a multi-device parallel computing system, efficient interconnection between devices is particularly important, which directly affects the efficiency of data parallelism and task parallelism, and thus affects the performance of the entire system.

[0003] Currently, in order to achieve interconnection between multiple devices, a script configuration method is usually adopted. This method allows developers to flexibly define the connection relationships and network topologies between devices according to specific hardware environments and application requirements. However, with the expansion of system scale and the increase in complexity, this decentralized script configuration method has defects such as a cumbersome configuration process, lack of centralized management, and high restart costs. Summary of the Invention

[0004] The present invention provides a topology management method, system, electronic device, and storage medium to solve the defects of cumbersome configuration process, lack of centralized management, and high restart costs in the topology management of multi-device interconnection in related technologies.

[0005] The present invention provides a topology management method, which is applied to a server side, and the method includes: Receiving device connection information and switch port information reported by each client; Generating device external interconnection topology information based on the model distributed training mode and the device connection information and switch port information of each client; Sending the device external interconnection topology information to each client, so that each client parses the device external interconnection topology information and performs topology configuration on local interconnected devices based on the parsing result, where the interconnected devices refer to devices connected to a switch cabinet.

[0006] According to a topology management method provided by the present invention, the generating device external interconnection topology information based on the model distributed training strategy and the device connection information and switch port information of each client includes: Based on the model distributed training mode, as well as the device connection information and switch port information of each client, allocate memory area identifiers for each interconnected device on each client, and determine the group identifiers of each interconnected device; Based on the device connection information and switch port information of each client, as well as the memory area identifiers and group identifiers of each interconnected device on each client, generate device external interconnection topology information.

[0007] According to a topology management method provided by the present invention, the device connection information of any client includes the connection information of each interconnected device on the any client, and the connection information of each interconnected device includes a domain identifier, a bus identifier, a device identifier, a function identifier, and a device number.

[0008] According to a topology management method provided by the present invention, it further includes: Monitor the device connection status of each client, and each client regularly reads the connection status of each local interconnected device and reports it; In the case where the device connection status of any client is monitored to be abnormal, send a connection recovery instruction to the any client, so that the any client re - establishes a device connection based on the connection recovery instruction and reports the result status of the re - connection.

[0009] According to a topology management method provided by the present invention, each client regularly reads the local device connection information and sends the device connection information to the server as heartbeat information.

[0010] According to a topology management method provided by the present invention, it further includes: Receive a mode switching instruction; In response to the mode switching instruction, determine a new training mode, and based on the new training mode, as well as the device connection information and switch port information of each client, re - allocate memory area identifiers and group identifiers to generate new device external interconnection topology information; Send the new device external interconnection topology information to each client, so that each client re - configures the topology of each local interconnected device based on the new device external interconnection topology information.

[0011] The present invention also provides a topology management method, which is applied to a client, and the method includes: Obtain the local device connection information and the switch port information of the switches connected to each local interconnected device, where the interconnected device refers to a device connected to a switch cabinet; Report the device link information and the switch port information to the server side, so that the server side can generate the external device interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; Receive the external device interconnection topology information sent by the server side, parse the external device interconnection topology information, and perform topology configuration on each interconnected device based on the parsing result.

[0012] According to a topology management method provided by the present invention, the parsing of the external device interconnection topology information and the topology configuration of each interconnected device based on the parsing result include: Parse the external device interconnection topology information to obtain the memory area identifier and group identifier of each local interconnected device; Based on the device management library, transparently transmit the memory area identifier and group identifier of each interconnected device to the kernel driver, so that the kernel driver performs topology configuration based on the memory area identifier and group identifier of each interconnected device.

[0013] The present invention also provides a topology management device, which is applied to the server side, and the device includes: A receiving unit, configured to receive the device link information and switch port information reported by each client; A generating unit, configured to generate the external device interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; A sending unit, configured to send the external device interconnection topology information to each client, so that each client parses the external device interconnection topology information and performs topology configuration on each local interconnected device, where the interconnected device refers to a device connected to the switch cabinet.

[0014] The present invention also provides a topology management device, which is applied to the client side, and the device includes: An obtaining unit, configured to obtain the local device link information and the switch port information of the switches connected to each local interconnected device, where the interconnected device refers to a device connected to the switch cabinet; A reporting unit, configured to report the device link information and the switch port information to the server side, so that the server side can generate the external device interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; A configuration unit, configured to receive the external device interconnection topology information sent by the server side, parse the external device interconnection topology information, and perform topology configuration on each interconnected device based on the parsing result.

[0015] The present invention also provides a topology management system, which includes a server side, a communication module, and at least one client; The client is used to obtain the local device link information and the switch port information of each local interconnected device, and transmit the device link information and the switch port information to the server side through the communication module, where the interconnected device refers to a device connected to a switch cabinet; The server side is used to generate device external interconnection topology information based on the device link information and switch port information of each client, and send the device external interconnection topology information to each client through the communication module; The client is also used to parse the device external interconnection topology information and perform topology configuration on each local interconnected device based on the parsing result.

[0016] The present invention also provides an electronic device, which includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the topology management method described in any one of the above is implemented.

[0017] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the topology management method described in any one of the above is implemented.

[0018] The present invention also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the topology management method described in any one of the above is implemented.

[0019] The topology management method, system, electronic device, and storage medium provided by the present invention enable each client to automatically report local device link information and switch port information. The server side automatically generates device external interconnection topology information and sends it to each client, so that each client can perform topology configuration on each local interconnected device according to the device external interconnection topology information, realizing automatic configuration. There is no need to manually execute different scripts for configuration on each client, greatly simplifying the configuration process and improving the configuration efficiency. In addition, by generating device external interconnection topology information on the server side, centralized management of the external interconnection topology of all interconnected devices can be achieved, reducing the management complexity. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following descriptions are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic structural diagram of multi-machine and multi-GPU interconnection provided by the present invention; Figure 2 It is one of the schematic flow diagrams of the topology management method provided by the present invention; Figure 3 It is the second schematic flow diagram of the topology management method provided by the present invention; Figure 4 It is the overall schematic flow diagram of multi-machine and multi-GPU interconnection topology configuration and monitoring provided by the present invention; Figure 5 It is the schematic diagram of the server-side and client-side deployment provided by the present invention; Figure 6 It is one of the schematic structural diagrams of the topology management device provided by the present invention; Figure 7 It is the second schematic structural diagram of the topology management device provided by the present invention; Figure 8 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0023] With the increase in model complexity and the explosive growth of data volume, the computing power of a single device has gradually been unable to meet the requirements. Therefore, multi-device parallel computing has become a key means to improve the efficiency and performance of model training. To achieve multi-device parallel computing, efficient interconnection between devices is particularly important. Here, the device can be a GPU (Graphics Processing Unit), GPGPU (General-purpose computing on Graphics Processing Units), TPU (Tensor Processing Unit), etc., and the present invention does not make specific limitations in this regard.

[0024] Figure 1 It is the schematic structural diagram of multi-machine and multi-GPU interconnection provided by the present invention, as Figure 1As shown, multi-machine and multi-GPU interconnection means connecting multiple GPUs on multiple computers (or servers) through specific communication technologies to form a unified computing resource pool for achieving larger-scale and higher-efficiency computing tasks. For example, the specific communication technology can be PCIE (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) technology, and multi-machine and multi-GPU interconnection can be to interconnect multiple GPUs on 2 machines or 4 machines.

[0025] Figure 1 Taking 2 machines as an example, each machine includes a total of 8 GPUs from GPU0 to GPU7. The GPUs within a single machine are in an asymmetric structure. On each machine, 4 GPUs (such as Figure 1 GPU0, GPU2, GPU5, and GPU7 shown in the figure) are interconnected at high speed with the GPUs on the external machine through the PCIE switch in the lower switch box. Here, the fact that the GPUs within a single machine are in an asymmetric structure means that the connections of the GPUs within each machine are not completely the same, that is, some GPUs will be connected to the external machine, while some GPUs will not be connected to the external. It should be noted that the switch box is a hardware device that allows multiple PCIE devices (such as GPUs, etc.) to be interconnected with each other through single or multiple PCIE ports. It can transfer data between multiple PCIE devices, allowing direct communication between devices without going through the host processor, thereby improving data transfer efficiency.

[0026] In Figure 1 the multi-machine and multi-GPU interconnection topology shown, multiple topological links are formed, which can thus support communication operations between multiple GPUs, such as reduce, gather, broadcast, etc. These operations are very important in distributed training. It should be understood that in addition to the interconnection of the GPUs within the two machines, these two machines can also be connected through a network card and the upper RoCE (RDMA over Converged Ethernet, a remote direct memory access protocol based on Ethernet) switch. Through the RoCE switch, the two machines can efficiently exchange data and support larger-scale distributed training. In addition, within each machine, the GPU can be communicatively connected to the CPU (Central Processing Unit) through the PCIE switch to improve the data transfer efficiency between the GPU and the CPU.

[0027] In the distributed training of models, different models may adopt different distributed training modes, such as TP8, TP16, TP32, etc. Here, TP8, TP16, and TP32 refer to the tensor parallelism mode in the model distributed training mode, which represent different segmentation granularities and parallel degrees. Specifically, the numbers after TP (such as 8, 16, 32) usually refer to the number of devices to which the model parameters or computing tasks are segmented during tensor parallelism. It should be understood that tensor parallelism means that the tensors in the model (i.e., the basic data units of the model, such as weight matrices) are segmented onto different devices, and each device is only responsible for processing the tensor fragments assigned to it. After the calculation is completed, the results of each fragment need to be synchronized to maintain the integrity of the model.

[0028] Since different distributed training modes have different requirements for communication and data transmission between GPUs, the interconnection topology configuration between GPUs also needs to change with different training models. However, the current interconnection topology between multiple GPUs is usually implemented by script configuration. Since the configuration commands for each machine are different, the administrator needs to log in to each machine layer by layer for configuration, which greatly increases the configuration time and labor costs. When the training mode of the model changes, each machine in the interconnected multi-machines needs to be restarted and the script configuration needs to be re-executed. After the configuration is completed, a test program needs to be run to verify the correctness of the topology configuration, resulting in a cumbersome configuration process and high restart costs. In addition, since different scripts need to be executed on each machine for configuration, the centralized management of the interconnection topology of multiple GPUs is lacking, which not only increases the management complexity but also easily leads to configuration errors and inconsistencies.

[0029] In view of this, the present invention provides a topology management method. By automatically reporting the local device link information and switch port information by each client, the server side automatically generates the external interconnection topology information of the devices and distributes it to each client, realizing automatic topology configuration, greatly simplifying the configuration process, and improving the configuration efficiency. At the same time, by generating the external interconnection topology information of the devices on the server side, the centralized management of the external interconnection topology of all interconnected devices can be realized, reducing the management complexity, thereby overcoming the above defects.

[0030] Based on the above embodiments, Figure 2 is one of the schematic flowcharts of the topology management method provided by the present invention. As Figure 2 shown, this method is applied to the server side, and this method includes: Step 210, receiving the device link information and switch port information reported by each client.

[0031] It should be noted that the method provided in the embodiments of the present invention can be applied to the server side. Here, the server side refers to a computer or program that provides services in a network. The client side refers to a computer or program that requests services in a network. For example, in a scenario of multi-machine and multi-GPU interconnection, assuming that there are 4 machines (machines 0 to 3) interconnected, that is, multiple GPUs on these 4 machines are interconnected at high speed through a PCIE switch in the switch box, one of them can be selected as both the server side and the client side, while the other 3 are only used as client sides.

[0032] In the embodiments of the present invention, the client side is responsible for obtaining local device connection information and switch port information, and reporting this information to the server side; the server side is responsible for receiving the device connection information and switch port information from each client side, generating external device interconnection topology information based on this information, and then sending the external device interconnection topology information to each client side. In addition, each client side also needs to receive the external device interconnection topology information sent by the server side and perform topology configuration of local interconnected devices based on this information. The technical solutions of the embodiments of the present invention will be introduced in detail below.

[0033] Specifically, each client side can read local device connection information and switch port information through the device management library on the machine where it is located, and then can use a specific API (Application Programming Interface) or message format to encapsulate this information, and send the encapsulated device connection information and switch port information to the server side through a network communication protocol (such as TCP / IP). Here, the device management library refers to a software library deployed on the machine for managing and monitoring various devices on the machine. For example, the device management library can be a GPU Management Library (i.e., GPU management library).

[0034] It can be understood that the device link information refers to the number of all interconnected devices (such as GPUs) on the machine where the client is located and the link information of each interconnected device, etc. The switch port information refers to the information of the switch ports used when all interconnected devices on the machine where the client is located are interconnected with external devices (such as GPUs of other machines) through the switch box. Here, the interconnected devices refer to the devices on the machine where the client is located that are connected to external devices (such as GPUs of other machines) through the switch box, and these interconnected devices are the objects to be considered when the server generates the external interconnection topology information of the device. It should be understood that the link information of each interconnected device refers to a set of parameters that can uniquely identify and describe the location and connection status of the device in the network. For example, taking the interconnected device as a GPU, the link information of each interconnected device may include the domain identifier, bus identifier, device identifier, and GPU ID of the device on the PCIE bus, etc. The embodiments of the present invention do not make specific limitations on this.

[0035] Step 220: Generate external device interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client.

[0036] Specifically, after the server receives the device link information and switch port information from each client, it will construct an external device interconnection topology information that describes the interconnection relationship between all interconnected devices according to the requirements of the model distributed training mode and the device link information and switch port information of each client. Here, the external device interconnection topology information refers to an interconnection relationship data structure that describes the high-speed interconnection between all interconnected devices through an external switch box, and it may include key information such as the number of nodes (or the number of machines, each machine is a node), the unique identifier of each node, and the set of external links of each node. Among them, the set of external links of each node may include information such as the identifier, location, and connection relationship with other devices of each interconnected device on the node.

[0037] Specifically, the server can parse the specific requirements for the external interconnection relationship of devices (such as GPUs) according to the model distributed training mode selected by the user (such as any one of TP8, TP16, and TP32). For example, in the scenario of 4 machines with 32 GPU cards, the TP8 mode means that the tensor is split into 8 parts, and each part is assigned to a GPU, which means that each group (i.e., clique family) requires 8 GPUs to complete the parallel calculation of a tensor. Therefore, 32 GPUs can be divided into 4 (i.e., 32 / 8 = 4) clique families; while the TP16 mode means that each tensor is split into 16 parts, that is, each clique family requires 16 GPUs to complete the parallel calculation of a tensor. Therefore, 32 GPUs can be divided into 2 clique families.

[0038] Based on the received device connection information, switch port information, and the requirements of the parsed model distributed training mode, the server starts to construct the external interconnection topology information of the devices. During the construction process, the server assigns a unique node identifier (such as self_node_id) to each node (i.e., each machine), and creates a NodeConn array to store the set of external connections of the node. For each connection in the NodeConn array, the server records information such as the source GPU device ID of the connection (denoted as src_gpu_id), the destination node ID (denoted as dst_node_id), the destination GPU device ID (denoted as dst_gpu_id), the switch port address of the destination GPU (denoted as dst_port), and the connection type (denoted as link_type). Based on the above information, the external interconnection topology information of the devices can be constructed.

[0039] Step 230: Send the external interconnection topology information of the devices to each of the clients, so that each of the clients can parse the external interconnection topology information of the devices and perform topology configuration on each local interconnected device based on the parsing result. The interconnected devices refer to the devices connected to the switch box.

[0040] Specifically, after generating the external interconnection topology information of the devices, the server can also send the external interconnection topology information of the devices to each client through the network communication protocol. For example, the server can use the same API or message format as the client to encapsulate the external interconnection topology information of the devices and send it to the listening ports of each client through the network.

[0041] After each client receives the external interconnection topology information of the devices, it will parse it to extract useful information. For example, operations such as decoding, format conversion, and extraction can be performed on the external interconnection topology information of the devices. The parsed information will be used in the subsequent topology configuration process. It should be understood that the parsing result refers to the useful information obtained by the client after parsing the external interconnection topology information of the devices. These information usually include the identifiers, locations, and connection relationships with other devices of each interconnected device. These information will be used as the basis for the client to perform topology configuration.

[0042] Based on the information in the parsing result, the client can perform topology configuration on each local interconnected device. Specifically, the client can configure network parameters such as the routing table of each local interconnected device, as well as the connection relationships and communication protocols between devices. Through topology configuration, the client can ensure that each local interconnected device can be interconnected and communicate with other devices according to the requirements in the external interconnection topology information of the devices.

[0043] It can be understood that performing topology configuration means configuring the network parameters and connection relationships of each local interconnected device according to the requirements in the external interconnected topology information of the device. Through topology configuration, it can be ensured that each device in the distributed training task can be correctly interconnected and communicate, so as to achieve efficient distributed training.

[0044] In the method provided by the embodiments of the present invention, each client automatically reports the local device link information and switch port information, and the server side automatically generates the external interconnected topology information of the device and distributes it to each client, so that each client can perform topology configuration on each local interconnected device according to the external interconnected topology information of the device, realizing automatic configuration, without manually executing different scripts for configuration on each client, greatly simplifying the configuration process and improving the configuration efficiency. In addition, by generating the external interconnected topology information on the server side, centralized management of the external interconnected topologies of all interconnected devices can be realized, reducing the management complexity.

[0045] Based on any of the above embodiments, the device link information of any client includes the link information of each interconnected device on the any client, and the link information of each interconnected device includes a domain identifier, a bus identifier, a device identifier, a function identifier, and a device number.

[0046] Specifically, the link information of each interconnected device refers to a set of parameters that can uniquely identify and describe the position and connection status of the device in the network. For example, taking the interconnected device as a GPU, the link information of each interconnected device may include the domain identifier, bus identifier, device identifier, function identifier, and device number of the device on the PCIE bus, etc.

[0047] Here, the domain identifier (i.e., the PCIE domain ID) refers to the unique identifier of the PCIE domain where the device is located. The PCIE domain is a logical division in the PCIE bus structure, used to distinguish different PCIE subsystems or domains. Each PCIE domain has its unique ID to ensure uniqueness in the system.

[0048] The bus identifier (i.e., the bus ID) is used to uniquely identify the PCIE bus to which the device is connected, and it is used to distinguish devices connected to different PCI buses.

[0049] The device identifier (i.e., the device ID) is used to uniquely identify the slot where the device is located, and it is used to distinguish different devices on the same bus. There can be multiple slots on one bus.

[0050] The function identifier (i.e., the function ID) represents the function of the device, and it is used to distinguish different functions implemented on the same slot. There can be multiple functions on one slot, and these functions usually correspond to different hardware units or logical units on the device.

[0051] The device number (i.e., GPU ID) refers to the number used to identify the GPU, which is usually a unique identifier assigned to the GPU in the system or software environment so that the system or software can correctly select and use GPU resources.

[0052] In the embodiments of the present invention, for each GPU, by combining Dbdf (i.e., PCIE domain ID, bus ID, device ID, function ID) and GPU ID, etc., a GPU device can be uniquely identified, ensuring the uniqueness and identifiability of the GPU device in the network, and providing a basis for generating external interconnection topology information and local topology configuration of the device.

[0053] Based on any of the above embodiments, step 220 specifically includes: Step 221, based on the model distributed training mode and the device link information and switch port information of each client, allocate memory area identifiers for each interconnected device on each client, and determine the group identifiers of each interconnected device.

[0054] It should be noted that when the server generates external interconnection topology information of the device, in addition to generating information such as the connection relationship between each interconnected device and other devices, it will also allocate memory area identifiers and group identifiers for each interconnected device, etc., further ensuring more efficient and stable communication between all interconnected devices.

[0055] Specifically, after receiving the device link information and switch port information from each client, the server analyzes the roles and tasks of each device in the training process according to the selected model distributed training mode (such as any one of TP8, TP16, TP32). To avoid conflicts when accessing remote devices, the server will allocate a unique memory area identifier for each interconnected device, and this identifier is used to identify and access the memory space of the device during the distributed training process. Here, the memory area identifier is a unique identifier used to identify and access the memory space of the device during the distributed training process. It ensures that each device has an independent memory area and avoids conflicts when accessing remote devices.

[0056] In addition, according to the device link information and switch port information, determine the connection relationship of each interconnected device. According to the connection relationship of the devices and the requirements of the model distributed training mode, divide the devices in the same communication group into a group, and allocate a unique group identifier for each group. This helps to more clearly represent the connection relationship and group relationship of the devices in the external interconnection topology information of the device. Here, the group identifier (i.e., clique id) refers to a unique identifier used to represent the group relationship of the devices, and the interconnected devices with the same group identifier can communicate with each other.

[0057] Step 222: Generate external device interconnection topology information based on the device link information and switch port information of each client, as well as the memory area identifiers and group identifiers of each interconnected device on each client.

[0058] Specifically, based on the device link information and switch port information of each client, the server can construct a data structure describing the connection relationship between devices. According to the allocated memory area identifier and group identifier of each interconnected device, the server can add the memory area identifier and group identifier for each device in the data structure. Based on the constructed device connection relationship data structure and the added memory area identifier, group identifier and other information, the external device interconnection topology information can be generated. This information includes complete information such as device node information, connection relationship, group relationship, memory area identifier, etc., which can be used for subsequent distributed training tasks.

[0059] It can be understood that the memory area identifier and group identifier are allocated by the server based on the model task to be trained. According to these allocated identifier information, the server will generate the external device interconnection topology information and send it to the client for corresponding configuration. For example, if the training mode of a large model is TP8, which means a tensor is split into 8 parts and each part is allocated to a GPU for processing, then these 8 GPUs will form a group (i.e., clique family). The server will allocate the same clique id to these 8 GPUs and then send it to the corresponding client, and the client will call the driver to implement the configuration.

[0060] In the embodiment of the present invention, steps 221 and 222 together implement the process of the server generating the external device interconnection topology information based on the device link information and switch port information of each client. Among them, the introduction of the memory area identifier and group identifier provides an effective means for device identification and group division, which helps to improve the efficiency and accuracy of distributed training.

[0061] Based on any of the above embodiments, each client periodically reads the local device link information and sends the device link information to the server as heartbeat information.

[0062] It should be noted that the heartbeat mechanism is a technology used to ensure that the connection between the client and the server remains active. Due to the instability of the network condition or the failure of intermediate devices, the connection may be interrupted unexpectedly. By periodically sending heartbeat information, the server can confirm that the client is still online and the connection is normal.

[0063] Specifically, each client can periodically read the device link information of the machine it is on through the GPU management library and report this information as heartbeat information to the server side. If the server side does not receive heartbeat information from a certain client within a certain period of time, it can infer that the client may have dropped the line or there is a problem with the connection. This mechanism allows the server side to detect and respond to connection failures in a timely manner, and thus take corresponding measures (such as re - establishing the connection or notifying the administrator).

[0064] In addition, including the device link information in the heartbeat information can ensure that the server side always has the latest status of the client devices at hand. This is very useful for monitoring the device health status, managing device resources, and making decisions based on the device status. If the status or configuration of the client changes during operation (for example, IP address change, device upgrade, etc.), these changes can be notified to the server side in a timely manner through the heartbeat information, which helps the server side maintain accurate client status information and ensure data consistency.

[0065] In the embodiment of the present invention, the client periodically reports the device link information of its own machine as heartbeat information to the server side, which is an effective network management strategy. It can not only keep the connection alive, but also help the server side detect and handle connection failures in a timely manner, while ensuring data accuracy and consistency.

[0066] Based on any of the above - mentioned embodiments, the method further includes: Monitoring the device link status of each client, where each client periodically reads the link status of each interconnected device locally and reports it; In the case where it is monitored that the device link status of any client is abnormal, sending a link recovery instruction to the any client, so that the any client re - establishes the device link based on the link recovery instruction and reports the result status of the re - link.

[0067] It should be noted that considering that the current GPU interconnection topology configuration method lacks real - time monitoring functions, once a certain GPU interconnection link fails, the system cannot issue an alarm or provide fault information in a timely manner. Usually, the administrator can only discover the problem after the model reports an error, which not only affects the stability and reliability of the system, but also increases the difficulty and time cost of troubleshooting. In response to this, the embodiment of the present invention can detect interconnection link failures in a timely manner and take corresponding measures by the server side monitoring the device link status of each client in real time, thereby overcoming the above - mentioned defects.

[0068] Specifically, the device connection status of the client refers to the current connection status of all GPUs (i.e., interconnected devices) connected to the external switch box on the client. This includes status information such as whether the device is online, whether the connection is stable, and whether there are communication failures. The device connection status is one of the key factors to ensure the smooth progress of distributed training tasks, because the disconnection or communication failure of any device may cause the interruption or failure of the training task.

[0069] Each client can regularly read the connection status of local interconnected devices through a built-in status monitoring module (such as a GPU management library). Once a status change is detected, the client will immediately encapsulate this information into a specific message format and report it to the server through a network communication protocol (such as TCP / IP). The server will receive and process these messages to update the external interconnected topology information and device connection status of the devices in real time.

[0070] When the server monitors that the device connection status of any client is abnormal, it indicates that there is a connection problem with one or some of the interconnected devices on that client (for example, a certain link is disconnected). This may be caused by network failures, device failures, configuration errors, etc. The occurrence of the abnormal status will directly affect the smooth progress of the distributed training task, so the server needs to immediately take measures to restore the device connection status.

[0071] After the server monitors the abnormal device connection status, it will determine the problematic client and interconnected devices based on the external interconnected topology information and device connection status of the devices. Then, the server will generate a link recovery instruction, which contains all the necessary information required to restore the device connection (such as device identification, connection parameters, etc.). Finally, the server will send the link recovery instruction to the problematic client through the network communication protocol.

[0072] It can be understood that the link recovery instruction is an instruction generated by the server to guide the client to restore the device connection status. It contains all the necessary information required to restore the device connection, such as device identification, connection parameters, number of retries, etc. After receiving the link recovery instruction, the client will re-establish the connection with the interconnected device according to the information in it.

[0073] After receiving the link recovery instruction, the client will immediately start to perform the operation of restoring the device connection. For example, it will reconfigure the network link parameters, restart the device driver, and try to re-establish the connection, etc. Once the device connection is successfully restored, the client will immediately encapsulate the result status of the reconnection into a specific message format and report it to the server through the network communication protocol. The server will receive and process these messages to update the external interconnected topology information and device connection status of the devices.

[0074] Here, the result status of re-linking refers to the result obtained by the client after attempting to re-establish the device link, which can include information such as whether the link is successful, whether there are any errors or warnings, etc. The result status of re-linking is crucial for the server side because it can help the server side determine whether the device connection has been fully restored and accordingly decide whether to continue executing the distributed training task.

[0075] In the embodiments of the present invention, by steps such as the server side monitoring the device link status of each client, sending a link recovery instruction, re-establishing the device link, and reporting the result status, it is ensured that the distributed training task can quickly resume and continue to execute when the device link is abnormal, greatly improving the reliability and stability of the distributed training task.

[0076] Based on any of the above embodiments, the method further includes: Receiving a mode switching instruction; In response to the mode switching instruction, determining a new training mode, and based on the new training mode, as well as the device link information and switch port information of each client, reallocating the memory area identifier and group identifier to generate new device external interconnection topology information; Sending the new device external interconnection topology information to each client, so that each client reconfigures the topology of each local interconnection device based on the new device external interconnection topology information.

[0077] Specifically, the mode switching instruction refers to an instruction generated by a user or a system for instructing the server side to change the current distributed training mode. In distributed training, different training modes may require different device configurations and connection methods. For example, TP8 (tensor parallel training mode with 8 GPUs) and TP16 (tensor parallel training mode with 16 GPUs) may require different device interconnection topologies to optimize the training performance and efficiency. Therefore, the mode switching instruction is used to trigger the server side to regenerate the device external interconnection topology information according to the new training mode and notify each client to make corresponding configuration changes. It should be understood that the server side usually receives the corresponding mode switching instruction according to different model training tasks.

[0078] It can be understood that the server side can receive the mode switching instruction through a specific interface or API. For example, this interface can be an HTTP-based RESTful API or an interface of other network communication protocols (such as TCP / IP, etc.). The user or the system can send the mode switching instruction by calling this interface and passing the corresponding parameters (such as the identifier of the new training mode). After receiving the instruction, the server side will parse the parameters and perform corresponding operations according to the instruction.

[0079] Specifically, after receiving the mode switching instruction, the server side will first parse the instruction to determine the new training mode. Then, based on the new training mode, as well as the current latest device connection information and switch port information of each client, the server side will reassign memory identifiers to each interconnected device of each client and determine the group identifiers of each interconnected device. Next, based on the newly assigned memory area identifiers and group identifiers, new external device interconnection topology information will be generated and sent to each client. This information describes the connection relationship and configuration requirements for high-speed interconnection between all interconnected devices of each client through an external switch in the new training mode.

[0080] After each client receives the new external device interconnection topology information, it will parse it to extract useful information. Based on the parsed information, each client can reconfigure the local interconnection topology of each device to switch to running in the new training mode, without the need to restart the machine or the corresponding software service in the middle, realizing the online switching of the training mode and meeting the requirements of model training.

[0081] Based on any of the above embodiments, Figure 3 is the second flowchart of the topology management method provided by the present invention. As Figure 3 shown, this method is applied to the client side, and this method includes: Step 310, obtain the local device connection information and the switch port information of the switches connected to each local interconnected device, where the interconnected device refers to the device connected to the switch box; Step 320, report the device connection information and the switch port information to the server side, so that the server side generates external device interconnection topology information based on the model distributed training mode and the device connection information and switch port information of each client.

[0082] It should be noted that the method provided by the embodiments of the present invention can be applied to the client side. Here, the client side refers to a computer or program that requests services in the network, and the server side refers to a computer or program that provides services in the network. For example, in the scenario of multi-machine and multi-GPU interconnection, assuming that there are 4 machines (machines 0 to 3) interconnected, that is, multiple GPUs on these 4 machines are interconnected at high speed through the PCIE switch in the switch box, one of them can be selected as both the server side and the client side, and the other 3 are only used as client sides.

[0083] Specifically, each client can read the local device link information and switch port information through the device management library on the machine where it is located. Then, it can use a specific API or message format to encapsulate this information and send the encapsulated device link information and switch port information to the server side through a network communication protocol (such as TCP / IP). Here, the device management library refers to a software library deployed on the machine for managing and monitoring various devices on the machine. For example, the device management library can be the GPU Management Library (i.e., the GPU management library).

[0084] It can be understood that the device link information refers to the number of all interconnected devices (such as GPUs) on the machine where the client is located and the link information of each interconnected device, etc. The switch port information refers to the information of the switch ports used when all interconnected devices on the machine where the client is located are interconnected with external devices (such as GPUs of other machines) through a switch box. Here, the interconnected devices refer to the devices on the machine where the client is located that are connected to external devices (such as GPUs of other machines) through a switch box, and these interconnected devices are the objects that need to be considered when the server side generates the device external interconnection topology information. It should be understood that the link information of each interconnected device refers to a set of parameters that can uniquely identify and describe the location and connection status of the device in the network. For example, taking the interconnected device as a GPU, the link information of each interconnected device can include the domain identifier, bus identifier, device identifier, and GPU ID of the device on the PCIE bus, etc. The embodiments of the present invention do not make specific limitations on this.

[0085] After receiving the device link information and switch port information from each client, the server side will construct a device external interconnection topology information that describes the interconnection relationship between all interconnected devices according to the requirements of the model distributed training mode and the device link information and switch port information of each client. Here, the device external interconnection topology information refers to an interconnection relationship data structure that describes the high-speed interconnection between all interconnected devices through an external switch box, and it can include key information such as the number of nodes (or the number of machines, each machine is a node), the unique identifier of each node, and the set of external links of each node. Among them, the set of external links of each node can include information such as the identifier, location, and connection relationship with other devices of each interconnected device on the node.

[0086] Specifically, based on the model distributed training mode selected by the user (such as any one of TP8, TP16, and TP32), the server side can parse out the specific requirements for the external interconnection relationship of devices (such as GPUs). For example, the TP8 mode may require dividing 8 GPU devices into two groups, with 4 GPU devices in each group for data parallel training; while the TP16 mode may require dividing 16 GPU devices into four groups, with 4 GPU devices in each group for model parallel training.

[0087] Based on the received device link information, switch port information, and the parsed requirements of the model distributed training mode, the server side starts to construct the external interconnection topology information of the devices. During the construction process, the server side assigns a unique node identifier (such as self_node_id) to each node (i.e., each machine), and creates a NodeConn array to store the set of external links of the node. For each link in the NodeConn array, the server side records information such as the source GPU device ID of the link (denoted as src_gpu_id), the destination node ID (denoted as dst_node_id), the destination GPU device ID (denoted as dst_gpu_id), the switch port address of the destination GPU (denoted as dst_port), and the link type (denoted as link_type). Based on the above information, the external interconnection topology information of the devices can be constructed.

[0088] Step 330: Receive the external interconnection topology information of the devices sent by the server side, parse the external interconnection topology information of the devices, and perform topology configuration on each interconnected device based on the parsing result.

[0089] Specifically, after generating the external interconnection topology information of the devices, the server side can also send the external interconnection topology information of the devices to each client through the network communication protocol. For example, the server side can use the same API or message format as the client to encapsulate the external interconnection topology information of the devices and send it to the listening ports of each client through the network.

[0090] After each client receives the external interconnection topology information of the devices, it will parse it to extract useful information. For example, operations such as decoding, format conversion, and extraction can be performed on the external interconnection topology information of the devices. The parsed information will be used in the subsequent topology configuration process. It should be understood that the parsing result refers to the useful information obtained by the client after parsing the external interconnection topology information of the devices. This information usually includes the identifiers, locations, and connection relationships with other devices of each interconnected device. This information will be used as the basis for the client to perform topology configuration.

[0091] Based on the information in the parsing result, the client can perform topology configuration on each local interconnected device. Specifically, the client can configure network parameters such as the routing table of each local interconnected device, as well as the connection relationships and communication protocols between devices. Through topology configuration, the client can ensure that each local interconnected device can be interconnected and communicate with other devices according to the requirements in the external interconnected topology information of the device.

[0092] It can be understood that performing topology configuration means configuring the network parameters and connection relationships of each local interconnected device according to the requirements in the external interconnected topology information of the device. Through topology configuration, it can be ensured that each device in the distributed training task can be correctly interconnected and communicate, so as to achieve efficient distributed training.

[0093] In the method provided by the embodiments of the present invention, by automatically reporting the local device link information and switch port information by each client, and the server side automatically generating the external interconnected topology information of the device and distributing it to each client, each client can perform topology configuration on each local interconnected device according to the external interconnected topology information of the device, realizing automatic configuration, without manually executing different scripts for configuration on each client, greatly simplifying the configuration process and improving the configuration efficiency. In addition, by generating the external interconnected topology information by the server side, centralized management of the external interconnected topology of all interconnected devices can be realized, reducing the management complexity.

[0094] It should be noted that other embodiments of the topology management method applied to the client of the present invention can refer to the respective embodiments of the topology management method applied to the server side, which will not be elaborated here.

[0095] Based on any of the above embodiments, in step 330, the parsing of the external interconnected topology information of the device and the topology configuration of each interconnected device based on the parsing result include: Step 331, parse the external interconnected topology information of the device to obtain the memory area identifier and group identifier of each local interconnected device.

[0096] It should be noted that when the server side generates the external interconnected topology information of the device, in addition to generating information such as the connection relationship between each interconnected device and other devices, it will also assign a memory area identifier and group identifier to each interconnected device, etc., further ensuring that the communication between all interconnected devices is more efficient and stable. Therefore, the generated external interconnected topology information of the device can also include the memory area identifier and group identifier of each interconnected device.

[0097] Specifically, after receiving the device external interconnection topology information sent by the server, the client can parse it, convert the device external interconnection topology information into a data structure that can be understood by the device management library, and then transparently transmit it to the kernel driver through the device management library, and the kernel driver performs topology configuration. Here, It can be understood that the parsed information includes the memory area identifier and group identifier of each interconnected device. Among them, the memory area identifier is a unique identifier used to identify and access the device memory space during the distributed training process. It ensures that each device has an independent memory area and avoids conflicts when accessing remote devices. The group identifier (i.e., clique id) is a unique identifier used to represent the device group relationship. Interconnected devices with the same group identifier can communicate with each other.

[0098] Step 332: Based on the device management library, transparently transmit the memory area identifier and group identifier of each interconnected device to the kernel driver, so that the kernel driver performs topology configuration based on the memory area identifier and group identifier of each interconnected device.

[0099] Specifically, the client first extracts the memory area identifier and group identifier of each local interconnected device from the device external interconnection topology information. These information have been allocated and determined by the server when generating the device external interconnection topology information. Then, the client calls the API or interface provided by the device management library and transmits the extracted information (such as the memory area identifier and group identifier) to the device management library.

[0100] After receiving the information, the device management library communicates with the kernel driver. This is usually achieved through the kernel-mode interface or specific communication mechanism provided by the device management library. The device management library transmits the information to the kernel driver so that the kernel driver can perform topology configuration on the interconnected devices based on this information. After receiving the device information, the kernel driver configures all the interconnected devices on the local machine according to this information. For example, set the memory area identifier and group identifier of each interconnected device into the hardware register.

[0101] Based on any of the above embodiments, an embodiment of the present invention provides a topology management system, including a server, a communication module, and at least one client; The client is used to obtain the local device link information and the switch port information linked to each local interconnected device, and transmit the device link information and the switch port information to the server through the communication module. The interconnected device refers to a device connected to the switch cabinet; The server is used to generate external device interconnection topology information based on the device connection information and switch port information of each client, and send the external device interconnection topology information to each client through the communication module; The client is also used to parse the external device interconnection topology information and perform topology configuration on each local interconnected device based on the parsing result.

[0102] Specifically, the system provided by the embodiments of the present invention mainly includes a server, a communication module and a client. Among them, the server manages the global multi-machine and multi-device interconnection topology configuration, such as the GPU topology link configuration of 2 machines with 16 cards (i.e., 16 GPU devices on 2 machines) or 4 machines with 32 cards (i.e., 32 GPU devices on 4 machines). The communication module realizes the communication and information interaction between each client and the server through the HTTP protocol.

[0103] Figure 4 is a schematic diagram of the overall process of multi-machine and multi-GPU interconnection topology configuration and monitoring provided by the present invention. As Figure 4 shown, the client reads the device connection information of the local machine and the switch port information of the local connection through the GPU management library, and reports them to the server through the communication module. After receiving this information, the server combines the requirements of the model distributed training mode to ensure that there is no conflict when accessing remote GPUs, assigns a memory area identifier to the interconnected GPUs on each client's machine, and defines different group identifiers to organize them into a multi-GPU interconnection topology information. At the same time, each client will receive the multi-GPU interconnection topology information configured by the server, and convert this information into topology configuration information (or topology structure) that can be understood by the GPU management library, and pass it through the GPU management library to the GPU KMD (i.e., kernel driver) for setting. In addition, the client will regularly read the device connection information of the local machine through the GPU management library as heartbeat information and report it to the server, and at the same time regularly read the device connection status of the local machine and report the connection status to the server. When the server discovers that a certain device is disconnected, it sends a notification to the client through the communication module, and then the client issues a reset P2P port instruction (i.e., link recovery instruction) to the GPU KMD through the GPU management library, and the KMD attempts to re-establish the link and reports the result status of the reconnection.

[0104] It can be understood that the above method is based on the hardware architecture of multi-machine and multi-GPU interconnection and the underlying software hierarchical structure (such as GPU management library, GPU KMD, etc.) implemented by an external switch cabinet, providing tools for flexible online configuration of multi-machine and multi-GPU interconnection topologies and real-time monitoring of the link status of multiple GPUs for different training models. It realizes online configuration of GPU interconnection topologies, real-time monitoring of GPU interconnection status, flexible switching between different training modes, meets the topology configuration requirements of different model training, and improves efficiency.

[0105] Based on any of the above embodiments, Figure 5 is a schematic diagram of the deployment of the server side and the client side provided by the present invention. As Figure 5 shown, the deployed software package mainly includes two services, Fabric server and Fabric client. Suppose there are 4 interconnected servers (i.e., server 0 to server 3, and each server is a machine). After the basic software packages of the GPU management library and GPU KMD are installed, the Fabric server and Fabric client software packages can be installed in the 4 interconnected servers. Here, the GPU management library is deployed in the user space, while the GPU KMD is deployed in the kernel space. Among them, the user space and the kernel space are two different running environments in the Linux system, which represent different privilege levels and functions respectively. The user space is responsible for running ordinary applications, while the kernel space provides underlying support and system services. The two achieve efficient cooperation through system calls, jointly ensuring the stability and security of the system.

[0106] Select one of the servers (such as server 0) to install both Fabric server and Fabric client (i.e., acting as both the server side and the client side at the same time), and the other servers only install Fabric client (i.e., only acting as the client side). Start the services respectively through corresponding commands to configure the multi-GPU interconnection topology. Subsequently, Fabric server, as a monitoring program, real-time monitors the link status of the multi-GPU external interconnection. When the link status of a certain GPU is abnormal, it attempts to recover the real-time link, and Fabric client reports the status of the attempt to re-link. When a new model training task requires switching the training mode, run the command to switch the mode on the server installed with Fabric server. Fabric server receives the configuration requirements of the new training mode, reconfigures the multi-GPU interconnection topology, and sends the configuration topology command to each Fabric client (i.e., sends the multi-GPU interconnection topology information to each client), and switches to run in the new training mode. There is no need to restart the machine and no need to restart the software service in the middle, achieving online switching and meeting the requirements of model training.

[0107] The topological management device provided by the present invention will be described below. The topological management device described below can be mutually referred to the topological management method described above.

[0108] Based on any of the above embodiments, Figure 6 is one of the schematic structural diagrams of the topological management device provided by the present invention. As Figure 6 shown, the device is applied to the server side, and the device includes: A receiving unit 610, configured to receive device link information and switch port information reported by each client; A generating unit 620, configured to generate device external interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; A sending unit 630, configured to send the device external interconnection topology information to each client, so that each client parses the device external interconnection topology information, and performs topology configuration on each local interconnection device based on the parsing result, where the interconnection device refers to a device connected to the switch cabinet.

[0109] In the device provided by the embodiment of the present invention, by automatically reporting local device link information and switch port information by each client, the server side automatically generates device external interconnection topology information and sends it to each client, so that each client can perform topology configuration on each local interconnection device according to the device external interconnection topology information, realizing automatic configuration, without manually executing different scripts for configuration on each client, greatly simplifying the configuration process and improving the configuration efficiency. In addition, by generating device external interconnection topology information on the server side, centralized management of the external interconnection topology of all interconnection devices can be realized, reducing the management complexity.

[0110] Based on any of the above embodiments, the generating unit 620 is specifically configured to: Based on the model distributed training mode and the device link information and switch port information of each client, allocate memory area identifiers for each interconnection device on each client, and determine the group identifiers of each interconnection device; Based on the device link information and switch port information of each client, and the memory area identifiers and group identifiers of each interconnection device on each client, generate device external interconnection topology information.

[0111] Based on any of the above embodiments, the device link information of any client includes the link information of each interconnection device on the any client, and the link information of each interconnection device includes a domain identifier, a bus identifier, a device identifier, a function identifier, and a device number.

[0112] Based on any of the above embodiments, the device further includes a monitoring unit, and the monitoring unit is configured to: Monitor the device connection status of each client, where each client periodically reads the connection status of each interconnected device locally and reports it; In the case where the device connection status of any client is monitored to be abnormal, send a connection recovery instruction to the any client, so that the any client re - establishes a device connection based on the connection recovery instruction and reports the result status of the re - connection.

[0113] Based on any of the above embodiments, each client periodically reads the device connection information locally and sends the device connection information to the server as heartbeat information.

[0114] Based on any of the above embodiments, the device further includes a switching unit, and the switching unit is configured to: Receive a mode switching instruction; In response to the mode switching instruction, determine a new training mode, and based on the new training mode, the device connection information of each client, and the switch port information, re - allocate the memory area identifier and the group identifier to generate new device external interconnection topology information; Send the new device external interconnection topology information to each client, so that each client re - configures the topology of each local interconnected device based on the new device external interconnection topology information.

[0115] Based on any of the above embodiments, Figure 7 is the second structural schematic diagram of the topology management device provided by the present invention. As Figure 7 shown, the device is applied to a client, and the device includes: An acquisition unit 710, configured to acquire the device connection information locally and the switch port information of the switches connected to each local interconnected device, where the interconnected device refers to a device connected to a switch box; A reporting unit 720, configured to report the device connection information and the switch port information to the server, so that the server generates device external interconnection topology information based on the model distributed training mode, the device connection information of each client, and the switch port information; A configuration unit 730, configured to receive the device external interconnection topology information sent by the server, parse the device external interconnection topology information, and perform topology configuration on each interconnected device based on the parsing result.

[0116] The device provided by the embodiment of the present invention enables each client to automatically report local device link information and switch port information, and the server side automatically generates device external interconnection topology information and distributes it to each client, so that each client can perform topology configuration on local interconnected devices according to the device external interconnection topology information, realizing automatic configuration, without manually executing different scripts for configuration on each client, greatly simplifying the configuration process and improving the configuration efficiency. In addition, by generating device external interconnection topology information on the server side, centralized management of the external interconnection topology of all interconnected devices can be realized, reducing the management complexity.

[0117] Based on any of the above embodiments, the configuration unit 730 is specifically configured to: Parse the device external interconnection topology information to obtain the memory area identifiers and group identifiers of local interconnected devices; Based on the device management library, pass the memory area identifiers and group identifiers of the interconnected devices to the kernel driver transparently, so that the kernel driver performs topology configuration based on the memory area identifiers and group identifiers of the interconnected devices.

[0118] Figure 8 An entity structure diagram of an electronic device is exemplified, as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute a topology management method, which is applied to the server side. The method includes: receiving the device link information and switch port information reported by each client; generating device external interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; and distributing the device external interconnection topology information to each client, so that each client parses the device external interconnection topology information and performs topology configuration on local interconnected devices based on the parsing result. The interconnected devices refer to the devices connected to the switch cabinet.

[0119] The processor 810 can also call the logical instructions in the memory 830 to execute a topology management method, which is applied to the client side. The method includes: obtaining the local device link information and the switch port information of each interconnected device linked to the local area, where the interconnected device refers to the device connected to the switch cabinet; reporting the device link information and the switch port information to the server side, so that the server side generates the device external interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; receiving the device external interconnection topology information sent by the server side, parsing the device external interconnection topology information, and performing topology configuration on each interconnected device based on the parsing result.

[0120] In addition, when the logical instructions in the above-mentioned memory 830 are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the related technology, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs and other various media that can store program codes.

[0121] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the topology management method provided by the above-mentioned various methods. This method is applied to the server side and includes: receiving the device link information and switch port information reported by each client; generating the device external interconnection topology information based on the model distributed training mode and the device link information and switch port information of each client; sending the device external interconnection topology information to each client, so that each client parses the device external interconnection topology information and performs topology configuration on each local interconnected device based on the parsing result. The interconnected device refers to the device connected to the switch cabinet.

[0122] The computer is also capable of executing the topology management method provided by each of the above methods. This method is applied to the client side and includes: obtaining the local device link information and the switch port information of each interconnected device linked to the local area, where the interconnected device refers to the device connected to the switch cabinet; reporting the device link information and the switch port information to the server side, so that the server side generates the external interconnected topology information of the device based on the model distributed training mode and the device link information and switch port information of each client; receiving the external interconnected topology information of the device sent by the server side, parsing the external interconnected topology information of the device, and performing topology configuration on each interconnected device based on the parsing result.

[0123] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the topology management method provided by each of the above methods. This method is applied to the server side and includes: receiving the device link information and switch port information reported by each client; generating the external interconnected topology information of the device based on the model distributed training mode and the device link information and switch port information of each client; sending the external interconnected topology information of the device to each client, so that each client parses the external interconnected topology information of the device and performs topology configuration on each local interconnected device, where the interconnected device refers to the device connected to the switch cabinet.

[0124] When the computer program is executed by a processor, it is implemented to execute the topology management method provided by each of the above methods. This method is applied to the client side and includes: obtaining the local device link information and the switch port information of each interconnected device linked to the local area, where the interconnected device refers to the device connected to the switch cabinet; reporting the device link information and the switch port information to the server side, so that the server side generates the external interconnected topology information of the device based on the model distributed training mode and the device link information and switch port information of each client; receiving the external interconnected topology information of the device sent by the server side, parsing the external interconnected topology information of the device, and performing topology configuration on each interconnected device based on the parsing result.

[0125] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0126] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0127] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A topology management method, characterized in that, The method is applied to the server side, and the method includes: Receiving device connection information and switch port information reported by each client; Generating device external interconnection topology information based on the model distributed training mode and the device connection information and switch port information of each client; Sending the device external interconnection topology information to each client, so that each client parses the device external interconnection topology information and performs topology configuration on each local interconnection device based on the parsing result, where the interconnection device refers to a device connected to the switch cabinet.

2. The topology management method according to claim 1, characterized in that The generating device external interconnection topology information based on the model distributed training strategy and the device connection information and switch port information of each client includes: Based on the model distributed training mode and the device connection information and switch port information of each client, allocating memory area identifiers to each interconnection device on each client and determining the group identifiers of each interconnection device; Generating device external interconnection topology information based on the device connection information and switch port information of each client and the memory area identifiers and group identifiers of each interconnection device on each client.

3. The topology management method according to claim 1, wherein The device connection information of any client includes the connection information of each interconnection device on the any client, and the connection information of each interconnection device includes a domain identifier, a bus identifier, a device identifier, a function identifier, and a device number.

4. The topology management method according to any one of claims 1 to 3, characterized in that It further includes: Monitoring the device connection status of each client, where each client periodically reads the connection status of each local interconnection device and reports it; In the case of monitoring that the device connection status of any client is abnormal, sending a connection recovery instruction to the any client, so that the any client re - establishes a device connection based on the connection recovery instruction and reports the result status of the re - connection.

5. The topology management method according to any one of claims 1 to 3, characterized in that It further includes: Each client periodically reads the local device connection information and sends the device connection information to the server side as heartbeat information.

6. The topology management method according to any one of claims 1 to 3, characterized in that, It further includes: Receiving a mode switching instruction; In response to the mode switching instruction, determining a new training mode, and re - allocating memory area identifiers and group identifiers based on the new training mode and the device connection information and switch port information of each client to generate new device external interconnection topology information; Sending the new device external interconnection topology information to each client, so that each client re - performs topology configuration on each local interconnection device based on the new device external interconnection topology information.

7. A topology management method, characterized in that, The method is applied to the client side, and the method includes: Obtaining local device connection information and switch port information of switches connected to each local interconnection device, where the interconnection device refers to a device connected to the switch cabinet; Reporting the device connection information and the switch port information to the server side, so that the server side generates device external interconnection topology information based on the model distributed training mode and the device connection information and switch port information of each client. Receive the device external interconnection topology information sent by the server side, parse the device external interconnection topology information, and perform topology configuration on each interconnection device based on the parsing result.

8. The topology management method according to claim 7, characterized in that The parsing of the device external interconnection topology information and the topology configuration of each interconnection device based on the parsing result include: Parse the device external interconnection topology information to obtain the memory area identifiers and group identifiers of each local interconnection device; Based on the device management library, transmit the memory area identifiers and group identifiers of each interconnection device to the kernel driver, so that the kernel driver performs topology configuration based on the memory area identifiers and group identifiers of each interconnection device.

9. A topology management system, characterized in that, It includes a server side, a communication module, and at least one client; The client is used to obtain the local device link information and the switch port information linked to each local interconnection device, and transmit the device link information and the switch port information to the server side through the communication module. The interconnection device refers to the device connected to the switch cabinet; The server side is used to generate device external interconnection topology information based on the device link information and switch port information of each client, and send the device external interconnection topology information to each client through the communication module; The client is also used to parse the device external interconnection topology information and perform topology configuration on each local interconnection device based on the parsing result.

10. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the topology management method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the topology management method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the topology management method according to any one of claims 1 to 8.

Citation Information

Cited By

  • Method and device for constructing API topological relation graph, equipment and medium

    CN121350516A

  • A method, apparatus, device, and medium for constructing API topology graphs

    CN121350516B