Data transmission method, system, device and storage medium
Patent Information
- Application Number
- CN202610786821.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]然而,随着大模型的数据规模增大,节点对共享存储器的读写操作会导致模型训练的时间开销增加,且可能造成共享存储器的网络带宽饱和,进而导致模型训练的训练效率下降
通过管理节点向多个计算节点中的目标计算节点发送数据同步指令,目标计算节点可以根据上述数据同步指令,将其本地存储器中存储的模型数据直接发送至新节点的本地存储器。上述过程中,目标计算节点无需将模型数据写入共享存储器,新节点也无需从共享存储器中读取模型数据,避免了对于共享存储器的读写操作,从而减少了时间开销。同时,由于数据模型直接在目标计算节点与新节点之间传输,避免了多个节点同时读写共享存储器所导致的网络带宽饱和。综上,本申请实施例提供的技术方案,提高了模型训练的训练效率。
Smart Images

Figure CN122802514A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a data transmission method, system, device, and storage medium. Background Technology
[0002] In the field of artificial intelligence, large model training is typically performed in a distributed training system containing multiple nodes, and the number of nodes can vary. In related technologies, when a new node is waiting to join, all currently training nodes are paused, and the model data of those nodes is written to shared memory. Subsequently, the new node reads the model data from the shared memory and stores it in its corresponding local memory for subsequent model training.
[0003] However, as the data scale of large models increases, the read and write operations of nodes on shared memory will increase the time overhead of model training and may cause network bandwidth saturation of shared memory, thereby reducing the training efficiency of model training. Summary of the Invention
[0004] This application provides a data transmission method, system, device, and storage medium. The technical solutions provided by this application are as follows.
[0005] According to one aspect of the embodiments of this application, a data transmission method for a distributed training system is provided, the distributed training system including a management node and multiple computing nodes, the method comprising: In the case that there are new nodes waiting to be added in the distributed training system, the management node sends a data synchronization instruction to the target computing node among the multiple computing nodes. The data synchronization instruction is used to instruct the model data stored on the target computing node to be synchronized to the new node. According to the data synchronization instruction, the target computing node sends the model data stored in its local memory to the local memory of the new node.
[0006] According to one aspect of the embodiments of this application, a distributed training system is provided, the distributed training system including a management node and multiple computing nodes: The management node is used to send a data synchronization instruction to the target computing node among the multiple computing nodes when there are new nodes waiting to be added in the distributed training system. The data synchronization instruction is used to instruct the model data stored on the target computing node to be synchronized to the new node. The target computing node is configured to send the model data stored in its local memory to the local memory of the new node according to the data synchronization instruction.
[0007] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the above-described data transmission method.
[0008] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the above-described data transmission method.
[0009] According to one aspect of the embodiments of this application, a chip is provided, the chip including programmable logic circuits and / or program instructions, which, when the chip is running, are used to implement the above-described data transmission method.
[0010] According to one aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program, the computer program being loaded and executed by a processor to implement the above-described data transmission method.
[0011] The technical solution provided in this application can bring the following beneficial effects: By sending data synchronization commands from the management node to the target computing node among multiple computing nodes, the target computing node can directly send the model data stored in its local memory to the local memory of the new node according to the data synchronization commands. In this process, the target computing node does not need to write model data to shared memory, and the new node does not need to read model data from shared memory, avoiding read and write operations on shared memory and thus reducing time overhead. Simultaneously, since the data model is directly transmitted between the target computing node and the new node, network bandwidth saturation caused by multiple nodes simultaneously reading and writing to shared memory is avoided. In summary, the technical solution provided by the embodiments of this application improves the training efficiency of model training. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of a distributed training system provided in one possible implementation of this application; Figure 2 This is a flowchart of a data transmission method provided in one possible implementation of this application; Figure 3This is a schematic diagram of the node ready process provided in one possible implementation of this application; Figure 4 This is a schematic diagram illustrating the determination of the target computing node in one possible implementation of this application; Figure 5 This is a schematic diagram of a data transmission process provided in one possible implementation of this application; Figure 6 This is a flowchart of a data transmission method provided in another possible implementation of this application; Figure 7 This is a block diagram of the distributed training system provided in one possible implementation of this application; Figure 8 This is a structural block diagram of a computer device provided in one possible implementation of this application. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] Before introducing the technical solution proposed in this application, the relevant technical background and related technologies will be briefly described below.
[0016] With the rapid development of artificial intelligence (AI) technology, the number of parameters in deep learning models, represented by large language models (LLM), has reached hundreds of billions or even trillions. In order to complete the training of large-scale models within a reasonable time, it is usually necessary to use computing clusters containing a large number of computing nodes (such as graphics processing unit (GPU) accelerator cards) and to adopt distributed training strategies such as data parallelism and model parallelism.
[0017] In practical applications of distributed training, training tasks often require "elastic training" capabilities, meaning the number of computing nodes participating in the computation can be dynamically increased or decreased during the training process. This requirement stems primarily from two aspects: first, to improve the utilization of cluster resources, training tasks can dynamically expand by adding new computing nodes when resources are idle to accelerate training; second, to cope with computing node failures. Specifically, training tasks need to dynamically remove computing nodes when they fail to ensure the continuity of training. Therefore, how to efficiently and reliably synchronize model data in scenarios where the number of nodes changes dynamically has become a pressing technical problem in the field of distributed training.
[0018] In related technologies, when a new node is waiting to join in a distributed training system, all currently training computing nodes are paused, and the complete model data held by these nodes is written to shared storage, such as a distributed file system. Subsequently, the new node reads the model data from this shared storage and loads it into its local storage. Finally, the system resumes the training tasks of all nodes (including existing computing nodes and the new node) and continues subsequent model training.
[0019] In some embodiments, the above process can be implemented as follows. Exemplarily, a computing node refers to a physical or virtual device that actually performs the model training task. Each computing node is configured with at least one GPU, whose video memory serves as the computing node's local storage for storing model data. A new node refers to a new computing node waiting to join the distributed training system. The new node does not yet possess the currently trained model data and needs to obtain the model data from shared memory before participating in training. Based on the above definitions, the specific implementation steps of this embodiment are as follows: When a new node is waiting to join in a distributed training system, the system first pauses all currently training compute nodes. These nodes were previously working collaboratively in a data-parallel manner, each storing complete model data in its local memory. After all compute nodes have finished pausing, the system writes the model data stored in the compute nodes' local memory to a shared memory accessible to all nodes. For example, the shared memory can be a distributed file system (such as HDFS (Hadoop Distributed File System)), a network file system (NFS), or other types of shared storage devices. Subsequently, the new node reads the previously written model data from the shared memory and stores the read model data in its own local memory (e.g., the new node's GPU memory). Finally, the system resumes the training task for all nodes (including existing compute nodes and the newly joined node). The new node can then participate in subsequent model training along with other compute nodes based on the model data in its local memory.
[0020] However, the aforementioned technologies have the following problems: As the data scale of large models increases, the volume of model data can reach the terabyte (TB) level. In this case, node read / write operations on shared memory generate significant disk input / output (I / O) overhead, leading to increased model training time. Simultaneously, when multiple computing nodes read and write data to shared memory concurrently, a large amount of network traffic is generated instantaneously, easily causing network bandwidth saturation of the shared memory. Furthermore, related technologies are often unaware of the physical network topology during data transmission; data transmission paths are often random, potentially traversing core layer switches or even different data center racks, further exacerbating network congestion and transmission latency. In summary, all these factors contribute to a decrease in model training efficiency.
[0021] To address the aforementioned problems, this application provides a data transmission method for distributed training systems. Before introducing the data transmission method, the distributed training system provided in this application will be described first.
[0022] In some embodiments, the distributed training system includes a management node and multiple computing nodes.
[0023] In one aspect of this embodiment, a management node refers to a control device or control process responsible for coordinating and managing the entire distributed training system. Exemplarily, the main functions of a management node may include: maintaining the status information of each computing node in the system, sensing the physical network topology of the cluster, deciding on the optimal target computing node for new nodes, sending control commands (such as pause, synchronization, resumption, etc.) to each computing node, and triggering fault-tolerance mechanisms in abnormal situations. It is worth noting that the specific form of a management node can be diverse; as long as it can achieve the coordination and management functions of each computing node in the distributed training system, it falls within the protection scope of this application. Exemplarily, the following describes several possible forms of a management node: 1. Independent Deployment: In some embodiments, the management node can be deployed independently on a dedicated physical server. This physical server does not participate in specific model training computation tasks, but is only responsible for executing management functions. The advantages of independent deployment are: resource isolation between the management node and the computing nodes, the execution of management tasks will not be interfered with by training tasks, and the failure of the management node will not directly affect the ongoing training computation.
[0024] 2. Shared Deployment: In some embodiments, the management node can share the same physical device with one or more compute nodes. For example, in scenarios with a small cluster size or limited resources, the management node's processes can be deployed on one of the compute nodes, which then handles both management and training computation functions. The advantages of shared deployment are: saving hardware resources, reducing deployment costs, and suitability for development and testing environments or small-scale production environments.
[0025] 3. Clustered Deployment: In some embodiments, management nodes can be deployed in a clustered manner to achieve high availability. For example, in an Active-Standby mode: one primary management node handles all control requests, and one or more standby management nodes synchronize the status information of the primary node. When the primary node fails, the standby nodes automatically take over management responsibilities. Another example is an Active-Active mode: multiple management nodes provide services simultaneously, using a distributed coordination service for leader election and status synchronization, achieving load balancing and fault tolerance. Clustered deployment is suitable for large-scale production environments and can meet the business requirements of high concurrency and high availability.
[0026] It is worth noting that the above methods can be used individually or in combination. In summary, this application does not limit the specific form of the management node, and any deployment method that can realize the function of the management node falls within the protection scope of this application.
[0027] In one aspect of this embodiment, a computing node refers to a physical or virtual device that actually performs the model training task. For example, each computing node may be configured with at least one GPU, whose video memory can serve as the computing node's local storage for storing model data (such as model weight parameters, gradients, etc.). In data-parallel training mode, each computing node can hold a complete copy of the model and collaboratively optimize the model by computing different batches of data. It should be noted that the specific forms of computing nodes are diverse; several possible forms of computing nodes are described below: 1. Physical Server Approach: In some embodiments, the compute node can be a single-GPU server, meaning one physical server is configured with one GPU accelerator card. This configuration is simple and suitable for small-scale training tasks or prototype verification scenarios. In other embodiments, the compute node can also be a multi-GPU server, meaning one physical server is configured with multiple GPU accelerator cards (e.g., 4, 8, or even more). With this configuration, multiple GPUs within the same server can communicate via high-speed interconnect technology (such as PCIe (Peripheral Component Interconnect express, a high-speed serial computer expansion bus standard)) to achieve model parallelism or data parallelism within the server. Multi-GPU servers offer stronger computing power and are suitable for large-scale model training tasks.
[0028] 2. Virtualization: In some embodiments, the computing node can be a virtual machine instance. Through virtualization technology, a physical server can be divided into multiple virtual machines, each participating in distributed training as an independent computing node. The advantages of virtualization are: good resource isolation, enabling secure isolation between different training tasks; and the ability for virtual machines to be easily migrated and snapshotted, facilitating elastic scheduling of training tasks.
[0029] 3. Heterogeneous Computing Node Approach: In some embodiments, computing nodes can be configured with different types of accelerators. Besides GPUs, computing nodes can also be configured with other types of AI acceleration chips, such as TPUs (Tensor Processing Units), NPUs (Neural Processing Units), and FPGAs (Field-Programmable Gate Arrays). It is worth noting that regardless of the type of accelerator used, as long as the accelerator has local memory and can store model data, it can serve as a computing node in this application. This application does not limit the type of accelerator used in the computing nodes.
[0030] It should be noted that one or more computing nodes can run on a single physical device. For example, on an 8-GPU server, each GPU card can be considered an independent computing node, meaning one physical device corresponds to 8 computing nodes; alternatively, the entire server can be considered a single computing node, with multiple GPU cards collaborating to complete the training task of that node. The former is suitable for data-parallel scenarios, while the latter is suitable for model-parallel scenarios. This application does not limit the mapping relationship between computing nodes and physical devices.
[0031] In some embodiments, exemplarily, please refer to Figure 1 , Figure 1 This is a schematic diagram of a distributed training system provided in one possible implementation of this application. For example... Figure 1 As shown, the distributed training system includes a management node 110 and multiple computing nodes (such as...). Figure 1 The computing nodes 131 and 132 shown are illustrated. On the other hand, the system also includes local controllers deployed on each computing node side (such as...). Figure 1 (Local controller 121 and local controller 122 shown).
[0032] In this context, the local controller refers to the control agent module deployed corresponding to the compute node. For example, each compute node can be configured with a corresponding local controller. A control channel exists between the local controller and the corresponding compute node, used to issue control commands to the compute node (such as pausing training, starting data transmission, resuming training, etc.) and to collect status information from the compute node (such as training progress, node health status, etc.). For example, the functions of the local controller may include: receiving and executing control instructions issued by the management node; reporting the status information of the compute node it resides in (such as registration information, ready signals, etc.); during data transmission, controlling the compute node to start a server or client and establish connections with other compute nodes; monitoring the health status of the compute node it resides in, and reporting fault information to the management node when a fault occurs.
[0033] It should be noted that the correspondence between the local controller and the computing node can take various forms. In some embodiments, the local controller can be directly deployed in the corresponding computing node, for example, running as a system service on the computing node's operating system and bound to the training process. In other embodiments, the local controller can also be deployed independently of the computing node, for example, deployed on the management unit of the physical server where the computing node is located, or deployed on other devices connected to the computing node via a network. As long as the local controller can implement control functions for the corresponding computing node (including issuing commands, collecting status, establishing network connections, etc.), it falls within the protection scope of this application. In other words, the deployment relationship between the local controller and the computing node can be tightly coupled (e.g., built into the computing node) or loosely coupled (e.g., deployed independently but network reachable), and this application does not limit this.
[0034] like Figure 1 As shown, management node 110 and local controller (such as...) Figure 1 A control channel exists between the local controllers 121 and 122 shown. The management node 110 sends control commands to the local controllers through the control channel and receives various signals and status information reported by the local controllers. Each local controller is connected to its corresponding computing node (e.g., Figure 1The local controller 121 (corresponding to computing node 131), local controller 122 (corresponding to computing node 132), etc., shown in the figure have a control channel between them, which is used to issue specific execution instructions to the computing nodes.
[0035] It is worth emphasizing that, such as Figure 1 As shown, direct communication connections exist between the computing nodes. Specifically, computing node 131 and computing node 132 can directly transmit data without going through management node 110 or other intermediate storage devices (such as the shared memory mentioned above). Exemplarily, the direct communication connection between computing nodes can be a peer-to-peer (P2P) connection. Through a P2P connection, one computing node can directly send model data stored in its local memory to another computing node, which then stores the received model data in its local memory. Unlike other computing nodes that need to read and write data through shared memory, this application allows direct data channels to be established between computing nodes, thereby achieving efficient migration of model data. This direct transmission mechanism not only eliminates the dependence on shared memory read and write operations, avoiding disk I / O bottlenecks, but also further optimizes the transmission path by combining physical network topology information, reducing network congestion and improving overall training efficiency.
[0036] from Figure 1 As can be seen, the distributed training system provided in this application achieves the separation of control channels and data channels. Specifically, a control channel exists between the management node and each local controller. The management node sends control commands to the local controllers and receives various signals and status information reported by the local controllers through this control channel. Each local controller, in turn, has a control channel with its corresponding computing node, used to issue specific execution commands to the computing node. On the other hand, a direct data channel exists between each computing node for point-to-point transmission of model data. This data channel is independent of the control channel, does not pass through the management node, and does not pass through shared memory. This separation of control and data channels decouples the control logic and data transmission logic of the distributed training system. The management node can focus on global scheduling and status management, while computing nodes can efficiently migrate data, avoiding interference between control and data traffic, and further improving the overall performance of the system.
[0037] Next, the data transmission method provided in this application for the aforementioned distributed training system will be described. For example, please refer to... Figure 2 , Figure 2 This is a flowchart of a data transmission method provided in one possible implementation of this application. The data transmission method may include at least one of the following steps 210-220.
[0038] Step 210: In the case of new nodes waiting to be added in the distributed training system, the management node sends a data synchronization instruction to the target computing node among multiple computing nodes. The data synchronization instruction is used to instruct the model data stored on the target computing node to be synchronized to the new node.
[0039] In one aspect of this embodiment, a new node refers to a new computing node waiting to join the distributed training system. The new node does not yet possess the currently trained model data and needs to obtain model data from existing computing nodes in the system before participating in training. It is important to emphasize that the terms "computing node" and "new node" in this application do not refer to two different types of physical devices, but rather are a distinction based on the stage and role of the node in the training process. Specifically: 1. A computing node refers to a node in a distributed training system that has completed loading model data and is participating in the model training task. In other words, a computing node holds the current model data and continues to perform iterative training.
[0040] 2. A new node refers to a node in the distributed training system that does not yet hold the current model data and is waiting to be added to the training task. A new node may be: a computing resource newly started in response to an elastic scaling command (i.e., a physically new node); a computing node that went offline due to a failure and is ready to be rejoined for training after recovery; or a training process newly started on the same physical device (for example, starting another training process on the same server to participate in training when one training process is already running).
[0041] In summary, "compute node" and "new node" can be different names for the same physical device at different times. Once a node completes model data synchronization and begins training, it transitions from a new node to a compute node. Conversely, when a compute node rejoins after going offline due to a fault, it is also considered a new node during the model data resynchronization process. Therefore, these names are only used to distinguish the role of a node in the synchronization process and do not constitute a limitation on the physical form or type of the node.
[0042] In one aspect of this embodiment, the target computing node refers to the computing node selected by the management node from among multiple existing computing nodes in the system, used to provide model data to the new node. In other words, for any new node waiting to join, the management node needs to determine a corresponding data source node for it, which is the target computing node. The target computing node sends the model data stored in its local memory to the new node, and the new node stores the received model data into its own local memory, thereby completing the model data synchronization.
[0043] It should be noted that a target computing node can send model data to a new node, or simultaneously send model data to multiple new nodes; similarly, a new node can also obtain model data from a single target computing node, or obtain model data in chunks from multiple target computing nodes (e.g., different target computing nodes send different model weight fragments). This application does not limit this, and any correspondence that enables the transmission of model data from existing computing nodes to new nodes falls within the protection scope of this application.
[0044] In one aspect of this embodiment, this application does not limit the method by which the management node determines the target computing node from multiple computing nodes. As long as the management node can select one or more computing nodes as the target computing node from multiple computing nodes based on a certain selection strategy, it falls within the protection scope of this application. For example, the management node can determine the target computing node based on at least one or more of the following factors: 1. Based on physical network distance: In some embodiments, the management node can obtain the physical network topology information of the distributed training system. This physical network topology information indicates the location of each computing node in the physical network (e.g., its hierarchical relationship with servers, top-of-rack switches (ToR switches), aggregation switches, core switches, etc.). The management node selects the computing node with the closest physical network distance to the new node as the target computing node based on the physical network distance between the new node and other computing nodes.
[0045] 2. Based on node load status: In some embodiments, the management node can obtain the current load status information of each computing node, such as CPU (Central Processing Unit) utilization, GPU utilization, network bandwidth usage, and memory usage. The management node selects computing nodes with lower loads as target computing nodes to avoid excessive interference to computing nodes undergoing training due to data transmission.
[0046] 3. Based on data consistency: In some embodiments, the management node can obtain model data version information or iteration step information from each computing node. In distributed training, there may be slight differences in model data between different computing nodes (e.g., different computing nodes may be at different iteration steps). The management node can select the computing node with the latest model data version or the highest iteration step as the target computing node to ensure that the model data obtained by the new node has a high degree of consistency.
[0047] 4. Based on historical transmission records: In some embodiments, the management node can record historical information about each computing node's participation in model data synchronization, such as historical transmission success rate, historical transmission latency, and historical failure rate. The management node can select computing nodes with better historical transmission performance as target computing nodes.
[0048] 5. Comprehensive Decision-Making Based on Multiple Factors: In some embodiments, the management node can make a weighted decision by combining multiple factors mentioned above. For example, the management node can simultaneously consider physical network distance, node load status, and data consistency, and assign different weights to each factor to determine the comprehensive score of the computing node, and finally select the computing node with the higher score as the target computing node.
[0049] It should be noted that the above-described methods for determining the target computing node are merely illustrative examples and are not intended to limit this application. Any strategy capable of selecting a target computing node from multiple computing nodes falls within the protection scope of this application. Specific implementation methods for determining the target computing node will be described in detail in subsequent embodiments.
[0050] In one aspect of this embodiment, a data synchronization command refers to a control command sent by the management node to the target computing node, used to trigger and coordinate the synchronization process of model data from the target computing node to the new node. The data synchronization command is one of the core message types in the control channel communication of this application. By sending a data synchronization command, the management node informs the relevant nodes to begin the transmission operation of model data. It should be noted that this application does not limit the specific form of the data synchronization command; as long as it can realize the control intent of the management node to convey "model data synchronization" to the target computing node and the new node, it falls within the protection scope of this application. For example, a data synchronization command can be an independent control message. The management node sends this control message to the target computing node, and the message carries relevant parameters for this synchronization task. In some embodiments, the data synchronization command can be nested within other control signaling. For example, the management node can send the data synchronization command along with other control information such as pausing training or starting the server to the target computing node. In another aspect of this embodiment, regarding the sending method of the data synchronization command, the management node can use unicast to send the data synchronization command to multiple target computing nodes separately, or it can use broadcast to send the data synchronization command to multiple target computing nodes simultaneously. For example, in the presence of multiple new nodes and multiple corresponding target computing nodes, the management node can send a data synchronization command to all participating target computing nodes through a single broadcast. This application does not limit the method of sending the data synchronization command.
[0051] It should be noted that the above description of the target computing node as the recipient of the data synchronization command is merely illustrative. In some embodiments, the management node may also send the data synchronization command to the new node, or simultaneously to both the target computing node and the new node. Specifically: when the data synchronization command is sent to the target computing node, the command instructs the target computing node to start data transmission preparation (e.g., start the server and listen on the port), provide the address information of the peer node (new node), and inform the transmission configuration parameters of this transmission task; when the data synchronization command is sent to the new node, the command instructs the new node to start data reception preparation (e.g., start the client), provide the address information of the peer node (target computing node) to initiate a connection request, and inform the transmission configuration parameters of this synchronization task. These two sending methods can be used individually or in combination. For example, the management node may send data synchronization instructions only to the target computing node, which then prepares to send the instructions and informs the new node of the connection information through other means (such as broadcast notification, new node polling, etc.); alternatively, it may send data synchronization instructions only to the new node, which then actively requests to establish a connection with the target computing node; or it may send data synchronization instructions to both nodes simultaneously to achieve synchronized instruction delivery and efficient coordination of the transmission process. This application does not limit the specific recipients or methods of sending data synchronization instructions; any method that can trigger the establishment of data transmission between the target computing node and the new node falls within the protection scope of this application.
[0052] In one aspect of this embodiment, regardless of whether the data synchronization instruction is sent to a target computing node or a new node, the data synchronization instruction can carry various information to guide the synchronization of model data. For example, the data synchronization instruction may include, but is not limited to, the following: 1) Role identification information: The data synchronization instruction may include role identification information to indicate the role played by the receiving node in this synchronization process. For example, the instruction may include a role field, the value of which can be "sender" (indicating that the node is the target computing node and needs to send model data) or "receiver" (indicating that the node is the new node and needs to receive model data). The receiving node can determine the operation it should perform based on the role identification information; 2) Peer node address information: The data synchronization instruction may include the network address information of the peer node to indicate the address required to establish a point-to-point connection between the target computing node and the new node. For example, for the target computing node, the instruction may include the network address of the new node (such as address and port number), indicating which address the target computing node needs to send data to. For the new node, the instruction may include the network address of the target computing node, indicating which address the new node needs to connect to to receive data. Through the above methods, the two nodes can directly establish a communication connection without the need for a management node as an intermediary; 3) Model data description information: The data synchronization instruction can include model data description information to indicate the specific content of the model data to be synchronized this time. For example, the instruction can include the name of the model parameters, hierarchical structure, tensor shape, data type, data size, etc. The above information can help the receiving node correctly parse and process the received model data. In some embodiments, if only a portion of the model data needs to be synchronized (e.g., in a model parallel scenario), the data synchronization instruction can also include the parameter range or parameter index to be synchronized; 4) Transmission configuration parameters: The data synchronization instruction can include transmission configuration parameters to control the behavior of data transmission. For example, the transmission configuration parameters can include at least one of the following: transmission timeout (used to set the maximum waiting time for data transmission), number of retries (used to set the maximum number of retries after transmission failure), transmission block size (used to set the block size when data is transmitted in blocks), compression flag (used to indicate whether to compress the transmitted data), encryption flag (used to indicate whether to encrypt the transmitted data), etc. By transmitting configuration parameters, the management node can flexibly adjust the transmission strategy according to network conditions and security requirements; 5) Synchronization identification information: The data synchronization command can include synchronization identification information to uniquely identify this synchronization task. For example, the management node can assign a globally unique synchronization ID (Identifier) to each data synchronization operation. The target computing node and the new node can carry the above synchronization ID in subsequent aggregation signals so that the management node can distinguish and track different synchronization tasks.Synchronization identification information can also be used in fault recovery scenarios, allowing management nodes to determine which synchronization tasks have been completed and which have not, based on the synchronization ID.
[0053] It should be noted that the data content that the above data synchronization instructions may include is merely illustrative and not intended to limit this application. In practical applications, data synchronization instructions may include any one or more combinations of the above content, or other content not listed, and this application does not impose any limitations on this.
[0054] In some embodiments, the management node can simultaneously send data synchronization instructions to both the target computing node and the new node to achieve synchronized instruction delivery. In other embodiments, the management node can prioritize sending data synchronization instructions to the target computing node, instructing it to start the server and listen on the corresponding port. Once the target computing node is ready, the management node then sends data synchronization instructions to the new node, instructing it to connect to the corresponding port of the target computing node. It is worth noting that this application does not limit the timing of data synchronization instruction transmission; any transmission method that can trigger the establishment of data transmission between the target computing node and the new node falls within the protection scope of this application.
[0055] Step 220: The target computing node sends the model data stored in its local memory to the local memory of the new node according to the data synchronization instruction.
[0056] In one aspect of this embodiment, the target computing node, upon receiving a data synchronization instruction, directly sends the model data stored in its local memory to the local memory of the new node. It is important to emphasize that during this transmission process, the model data does not pass through shared storage (such as a distributed file system, network file system, etc.), does not pass through a management node, and does not pass through any other relay devices or intermediate storage media. The model data originates directly from the target computing node's local memory, is transmitted over the network, and arrives at the new node's local memory, thus achieving end-to-end direct transmission.
[0057] In one aspect of this embodiment, direct transfer between the target computing node and the new node can be implemented in various forms. As long as model data can be directly transferred from the source node's local memory to the target node's local memory without passing through shared memory or other intermediary devices, it falls within the protection scope of this application. Several possible forms of direct transfer are described below; these methods can be used individually or in combination as needed: 1. Point-to-point direct transmission: In some embodiments, the target computing node and the new node adopt a point-to-point direct transmission method. Specifically, the target computing node acts as the server and the new node acts as the client. The two can establish a direct communication connection (such as a TCP (Transmission Control Protocol) connection). Through the above communication connection, the target computing node sends the model data in its local memory to the new node, and the new node directly writes the received data into its local memory.
[0058] 2. Remote Direct Memory Access (RDMA) Transfer: In some embodiments, RDMA technology can be used for direct transfer between the target computing node and the new node. RDMA is a technology that allows a computer to directly access the memory of a remote computer without the involvement of the kernels of the operating systems of both nodes. Specifically, the RDMA network card of the target computing node can directly read model data from its local memory (such as GPU memory) and send it to the RDMA network card of the new node via the network. The RDMA network card of the new node then directly writes the data to its local memory. The advantages of the above RDMA transfer are: the data transfer does not involve the CPU, saving CPU resources; the data copy path is short; the latency is extremely low; and the throughput is high, making it suitable for high-speed transmission scenarios of large-scale data.
[0059] 3. High-speed interconnect-based transmission: In some embodiments, when the target computing node and the new node are located on the same physical device and both are equipped with GPUs, direct transmission can be achieved using high-speed interconnect technology between GPUs. For example, NVLink (a high-speed interconnect technology developed by NVIDIA) or PCIe can be used to directly transmit model data between multiple GPUs within the same server. This transmission method does not pass through host memory or the network protocol stack, resulting in extremely high transmission bandwidth and extremely low transmission latency.
[0060] In one aspect of this embodiment, the specific implementation of the direct transmission described above can be selected and optimized based on the physical location relationship between the target computing node and the new node. For example, when the target computing node and the new node are located on different GPUs of the same physical device, high-speed interconnect-based transmission can be prioritized to achieve the highest transmission efficiency. When the target computing node and the new node are located in the same rack (e.g., connected under the same rack switch), RDMA transmission can be prioritized to utilize the high bandwidth and low latency characteristics of the rack network. When the target computing node and the new node are located in different racks, point-to-point TCP transmission or RDMA transmission can be used. Although the transmission latency is relatively high, it still avoids the read / write overhead of shared memory. It is worth noting that the management node can select the optimal direct transmission method for node pairs with different location relationships based on physical network topology information. The specific implementation of topology-aware selection of the target computing node will be described in detail in subsequent embodiments and will not be repeated here.
[0061] It should be emphasized that the direct transmission methods listed above are merely illustrative and not intended to limit this application. Any technical solution that enables the direct transmission of model data from the local memory of the target computing node to the local memory of the new node, without passing through shared memory or other relay devices, falls within the protection scope of this application. Furthermore, this application does not limit the specific protocol, specific technology, or specific hardware platform used for direct transmission; as long as it achieves the core feature of "direct transmission," it falls within the protection scope of this application.
[0062] By sending data synchronization commands from the management node to the target computing node among multiple computing nodes, the target computing node can directly send the model data stored in its local memory to the local memory of the new node according to the data synchronization commands. In this process, the target computing node does not need to write model data to shared memory, and the new node does not need to read model data from shared memory, avoiding read and write operations on shared memory and thus reducing time overhead. Simultaneously, since the data model is directly transmitted between the target computing node and the new node, network bandwidth saturation caused by multiple nodes simultaneously reading and writing to shared memory is avoided. In summary, the data transmission method provided in this application improves the training efficiency of model training.
[0063] In some embodiments, before the management node sends a data synchronization instruction to the target computing node among multiple computing nodes, the method further includes: The management node sends a pause signal to multiple compute nodes, which instructs the compute nodes to pause training after completing the current training task, so that the model data of the multiple compute nodes are in a consistent state. After completing the current training task, the compute nodes pause training and send a ready signal to the management node, which instructs the compute nodes to complete the current training task and pause training. After receiving the ready signals sent by the multiple compute nodes, the management node executes the step of sending a data synchronization command to the target compute node among the multiple compute nodes.
[0064] In one aspect of this embodiment, the pause signal is a control signal sent by the management node to the computing nodes, instructing the computing nodes to pause training after completing the current training task. Specifically, the pause signal ensures that each computing node places its model data in a consistent state before transmitting the model data, i.e., the state after completing the gradient update of the current iteration step, avoiding data transmission under inconsistent conditions. It should be noted that the specific implementation of the pause signal can be varied. In some embodiments, the pause signal can be a separate control message. In other embodiments, the pause signal can be sent in conjunction with other control information (such as a synchronization preparation command). The pause signal can be sent to each computing node individually via unicast or simultaneously to all computing nodes via broadcast. This application does not limit this.
[0065] In one aspect of this embodiment, a ready signal refers to a status report signal sent by a computing node to a management node, indicating that the computing node has completed its current training task and paused training. By receiving the ready signal, the management node can ascertain the current status of each computing node and thus determine whether it can proceed to the next stage (i.e., send a data synchronization command to the target computing node). It should be noted that the specific implementation of the ready signal can vary. In some embodiments, the ready signal can be an independent status message containing the identification information of the computing node sending the signal. In other embodiments, the ready signal can carry other status information of the computing node, such as the current iteration step number, model data version, timestamp, etc. This application does not limit this.
[0066] In one aspect of this embodiment, completing the current training task means that the computing node completes a full iteration step it is currently executing. In distributed training, an iteration step typically includes operations such as forward propagation to calculate the loss, backpropagation to calculate the gradient, and updating model parameters. Here, "Step" and "Iteration" have the same meaning in the context of distributed training, both representing the complete process of the model processing a batch of data and completing a parameter update. After completing the current iteration step, the model data held by the computing node is in a state where the parameter update has just been completed, and at this time, the model data of all computing nodes is consistent (i.e., the aforementioned consistent state). Pausing training at the aforementioned point in time ensures that the model data subsequently synchronized to new nodes is valid and consistent. It should be noted that the specific definition of the current training task can be adjusted according to the actual training strategy. For example, in some embodiments, the current training task can be a full iteration step; in some embodiments, the current training task can be a batch composed of multiple iteration steps; in some embodiments, the current training task can be a synchronization point after gradient accumulation. This application does not limit this; any task boundary that enables the model data of computing nodes to reach a consistent state is applicable.
[0067] In one aspect of this embodiment, the management node needs to wait for all computing nodes to complete their pause and report ready signals, rather than starting the subsequent process after receiving ready signals from only a portion of the computing nodes. This mechanism ensures that the model data of all computing nodes is in a consistent state before model data synchronization, and avoids the following problem: if a computing node has not yet paused, its model data may be being updated. Sending data to a new node using this node as the target computing node might result in intermediate or inconsistent data, thus affecting the correctness of the training task. It should be noted that in some embodiments, the management node can also set a timeout mechanism. For example, if most computing nodes have sent ready signals, and only a few computing nodes fail to respond due to faults or other reasons, the management node can mark the unresponsive computing nodes as faulty nodes and remove them after the timeout, and continue executing subsequent steps based on the ready computing nodes.
[0068] In some embodiments, before the management node sends a data synchronization instruction to the target computing node, a node readiness process, such as a new node initiating a registration request, is also included. The node readiness process described above is explained below with reference to the accompanying drawings. For example, please refer to... Figure 3 , Figure 3 This is a schematic diagram of a node ready process provided in one possible implementation of this application. The node ready process may include the following steps S1 to S5: Step S1: The new node initiates a registration request to the management node. In one aspect of this embodiment, when an expansion command is issued, a new computing node (i.e., the new node) starts up. After completing initialization, the new node sends a registration request to the management node. It should be noted that the registration request may contain relevant information about the new node so that the management node can identify and manage the new node. For example, the registration information may include at least one of the following: the new node's network address (such as an Internet Protocol address (IP address)), physical location identifier (such as rack ID, server ID, switch port identifier, etc.), hardware configuration information (such as GPU model, number of GPUs, memory size, etc.), and software environment information (such as operating system version, deep learning framework version, etc.). After receiving the registration request, the management node can include the new node in the system's management scope according to the registration information.
[0069] Step S2: The management node marks the new node as a waiting state. In one aspect of this embodiment, after receiving a registration request from a new node, the management node marks the new node as a waiting state (e.g., marked as PENDING). New nodes in the waiting state are temporarily suspended and do not participate in model training computation. At this time, the new node does not yet possess the currently trained model data and needs to wait for further instructions from the management node. It should be noted that the management node can maintain a state machine for each node (including computing nodes and new nodes), and the states can include: unregistered, registered and waiting, synchronizing, ready, training, faulty, etc. Marking new nodes as waiting states helps the management node understand the current stage of each node, facilitating subsequent scheduling and fault tolerance.
[0070] Step S3: The management node sends a pause signal to the compute nodes. In one aspect of this embodiment, after completing the registration and marking of new nodes, the management node sends a pause signal to all running compute nodes. The pause signal instructs the compute nodes to pause training after completing the current training task. It should be noted that after receiving the pause signal, the compute nodes do not immediately pause training, but continue to complete a full iteration step that is currently being executed.
[0071] Step S4: The compute node sends a ready signal to the management node. In one aspect of this embodiment, after completing the current iteration step and pausing training, the compute node sends a ready signal to the management node to indicate that it has completed the current training task and paused training. At this time, the model data is ready and can be used to synchronize with new nodes. It should be noted that the ready signal may carry relevant information about the compute node, such as node identifier, current iteration step number, model data version, timestamp, etc. By receiving the ready signals from each compute node, the management node can determine which compute nodes are ready.
[0072] Step S5: The new node sends a ready signal to the management node. In one aspect of this embodiment, after completing its own initialization preparation, the new node also sends a ready signal to the management node to indicate that the new node is ready and can act as a receiver waiting for model data synchronization. It should be noted that the ready signals sent by the new node and the compute nodes are functionally different: the ready signal of the compute node indicates that "the data source is ready" (i.e., the model data of the compute node can be sent), while the ready signal of the new node indicates that "the receiver is ready" (i.e., the new node is ready to receive model data). By receiving the ready signals from all nodes (including compute nodes and new nodes), the management node confirms that the entire system is in a state where data synchronization can begin.
[0073] On the other hand, after sending the pause signal, the management node continues to wait and collect the ready signals sent by each node. Once the management node receives the ready signals from all nodes (including compute nodes and new nodes), it confirms that all nodes are ready. At this point, the management node can execute the next step—sending a data synchronization command to the target compute node.
[0074] It should be noted that, Figure 3 Steps S1 to S5 shown and the specific description above are merely an exemplary implementation of this application and do not constitute a limitation on the scope of protection of this application. Steps such as registering new nodes, marking their status, and sending ready signals by new nodes are optional or exemplary implementation details and are not essential features as defined in these claims.
[0075] By first sending pause signals to multiple computing nodes, each node pauses training and sends a ready signal after completing its current training task. The management node only executes the step of sending a data synchronization command to the target computing node after receiving ready signals from all computing nodes. This mechanism ensures that all computing nodes have paused training and that the model data on all computing nodes is in a consistent state before model data synchronization, avoiding training errors and other problems caused by synchronizing inconsistent model data to a new node. In summary, the technical solution provided in this embodiment provides a reliable consistency guarantee for the model data synchronization process, thereby improving the stability of the distributed training system and the accuracy of the training results.
[0076] In some embodiments, the method further includes: a management node determining the node distance between the new node and each computing node based on physical network topology information, wherein the physical network topology information is used to indicate the location of the multiple computing nodes in the physical network, and the node distance is used to indicate the physical network distance between the new node and the computing nodes; and determining the target computing node from the multiple computing nodes based on the node distance between the new node and each computing node.
[0077] In one aspect of this embodiment, physical network topology information refers to information used to indicate the location of multiple computing nodes in a physical network. Specifically, the physical network refers to the network infrastructure that actually connects the various computing nodes in a distributed training system, including but not limited to: server internal buses (e.g., PCIe, NVLink, etc.), rack internal networks (rack switches connecting servers within the same rack), aggregation layer networks (aggregation switches connecting multiple rack switches), core layer networks (core switches connecting multiple aggregation switches), and other hierarchical structures. Physical network topology information may include at least one of the following: the server identifier of each computing node, the rack identifier of each computing node, the rack switch identifier connected to each computing node, the aggregation switch identifier connected to each computing node, the core switch identifier connected to each computing node, the network hop count between computing nodes, the network latency between computing nodes, and the available bandwidth between computing nodes. The management node can obtain physical network topology information by reading cluster configuration information, network management protocols (such as SNMP (Simple Network Management Protocol)), Link Layer Discovery Protocol (LLDP), etc., and this application does not limit this.
[0078] In one aspect of this embodiment, node distance is used to indicate the physical network distance between a new node and a computing node. Node distance is a relative metric used to characterize the proximity of two nodes during network transmission. Generally, the smaller the node distance, the shorter the physical network path between the two nodes, resulting in lower data transmission latency and higher available bandwidth; conversely, the larger the node distance, the longer the physical network path between the two nodes, potentially requiring more network switching equipment, leading to higher data transmission latency and lower available bandwidth.
[0079] In one aspect of this embodiment, node distance can be measured in various ways. For example, node distance can be determined based on the number of network switching devices traversed between two nodes (e.g., network hop count). For instance, two nodes located within the same server do not require data transmission through any external network switches, resulting in the shortest node distance; two nodes under the same rack switch require one rack switch, resulting in the next longest node distance; two nodes under the same aggregation switch require both rack and aggregation switches, resulting in a longer node distance; and two nodes spanning aggregation switches also require a core switch, resulting in the longest node distance. Alternatively, node distance can also be quantified based on network latency (e.g., round-trip time (RTT)) or network bandwidth between two nodes, which this application does not limit.
[0080] In one aspect of this embodiment, the management node determines the target computing node from multiple computing nodes based on the node distances between the new node and each computing node. Specifically, the management node can select the computing node with the smallest node distance to the new node as the target computing node based on the node distance. Since a smaller node distance results in a shorter data transmission path, lower transmission latency, and less bandwidth consumption on the backbone network, selecting the computing node with the smallest node distance as the data source can achieve optimal data transmission efficiency. It should be noted that there are multiple ways to determine the target computing node based on node distance; several possible implementations are listed below: 1. Selecting the node with the smallest distance: In some embodiments, the management node calculates the node distance between the new node and each computing node, and then selects the computing node with the smallest node distance as the target computing node. If there are multiple computing nodes with the same node distance, a secondary selection can be made based on other factors (such as node load, historical transmission success rate, etc.), or one of them can be randomly selected.
[0081] 2. Select a node whose distance is less than a threshold: In some embodiments, the management node can set a node distance threshold and select one as the target computing node from all computing nodes whose distance is less than the threshold.
[0082] 3. Hierarchical selection based on node distance: In some embodiments, the management node can categorize compute nodes into different levels based on node distance (Level 1 within the same server, Level 2 within the same rack, and Level 3 within the same aggregation switch), and then prioritize the highest-level (i.e., the closest) compute node as the target compute node. If no compute node is available in the highest level, the next highest-level compute node is selected in turn.
[0083] 4. Comprehensive selection based on other factors: In some embodiments, the management node can use node distance as a factor in the comprehensive decision-making process, and calculate a comprehensive score by weighting it with other factors (such as node load, data consistency, historical transmission records, etc.), and finally select the computing node with the highest comprehensive score as the target computing node.
[0084] It should be noted that the above implementation methods are merely illustrative examples and are not intended to limit this application. Any implementation method that determines the target computing node based on node distance falls within the protection scope of this application.
[0085] The management node determines the distance between the new node and each compute node based on the physical network topology information, and selects the target compute node according to the distance, so that model data can be transmitted from the compute node with the closest physical network distance to the new node. This scheme shortens the physical path of data transmission, reduces transmission latency, and reduces the bandwidth consumption of data traffic on the backbone network.
[0086] In some embodiments, determining a target computing node from a plurality of computing nodes based on the node distances between the new node and the various computing nodes includes: Based on the node distance between the new node and each compute node, if it is determined that there is a compute node on the same server as the new node, then the compute node on the same server as the new node is selected as the target compute node; if it is determined that there is no compute node on the same server as the new node, but there is a compute node on the same rack switch as the new node, then the compute node on the same rack switch as the new node is selected as the target compute node; if there is no compute node on the same rack switch as the new node, then the compute node on the same aggregation switch as the new node is selected as the target compute node.
[0087] In one aspect of this embodiment, the management node determines the target computing node from multiple computing nodes based on the node distance between the new node and each computing node, specifically employing a priority-based topology-aware selection strategy. The core of this strategy lies in: setting priorities for different network distances according to the hierarchical structure of the physical network topology, and preferentially selecting the computing node with the closest physical network distance to the new node as the target computing node.
[0088] In one aspect of this embodiment, the hierarchical structure of the physical network can be divided into multiple levels from near to far, and correspondingly, the priority of node distance is also divided in order from near to far. Specifically, this embodiment involves the following three priority levels: 1. Same Server Tier (Highest Priority): "Same server" means the new node and the compute node reside on the same physical server. In this scenario, the new node and the compute node share the same physical server's hardware resources (such as CPU, memory, motherboard, PCIe bus, etc.), and data transmission between them can occur without any external network devices. For example, within the same server, compute nodes can be different GPUs on the same server, or different training processes on the same server (e.g., training processes running on different GPUs, or different computational streams running on the same GPU). When the new node and the compute node reside on the same server, the physical network distance between them is minimized, data transmission latency is reduced, and available bandwidth is maximized.
[0089] 2. Same Rack Switch Level (Second Highest Priority): Same rack switch refers to a new node and a compute node located on different servers, but these servers are connected to the same rack switch. In a data center network architecture, multiple servers typically reside within a rack, interconnected via switches at the top of the rack (i.e., rack switches). When a new node and a compute node are located under the same rack switch, data transmission between them requires passing through a single rack switch, but not through higher-level aggregation or core switches. Compared to the same server level, the same rack switch level has slightly higher latency and potentially slightly lower bandwidth, but it still represents a shorter distance transmission and will not consume backbone network bandwidth.
[0090] 3. Same Aggregation Switch Level (Lower Priority): This refers to a new node and a compute node located in different racks, but the rack switches connected to these racks belong to the same aggregation switch. In a data center network architecture, the uplinks of multiple rack switches connect to an aggregation switch, which is responsible for data exchange within its jurisdiction. When a new node and a compute node are under the same aggregation switch, data transmission between them must pass through both the rack switches and the aggregation switch, but not the core switch. Compared to the same rack switch level, the data transmission path at the same aggregation switch level is longer and has higher latency, but it remains confined within the aggregation switch and does not consume core network bandwidth.
[0091] In one aspect of this embodiment, the management node makes decisions based on the node distances between the new node and each compute node, according to the aforementioned priority order. The specific decision-making logic is as follows: First priority judgment: The management node first determines whether there is a compute node located on the same server as the new node among the multiple compute nodes. If so, it selects the compute node located on the same server as the new node as the target compute node. If there are multiple candidate compute nodes on the same server, the management node can further select one or more optimal ones based on other factors (such as node load, GPU utilization, etc.), or it can make a random selection. Second priority judgment: If the management node determines that there is no compute node located on the same server as the new node among the multiple compute nodes, it continues to determine whether there is a compute node located on the same rack switch as the new node. If so, it selects the compute node located on the same rack switch as the new node as the target compute node. When multiple candidate compute nodes exist within the same rack switch, the management node can further select the best one. In the third priority decision, if the management node determines that none of the multiple compute nodes reside on the same server or the same rack switch as the new node, then it selects the compute node located on the same aggregation switch as the new node as the target compute node. If multiple candidate compute nodes also exist within the same aggregation switch, the management node can similarly select the best one.
[0092] The decision-making process of this embodiment will now be described in detail with reference to the accompanying drawings. For example, please refer to... Figure 4 , Figure 4 This is a schematic diagram illustrating the determination of the target computing node in one possible implementation of this application.
[0093] like Figure 4 As shown, the physical network topology of the distributed training system includes an aggregation switch 410, a rack 421, and a rack 422. Rack 421 contains a rack switch 431, which connects to compute nodes 441 and new nodes 450. Rack 422 contains a rack switch 432, which connects to compute nodes 442 and 443.
[0094] New node 450, located within rack 421 and connected to rack switch 431, is the new node awaiting training. The management node needs to select a target compute node from among the existing compute nodes to transfer model data to new node 450. Currently, several candidate compute nodes exist in the system: compute node 441 is also located within rack 421 and connected to the same rack switch 431; compute nodes 442 and 443 are located within rack 422 and connected to another rack switch 432, separated from new node 450 by aggregation switch 410.
[0095] The management node first obtains the physical network topology information and identifies the physical locations of each compute node and the new node. Specifically, the management node identifies that the new node 450 is located in rack 421 and connected to rack switch 431; compute node 441 is located in the same rack (rack 421) and under the same rack switch (rack switch 431) as the new node 450; compute nodes 442 and 443 are located in another rack (rack 422) and need to communicate with the new node 450 through at least aggregation switch 410.
[0096] The management node performs matching based on a preset physical network distance priority policy. This priority policy, from highest to lowest, is: within the same server, under the same rack switch, and under the same aggregation switch. In this embodiment, since the new node 450 and compute node 441 are located in the same rack (rack 421) and connected to the same rack switch (rack switch 431), the network relationship between compute node 441 and the new node 450 falls under the "same rack switch" priority. Compute nodes 442 and 443, however, require at least traversing aggregation switch 410 to reach the new node 450, thus their network distance falls under a lower priority (under the same aggregation switch).
[0097] Based on the above priority determination, the management node makes a clear decision: select compute node 441 as the target compute node. This is because compute node 441 has the shortest physical network distance to the new node 450, allowing data to be transmitted within the rack without needing to go up to the aggregation switch or cross the core switch. If compute node 442 or compute node 443 were selected as the data source, the data would have to go up to the aggregation switch 410 and then down to the rack 421, resulting in a longer transmission path, greater network distance, higher latency, and consuming bandwidth from the aggregation switch and even the core network.
[0098] In summary, the management node, through topology awareness, uses compute node 441 as the target compute node to transmit model data to the new node 450, thereby achieving the goal of minimizing network overhead and transmission latency.
[0099] It should be noted that, Figure 4 The topology and decision-making process shown are merely an exemplary implementation of this application and do not constitute a limitation on the scope of protection of this application. In practical applications, the physical network topology may include more layers (such as the core switch layer) and more devices, and the priority strategy can also be extended accordingly based on the actual network architecture (for example, adding priority levels such as "under the same core switch" or "within the same server"). As long as the core principle of "preferentially selecting the computing node with the closest physical network distance as the target computing node" is met, any specific topology and decision-making logic falls within the scope of protection of this application.
[0100] In one aspect of this embodiment, different data transmission methods can be employed based on the physical location relationship between the new node and the target computing node to fully utilize the characteristics of each layer and obtain optimal transmission efficiency. Specifically: 1. When the new node and the target compute node are located within the same server, high-speed interconnect technology within the server can be used for data transmission. For example, at least one of the following transmission methods can be used: 1) NVLink transmission: If the new node and the target compute node are two GPUs within the same server, and these two GPUs support NVLink interconnection, data transmission can be performed directly via NVLink; 2) PCIe transmission: If the new node and the target compute node are two GPUs within the same server, but do not support NVLink, data transmission can be performed via the PCIe bus; 3) Shared memory transmission: If the new node and the target compute node are two training processes on the same server (possibly sharing the same GPU or using different GPUs), data transmission can be performed via shared memory. Shared memory transmission does not go through the network protocol stack and has extremely low latency.
[0101] 2. When the new node and the target compute node are located under the same rack switch, data needs to be transmitted through the internal rack network. For example, at least one of the following transmission methods can be used: 1) RDMA transmission: If the network cards of both the new node and the target compute node support RDMA, RDMA can be used for data transmission. RDMA allows data to be transferred directly from the source node's memory to the target node's memory without the involvement of the operating system kernel, featuring high throughput and low latency; 2) GPUDirect transmission: If both the GPU and network card of the new node and the target compute node support GPUDirect technology, GPUDirect can be used for data transmission. GPUDirect allows model data in the GPU's video memory to be sent directly through the network card, or model data received by the network card to be directly written to the GPU's video memory, completely bypassing the host memory and CPU, further reducing transmission latency; 3) TCP / IP transmission: If the above high-speed transmission technologies are unavailable, the traditional TCP protocol can also be used for data transmission.
[0102] 3. When the new node and the target compute node are located under the same aggregation switch, data needs to pass through two layers of network devices: the rack switch and the aggregation switch. Although the transmission path is longer, a similar transmission method as under the same rack switch can still be used. For example, RDMA, GPUDirect, or TCP transmission can be used. It should be noted that because the transmission path passes through more network devices, the transmission latency will increase accordingly, but it still has a significant advantage compared to transmission across aggregation switches (which requires passing through the core switch).
[0103] By employing a priority-based selection strategy, the target compute node is chosen by sequentially assessing the priorities of the same server, rack switch, and aggregation switch, selecting the compute node with the closest physical network distance to the new node. This strategy restricts the transmission path of model data to a minimum network range, reducing the data traffic's bandwidth consumption on the backbone network and further lowering transmission latency.
[0104] In some embodiments, the target computing node sends the model data stored in its local memory to the local memory of the new node according to a data synchronization instruction, including: The management node sends a data synchronization command to the new node; the new node sends a connection request to the target computing node according to the data synchronization command; the target computing node receives the connection request according to the data synchronization command and establishes a communication connection with the new node; the target computing node sends the model data stored in its local memory to the new node's local memory through the communication connection.
[0105] In this embodiment, the timing relationship between the steps "the management node sends a data synchronization command to the new node" and "the management node sends a data synchronization command to the target computing node" is not limited. For example, the management node can send data synchronization commands to both the target computing node and the new node simultaneously. This method has the highest distribution efficiency and is suitable for scenarios with good network conditions and both nodes are ready. Secondly, the management node can first send a data synchronization command to the target computing node, and then send the data synchronization command to the new node after the target computing node has completed server startup and is ready. This method ensures that the server is in a receptive state when the new node initiates a connection, avoiding connection failure. Thirdly, the management node can first send a data synchronization command to the new node and then send the data synchronization command to the target computing node. This method is suitable for special scenarios where the new node requires a longer initialization time.
[0106] In one aspect of this embodiment, the establishment of a communication connection between the target computing node and the new node can adopt a server-client model. In this model, the two communicating parties assume different roles. The server is the party passively waiting for connection requests. In this embodiment, the target computing node acts as the server, starting a server program according to the data synchronization instruction, listening on a specified port, and waiting for connection requests from the new node. Upon receiving a connection request, the server accepts the request and establishes a communication connection. The client is the party actively initiating the connection request. In this embodiment, the new node acts as the client, actively initiating a connection request to the target computing node based on the target computing node's address information (such as IP address and port number) carried in the data synchronization instruction. The server-client model is the most basic and widely used connection establishment model in network communication, characterized by its simplicity, good compatibility, and ease of management.
[0107] Accordingly, in the server-client model, the data synchronization instructions sent by the management node to the target compute node and to the new node differ clearly in content to accommodate the needs of different roles. Specifically, the data synchronization instructions sent to the target compute node (server) are mainly used to instruct it to start the server, and may include a role identifier (identified as "server" or "sender"), server startup parameters (such as listening port range, maximum number of connections, timeout, etc.), transmission configuration (such as data block size, compression flags, etc.), and a unique identifier for this synchronization task. The data synchronization instructions sent to the new node (client) are mainly used to instruct it to initiate a connection, and may include a role identifier (identified as "client" or "receiver"), the target compute node's address information (IP address and listening port number), client connection parameters (such as connection timeout, number of retries, etc.), and the same synchronization identifier, so that the new node can match the synchronization task. Through this differentiated instruction design, the management node can precisely control the behavior of nodes with different roles, enabling the server and client to efficiently establish connections and complete data transmission according to the agreement.
[0108] In one aspect of this embodiment, the specific process of establishing a communication connection between the target computing node and the new node is as follows: First, after receiving the data synchronization instruction sent by the management node, the target computing node starts a server according to the instruction content. Specifically, the target computing node creates a socket, binds it to an available network port, and begins listening for connection requests on that port. The target computing node can inform the new node of the address information of the listening port (such as IP address and port number) through the management node or directly. Second, after receiving the data synchronization instruction sent by the management node, the new node starts a client according to the address information of the target computing node carried in the instruction and sends a connection request to the target computing node. Finally, after receiving the connection request sent by the new node, the target computing node accepts the connection request, and the two parties establish a communication connection. After the communication connection is successfully established, the target computing node sends the model data stored in its local memory to the new node via the network through this connection. After receiving the data, the new node stores the data in its local memory. In the above transmission process, the model data does not pass through shared memory or the management node, realizing direct end-to-end transmission.
[0109] It should be noted that the above-described timing relationships, transmission methods, and instruction content are merely illustrative examples and are not intended to limit this application. The server-client mode is only one exemplary implementation of this embodiment; alternative implementations may employ other connection establishment methods such as peer-to-peer connection mode and reverse connection mode. Any method capable of establishing a communication connection and transmitting data between the target computing node and the new node falls within the protection scope of this application.
[0110] Based on data synchronization instructions, the new node sends a connection request to the target computing node. The target computing node receives the connection request and establishes a communication connection with the new node. Both parties then directly transmit model data through this connection. This scheme establishes a point-to-point communication channel between the target computing node and the new node, allowing model data to be transmitted directly between them without passing through shared memory or a management node. Furthermore, the establishment of the communication connection is triggered by data synchronization instructions, a simple, efficient, and secure process that further improves the efficiency and reliability of model data transmission.
[0111] In some embodiments, please refer to Figure 5 , Figure 5 This is a schematic diagram illustrating the data transmission process provided in one possible implementation of this application. For example... Figure 5 As shown, the above data transmission process may include the following steps S6 to S10.
[0112] Step S6: The management node determines the target compute node. Based on the physical network topology information, the management node determines the target compute node from multiple compute nodes. For details on the specific implementation of the management node determining the target compute node, please refer to [link to relevant documentation]. Figure 4 The detailed descriptions in the foregoing embodiments are not repeated here.
[0113] It should be noted that in this embodiment, step S6 is shown to be executed after the node ready process (i.e., steps S1~S5), but this is only an exemplary implementation of this application. In practical applications, there is no necessary temporal dependency between the process of the management node determining the target computing node and the node ready process. For example, the management node can determine the target computing node immediately after receiving the registration request of a new node, and then send a pause signal to the computing node; or it can determine the target computing node after the computing node has completed the pause and sent a ready signal. Both of the above methods can achieve the technical effects of this application, and this application does not limit the execution order between "determining the target computing node" and "the node ready process".
[0114] Step S7: The management node sends a data synchronization command to the target computing node. After determining the target computing node, the management node sends a data synchronization command to it. This command instructs the target computing node to prepare to send model data. Upon receiving the data synchronization command, the target computing node starts the server according to the command content, listens on the specified port, and waits for connection requests from new nodes.
[0115] Step S8: The management node sends a data synchronization command to the new node. This command instructs the new node to prepare to receive model data. Upon receiving the data synchronization command, the new node initiates a connection request to the target computing node based on the address information of the target computing node carried in the command.
[0116] Step S9: The target computing node sends model data. After receiving the connection request from the new node, the target computing node accepts the request, and a communication connection is established. Subsequently, the target computing node sends the model data stored in its local memory over the network. This sending process does not involve writing the model data to shared memory; the model data directly starts from the target computing node's local memory and enters the network for transmission.
[0117] Step S10: The new node receives model data. The new node receives the model data sent by the target computing node via the network and stores the received model data directly into its local memory. This receiving process does not involve reading data from shared memory; the model data is written directly from the network to the new node's local memory.
[0118] Through steps S6 to S10, model data originates from the target computing node's local memory, travels via the network, and directly reaches the new node's local memory, achieving end-to-end direct transmission. The entire process avoids both writing model data to disk and reading it from shared memory, thus avoiding disk I / O overhead and the network bandwidth saturation problem of shared memory.
[0119] It should be noted that, Figure 5 Steps S6 to S10 shown and the specific description above are merely an exemplary implementation of this application and do not constitute a limitation on the scope of protection of this application. In practical applications, the execution order of each step can be adjusted according to the actual situation (for example, steps S7 and S8 can be executed simultaneously or sequentially). As long as the core feature of "after the management node sends the data synchronization instruction, the target computing node establishes a connection with the new node and directly transmits model data" is satisfied, any specific implementation falls within the scope of protection of this application.
[0120] In some embodiments, the distributed training system further includes a monitoring node, and the method further includes: During the process of the target computing node sending model data to the new node, the monitoring node monitors both the target computing node and the new node. If a fault is detected in the target computing node and / or the new node, the monitoring node reports the fault information to the management node. The management node triggers the fault tolerance mechanism based on the fault information.
[0121] In one aspect of this embodiment, a monitoring node refers to a control device or control process responsible for monitoring the operating status of each node and the data transmission process in a distributed training system. The main functions of a monitoring node include: real-time monitoring of the health status of target computing nodes and new nodes (such as whether processes are running normally, whether resource utilization is normal, etc.), monitoring the connectivity and transmission quality of data transmission links, and timely reporting of fault information when an anomaly is detected. Exemplarily, a monitoring node can include the following forms: 1. As a built-in module of a management node: In some embodiments, the monitoring node can be a functional module within the management node, with the management node uniformly responsible for monitoring and scheduling. The advantage of this approach is its simple system architecture, requiring no additional deployment; 2. Deployed as an independent node: In some embodiments, the monitoring node can be deployed independently of the management node, operating as an independent node in the system. The advantage of this approach is the decoupling of monitoring and management functions, avoiding the impact of excessive load on the management node on the real-time performance of monitoring; 3. Distributed deployment: In some embodiments, the monitoring node can be deployed in a distributed manner, for example, deploying a monitoring agent on each computing node, with multiple monitoring agents collaboratively completing the global monitoring task. The advantage of this approach is its wide monitoring coverage and strong fault tolerance. This application does not limit the specific form of the monitoring node; any implementation method that can realize the monitoring function and report fault information falls within the protection scope of this application.
[0122] In one aspect of this embodiment, during the process of the target computing node sending model data to the new node, the monitoring node monitors both the target computing node and the new node. It should be noted that, in addition to directly monitoring the target computing node and the new node themselves, the monitoring node can also monitor other elements involved in the data transmission process. For example, the monitoring objects may include at least one of the following: 1. The target computing node: monitoring whether its processes are running normally, whether its CPU / GPU usage is abnormal, whether its memory / video memory is sufficient, and whether its network interface is working properly; 2. The new node: monitoring whether its processes are running normally, whether its local storage has sufficient space, and whether its network interface is working properly; 3. The health status of the data transmission link: monitoring whether the network connection between the target computing node and the new node is smooth, whether the transmission latency is within the normal range, whether the packet loss rate is too high, and whether the bandwidth fluctuates abnormally.
[0123] It should be noted that in practical applications, link failures (such as network switch failures, disconnected network cables, network congestion leading to timeouts, etc.) often manifest as communication failures between nodes or data transmission interruptions. From a system perspective, the end result of link failures and node failures is the same—data synchronization failure. Therefore, in the context of this application, whether a node itself fails or the link between nodes fails, it can be regarded as a "node failure" or collectively referred to as a "synchronization failure," triggering the corresponding fault tolerance mechanism.
[0124] In one aspect of this embodiment, if a monitoring node detects a failure in a target computing node and / or a new node, the monitoring node reports the failure information to the management node. The failure information is a dataset describing the failure situation. The management node determines the failure type, scope, and appropriate fault-tolerance measures based on the failure information. For example, the failure information may include at least one of the following: 1. Failure node identifier, i.e., a unique identifier of the node that failed (such as node ID, IP address, etc.), used by the management node to locate the faulty node; 2. Failure type, describing the specific type of failure, such as no response, connection timeout, insufficient resources, network unreachable, etc.; 3. Failure occurrence time, i.e., the timestamp of the failure occurrence, used by the management node for failure analysis and log recording; 4. Failure details, i.e., a detailed description of the failure, such as error codes, error logs, exception stacks, etc., used for subsequent fault investigation and recovery; 5. Transmission progress, i.e., the amount of data transmitted or the percentage of progress completed at the time of the failure, used to determine whether retransmission is necessary and from where to resume transmission.
[0125] In one aspect of this embodiment, the management node triggers a corresponding fault tolerance mechanism based on the fault information reported by the monitoring node. The fault tolerance mechanism isolates the faulty node when it fails, ensuring that the unfailed nodes (including the target computing node and new nodes) can continue to complete the data synchronization task. It should be noted that there can be various specific implementations of the fault tolerance mechanism, and this application does not impose any specific limitations on them. For example, the fault tolerance mechanism may include at least one of the following measures: canceling the model data synchronization process in which the faulty node participates, removing the faulty node, recording the number of faulty nodes and determining whether it exceeds a threshold, re-synchronizing data after the fault is recovered, adjusting the global parallel configuration, etc. Detailed implementations of the fault tolerance mechanism will be described in subsequent embodiments.
[0126] By introducing monitoring nodes to monitor the target computing nodes and new nodes in real time, and reporting fault information to the management node when a fault is detected, the management node triggers a fault tolerance mechanism, thus realizing fault detection and automatic recovery during data transmission. This scheme can promptly detect and handle abnormal situations during transmission, avoiding data synchronization task failure or training interruption due to node failure, and improving the reliability and stability of the distributed training system in dynamic environments.
[0127] In some embodiments, when multiple new nodes exist, the management node determines a corresponding target computing node for each new node to synchronize model data. The management node triggers a fault tolerance mechanism based on fault information, including: The management node identifies the faulty node based on the fault information. If the faulty node is a new node, the management node controls the target computing node corresponding to the faulty node to cancel model data synchronization. If the faulty node is a target computing node, the management node records the number of faulty nodes. If the number of faulty nodes does not exceed a preset threshold, the management node controls the faulty nodes to cancel model data synchronization. If the number of faulty nodes exceeds the preset threshold, the management node controls all target computing nodes to cancel model data synchronization.
[0128] In one aspect of this embodiment, when multiple new nodes need to join the distributed training system simultaneously, the management node will assign a corresponding target computing node to each new node. In other words, multiple pairs of "target computing node-new node" model data synchronization tasks can exist simultaneously in the system. For example, when the system needs to expand from 4 computing nodes to 8 computing nodes, there are 4 new nodes. The management node will match 4 (or more) target computing nodes to each of these 4 new nodes and perform model data synchronization in parallel. In the above multi-node scenario, one or more pairs of data transmission tasks may fail. The management node needs to adopt differentiated fault tolerance strategies based on the type and number of failed nodes to complete as many data transmission tasks as possible while ensuring the stability of the distributed training system.
[0129] In one aspect of this embodiment, the management node identifies faulty nodes based on fault information reported by the monitoring nodes. A faulty node is a node that has experienced a failure; it can be a new node or a target computing node. Specifically, a faulty node that is a new node waiting to receive model data fails (e.g., the new node process crashes, the network disconnects, or there are insufficient resources). A faulty node that is a target computing node, serving as the data source, fails (e.g., the target computing node process crashes, the GPU fails, the network disconnects, or there is no network). The management node can determine the specific identity and type of the faulty node by parsing fields such as "faulty node identifier" and "fault type" in the fault information.
[0130] In one aspect of this embodiment, when the faulty node is a new node, the management node controls the target computing node corresponding to the faulty node to cancel model data synchronization. Specifically, each new node corresponds to at least one target computing node (i.e., the data source determined by the management node for that new node). When a new node fails, it cannot receive model data normally, and its corresponding target computing node no longer needs to send data to it. Therefore, the management node controls the corresponding target computing node to cancel the current data synchronization task. The target computing node can release relevant resources (such as shutting down the server, releasing memory buffers, etc.) and wait for subsequent instructions from the management node. It should be noted that canceling the data synchronization task corresponding to the faulty new node does not affect the synchronization tasks between other new nodes and their respective target computing nodes. The above fault tolerance mechanism embodies the design concept of fault isolation and improves the fault tolerance capability of the system.
[0131] In one aspect of this embodiment, when the faulty node is a target computing node, the management node's processing strategy is more complex, requiring it to distinguish whether the number of faulty target computing nodes exceeds a preset threshold. Specifically, the management node first records the number of faulty target computing nodes. Since each target computing node may correspond to one or more new nodes (for example, a target computing node can send model data to multiple new nodes), the fault of a target computing node will affect the data synchronization of all new nodes associated with it.
[0132] In one aspect of this embodiment, if the number of faulty target computing nodes does not exceed a preset threshold, the management node controls the faulty nodes to cancel model data synchronization, while other normal target computing nodes continue to execute their data synchronization tasks. The preset threshold is a pre-configured value representing the upper limit of the number of faulty nodes the system can tolerate. For example, the preset threshold can be set based on factors such as system size, the importance of the training task, and available resources. For instance, with a total of 10 target computing nodes, the threshold can be set to 2, allowing a maximum of 2 target computing nodes to fail without affecting the continuation of the overall data synchronization task. When the number of faulty nodes does not exceed the threshold, it indicates that the fault range is within the system's tolerance range. The management node only cancels the data synchronization tasks participated in by the faulty target computing nodes, while other normal target computing nodes continue to send model data to their respective new nodes. This mechanism can maximize data synchronization and reduce the impact of faulty nodes on the overall training task.
[0133] In one aspect of this embodiment, if the number of faulty target computing nodes exceeds a preset threshold, the management node controls all target computing nodes to cancel model data synchronization. When the number of faulty nodes exceeds the threshold, it indicates that the scope of the fault has exceeded the system's tolerance range. At this time, even if some target computing nodes are still normal, the overall data synchronization task success rate is already low due to the unavailability of a large number of data sources, or continued synchronization may lead to inconsistent system states. In the above situation, the management node can adopt a more conservative strategy: cancel the data synchronization task of all target computing nodes, that is, suspend the overall data synchronization process. Furthermore, after the overall data synchronization process is suspended, the management node can take further recovery measures, such as: notifying all nodes to restore to the state before the fault, waiting for the faulty nodes to recover and then re-initiating data synchronization, or removing the faulty nodes and continuing training with a reduced cluster size, etc. The above recovery measures can be flexibly configured according to the actual situation, and this application does not limit them.
[0134] In one aspect of this embodiment, the preset threshold is a key parameter used by the management node to determine the severity of a fault. The preset threshold can be set based on at least one of the following factors: 1. System size: The larger the system size, the greater the number of faulty nodes it can tolerate, and the higher the threshold can be set accordingly; 2. Redundancy strategy: If the system adopts a redundancy backup strategy (e.g., each new node is configured with multiple backup data sources), it can tolerate more faulty nodes, and the threshold can be set higher; 3. Historical fault statistics: The threshold can be dynamically adjusted based on historical fault statistics during system operation. It is worth noting that the preset threshold can be statically configured (e.g., specified through a configuration file) or dynamically adjusted (e.g., automatically calculated by the management node based on real-time status). This application does not limit the specific value and setting method of the preset threshold.
[0135] By employing differentiated fault-tolerance strategies based on the type of the faulty node (new node or target compute node) and whether the number of faulty target compute nodes exceeds a preset threshold, a fault-tolerance strategy is adopted. For a new node failure, only the data synchronization tasks related to that faulty node are cancelled, achieving fault isolation. For a target compute node failure, the strategy determines whether to cancel only the data synchronization tasks of the faulty node or all data synchronization tasks, depending on whether the number of faulty nodes exceeds a preset threshold. This approach, while ensuring system stability, completes as many data synchronization tasks as possible, avoiding the failure of the entire data synchronization task due to a local node failure. In summary, this embodiment improves the stability and resource utilization of the distributed training system during the execution of data synchronization tasks.
[0136] In some embodiments, after the target computing node sends the model data stored in its local memory to the local memory of the new node according to the data synchronization instruction, the method further includes: The management node determines the set of valid nodes, which includes computing nodes that store model data in their local memory and / or new nodes. The management node sends a recovery instruction to the computing nodes and / or new nodes in the set of valid nodes. The recovery instruction is used to instruct the start of the next round of training tasks.
[0137] In one aspect of this embodiment, during model data synchronization, each participating node reports a completion signal to the management node after completing its task, indicating that it has completed the relevant tasks for model data synchronization. For example, when a target computing node completes the transmission of model data, it can send a completion signal to the management node to report that the target computing node has completed the data transmission task; similarly, a new node can also send a completion signal to the management node after successfully receiving model data and writing it to its local memory, reporting that the new node has successfully obtained the model data and is ready to participate in training. By receiving these completion signals, the management node can monitor the synchronization progress and completion status of each node in real time and determine the set of valid nodes based on the received completion signals. It should be noted that the completion signals are an important basis for the management node to determine whether a node is healthy and ready to participate in subsequent training.
[0138] In one aspect of this embodiment, the management node determines the set of valid nodes. The set of valid nodes refers to the collection of all nodes whose local memory already contains model data after the completion of this data synchronization process, and whose nodes are qualified to participate in subsequent training tasks. Specifically, the set of valid nodes may include the following two types of nodes: 1. Existing computing nodes (including the target computing node as the data source and other computing nodes that did not participate in data synchronization or have completed data synchronization) already possess model data and are therefore valid nodes; 2. New nodes that have successfully completed model data reception, whose local memory already contains model data synchronized from the target computing node, are also valid nodes.
[0139] It should be noted that there are multiple ways for the management node to determine the set of valid nodes. The following are some examples of implementation methods: Method 1 (Determination Based on Completion Signal): The management node determines the set of valid nodes based on the received completion signals. Specifically, the management node includes the target computing nodes and new nodes that have successfully sent completion signals into the set of valid nodes. Since only nodes that have successfully completed model data synchronization can send completion signals, this method can accurately identify healthy nodes that are ready for training.
[0140] Method Two (Determination Based on Fault Information): When determining the set of valid nodes, the management node can make a comprehensive judgment based not only on the received completion signals but also on the fault information reported by the monitoring nodes. The following is a specific analysis based on different types of faulty nodes: When a faulty node is a new node, it cannot receive model data and therefore will not send a completion signal. The management node can determine that this new node has not been included in the set of valid nodes based on the received completion signal. Simultaneously, fault information reported by monitoring nodes can further confirm the node's fault status (e.g., node crash, network disconnection, insufficient resources). The management node can then mark this node as offline based on the fault information and exclude it from subsequent scheduling.
[0141] When the faulty node is a target compute node, it cannot complete the transmission of model data, or it interrupts transmission after sending partial data. In such cases, the management node may not receive the completion signal sent by the target compute node and therefore will not include it in the set of valid nodes. Furthermore, the management node can determine the identity and fault type of the faulty node based on the fault information reported by the monitoring nodes. If the number of faulty target compute nodes does not exceed a preset threshold, the management node only excludes the faulty node from the set of valid nodes, while other unaffected target compute nodes and new nodes continue to complete the data synchronization task and are eventually included in the set of valid nodes. If the number of faulty target compute nodes exceeds the preset threshold, the management node can cancel all data synchronization tasks. In this case, the set of valid nodes only includes some of the original compute nodes (and these compute nodes may already be in a paused state), or the management node decides not to send a recovery command and re-initiates the data synchronization process.
[0142] It should be noted that the above method for determining the effective node set is merely an illustrative example. This application does not limit the specific strategy for determining the effective node set. Any implementation that can determine the set of nodes that can participate in model training falls within the protection scope of this application.
[0143] In some embodiments, after determining the set of valid nodes, the management node may also perform at least one of the following operations to prepare for subsequent training tasks: 1. Update Global Parallel Configuration: The global parallel configuration refers to the configuration information used to describe the overall size of the distributed training cluster and the collaborative relationships between nodes, which includes at least the world size. The world size indicates the total number of currently active nodes participating in training. The management node updates the world size in the global parallel configuration based on the total number of nodes in the active node set. Newly added nodes and existing compute nodes need to know the new cluster size in order to correctly initialize the process group.
[0144] 2. Determine new physical network topology information: The cluster's physical network topology may change when a faulty node is removed (e.g., the number of nodes in certain racks or switches decreases). The management node can update the physical network topology information based on the set of valid nodes, so that subsequent task scheduling and data transmission can be optimized based on the latest topology.
[0145] 3. Other configuration information required to generate recovery instructions: The management node can also prepare node role information (i.e., the role assignment of each node in training, such as the node ID for data parallelism), training status information (the iteration step position, learning rate, optimizer status, etc., when the next round of training task should start), and synchronization point information (the synchronization point that all nodes need to reach to ensure training consistency), so that they can be carried in the recovery instructions and sent to each node.
[0146] In one aspect of this embodiment, after determining the set of valid nodes, the management node sends a recovery instruction to all nodes in the set (including compute nodes and / or new nodes). The recovery instruction is a control signaling issued by the management node to instruct the receiving node to resume training. It should be noted that sending a recovery instruction to nodes in the set of valid nodes essentially means issuing the instruction only to healthy nodes that are ready for training. In other words, determining the set of valid nodes is the process of removing faulty nodes. In practical implementation, this operation can be implemented in various specific ways. For example, if no faulty nodes are detected, the management node can directly broadcast the recovery instruction to all nodes in the system. Since all nodes are healthy at this time, this broadcast method has the highest efficiency. If a faulty node is detected, the management node can remove the faulty node based on the fault information and only send the recovery instruction to the remaining healthy nodes (this can be done by unicasting or by updating the set of valid nodes and then broadcasting). The essence of the above implementation methods is the same as "determining the set of valid nodes and sending recovery instructions to the nodes in the set". As long as the technical effect of sending recovery instructions only to healthy nodes can be achieved, regardless of whether broadcast, unicast or other sending methods are used, they all fall within the protection scope of this application.
[0147] In one aspect of this embodiment, after receiving a recovery instruction, the compute nodes and / or new nodes in the effective node set resume training according to the instruction content. For example, a node can parse the updated global parallel configuration from the recovery instruction and reset its local communication group according to this configuration. The communication group is a logical group used for inter-node communication in distributed training. Each node needs to know the total number of effective nodes participating in training in the cluster and the communication addresses of each node in order to correctly establish connections with other nodes and perform communication operations such as gradient exchange and parameter synchronization. When the cluster size changes (e.g., a new node is added or a faulty node is removed), the communication group needs to be reset according to the new world size to ensure that nodes can correctly communicate with other healthy nodes. After the communication group is reset, the node can begin executing the next round of training tasks. For example, existing compute nodes resume from a paused state to a working state and continue training for the next iteration based on the model data already in their local memory. New nodes that successfully receive model data participate in training for the first time and, based on the model data received in their local memory, begin executing training tasks together with other nodes.
[0148] For nodes that do not receive a recovery instruction, the subsequent handling varies depending on the situation: If the node is a faulty node (such as a node crash, network disconnection, etc.), the management node can mark it as offline and re-initiate a registration request after the fault is recovered, and re-execute the model data synchronization process to join the training cluster; If the node is a healthy node whose synchronization task was canceled because the number of faulty nodes exceeded a preset threshold (for example, the management node canceled all data synchronization tasks, causing even healthy nodes to fail to complete data synchronization), the management node can re-initiate a complete data synchronization process after the faulty node is repaired, or try to synchronize again after adjusting the strategy according to the actual situation.
[0149] The system determines the set of valid nodes by managing the node and sends recovery commands to these nodes to resume training. This approach ensures that, regardless of whether the system is in a normal or faulty state, only nodes that successfully hold model data are included in subsequent training tasks; faulty nodes are automatically removed, preventing the entire training task from stalling due to waiting for faulty nodes. Furthermore, by determining the set of valid nodes, the system can continue subsequent model training with more appropriately allocated computer resources, improving the utilization rate of computer resources.
[0150] In some embodiments, before the management node determines the set of valid nodes, the method further includes: the target computing node sending a transmission completion signal to the management node, the transmission completion signal indicating that the target computing node has completed sending model data; and the new node sending a reception completion signal to the management node, the reception completion signal indicating that the new node's local memory has stored model data. The management node determines the set of valid nodes by: the management node determining the set of valid nodes based on the transmission completion signal sent by at least one target computing node and the reception completion signal sent by at least one new node.
[0151] In one aspect of this embodiment, the transmission completion signal refers to the completion signal reported by the target computing node to the management node after completing the model data transmission task. This signal indicates that the target computing node has successfully transmitted all the model data stored in its local memory, and its task as the data sender has been completed. For example, the transmission completion signal may include at least one of the following information: the sender node identifier (such as node ID, IP address, etc.), the unique identifier of this synchronization task, the number of bytes of data transmitted, and the timestamp of transmission completion.
[0152] In one aspect of this embodiment, the "receive completion signal" refers to the completion signal reported by a new node to the management node after it has successfully received the model data and written it to its local memory. This signal indicates that the new node has successfully received the complete model data, and its local memory contains model data that can be used for subsequent training; that is, the new node is ready to participate in training. For example, the "receive completion signal" may include at least one of the following: the receiver node identifier, a unique identifier for this synchronization task, the number of bytes of data received, a timestamp indicating completion of reception, and a data integrity verification result (such as a hash value).
[0153] In one aspect of this embodiment, the management node jointly determines the set of valid nodes based on the transmission completion signal and the reception completion signal. It should be noted that model data synchronization is a complete end-to-end process involving both the sender (target computing node) and the receiver (new node). If only the transmission completion signal is used, the management node can only know that the sender has completed data transmission, but cannot confirm whether the receiver has successfully received the data and correctly written it to local memory. Conversely, if only the reception completion signal is used, the management node can only know that the receiver has completed reception, but cannot confirm whether the data comes from a healthy target computing node, nor can it confirm the synchronization status of other target computing nodes. Therefore, the management node needs to simultaneously receive and correlate the transmission completion signal and the reception completion signal to accurately determine whether the synchronization task between a corresponding set of target computing nodes and the new node has truly been completed.
[0154] Specifically, for a corresponding set of target compute nodes and new nodes, the management node can only confirm that the synchronization task has been successfully completed when it receives both a transmission completion signal from the target compute node and a receive completion signal from the new node. At this point, both the target compute node and the new node can be included in the set of valid nodes. If only a transmission completion signal is received but the corresponding receive completion signal is not, it indicates a possible fault on the new node side (e.g., reception failure, node crash, network disconnection, etc.), and the new node will not be included in the set of valid nodes. If only a receive completion signal is received but the corresponding transmission completion signal is not, it indicates a possible fault or signal loss on the target compute node side, and the target compute node will not be included in the set of valid nodes.
[0155] Through the above method, the management node can accurately identify healthy nodes that have truly completed model data synchronization based on the matching and verification of signals from both ends, thereby ensuring that every node in the effective node set holds correct and complete model data. This provides a reliable guarantee for the consistency and correctness of subsequent training tasks. It should be noted that in scenarios where multiple new nodes and multiple target computing nodes are synchronized in parallel, the management node can associate and match the transmission completion signal and the reception completion signal according to the synchronization task identifier, determine the completion status of each group of synchronization tasks, and then summarize them to obtain a complete set of effective nodes.
[0156] In some embodiments, the new node may perform a health check simultaneously with or before sending a reception completion signal, and report the health check results to the management node. A health check is a process by which a new node performs a self-check of its own operational status to confirm that the node possesses the basic conditions to participate in subsequent training tasks. Exemplarily, a health check may include at least one of the following: 1. Local storage status check: Check whether the new node's local storage (such as GPU memory) has enough free space to store model data, and whether the model data writing operation has been successfully completed and the data is complete and error-free.
[0157] 2. Network connection status check: Check whether the network interface of the new node is working properly, whether the network connection between the new node and the target computing node is stable, and whether the control channel between the new node and the management node is unobstructed.
[0158] 3. Hardware resource status check: Check whether the hardware resources of the new node, such as GPU, CPU, and memory, are in normal condition and whether there are any abnormalities such as overheating, memory errors, or ECC (Error Correcting Code) errors.
[0159] 4. Model data integrity verification: Verify the received model data (e.g., calculate the hash value and compare it with the expected value) to ensure that the data has not been damaged or lost during transmission.
[0160] New nodes can include the health check results in their reception completion signal and report them to the management node, or they can send a health check report separately before or after sending the reception completion signal. The health check results can include a health status indicator (such as "healthy" or "faulty"), detailed check items, and a description of any anomalies. Upon receiving the health check results, the management node can use these results to determine whether the new node is eligible to participate in training. For example, if a new node's reception completion signal indicates successful data reception, but the health check results indicate a write error or hardware failure in its local memory, the management node can still classify it as a faulty node and exclude it from the set of valid nodes.
[0161] Similarly, the target compute node can perform a health check while or before sending the transmission completion signal and report the results to the management node. The health check of the target compute node can include local memory status checks (whether the model data is complete and undamaged), network connectivity status checks, and hardware resource status checks. The management node can combine the health check results of the target compute node and the new node, as well as the transmission completion signal and the reception completion signal, to jointly determine the valid set of nodes.
[0162] Through the aforementioned health check mechanism, the management node can not only know whether the node has completed the model data synchronization task, but also further confirm the health status of the node, avoiding the inclusion of potentially faulty nodes into the effective node set, thereby improving the stability and reliability of subsequent training tasks.
[0163] By introducing a two-way confirmation mechanism—transmission completion signal and reception completion signal—the management node needs to simultaneously receive both the transmission completion signal from the target computing node and the reception completion signal from the new node to confirm the completion of a set of model data synchronization tasks and add the corresponding node to the valid node set. This two-way confirmation mechanism avoids erroneous inclusions based solely on single-end signals (either transmission completion or reception completion only), ensuring that each node in the valid node set holds the correct model data. In summary, this embodiment improves the accuracy of the valid node set, providing a reliable guarantee for the consistency and stability of subsequent training tasks.
[0164] The following description, in conjunction with the accompanying drawings, illustrates a specific embodiment of the data transmission method provided in this application. For example, please refer to... Figure 6 , Figure 6 This is a flowchart of a data transmission method provided in another possible implementation of this application. For example... Figure 6As shown, the data transmission method provided in this embodiment can be divided into three stages: node readiness process, data transmission process, and training recovery process.
[0165] In one aspect of this embodiment, the node ready process corresponds to... Figure 6 Steps S1 to S5 in the process. In summary, the node ready process includes: Step S1: The new node sends a registration request to the management node.
[0166] In step S2, the management node marks the new node as waiting.
[0167] Step S3: The management node sends a pause signal to the computing node.
[0168] Step S4: After completing the current training task, the computing node sends a ready signal to the management node.
[0169] Step S5: The new node sends a ready signal to the management node.
[0170] Through the above process, all computing nodes pause training and reach a consistent state, and new nodes are ready to receive data. It is worth noting that the specific implementation of the node readiness process can be found in [link to documentation]. Figure 3 The detailed descriptions in the foregoing embodiments are not repeated here.
[0171] In one aspect of this embodiment, step S6 involves the management node determining the target computing node. In general, the management node selects the computing node with the closest physical network distance as the target computing node based on the physical network topology information and according to a priority strategy (same server, same rack switch, same aggregation switch). For a detailed implementation of the above process, please refer to [link to relevant documentation]. Figure 4 The detailed descriptions in the foregoing embodiments are not repeated here.
[0172] In one aspect of this embodiment, the data transmission process corresponds to... Figure 6 Steps S7 to S10 in the process. In summary, the data transmission process includes: Step S7: The management node sends a data synchronization command to the target computing node.
[0173] Step S8: The management node sends a data synchronization command to the new node.
[0174] Step S9: The target computing node sends model data.
[0175] Step S10: The new node receives model data.
[0176] Through the above process, model data is transferred directly from the target computing node's local storage to the new node's local storage via the network, without passing through shared storage. It is worth noting that the specific implementation details of the data transfer process can be found in [link to relevant documentation]. Figure 5 The detailed descriptions in the foregoing embodiments are not repeated here.
[0177] In one aspect of this embodiment, the training recovery process corresponds to... Figure 6 Steps S11 to S16 in the process. Specifically: Step S11: The target computing node sends a transmission completion signal to the management node. After the target computing node completes the transmission of model data, it sends a transmission completion signal to the management node to report that the data transmission task has been completed.
[0178] Step S12: The new node sends a reception completion signal to the management node. After the new node completes the reception of the model data and successfully writes the model data into its local memory, the new node sends a reception completion signal to the management node to report that it has successfully received the model data and is ready to participate in training.
[0179] Step S13: The management node determines the set of valid nodes. After receiving completion signals from each node (i.e., the aforementioned transmission completion signal and reception completion information), the management node determines the set of valid nodes. Specifically, for any pair of corresponding target computing nodes and new nodes, the management node only confirms the synchronization of the model data for that pair is complete and includes the target computing node and new node in the set of valid nodes when it receives both the transmission completion signal from the target computing node and the reception completion signal from the new node. The set of valid nodes includes computing nodes and / or new nodes that store model data in their local memory; that is, all nodes that successfully hold valid model data and can participate in subsequent training. When a failure occurs during data transmission, the management node can exclude the faulty node from the set of valid nodes.
[0180] Step S14: The management node confirms the new physical network topology information. In the presence of faulty nodes, the management node needs to confirm the new physical network topology information. Since some nodes may be removed due to failure, the cluster's physical network topology changes (e.g., the number of nodes in certain racks or switches decreases), and the management node needs to update the topology information for subsequent scheduling and training.
[0181] Step S15: The management node sends a recovery command to the compute nodes in the set of valid nodes. After determining the set of valid nodes, the management node sends a recovery command to the compute nodes in the set of valid nodes, instructing them to start executing the next round of training tasks.
[0182] Step S16: The management node sends a recovery command to the new nodes in the set of valid nodes. Simultaneously, the management node sends a recovery command to the new nodes in the set of valid nodes, instructing them to begin the next round of training.
[0183] It is important to emphasize that step S14 is only required when a faulty node exists. Specifically: if no fault occurs during data transmission, the management node can directly send a recovery command to all nodes after receiving transmission completion signals from all target computing nodes and reception completion signals from all new nodes, without needing to perform steps such as faulty node statistics and topology information confirmation; if a fault occurs during data transmission (e.g., some new nodes or target computing nodes fail), the management node needs to execute step S14 and then send the recovery command only to nodes in the valid node set. Furthermore, regardless of whether a faulty node exists, after determining the valid node set, the management node can update the world size in the global parallel configuration based on the total number of nodes in the valid node set, so that each node can correctly reset the communication group after receiving the recovery command. More implementation details regarding the above training recovery process can be found in the detailed description in the foregoing embodiments, and will not be repeated here.
[0184] It is worth noting that, Figure 6 The stage divisions and step numbers shown are merely an exemplary implementation of this application and do not constitute a limitation on the scope of protection of this application. In practical applications, the execution order of each stage can be adjusted according to the actual situation. As long as the core process of "under the coordination of the management node, the target computing node establishes a connection with the new node and directly transmits model data, and after synchronization is completed, the next round of training tasks begins" is met, any specific implementation method falls within the scope of protection of this application.
[0185] The following are system embodiments of this application, which can be used to execute the method embodiments of this application. For details not disclosed in the system embodiments of this application, please refer to the method embodiments of this application.
[0186] For example, please refer to Figure 7 , Figure 7 This is a block diagram of a distributed training system provided in one possible implementation of this application. The distributed training system 700 includes a management node 710 and multiple computing nodes 720.
[0187] Management node 710 is used when there are new nodes waiting to join in the distributed training system. Figure 7 In the case of (not shown), to the target computing node among multiple computing nodes 720 ( Figure 7 (Not shown in the image) Sends a data synchronization command, which instructs the model data stored on the target computing node to be synchronized to the new node.
[0188] The target computing node is used to send the model data stored in the target computing node's local memory to the new node's local memory according to the data synchronization command.
[0189] In some embodiments, the management node 710 is further configured to: determine the node distance between the new node and each computing node based on physical network topology information, wherein the physical network topology information is used to indicate the location of the multiple computing nodes in the physical network, and the node distance is used to indicate the physical network distance between the new node and the computing nodes; and determine the target computing node from the multiple computing nodes based on the node distance between the new node and each computing node.
[0190] In some embodiments, the management node 710 is further configured to: based on the node distance between the new node and each computing node, if it is determined that there is a computing node located on the same server as the new node among the multiple computing nodes, then select the computing node located on the same server as the new node as the target computing node; if it is determined that there is no computing node located on the same server as the new node among the multiple computing nodes, and there is a computing node located on the same rack switch as the new node among the multiple computing nodes, then select the computing node located on the same rack switch as the new node as the target computing node; if there is no computing node located on the same rack switch as the new node among the multiple computing nodes, then select the computing node located on the same aggregation switch as the new node as the target computing node.
[0191] In some embodiments, the management node 710 is further configured to send a pause signal to multiple computing nodes, the pause signal being used to instruct the computing nodes to pause training after completing the current training task, so that the model data of the multiple computing nodes 720 are in a consistent state.
[0192] Compute node 720 is also used to pause training after completing the current training task and send a ready signal to the management node. The ready signal is used to indicate that the compute node has completed the current training task and paused training.
[0193] The management node 710 is also configured to, after receiving ready signals from multiple computing nodes, execute the step of sending a data synchronization instruction to a target computing node among the multiple computing nodes 720.
[0194] In some embodiments, the management node is used to send data synchronization instructions to the new node.
[0195] The new node is used to send a connection request to the target computing node according to the data synchronization instructions.
[0196] The target computing node is used to receive connection requests and establish communication connections with new nodes according to data synchronization instructions.
[0197] The target computing node is also used to send model data stored in the target computing node's local memory to the new node's local memory via a communication connection.
[0198] In some embodiments, the distributed training system 700 further includes a monitoring node ( Figure 7 (Not shown in the image).
[0199] The monitoring node is used to: monitor the target computing node and the new node during the process of the target computing node sending model data to the new node; if a fault is detected in the target computing node and / or the new node, the monitoring node reports the fault information to the management node.
[0200] Management node 710 is used to trigger the fault tolerance mechanism based on fault information.
[0201] In some embodiments, when multiple new nodes exist, the management node determines a corresponding target computing node for each new node to synchronize model data.
[0202] Management node 710 is used for: identifying faulty nodes based on fault information; controlling the target computing nodes corresponding to the faulty nodes to cancel model data synchronization when the faulty node is a new node; recording the number of faulty nodes when the faulty node is a target computing node; controlling the faulty nodes to cancel model data synchronization if the number of faulty nodes does not exceed a preset threshold; and controlling all target computing nodes to cancel model data synchronization if the number of faulty nodes exceeds the preset threshold.
[0203] In some embodiments, the management node 710 is further configured to: determine a set of valid nodes, the set of valid nodes including computing nodes and / or new nodes that store model data in local memory; and send a recovery instruction to the computing nodes and / or new nodes in the set of valid nodes, the recovery instruction being used to instruct the start of the next round of training tasks.
[0204] In some embodiments, the target computing node is further configured to send a transmission completion signal to the management node, the transmission completion signal indicating that the target computing node has completed sending the model data.
[0205] The new node is also used to send a reception completion signal to the management node, which indicates that the model data has been stored in the new node's local memory.
[0206] The management node is also used to determine the set of valid nodes based on the transmission completion signal sent by at least one target computing node and the reception completion signal sent by at least one new node.
[0207] The distributed training system provided in this application sends data synchronization instructions from a management node to a target computing node among multiple computing nodes. The target computing node can then directly send the model data stored in its local memory to the local memory of the new node according to these instructions. In this process, the target computing node does not need to write model data to shared memory, and the new node does not need to read model data from shared memory, avoiding read / write operations on shared memory and thus reducing time overhead. Simultaneously, since the data model is directly transmitted between the target computing node and the new node, network bandwidth saturation caused by multiple nodes simultaneously reading and writing to shared memory is avoided. In summary, the distributed training system provided in this application improves the training efficiency of model training.
[0208] It should be noted that the system provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0209] For example, please refer to Figure 8 , Figure 8 This is a structural block diagram of a computer device provided in one possible implementation of this application. The computer device 800 can be any electronic device with data computing, processing, and storage functions. The computer device 800 can be used to implement the data transmission method provided in the above embodiments.
[0210] Typically, computer device 800 may include a processor 810 and a memory 820.
[0211] Processor 810 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 810 may be implemented using at least one hardware form selected from DSP (Digital Signal Processor), FPGA, and PLA (Programmable Logic Array). Processor 810 may also include a main processor and a coprocessor. The main processor, also known as the CPU, is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 810 may integrate a GPU, which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 810 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0212] The memory 820 may include one or more computer-readable storage media, which may be non-transitory. The memory 820 may also include high-speed random access memory and NVM (Non-Virtual Machine). Volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage medium in memory 820 is used to store a computer program configured to be executed by one or more processors to implement the data transfer method described above.
[0213] Those skilled in the art will understand that Figure 8 The structure shown does not constitute a limitation on the computer device 800, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0214] In an illustrative embodiment, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor of a computer device, implements the aforementioned data transmission method. Optionally, the computer-readable storage medium may be a ROM (Read-Only Memory), RAM (Random Access Memory), CD-ROM (Compact Disc Read-Only Memory), magnetic tape, floppy disk, or optical data storage device, etc.
[0215] In an exemplary embodiment, a chip is also provided, the chip including programmable logic circuitry and / or program instructions stored in a computer-readable storage medium. A processor of a computer device reads the programmable logic circuitry and / or program instructions from the computer-readable storage medium, and executes the programmable logic circuitry and / or program instructions, causing the computer device to perform the data transmission method described above.
[0216] In an exemplary embodiment, a computer program product is also provided, which includes a computer program that is loaded and executed by a processor to implement the data transmission method described above.
[0217] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0218] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data transmission method applied to a distributed training system, characterized in that, The distributed training system includes a management node and multiple computing nodes, and the method includes: In the case that there are new nodes waiting to be added in the distributed training system, the management node sends a data synchronization instruction to the target computing node among the multiple computing nodes. The data synchronization instruction is used to instruct the model data stored on the target computing node to be synchronized to the new node. According to the data synchronization instruction, the target computing node sends the model data stored in its local memory to the local memory of the new node.
2. The method according to claim 1, characterized in that, The method further includes: The management node determines the node distance between the new node and each of the computing nodes based on the physical network topology information. The physical network topology information is used to indicate the position of the plurality of computing nodes in the physical network, and the node distance is used to indicate the physical network distance between the new node and the computing nodes. The target computing node is determined from the plurality of computing nodes based on the node distance between the new node and each of the computing nodes.
3. The method according to claim 2, characterized in that, The step of determining the target computing node from the plurality of computing nodes based on the node distance between the new node and each of the computing nodes includes: Based on the node distance between the new node and each of the computing nodes, if it is determined that there is a computing node on the same server as the new node among the plurality of computing nodes, then the computing node on the same server as the new node is selected as the target computing node. If it is determined that there is no computing node on the same server as the new node among the plurality of computing nodes, and there is a computing node on the same rack switch as the new node among the plurality of computing nodes, then the computing node on the same rack switch as the new node is selected as the target computing node. If none of the plurality of computing nodes is located on the same rack switch as the new node, then the computing node located on the same aggregation switch as the new node is selected as the target computing node.
4. The method according to claim 1, characterized in that, Before the management node sends a data synchronization command to the target computing node among the plurality of computing nodes, the method further includes: The management node sends a pause signal to the plurality of computing nodes. The pause signal is used to instruct the computing nodes to pause training after completing the current training task, so that the model data of the plurality of computing nodes are in a consistent state. After completing the current training task, the computing node pauses training and sends a ready signal to the management node. The ready signal is used to indicate that the computing node has completed the current training task and paused training. After receiving the ready signals sent by the plurality of computing nodes, the management node executes the step of sending a data synchronization command to the target computing node among the plurality of computing nodes.
5. The method according to claim 1, characterized in that, The target computing node, according to the data synchronization instruction, sends the model data stored in its local memory to the local memory of the new node, including: The management node sends the data synchronization command to the new node; The new node sends a connection request to the target computing node according to the data synchronization instruction; The target computing node receives the connection request according to the data synchronization instruction and establishes a communication connection with the new node; The target computing node sends the model data stored in its local memory to the local memory of the new node through the communication connection.
6. The method according to claim 1, characterized in that, The distributed training system also includes a monitoring node, and the method further includes: During the process of the target computing node sending the model data to the new node, the monitoring node monitors both the target computing node and the new node; If a fault is detected in the target computing node and / or the new node, the monitoring node reports the fault information to the management node; The management node triggers a fault tolerance mechanism based on the fault information.
7. The method according to claim 6, characterized in that, In the presence of multiple new nodes, the management node determines a corresponding target computing node for each new node to synchronize model data; The management node triggers a fault tolerance mechanism based on the fault information, including: The management node determines the faulty node based on the fault information; If the faulty node is the new node, the management node controls the target computing node corresponding to the faulty node to cancel the model data synchronization. If the faulty node is the target computing node, the management node records the number of faulty nodes; if the number of faulty nodes does not exceed a preset threshold, the management node controls the faulty node to cancel model data synchronization; if the number of faulty nodes exceeds the preset threshold, the management node controls all target computing nodes to cancel model data synchronization.
8. The method according to any one of claims 1 to 7, characterized in that, After the target computing node sends the model data stored in its local memory to the local memory of the new node according to the data synchronization instruction, the method further includes: The management node determines a set of valid nodes, which includes computing nodes and / or new nodes that store the model data in the local memory. The management node sends a recovery instruction to the compute nodes and / or new nodes in the set of valid nodes. The recovery instruction is used to indicate the start of the next round of training tasks.
9. The method according to claim 8, characterized in that, Before the management node determines the set of valid nodes, the method further includes: The target computing node sends a transmission completion signal to the management node, the transmission completion signal being used to indicate that the target computing node has completed sending the model data; The new node sends a reception completion signal to the management node, the reception completion signal indicating that the model data has been stored in the new node's local memory; The management node determines the set of valid nodes, including: The management node determines the set of valid nodes based on the transmission completion signal sent by at least one target computing node and the reception completion signal sent by at least one new node.
10. A distributed training system, characterized in that, The distributed training system includes a management node and multiple computing nodes: The management node is used to send a data synchronization instruction to the target computing node among the multiple computing nodes when there are new nodes waiting to be added in the distributed training system. The data synchronization instruction is used to instruct the model data stored on the target computing node to be synchronized to the new node. The target computing node is configured to send the model data stored in its local memory to the local memory of the new node according to the data synchronization instruction.
11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing a computer program that is loaded and executed by the processor to implement the method as claimed in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method as described in any one of claims 1 to 9.
13. A chip, characterized in that, The chip includes programmable logic circuitry and / or program instructions, which, when the chip is running, are used to implement the method as described in any one of claims 1 to 9.
14. A computer program product, characterized in that, The computer program product includes a computer program that is loaded and executed by a processor to implement the method as described in any one of claims 1 to 9.