Distributed computing system communication method and device
By dynamically adjusting the communication timeout threshold of the communication group in the distributed computing system, the communication timeout detection problem is solved, the communication efficiency and the stability of model training are improved, and training failure caused by communication exceptions is avoided.
Patent Information
- Application Number
- CN202510046589.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
In a distributed computing system, how to reasonably detect communication timeouts, improve communication efficiency and stability between computing nodes, improve model training efficiency, and avoid training failures and system instability caused by communication exceptions.
By determining the communication timeout thresholds for each communication group and dynamically adjusting these thresholds according to the synchronization wait time during the current training process, the accurate detection of the communication synchronization status is ensured.
Dynamic adjustment of communication timeout threshold during model training is realized, avoiding the system's premature judgment of communication failure and unnecessary waiting time, maximizing communication efficiency, reducing communication synchronization waiting time, and improving the efficiency and stability of model training.
Smart Images

Figure CN119945954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a distributed computing system communication method and device. Background Art
[0002] In a distributed computing system, model training is a complex and resource-intensive process. It relies on efficient collaboration between multiple computing nodes. These computing nodes may be distributed in different geographical locations, connected together through a network, and work together to complete the model training task. In this setting, data consistency and model state synchronization are crucial because they directly affect the accuracy and efficiency of model training. To achieve this, nodes need to communicate and synchronize data frequently. Communication anomalies will block the continuation of the entire model training task.
[0003] Related technologies generally set static timeout thresholds based on experience to detect communication timeouts in distributed computing systems. When the network is in good condition, setting a timeout threshold that is too large will result in unnecessary waiting time and reduce training efficiency. On the contrary, when the network is in poor condition, setting a timeout threshold that is too small may cause premature interruption of communication, thereby increasing the risk of training failure.
[0004] Therefore, how to reasonably detect communication timeouts in distributed computing systems, improve the communication efficiency and stability between computing nodes, and improve the training efficiency of models in distributed computing systems has become a technical problem that needs to be urgently solved in the industry. Summary of the invention
[0005] The present invention provides a distributed computing system communication method and device, which are used to solve the technical problems of how to reasonably detect communication timeouts in a distributed computing system, improve the communication efficiency and stability between computing nodes, and improve the training efficiency of models in a distributed computing system.
[0006] The present invention provides a distributed computing system communication method, comprising: Determine a communication timeout threshold for each communication group; the communication group is determined based on processes having data communication relationships in each computing node in the distributed computing system; the process is used to execute a training task for a target model; Determine the communication synchronization state of each communication group in the current training process based on a comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process; The communication timeout threshold of each communication group is adjusted based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
[0007] In some embodiments, determining the communication timeout threshold of each communication group includes: Before training the target model, an initial value of a communication timeout threshold of each communication group is set; the initial value of the communication timeout threshold is greater than the sum of the operator compilation time and the data set initialization time of the target model.
[0008] In some embodiments, adjusting the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group includes: Based on the synchronization waiting time of each communication group in the current training process, determining the maximum synchronization waiting time and the synchronization waiting time variance of each communication group in the current training process; Based on the maximum synchronization waiting time and the synchronization waiting time variance, a communication timeout threshold adjustment value of each communication group is determined.
[0009] In some embodiments, determining the communication timeout threshold adjustment value of each communication group based on the maximum synchronization waiting time and the synchronization waiting time variance includes: Compare the maximum synchronization waiting time in the current training process with the maximum synchronization waiting time in each training process before the current training process, and determine multiple maximum synchronization waiting time changes; When the maximum value among the multiple maximum synchronization waiting time changes is greater than the first preset threshold, determining the communication timeout threshold adjustment value of each communication group based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process; and / or, comparing the synchronization waiting time variance in the current training process with the synchronization waiting time variance in each training process before the current training process, to determine a plurality of synchronization waiting time variance changes; When the maximum value among the multiple synchronization waiting time variance changes is greater than the second preset threshold, the communication timeout threshold adjustment value of each communication group is determined based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process.
[0010] In some embodiments, the method further comprises: Acquire heartbeat information of each computing node in the distributed computing system; the heartbeat information is used to indicate the health status of the computing node; In the event that the heartbeat information of any computing node is abnormal, determining that any computing node is faulty; Resynchronize any of the computing nodes, or distribute the training task executed by any of the computing nodes to other computing nodes.
[0011] In some embodiments, after allocating the training task executed by any computing node to other computing nodes, the method further includes: Determine the faulty computing node based on the heartbeat information of each computing node, and remove the faulty computing node from the distributed computing system; Based on the remaining computing nodes in the distributed computing system and the current training status of the target model, continue to train the target model.
[0012] The present invention provides a distributed computing system communication device, comprising: A determination module, used to determine a communication timeout threshold of each communication group; the communication group is determined based on processes having data communication relationships in each computing node in the distributed computing system; the process is used to execute a training task of a target model; A detection module, configured to determine the communication synchronization state of each communication group in the current training process based on a comparison result of a communication timeout threshold of each communication group and a synchronization waiting time of each communication group in the current training process; An adjustment module is used to adjust the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
[0013] The present invention provides an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the distributed computing system communication method when executing the computer program.
[0014] The present invention provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the distributed computing system communication method is implemented.
[0015] The present invention provides a computer program product, comprising a computer program, wherein the computer program implements the distributed computing system communication method when executed by a processor.
[0016] The communication method and device of the distributed computing system provided by the present invention determine the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model; based on the comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process, the communication synchronization state of each communication group in the current training process is determined; based on the synchronization waiting time of each communication group in the current training process, the communication timeout threshold of each communication group is adjusted to obtain the communication timeout threshold adjustment value of each communication group; based on the communication timeout threshold adjustment value of each communication group, the communication synchronization state of each communication group in the next training process is detected; because the communication timeout threshold of each communication group is dynamically adjusted during the model training process, it is avoided that the system judges that the communication fails too early, and it is avoided that unnecessary waiting time is caused, so as to ensure the maximization of communication efficiency, reduce the communication synchronization waiting time, improve the training efficiency of the model in the distributed computing system, and at the same time reduce the communication interruption and system instability, and improve the continuity and stability of the model training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0019] Figure 1 It is a flow chart of the distributed computing system communication method provided by the present invention.
[0020] Figure 2 It is a schematic diagram of the hybrid parallel strategy for model training provided by the present invention.
[0021] Figure 3 It is a schematic diagram of the structure of the distributed computing system communication device provided by the present invention.
[0022] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the present invention are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units or modules is not necessarily limited to those steps or units or modules that are clearly listed, but may include other steps or units or modules that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] When using a distributed computing system for model training, nodes need to communicate and synchronize data frequently, which involves operations such as data distribution, aggregation, and updating model parameters. In model training, collective communication and synchronization operations are key parts of distributed training, which ensure data consistency and collaboration between different computing nodes. When performing these operations, the processor generally stops running and waits for synchronization until all communication participants complete the operation. The time setting for this synchronization wait will affect the success rate of the task and the utilization of resources.
[0026] However, this synchronization process can become complicated and time-consuming for a variety of reasons. For example, from the perspective of distributed computing systems, network delays, node failures, security attacks, and power outages can all lead to communication interruptions or delays. From the perspective of model training, unexpected process crashes and excessive amounts of communication data can also lead to communication delays.
[0027] In order to solve the shortcomings of related technologies, Figure 1 FIG. 1 is a flow chart of a distributed computing system communication method provided by the present invention, such as Figure 1 As shown, the method includes step 110 , step 120 and step 130 .
[0028] Step 110, determine the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model.
[0029] Specifically, the execution subject of the distributed computing system communication method provided in the embodiment of the present invention is a distributed computing system communication device. The device can be implemented by software, such as a distributed computing system communication program; or by hardware, such as a processor, computer or server that executes the distributed computing system communication method.
[0030] The application scenario of the method provided in the embodiment of the present invention is to perform distributed training on the target model. The target model refers to the neural network model that needs to be trained, for example, it can be various large language models. For example, when the neural network model is used to perform tasks such as image processing, speech processing or text processing, the neural network model can be trained using a distributed computing system. The neural network model can be trained using training data such as image samples, speech samples or text samples.
[0031] A distributed computing system is a system formed by multiple distributed computing nodes connected through a network. The computing nodes can be computers or servers. These computing nodes cooperate with each other to achieve one or more common goals, such as training or reasoning a large language model.
[0032] A process refers to an independent computing unit that performs each training task of the target model. Each process is responsible for handling a specific computing task, usually a part of the model (such as data parallelism in distributed training). Processes need to communicate with each other, especially when training on multiple nodes or across nodes, when processes need to synchronize model parameters (such as gradients) or share calculation results.
[0033] Multiple AI processors, such as a graphics processing unit (GPU), can be set up inside each computing node. A computing node is a unit of physically distributed computing resources. Each node may run multiple processes, and each process corresponds to an AI processor.
[0034] A communication group refers to multiple processes that have data communication relationships in each computing node in a distributed computing system. The data communication relationship can be determined based on the parallel strategy of model training. Figure 2 Schematic diagram of the hybrid parallel strategy for model training provided by the present invention, such as Figure 2As shown, it includes computing node Node0 and computing node Node1. Among them, computing node Node0 is equipped with 8 GPUs, respectively represented as GPU0 to GPU7; computing node Node1 is equipped with 8 GPUs, respectively represented as GPU8 to GPU15. The above two nodes use the distributed training strategy of DP2-TP2-PP4. DP (Data Parallel) is a data parallel strategy. TP (Tensor Parallel) is a tensor parallel strategy. PP (Pipeline Parallel) is a pipeline parallel strategy.
[0035] The two large rounded dashed boxes on the left and right represent two DP communication groups. The corresponding GPUs between the two communication groups (connected by thin black solid arrows, such as GPU0 and GPU2) generally perform all-reduce communication. The eight GPUs in the same DP communication group can also be labeled A to H.
[0036] There are 4 rows from top to bottom, representing 4 PP communication groups. The corresponding GPUs between two adjacent communication groups (such as GPU0->GPU4->GPU8->GPU12) perform send and receive (send / recv) communications.
[0037] The corresponding GPUs in the small dotted boxes in the same DP communication group and the same PP communication group act as a TP communication group to perform all gather or reduce scatter communication.
[0038] The communication timeout threshold refers to the time limit in a distributed computer system when a process's communication synchronization request is not responded to within a predetermined time, at which the system will consider the communication synchronization failure and take corresponding processing measures.
[0039] Step 120: Determine the communication synchronization state of each communication group in the current training process based on the comparison result between the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process.
[0040] Specifically, the communication synchronization state refers to the state of communication synchronization failure or communication synchronization success of processes in the communication group. The synchronization waiting time refers to the waiting time that a task to be executed by a process must wait for the completion of tasks to be executed by other processes before it can continue to execute.
[0041] The training of the target model is performed by iterating multiple training processes (steps). In each training process, the synchronization waiting time in each training process can be compared with the communication timeout threshold to detect the communication synchronization status of each communication group in the current training process.
[0042] Taking the current training process as an example, the synchronization waiting time of each process in the communication group is compared with the communication timeout threshold. If the synchronization waiting time is greater than the communication timeout threshold, the process communication synchronization state can be considered to be failed; if the synchronization waiting time is less than the communication timeout threshold, the process communication synchronization state can be considered to be successful.
[0043] Step 130: adjust the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group. The communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
[0044] Specifically, the statically set communication timeout threshold lacks flexibility and cannot adapt to dynamic changes in network conditions. When the network conditions are good, a timeout setting that is too long (the communication timeout threshold is too large) may cause unnecessary waiting time and reduce training efficiency. On the contrary, when the network conditions are poor, a timeout setting that is too short (the communication timeout threshold is too small) may cause premature interruption of communication, thereby increasing the risk of training failure. For example, in a large-scale distributed training task, the communication timeout threshold is set to 30 seconds. When the network conditions are good, this communication timeout threshold can accurately detect communication timeouts because communication is usually completed within a few seconds. However, when the network conditions suddenly deteriorate, such as due to network congestion or hardware failure, the communication time may be extended to more than 1 minute. In this case, the static setting of 30 seconds will cause the system to prematurely determine that the communication has failed, thereby interrupting the training process.
[0045] In each training process, the communication timeout threshold can be dynamically adjusted according to the synchronization waiting time. If the synchronization waiting time is long, it means that the network conditions have changed and more time is needed for communication synchronization. In this case, the communication timeout threshold can be increased to prevent the system from prematurely judging communication failure. If the synchronization waiting time is short, it means that the network conditions have changed and less time is needed for communication synchronization. In this case, the communication timeout threshold can be reduced to avoid unnecessary waiting time. After adjusting the communication timeout threshold, the communication timeout threshold adjustment value can be obtained.
[0046] Since the amount of data and the communication links that need to be synchronized for each communication group may not be the same, independent adjustments can be made for each communication group.
[0047] The communication timeout threshold adjustment value of each communication group may be used as the communication timeout threshold of each communication group in the next training process, and the communication synchronization state of each communication group in the next training process may be detected.
[0048] The communication method for a distributed computing system provided by an embodiment of the present invention determines the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model; based on the comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process, the communication synchronization state of each communication group in the current training process is determined; based on the synchronization waiting time of each communication group in the current training process, the communication timeout threshold of each communication group is adjusted to obtain the communication timeout threshold adjustment value of each communication group; based on the communication timeout threshold adjustment value of each communication group, the communication synchronization state of each communication group in the next training process is detected; since the communication timeout threshold of each communication group is dynamically adjusted during the model training process, it is avoided that the system judges that the communication fails too early and avoids causing unnecessary waiting time, thereby ensuring that the communication efficiency is maximized, reducing the communication synchronization waiting time, and improving the training efficiency of the model in the distributed computing system, while reducing the communication interruption and system instability, and improving the continuity and stability of the model training process.
[0049] It should be noted that each implementation of the present invention can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.
[0050] In some embodiments, determining a communication timeout threshold for each communication group includes: Before training the target model, an initial value of a communication timeout threshold of each communication group is set; the initial value of the communication timeout threshold is greater than the sum of the operator compilation time and the data set initialization time of the target model.
[0051] Specifically, at the beginning of training, the target model generally performs data set initialization and operator compilation, which is generally time-consuming. Operator compilation refers to the time required to convert operators (such as convolution, matrix multiplication, etc.) into machine code or low-level instructions that can be executed on a specific hardware architecture. Data set initialization refers to the time required for data to be loaded, preprocessed, batched, and other operations.
[0052] Before the target model is trained, the communication capability of the network, the computing capability of the nodes, the storage transmission bandwidth, etc. cannot be accurately understood, so the communication timeout threshold in the first training process cannot be accurately determined.
[0053] Therefore, when setting the initial value of the communication timeout threshold of each communication group, a larger value can be set so that the value is greater than the sum of the operator compilation time and the data set initialization time of the target model. For example, when the number of artificial intelligence processors in the distributed computing system is large (for example, greater than 1000) and the training parameters of the target model are large (for example, greater than 70 billion), the initial value of the communication timeout threshold can be set to 30 minutes.
[0054] In model hybrid parallel training, the amount of data transmitted and the communication links of different communication groups are not necessarily the same. Therefore, the communication timeout thresholds of different communication groups cannot use a unified parameter. Instead, each communication group has an independent communication timeout threshold. When creating distributed hybrid parallel communication groups such as DP, PP, and TP, you can pass in the corresponding timeout setting parameters. For the DP communication group, you can create an array that is used to store the communication timeout threshold (timeout) of each process. Similarly, for the PP and TP communication groups, there are corresponding arrays to store the communication timeout thresholds.
[0055] The distributed computing system communication method provided by an embodiment of the present invention sets the initial value of the communication timeout threshold to be greater than the sum of the operator compilation time and the data set initialization time of the target model at the beginning of model training, so as to avoid the system from judging the communication failure too early.
[0056] In some embodiments, adjusting the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group includes: Based on the synchronization waiting time of each communication group in the current training process, determining the maximum synchronization waiting time and the synchronization waiting time variance of each communication group in the current training process; Based on the maximum synchronization waiting time and the synchronization waiting time variance, a communication timeout threshold adjustment value of each communication group is determined.
[0057] Specifically, during each training process of the model, the communication time consumption and fluctuation of each communication group can be monitored in real time through communication operator performance statistics tools, etc., to obtain the synchronization waiting time, throughput, communication data volume, etc. of each communication group in the current training process. The synchronization waiting time of each communication group in the current training process may include the synchronization waiting time of each process.
[0058] By counting the synchronization waiting time of each communication group during the current training process, the maximum synchronization waiting time can be determined ( ) and synchronization waiting time variance ( ). The maximum synchronization waiting time may be the maximum synchronization waiting time of each process. The synchronization waiting time variance may be the variance of the synchronization waiting time of each process, and the variance is used to describe the degree of deviation of the synchronization waiting time from the average value.
[0059] Determine the communication timeout threshold adjustment value of each communication group according to the maximum synchronization waiting time and the synchronization waiting time variance ( ), which can be expressed as: .
[0060] in, It is a floating coefficient and can be any value from 1 to 3.
[0061] Each process has its own communication group. Based on the data collected by the performance statistics tool, the synchronization waiting time of each communication group is calculated by classifying them into communication groups ( )、Maximum synchronization waiting time( ) and synchronization waiting time variance ( ).
[0062] Because when training a model, the first training process generally has more additional calculations, such as compiling operators and preparing data sets. These additional calculations will be placed in the processes corresponding to some processors, while the communication on the processes in other processors needs to be in a waiting state, and this waiting time generally takes several minutes. However, after the first training process, the training performance will stabilize. In order to ensure the accuracy of fluctuation statistics, the synchronization waiting time of the first training process and the synchronization waiting time of subsequent training processes will be counted separately.
[0063] The distributed computing system communication method provided by the embodiment of the present invention determines the communication timeout threshold adjustment value of each communication group according to the maximum synchronization waiting time and the synchronization waiting time variance. It not only takes into account the maximum value of the synchronization waiting time, but also takes into account the fluctuation of the synchronization waiting time, thereby improving the accuracy of the communication timeout threshold adjustment value.
[0064] In some embodiments, determining the communication timeout threshold adjustment value of each communication group based on the maximum synchronization waiting time and the synchronization waiting time variance includes: Compare the maximum synchronization waiting time in the current training process with the maximum synchronization waiting time in each training process before the current training process, and determine multiple maximum synchronization waiting time changes; When the maximum value among the multiple maximum synchronization waiting time changes is greater than the first preset threshold, determining the communication timeout threshold adjustment value of each communication group based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process; and / or, comparing the synchronization waiting time variance in the current training process with the synchronization waiting time variance in each training process before the current training process, to determine a plurality of synchronization waiting time variance changes; When the maximum value among the multiple synchronization waiting time variance changes is greater than the second preset threshold, the communication timeout threshold adjustment value of each communication group is determined based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process.
[0065] Specifically, multiple training processes are performed, and the maximum synchronization waiting time and the synchronization waiting time variance in each training process are recorded.
[0066] On the one hand, the maximum synchronization waiting time in the current training process can be compared with the maximum synchronization waiting time in each training process before the current training process, and the comparison method can be a difference operation to obtain multiple maximum synchronization waiting time changes. If the maximum value of these maximum synchronization waiting time changes is greater than the first preset threshold, the communication timeout threshold adjustment value of each communication group can be determined based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process.
[0067] On the other hand, the synchronization waiting time variance in the current training process can also be compared with the synchronization waiting time variance in each training process before the current training process, and the comparison method can be a difference operation to obtain multiple synchronization waiting time variance changes. If the maximum value of these synchronization waiting time variance changes is greater than the second preset threshold, the communication timeout threshold adjustment value of each communication group can be determined based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process.
[0068] The above two aspects can be used alone or in combination to determine whether to adjust the communication timeout threshold of each communication group. The communication timeout threshold adjustment value will be used in the next training process.
[0069] It is also possible to determine whether to adjust the communication timeout threshold based on the change ratio of the maximum synchronization waiting time. For example, the maximum synchronization waiting time of the current training process (new ) and the maximum synchronization waiting time in the previous multiple training processes (old ) is changed in the ratio of ((new latency – old latency) / old latency). If the absolute value of this value is greater than 0.2, it can be considered that the network status has changed significantly and the communication timeout threshold needs to be adjusted.
[0070] The distributed computing system communication method provided by the embodiment of the present invention can avoid frequent adjustment of the communication timeout threshold, thereby improving the reliability and stability of the communication timeout threshold adjustment.
[0071] In some embodiments, the maximum synchronization waiting time and the synchronization waiting time variance in each training process are recorded. At the same time, the current model parameters of the target model, the training progress, and the state parameters of the optimizer in the distributed computing system and other parameters can be saved as the current training status (checkpoint) of the target model.
[0072] In some embodiments, training can be restarted based on the saved current training state (checkpoint). The communication timeout threshold of each communication group is updated when training is started. For the communication timeout threshold of the first training process after restarting training, the total initialization time of the first training process previously recorded (before restarting training) is multiplied by a coefficient greater than 1, such as 1.2, and this time is used as the communication timeout threshold of the first training process.
[0073] In some embodiments, the method further comprises: Obtain the heartbeat information of each computing node in the distributed computing system; the heartbeat information is used to indicate the health status of the computing node; In the event that the heartbeat information of any computing node is abnormal, it is determined that any computing node is faulty; Resynchronize any computing node, or distribute the training task executed by any computing node to other computing nodes; the heartbeat information of other computing nodes is normal.
[0074] Specifically, a computing node monitoring process (master process) can be set up in the distributed computing system, which is responsible for information exchange with the training processes in each artificial intelligence processor in all computing nodes and collecting heartbeat information of each computing node. Heartbeat information refers to a signal or message sent periodically by a server (computing node) to indicate to other servers, monitoring systems or clients that it is still operating normally and remains online. The heartbeat mechanism is widely used in distributed systems, database clusters, load balancing, fault-tolerant systems, and other applications that require real-time monitoring and management.
[0075] Heartbeat information is used to indicate the health status of the computing node. When the heartbeat of a computing node is not received for a long time, it means that the node is faulty and training task fault recovery is required. In addition, each computing node can also periodically call the artificial intelligence processor (such as GPU) and network status detection instructions to monitor the network communication status, including communication delay, packet loss rate and bandwidth usage. These indicators can be monitored in real time through network detection tools to identify communication anomalies. Simple operators can also be executed to verify whether the GPU is calculating normally. Then report the status of the current computing node to the master process. When the master process finds that the GPU of a computing node is in an unhealthy state, such as card drop, PCIE link speed drop, GPU error, operator error, etc., it will also initiate a training task fault recovery operation.
[0076] For any computing node, if the heartbeat information of the computing node is abnormal, it can be determined that the computing node is faulty.
[0077] Once a fault is detected, the system will automatically try to resynchronize or reallocate tasks to other healthy nodes based on the type of fault to minimize the impact of training interruption. For example, if the fault code of the GPU card determines that it is a restartable fault, the GPU will be reset on the computing node, and the training process will be restarted and resynchronized. If it is an unrecoverable error or the node heartbeat cannot be received, the task needs to be scheduled to other computing nodes and then the training task will be started.
[0078] The distributed computing system communication method provided by the embodiment of the present invention can quickly identify faults and automatically recover, and can keep the system running even when some nodes fail, thereby enhancing the robustness of the system.
[0079] In some embodiments, after allocating the training task executed by any computing node to other computing nodes, the method further includes: Based on the heartbeat information of each computing node, determine the faulty computing node and remove the faulty computing node from the distributed computing system; Based on the remaining computing nodes in the distributed computing system and the current training status of the target model, the target model continues to be trained.
[0080] Specifically, after the training task is restarted, each computing node can be checked, for example, the heartbeat information of each computing node can be obtained, the faulty computing node can be determined, and the faulty computing node can be removed from the distributed computing system. If there is a new computing node that can be replaced, each computing node needs to be checked again after the new computing node is added to the distributed computing system.
[0081] According to the current training status of the target model, the current model parameters, training progress, and the state parameters of the optimizer in the distributed computing system are determined, and the target model is continued to be trained using the remaining computing nodes in the distributed computing system.
[0082] The distributed computing system communication method provided by the embodiment of the present invention continues to train the target model, avoiding the need to start training from scratch, and saving time and computing resources.
[0083] The following describes a system provided by an embodiment of the present invention. The system described below and the method described above can be referenced to each other.
[0084] Figure 3 is a schematic diagram of the structure of the distributed computing system communication device provided by the present invention, such as Figure 3 As shown, the device comprises: Determination module 310, used to determine the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model; A detection module 320, configured to determine the communication synchronization state of each communication group in the current training process based on a comparison result of a communication timeout threshold of each communication group and a synchronization waiting time of each communication group in the current training process; The adjustment module 330 is used to adjust the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization status of each communication group in the next training process.
[0085] The communication device of the distributed computing system provided by the present invention determines the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model; based on the comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process, the communication synchronization state of each communication group in the current training process is determined; based on the synchronization waiting time of each communication group in the current training process, the communication timeout threshold of each communication group is adjusted to obtain the communication timeout threshold adjustment value of each communication group; based on the communication timeout threshold adjustment value of each communication group, the communication synchronization state of each communication group in the next training process is detected; because the communication timeout threshold of each communication group is dynamically adjusted during the model training process, it is avoided that the system judges that the communication fails too early, and it is avoided that unnecessary waiting time is caused, so as to ensure the maximization of communication efficiency, reduce the communication synchronization waiting time, improve the training efficiency of the model in the distributed computing system, and at the same time reduce the communication interruption and system instability, and improve the continuity and stability of the model training process.
[0086] Figure 4 is a schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 4 As shown, the electronic device may include: a processor (Processor) 410, a communication interface (Communications Interface) 420, a memory (Memory) 430 and a communication bus (Communications Bus) 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic command in the memory 430 to execute the method described in the above embodiment, for example: Determine the communication timeout threshold of each communication group; the communication group is determined based on the process having a data communication relationship in each computing node in the distributed computing system; the process is used to execute the training task of the target model; based on the comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process, determine the communication synchronization state of each communication group in the current training process; adjust the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
[0087] In addition, the logic commands in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several commands to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.
[0088] The processor in the electronic device provided in the embodiment of the present invention can call the logic instructions in the memory to implement the above method. Its specific implementation method is consistent with the implementation method of the aforementioned method and can achieve the same beneficial effects, which will not be repeated here.
[0089] An embodiment of the present invention further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.
[0090] Its specific implementation is consistent with the aforementioned method implementation and can achieve the same beneficial effects, so it will not be repeated here.
[0091] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described above is implemented.
[0092] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0093] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0094] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A distributed computing system communication method, characterized in that: include: Determine a communication timeout threshold for each communication group; The communication group is determined based on processes having data communication relationships in various computing nodes in the distributed computing system; The process is used to execute the training task of the target model; Determine the communication synchronization state of each communication group in the current training process based on a comparison result of the communication timeout threshold of each communication group and the synchronization waiting time of each communication group in the current training process; The communication timeout threshold of each communication group is adjusted based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
2. The distributed computing system communication method according to claim 1, characterized in that: The determining of the communication timeout threshold of each communication group includes: Before training the target model, an initial value of a communication timeout threshold of each communication group is set; the initial value of the communication timeout threshold is greater than the sum of the operator compilation time and the data set initialization time of the target model.
3. The distributed computing system communication method according to claim 1, characterized in that: The adjusting the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group includes: Based on the synchronization waiting time of each communication group in the current training process, determining the maximum synchronization waiting time and the synchronization waiting time variance of each communication group in the current training process; Based on the maximum synchronization waiting time and the synchronization waiting time variance, a communication timeout threshold adjustment value of each communication group is determined.
4. The distributed computing system communication method according to claim 3, characterized in that: The determining, based on the maximum synchronization waiting time and the synchronization waiting time variance, of the communication timeout threshold adjustment value of each communication group comprises: Compare the maximum synchronization waiting time in the current training process with the maximum synchronization waiting time in each training process before the current training process, and determine multiple maximum synchronization waiting time changes; When the maximum value among the multiple maximum synchronization waiting time changes is greater than the first preset threshold, determining the communication timeout threshold adjustment value of each communication group based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process; and / or, comparing the synchronization waiting time variance in the current training process with the synchronization waiting time variance in each training process before the current training process, to determine a plurality of synchronization waiting time variance changes; When the maximum value among the multiple synchronization waiting time variance changes is greater than the second preset threshold, the communication timeout threshold adjustment value of each communication group is determined based on the maximum synchronization waiting time and the synchronization waiting time variance in the current training process.
5. The distributed computing system communication method according to any one of claims 1 to 4, characterized in that: The method further comprises: Acquire heartbeat information of each computing node in the distributed computing system; the heartbeat information is used to indicate the health status of the computing node; In the event that the heartbeat information of any computing node is abnormal, determining that any computing node is faulty; Resynchronize any of the computing nodes, or distribute the training task executed by any of the computing nodes to other computing nodes.
6. The distributed computing system communication method according to claim 5, characterized in that: After allocating the training task executed by any computing node to other computing nodes, the method further includes: Determine the faulty computing node based on the heartbeat information of each computing node, and remove the faulty computing node from the distributed computing system; Based on the remaining computing nodes in the distributed computing system and the current training status of the target model, continue to train the target model.
7. A distributed computing system communication device, characterized in that: include: A determination module, used to determine a communication timeout threshold for each communication group; The communication group is determined based on processes having data communication relationships in various computing nodes in the distributed computing system; The process is used to execute the training task of the target model; A detection module, configured to determine the communication synchronization state of each communication group in the current training process based on a comparison result of a communication timeout threshold of each communication group and a synchronization waiting time of each communication group in the current training process; An adjustment module is used to adjust the communication timeout threshold of each communication group based on the synchronization waiting time of each communication group in the current training process to obtain the communication timeout threshold adjustment value of each communication group; the communication timeout threshold adjustment value of each communication group is used to detect the communication synchronization state of each communication group in the next training process.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the distributed computing system communication method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the distributed computing system communication method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the distributed computing system communication method according to any one of claims 1 to 6 is implemented.