Large language model distributed training method and system based on dynamic resource scheduling
Through dynamic resource scheduling and gradient compression algorithms, distributed training of large language models is solved, resource utilization imbalance and communication bottlenecks are achieved, and efficient and stable training of hundreds of billions of parameter model is achieved.
Patent Information
- Application Number
- CN202510764976.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Traditional distributed training methods have problems such as unbalanced resource utilization, communication bottlenecks and insufficient fault tolerance in large-scale pre-training models, which are difficult to meet the training needs of hyper-large-scale models.
A large language model distributed training method based on dynamic resource scheduling is adopted to periodically collect resource state data of computing nodes, and subtasks are allocated using reinforcement learning strategies, combining gradient compression algorithms and elastic fault tolerance mechanisms to optimize resource utilization and communication efficiency.
显著提升了大语言模型的训练效率和稳定性,支持千亿级参数模型的稳定训练,优化了资源利用率和容错能力。
Smart Images

Figure CN120278283A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and distributed computing, and particularly relates to a distributed training method and system for large language models based on dynamic resource scheduling. Background Art
[0002] With the rapid growth of the parameter scale of large language models, many problems and technical bottlenecks have gradually emerged in the practical application of traditional distributed training methods. Currently, large-scale pre-trained models often contain tens of billions or even trillions of parameters, which pose extremely high requirements on the underlying computing, storage, and communication systems. To support such a huge computing task, distributed training has become an inevitable choice, and various parallel strategies such as data parallelism, model parallelism, and tensor parallelism are widely adopted. However, in the actual training process, the design and execution methods of traditional distributed training architectures are difficult to meet the training requirements of ultra-large models.
[0003] First of all, in a heterogeneous computing environment, the problem of unbalanced resource utilization is particularly prominent. Distributed training is usually deployed on a cluster composed of different types and specifications of computing nodes (such as GPUs, TPUs, and high and low configuration CPU servers). There are often situations where some nodes are overloaded and some node resources are idle, resulting in the training speed being limited by the slowest computing node, thus significantly reducing the overall efficiency.
[0004] Secondly, as the model scale continues to expand, the communication cost of parameter synchronization also increases exponentially. In mainstream training frameworks, multiple computing nodes need to frequently perform gradient aggregation and parameter synchronization operations. Especially in the case of adopting the full synchronization mechanism, a large amount of data is frequently transmitted in the network, resulting in an obvious communication bottleneck and severely restricting the training performance.
[0005] In addition, the current distributed training systems generally have weak fault tolerance capabilities. During a long training process, phenomena such as node failures, GPU crashes, or network fluctuations are relatively common. However, most training frameworks lack automatic detection and fault isolation mechanisms, and once an anomaly occurs, it may lead to the interruption of training or even the overall failure of the task. Summary of the Invention
[0006] The present invention provides a distributed training method and system for large language models based on dynamic resource scheduling to solve the defect that the prior art is difficult to support the efficient training of ultra-large models.
[0007] The present invention provides a distributed training method for large language models based on dynamic resource scheduling, including: For a distributed training cluster including heterogeneous computing nodes, collecting resource status data of each computing node at periodic time intervals; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; In the current training batch, divide the training tasks of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type labels. Adopt a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node. The parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
[0008] According to a distributed training method for a large language model based on dynamic resource scheduling provided by the present invention, the step of allocating each type of subtask to the optimal computing node in an optimal proportion based on the resource status data of each computing node and the task description data of each type of subtask by using a reinforcement learning strategy includes: Construct a state space composed of the GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and task description vectors output by the graph neural network based on the task description data of each type of subtask. Construct an action space including a task allocation weight matrix; wherein, the task allocation weight matrix defines the task proportions of each type of subtask borne by each computing node in the corresponding training batch. Based on a preset reward function, the state space, and the action space, use a pre-trained deep Q-network to determine the types of subtasks borne by each computing node and the task proportions of the corresponding type of subtasks.
[0009] According to a distributed training method for a large language model based on dynamic resource scheduling provided by the present invention, the preset reward function is specifically: R = α × resource balance coefficient + β × single-iteration time shortening rate - γ × node overload penalty Wherein, α, β, and γ are respectively preset weight coefficients, and the resource balance coefficient is determined based on the reciprocal of the variance of the resource utilization rates of each computing node.
[0010] According to a distributed training method for a large language model based on dynamic resource scheduling provided by the present invention, the step of adopting a gradient compression algorithm to compress the gradient data generated on the computing node and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node includes: Sparsify the gradient data generated on the computing node based on the current compression rate, and generate a non-zero gradient index table; Perform dynamic quantization encoding on the sparsified gradient data, and map 32-bit floating-point values to an 8-bit integer space according to the gradient distribution range; Real-time monitor the network bandwidth utilization rate of the computing node. If the current bandwidth utilization rate exceeds the preset threshold, reduce the compression rate and enable quadratic Huffman coding to compress the non-zero gradient index table.
[0011] According to a large language model distributed training method based on dynamic resource scheduling provided by the present invention, the quantization range of the dynamic quantization encoding is determined by the following formula: scale_factor = 127 / max(|gradmin|, |gradmax|) where gradmin and gradmax are the minimum and maximum values of the sparsified gradients in the current batch.
[0012] According to a large language model distributed training method based on dynamic resource scheduling provided by the present invention, the method further includes: Perform fault detection on the computing node based on an elastic fault tolerance mechanism, and resume the training tasks responsible for the faulty node based on the incremental checkpoint snapshot and the task migration priority queue.
[0013] According to a large language model distributed training method based on dynamic resource scheduling provided by the present invention, performing fault detection on the computing node based on an elastic fault tolerance mechanism and resuming the training tasks responsible for the faulty node based on the incremental checkpoint snapshot and the task migration priority queue includes: Save snapshots of the full checkpoint and the incremental checkpoint at a preset time interval. The incremental checkpoint only stores the gradient difference data between the current snapshot and the previous snapshot; the gradient difference data is compressed by a differential coding algorithm; Monitor the survival status of the computing node through a heartbeat detection mechanism, and determine whether the computing node is a faulty node based on the number of consecutive unresponsive times; For the faulty node, restore the model parameters based on the snapshot of the latest full checkpoint and the snapshot of the latest incremental checkpoint corresponding to the faulty node, add the unfinished tasks of the faulty node to the task migration priority queue, migrate the unfinished tasks to the standby node based on the priority of the unfinished tasks, and verify the integrity of the model parameters through a consistency protocol.
[0014] According to a large language model distributed training method based on dynamic resource scheduling provided by the present invention, the method further includes: Dynamically adjust the parameter synchronization interval based on the network bandwidth of each computing node; Among them, when the remaining network bandwidth of each computing node is higher than the preset threshold, the parameter synchronization interval is shortened to the minimum allowable value; When the remaining network bandwidth of any computing node is lower than the preset threshold, the parameter synchronization interval is extended to the maximum allowable value.
[0015] According to a large language model distributed training method based on dynamic resource scheduling provided by the present invention, the periodic time interval is 1 second, and resource status data is collected through a lightweight proxy program. The proxy program is deployed on each computing node, and historical resource status data sequences are stored through a time series database.
[0016] The present invention also provides a large language model distributed training system based on dynamic resource scheduling, including: A resource monitoring module for collecting resource status data of each computing node at a periodic time interval for a distributed training cluster including heterogeneous computing nodes; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; A dynamic scheduling decision module for dividing the training tasks of the large language model into multiple types of subtasks in the current training batch, and based on the resource status data of each computing node and the task description data of each type of subtask, using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in the optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type tags; A gradient optimization communication module for compressing the gradient data generated on the computing node using a gradient compression algorithm, and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; A parameter update module for weighted fusion of the compressed gradient data of different computing nodes based on the parameter synchronization interval on the parameter server side, updating the global model parameters based on the fusion result, and broadcasting the global model parameters to each computing node.
[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the large language model distributed training method based on dynamic resource scheduling as described in any one of the above.
[0018] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the large language model distributed training method based on dynamic resource scheduling as described in any one of the above.
[0019] The present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the large language model distributed training method based on dynamic resource scheduling as described in any one of the above.
[0020] The large language model distributed training method and system based on dynamic resource scheduling provided by the present invention collect resource status data of each computing node at periodic time intervals, so as to divide the training task of the large language model into multiple types of subtasks in the current training batch, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; in addition, a gradient compression algorithm is used to compress the gradient data generated on the computing node, and the compression rate of the gradient data is dynamically adjusted in combination with the current network bandwidth utilization rate of the computing node; the parameter server then performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node, which can significantly improve the training efficiency of the large language model, achieve a greater improvement in resource utilization rate under the same hardware conditions, and support the stable training of a large language model with hundreds of billions of parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0022] Figure 1 is a flowchart of the large language model distributed training method based on dynamic resource scheduling provided by the present invention; Figure 2 is a flowchart of the task scheduling method provided by the present invention; Figure 3 is a flowchart of the gradient compression method provided by the present invention; Figure 4 is a structural diagram of the large language model distributed training system based on dynamic resource scheduling provided by the present invention; Figure 5 is a structural diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Figure 1 is a schematic flowchart of a large language model distributed training method based on dynamic resource scheduling provided by the present invention. As Figure 1 shown, the method includes: Step 110: For a distributed training cluster including heterogeneous computing nodes, collect resource status data of each computing node at periodic time intervals; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; Step 120: In the current training batch, divide the training tasks of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type labels; Step 130: Use a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; Step 140: The parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
[0025] Here, a distributed training cluster containing heterogeneous computing nodes can be constructed. The distributed training cluster includes at least one parameter server (equipped with a high-bandwidth network interface card) and multiple computing nodes. The computing nodes are of a hybrid GPU architecture (for example: 4 NVIDIA A100 nodes, 8 V100 nodes). The parameter server and the computing nodes communicate through the RDMA protocol, and the network topology is a Fat-Tree structure. In addition, a distributed training framework (such as PyTorch Distributed or Horovod) is deployed on both the parameter server and the computing nodes, and a dynamic resource monitoring agent is installed on each computing node to collect the resource status data of each computing node at periodic time intervals. The resource status data includes GPU computing power utilization, video memory occupancy, remaining network bandwidth, and gradient data distribution characteristics, etc. Among them, the periodic time interval can be 1 second, and the resource status data is collected through a lightweight agent program. The agent program is deployed on each computing node and can store the historical resource status data sequence through a time series database. In some embodiments, the monitoring agent of each computing node uploads the collected resource status data to the status database of the parameter server (for example, using Redis to implement real-time data caching). The status database stores the resource status data sequence of the most recent 10 minutes according to the timestamp for subsequent dynamic resource analysis and scheduling.
[0026] After collecting the resource status data of each computing node, the resource status data of each computing node can be encoded into a vector, for example: '[GPU utilization, video memory occupancy, remaining bandwidth]', and the vector is normalized to eliminate the dimension difference.
[0027] For the training task of large language models, training is usually carried out in batches. In the current training batch, the training task of the large language model can be divided into multiple types of subtasks, such as the forward propagation task (Forward), the backward propagation task (Backward), etc. Among them, the classification method of the training task can be determined based on the parallel training strategy supported by the current environment, and the embodiments of the present invention do not make specific limitations on this. Since a training batch contains multiple training samples (for example, 10,000), there are also multiple subtasks of the same type. In order to better improve the training efficiency and stability of large language models, based on the resource status data of each computing node and the task description data of various types of subtasks, the reinforcement learning strategy can be used to allocate various types of subtasks to the optimal computing nodes in the optimal proportion. Here, the task description data includes computational intensity, video memory requirements, communication dependencies, execution latency, and task type tags. It should be noted that after using the reinforcement learning strategy to allocate various types of subtasks to the optimal computing nodes in the optimal proportion, the same type or different types of subtasks can be executed on the same computing node, and multiple subtasks of the same type can be distributed to different computing nodes to improve the resource utilization rate of the computing nodes, thereby improving the training efficiency and stability of large language models.
[0028] In some embodiments, as Figure 2 shown, the using of the reinforcement learning strategy to allocate various types of subtasks to the optimal computing nodes based on the resource status data of each computing node and the task description data of various types of subtasks includes: Step 210, constructing a state space composed of the GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and task description vectors output by the graph neural network based on the task description data of various types of subtasks; Step 220, constructing an action space including a task allocation weight matrix; wherein, the task allocation weight matrix defines the task proportions of various types of subtasks borne by each computing node in the corresponding training batch; Step 230, based on a preset reward function, the state space, and the action space, using a pre-trained deep Q network to determine the types of subtasks borne by each computing node and the task proportions of the corresponding types of subtasks.
[0029] Among them, a pre-trained deep Q-network can be used to implement scheduling decisions based on reinforcement learning techniques, and the model inference frequency is synchronized with the periodic time interval (making a decision once per second). The state space of the deep Q-network covers the GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and the task description vectors output by the graph neural network based on the task description data of various subtasks. That is, the input of the deep Q-network can be a multi-dimensional vector composed of the current GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and the task description vectors of various subtasks in the current training batch output by the graph neural network. Here, a task node graph can be constructed based on the task description data of various subtasks and the temporal dependence relationship between various subtasks. Each node in the task node graph corresponds to a type of subtask, and the node attributes of each node correspond to the task description data of the corresponding subtask. Inputting this task node graph into the graph neural network can obtain the node vector of each node output by the graph neural network as the task description vector of the corresponding type of subtask.
[0030] The action space of the deep Q-network includes a task allocation weight matrix. Among them, the task proportion of various subtasks undertaken by each computing node in the corresponding training batch is defined in this task allocation weight matrix. That is, the output of the deep Q-network can be the task proportion of various subtasks undertaken by each computing node in the corresponding training batch. In some other embodiments, in order to take into account the communication overhead in the distributed training process, the action space can also include a parameter synchronization interval. That is, the output of the deep Q-network includes the task proportion of various subtasks undertaken by each computing node in the corresponding training batch and the parameter synchronization interval. This parameter synchronization interval determines the parameter synchronization frequency between each computing node and the parameter server. Therefore, by optimizing the parameter synchronization interval in real time through the deep Q-network, the task allocation and parameter synchronization strategies can be jointly optimized to further improve the training efficiency and stability of the large language model.
[0031] Based on the preset reward function, the above-designed state space and action space, the deep Q-network can output the types of subtasks undertaken by each computing node and the task proportion of the corresponding type of subtasks, realizing the efficient distributed training of the large language model.
[0032] In some embodiments, the preset reward function is specifically: R = α × resource balance coefficient + β × single-iteration time shortening rate - γ × node overload penalty Among them, α, β, and γ are respectively preset weight coefficients, and the resource balance coefficient is determined based on the reciprocal of the variance of the resource utilization rates of each computing node.
[0033] In some other embodiments, the parameter synchronization interval can also be dynamically adjusted based on the network bandwidth of each computing node. Among them, when the remaining network bandwidth of each computing node is higher than the preset threshold, the parameter synchronization interval is shortened to the minimum allowable value; when the remaining network bandwidth of any computing node is lower than the preset threshold, the parameter synchronization interval is extended to the maximum allowable value.
[0034] When the computing node executes subtasks, gradient data of the model will be generated, and a large amount of gradient data will bring great communication pressure and storage pressure. In this regard, a gradient compression algorithm can be used to compress the gradient data generated on the computing node, and the compression rate of the gradient data on the corresponding computing node can be dynamically adjusted in combination with the current network bandwidth utilization rate of the computing node, so as to avoid overloading the computing node due to the excessive amount of gradient data, thereby balancing the model training efficiency and training accuracy.
[0035] In some embodiments, as Figure 3 shown, using the gradient compression algorithm to compress the gradient data generated on the computing node and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node includes: Step 310, sparsify the gradient data generated on the computing node based on the current compression rate, and generate a non-zero gradient index table; Step 320, perform dynamic quantization encoding on the sparsified gradient data, and map 32-bit floating-point values to an 8-bit integer space according to the gradient distribution range; Step 330, monitor the network bandwidth utilization rate of the computing node in real time. If the current bandwidth utilization rate exceeds the preset threshold, reduce the compression rate, and enable secondary Huffman coding to compress the non-zero gradient index table.
[0036] Specifically, the gradient data (i.e., gradient tensor (Tensor)) generated on the computing node can be flattened into a one-dimensional vector, the absolute values of each gradient are calculated and sorted, and the top K% of the sorted gradient absolute values are retained based on the current compression rate (denoted as K%) and the remaining gradient absolute values are set to zero to achieve sparsification. At the same time, record the index positions of the non-zero gradient absolute values to generate a non-zero gradient index table. Subsequently, perform dynamic quantization encoding on the sparsified gradient data, and map 32-bit floating-point values to an 8-bit integer space according to the gradient distribution range.
[0037] For example, the quantization encoding can be implemented based on the following method: quantized_grad=np.round(sparse_grad×scale_factor).astype(np.int8) Among them, the quantization range scale_factor of the dynamic quantization encoding can be determined by the following formula: scale_factor = 127 * max(|gradmin|, |gradmax|) Here, gradmin and gradmax are the minimum and maximum values of the sparsified gradients in the current batch.
[0038] It should be noted that on the parameter server side, the quantized gradient data can be dequantized based on the following method: dequantized_grad = quantized_grad.astype(np.float32) / scale_factor In addition, the network bandwidth utilization rate of the computing nodes can be monitored in real time. When the current bandwidth utilization rate exceeds the preset threshold, the aforementioned compression rate is reduced, and the second Huffman coding is enabled to compress the non-zero gradient index table.
[0039] For example, if the bandwidth utilization rate > 80%, the compression rate is reduced from 10% to 5% to further reduce the data volume; if the bandwidth utilization rate < 50%, the compression rate is restored to 10% to give priority to ensuring the model accuracy.
[0040] The compressed gradient data of the computing nodes is periodically uploaded to the parameter server at the parameter synchronization interval. The parameter server then performs weighted fusion on the compressed gradient data of different computing nodes (the weights corresponding to each computing node can be preset in advance), thereby updating the global model parameters based on the fusion result, and broadcasting the global model parameters to each computing node as the starting point for the next training.
[0041] In some embodiments, in order to ensure the high availability of training and reduce the impact caused by the failure of computing nodes, the computing nodes can be fault-detected based on an elastic fault tolerance mechanism, and the training tasks responsible for the faulty nodes can be restored based on the incremental checkpoint snapshot and the task migration priority queue.
[0042] Specifically, snapshots of full checkpoints and incremental checkpoints can be saved at a preset time interval. For example, a full checkpoint save (storing complete model parameters and optimizer states) is triggered every 30 minutes, and an incremental checkpoint is saved every 5 minutes (only recording the differential gradients from the previous snapshot). Among them, the incremental checkpoint only stores the gradient difference data between the current snapshot and the previous snapshot, and the gradient difference data is compressed by a differential encoding algorithm. In some embodiments, in the incremental checkpoint, the difference ΔW between the current parameter Wt and the parameter Wt−1 saved in the previous checkpoint can be calculated, and ΔW is sparsified and the Zstandard compression algorithm is applied, so that the amount of compressed data is less than 30% of the full checkpoint. The alive state of the computing node is monitored through a heartbeat detection mechanism. For example, the parameter server can send a heartbeat packet to the computing node every 5 seconds, and determine whether the computing node is a faulty node based on the number of consecutive unresponses of the computing node. For example, if a computing node fails to respond continuously for 3 times (accumulative timeout of 15 seconds), it is marked as a faulty node.
[0043] For a faulty node, new tasks can be immediately stopped from being assigned to the faulty node, and its unfinished tasks are recorded. The model parameters are restored based on the snapshot of the latest full checkpoint and the snapshot of the latest incremental checkpoint corresponding to the faulty node, and the unfinished tasks of the faulty node are added to the task migration priority queue to migrate the unfinished tasks to the standby node based on the priority of the unfinished tasks. Among them, the unfinished tasks of the faulty node can be added to the task migration priority queue with a higher priority, and according to the current resource status, the standby node with the lowest load is selected to execute the unfinished tasks of the faulty node. In addition, the integrity of the model parameters needs to be verified through a consistency protocol, that is, comparing the restored model parameters with the latest parameters of the healthy nodes to ensure consistency. If the difference between the restored model parameters and the latest parameters of the healthy nodes exceeds the threshold, full synchronization is triggered to update the model parameters corresponding to the faulty node as the training starting point for the unfinished tasks.
[0044] In summary, the method provided in the embodiments of the present invention collects the resource status data of each computing node at periodic time intervals, thereby dividing the training task of the large language model into multiple types of subtasks in the current training batch, and based on the resource status data of each computing node and the task description data of each type of subtask, using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; in addition, a gradient compression algorithm is used to compress the gradient data generated on the computing node, and the compression rate of the gradient data is dynamically adjusted in combination with the current network bandwidth utilization rate of the computing node; the parameter server then performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node, which can significantly improve the training efficiency of the large language model, achieve a greater improvement in resource utilization under the same hardware conditions, and support the stable training of a large language model with hundreds of billions of parameters.
[0045] The distributed training system of the large language model based on dynamic resource scheduling provided by the present invention will be described below. The distributed training system of the large language model based on dynamic resource scheduling described below can be correspondingly referred to with the distributed training method of the large language model based on dynamic resource scheduling described above.
[0046] Based on any of the above embodiments, Figure 4 is a schematic structural diagram of the distributed training system of the large language model based on dynamic resource scheduling provided by the present invention. As Figure 4 shown, the system includes: A resource monitoring module 410, configured to collect the resource status data of each computing node at periodic time intervals for a distributed training cluster including heterogeneous computing nodes; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; A dynamic scheduling decision module 420, configured to divide the training task of the large language model into multiple types of subtasks in the current training batch, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type tags; A gradient optimization communication module 430, configured to use a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; A parameter update module 440, configured to perform weighted fusion on the compressed gradient data of different computing nodes at the parameter server end based on the parameter synchronization interval, update the global model parameters based on the fusion result, and broadcast the global model parameters to each computing node.
[0047] The system provided by the embodiments of the present invention collects the resource status data of each computing node at periodic time intervals, so that in the current training batch, the training task of the large language model is divided into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, the reinforcement learning strategy is used to allocate each type of subtask to the optimal computing node in the optimal proportion; in addition, the gradient compression algorithm is used to compress the gradient data generated on the computing node, and the compression rate of the gradient data is dynamically adjusted in combination with the current network bandwidth utilization rate of the computing node; the parameter server then performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node, which can significantly improve the training efficiency of the large language model, achieve a greater improvement in resource utilization under the same hardware conditions, and support the stable training of a large language model with hundreds of billions of parameters.
[0048] Based on any of the above embodiments, the using the reinforcement learning strategy to allocate each type of subtask to the optimal computing node in the optimal proportion based on the resource status data of each computing node and the task description data of each type of subtask includes: Construct a state space composed of the GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and task description vectors output by the graph neural network based on the task description data of each type of subtask; Construct an action space including a task allocation weight matrix; wherein, the task allocation weight matrix defines the task proportions of each type of subtask undertaken by each computing node in the corresponding training batch; Based on a preset reward function, the state space and the action space, use the pre-trained deep Q network to determine the types of subtasks undertaken by each computing node and the task proportions of the corresponding type of subtasks.
[0049] Based on any of the above embodiments, the preset reward function is specifically: R = α × resource balance coefficient + β × single iteration time shortening rate - γ × node overload penalty Wherein, α, β, and γ are respectively preset weight coefficients, and the resource balance coefficient is determined based on the reciprocal of the variance of the resource utilization rates of each computing node.
[0050] Based on any of the above embodiments, the using the gradient compression algorithm to compress the gradient data generated on the computing node and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node includes: Perform sparsification processing on the gradient data generated on the computing node based on the current compression rate, and generate a non-zero gradient index table; Perform dynamic quantization encoding on the sparsified gradient data, and map 32-bit floating-point values to an 8-bit integer space according to the gradient distribution range; Real-time monitor the network bandwidth utilization rate of the computing node. If the current bandwidth utilization rate exceeds the preset threshold, reduce the compression rate and enable quadratic Huffman coding to compress the non-zero gradient index table.
[0051] Based on any of the above embodiments, the quantization range of the dynamic quantization encoding is determined by the following formula: scale_factor = 127 / max(|gradmin|, |gradmax|) where gradmin and gradmax are the minimum and maximum values of the current batch of sparsified gradients.
[0052] Based on any of the above embodiments, the system further includes a fault tolerance module for: Perform fault detection on the computing node based on an elastic fault tolerance mechanism, and restore the training tasks responsible for the faulty node based on the incremental checkpoint snapshot and the task migration priority queue.
[0053] Based on any of the above embodiments, performing fault detection on the computing node based on an elastic fault tolerance mechanism and restoring the training tasks responsible for the faulty node based on the incremental checkpoint snapshot and the task migration priority queue includes: Save snapshots of the full checkpoint and the incremental checkpoint at a preset time interval. The incremental checkpoint only stores the gradient difference data between the current snapshot and the previous snapshot; the gradient difference data is compressed through a differential coding algorithm; Monitor the survival status of the computing node through a heartbeat detection mechanism, and determine whether the computing node is a faulty node based on the number of consecutive unresponsive times; For the faulty node, restore the model parameters based on the snapshot of the latest full checkpoint and the snapshot of the latest incremental checkpoint corresponding to the faulty node, add the unfinished tasks of the faulty node to the task migration priority queue, migrate the unfinished tasks to the standby node based on the priority of the unfinished tasks, and verify the integrity of the model parameters through a consistency protocol.
[0054] Based on any of the above embodiments, the system further includes a parameter synchronization optimization module for: Dynamically adjust the parameter synchronization interval based on the network bandwidth of each computing node; where when the remaining network bandwidth of each computing node is higher than the preset threshold, shorten the parameter synchronization interval to the minimum allowable value; When the remaining network bandwidth of any computing node is lower than the preset threshold, extend the parameter synchronization interval to the maximum allowable value.
[0055] Based on any of the above embodiments, the periodic time interval is 1 second, and the resource status data is collected through a lightweight agent program. The agent program is deployed on each computing node, and the historical resource status data sequence is stored in a time series database.
[0056] Figure 5 is a schematic structural diagram of an electronic device provided by the present invention. As Figure 5 shown, the electronic device may include: a processor 510, a memory 520, a communication interface 530, and a communication bus 540. Among them, the processor 510, the memory 520, and the communication interface 530 communicate with each other through the communication bus 540. The processor 510 can call the logical instructions in the memory 520 to execute a large language model distributed training method based on dynamic resource scheduling. The method includes: for a distributed training cluster including heterogeneous computing nodes, collecting the resource status data of each computing node at a periodic time interval; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; in the current training batch, dividing the training task of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type labels; using a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; the parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
[0057] In addition, when the logical instructions in the above-mentioned memory 520 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0058] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the large language model distributed training method based on dynamic resource scheduling provided by the above-mentioned various methods. The method includes: for a distributed training cluster including heterogeneous computing nodes, collecting resource status data of each computing node at periodic time intervals; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; in the current training batch, dividing the training task of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type labels; using a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; the parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
[0059] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the large language model distributed training method based on dynamic resource scheduling provided above. The method includes: for a distributed training cluster including heterogeneous computing nodes, collecting resource status data of each computing node at periodic time intervals; the resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; in the current training batch, dividing the training tasks of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirement, communication dependency, execution delay, and task type label; using a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjusting the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; the parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
[0060] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0061] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0062] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A distributed training method for large language models based on dynamic resource scheduling, characterized in that, Including: For a distributed training cluster containing heterogeneous computing nodes, collecting resource status data of each computing node at periodic time intervals; The resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; In the current training batch, divide the training task of the large language model into multiple types of subtasks, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type tags; Adopt a gradient compression algorithm to compress the gradient data generated on the computing node, and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node; The parameter server performs weighted fusion on the compressed gradient data of different computing nodes based on the parameter synchronization interval, updates the global model parameters based on the fusion result, and broadcasts the global model parameters to each computing node.
2. The large language model distributed training method based on dynamic resource scheduling according to claim 1, wherein, The step of using a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in an optimal proportion based on the resource status data of each computing node and the task description data of each type of subtask includes: Construct a state space composed of the GPU utilization rate, video memory occupancy rate, remaining network bandwidth, current batch gradient data volume of each computing node, and a task description vector output by the graph neural network based on the task description data of each type of subtask; Construct an action space including a task allocation weight matrix; wherein, the task allocation weight matrix defines the task proportion of each type of subtask undertaken by each computing node in the corresponding training batch; Based on a preset reward function, the state space, and the action space, use a pre-trained deep Q network to determine the subtask type undertaken by each computing node and the task proportion of the corresponding type of subtask.
3. The distributed training method of the large language model based on dynamic resource scheduling according to claim 2, wherein, The preset reward function is specifically: R = α × resource balance coefficient + β × single iteration time shortening rate - γ × node overload penalty Wherein, α, β, and γ are respectively preset weight coefficients, and the resource balance coefficient is determined based on the reciprocal of the variance of the resource utilization rates of each computing node.
4. The distributed training method of the large language model based on dynamic resource scheduling according to claim 1, characterized in that, The step of adopting a gradient compression algorithm to compress the gradient data generated on the computing node and dynamically adjust the compression rate of the gradient data in combination with the current network bandwidth utilization rate of the computing node includes: Perform sparsification processing on the gradient data generated on the computing node based on the current compression rate, and generate a non-zero gradient index table; Perform dynamic quantization encoding on the sparsified gradient data, and map 32-bit floating-point values to an 8-bit integer space according to the gradient distribution range; Real-time monitor the network bandwidth utilization rate of the computing node. If the current bandwidth utilization rate exceeds the preset threshold, reduce the compression rate and enable quadratic Huffman coding to compress the non-zero gradient index table.
5. The large language model distributed training method based on dynamic resource scheduling according to claim 4, wherein The quantization range of the dynamic quantization encoding is determined by the following formula: scale_factor = 127 / max(|gradmin|, |gradmax|) Among them, gradmin and gradmax are the minimum and maximum values of the sparse gradient of the current batch.
6. The method for distributed training of a large language model based on dynamic resource scheduling according to claim 1, wherein The method further includes: Performing fault detection on computing nodes based on an elastic fault tolerance mechanism, and restoring the training tasks responsible for the faulty nodes based on incremental checkpoint snapshots and a task migration priority queue.
7. The distributed training method for large language models based on dynamic resource scheduling according to claim 6, wherein, The performing fault detection on computing nodes based on an elastic fault tolerance mechanism, and restoring the training tasks responsible for the faulty nodes based on incremental checkpoint snapshots and a task migration priority queue includes: Saving snapshots of full checkpoints and snapshots of incremental checkpoints at a preset time interval, where the incremental checkpoint only stores gradient difference data between the current snapshot and the previous snapshot; the gradient difference data is compressed through a differential coding algorithm; Monitoring the survival status of computing nodes through a heartbeat detection mechanism, and determining whether a computing node is a faulty node based on the number of consecutive unresponses; For a faulty node, restoring model parameters based on the snapshot of the latest full checkpoint and the snapshot of the latest incremental checkpoint corresponding to the faulty node, adding the unfinished tasks of the faulty node to the task migration priority queue, migrating the unfinished tasks to standby nodes based on the priorities of the unfinished tasks, and verifying the integrity of the model parameters through a consistency protocol.
8. The distributed training method for large language models based on dynamic resource scheduling according to claim 1, wherein, The method further includes: Dynamically adjusting the parameter synchronization interval based on the network bandwidth of each computing node; Among them, when the remaining network bandwidth of each computing node is higher than a preset threshold, shortening the parameter synchronization interval to the minimum allowable value; When the remaining network bandwidth of any computing node is lower than a preset threshold, extending the parameter synchronization interval to the maximum allowable value.
9. The distributed training method of the large language model based on dynamic resource scheduling according to claim 1, wherein, The periodic time interval is 1 second, and resource status data is collected through a lightweight proxy program, which is deployed on each computing node, and historical resource status data sequences are stored through a time series database.
10. A distributed training system for large language models based on dynamic resource scheduling, characterized in that, It includes: A resource monitoring module, which is used to collect resource status data of each computing node at a periodic time interval for a distributed training cluster containing heterogeneous computing nodes; The resource status data includes GPU computing power utilization rate, video memory occupancy rate, remaining network bandwidth, and gradient data distribution characteristics; A dynamic scheduling decision module, which is used to divide the training tasks of a large language model into multiple types of subtasks in the current training batch, and based on the resource status data of each computing node and the task description data of each type of subtask, use a reinforcement learning strategy to allocate each type of subtask to the optimal computing node in the optimal proportion; the task description data includes computing intensity, video memory requirements, communication dependencies, execution latency, and task type tags; A gradient optimization communication module, which is used to compress the gradient data generated on computing nodes using a gradient compression algorithm, and dynamically adjust the compression ratio of the gradient data in combination with the current network bandwidth utilization rate of the computing nodes; A parameter update module, which is used to perform weighted fusion on the compressed gradient data of different computing nodes at the parameter server end based on the parameter synchronization interval, update the global model parameters based on the fusion result, and broadcast the global model parameters to each computing node.
Citation Information
Patent Citations
Asynchronous distributed deep learning training method, device and system
CN110245743A
Distributed training task scheduling method, system and device for intelligent computing
CN115248728A
Multitask processing method and system, computer equipment and storage medium
CN117035111A
Self-adaptive multi-level mixed index data storage system and storage method
CN118152407A
Data migration method and device
CN118550582A
Cited By
Inference request scheduling method and device based on reinforcement learning, equipment and medium
CN120950225A
Large model training video memory optimization method and system based on low-precision integer storage
CN121351902A
AI model intelligent training and reasoning integrated method and system
CN121352030A
Distributed model training method and device, server and storage medium
CN121433915A
Electric power large model distributed training method and system based on FSDP and gradient compression
CN121436095A