Heterogeneous cluster hybrid parallel training method and system combined with freezing mechanism

By optimizing the allocation of heterogeneous cluster resources through dynamic programming and freezing mechanisms, the inefficiency of static schemes in large-scale deep learning model training is solved, and training results with efficient utilization of heterogeneous hardware resources are achieved.

CN121145971AActive Publication Date: 2025-12-16SOUTHEAST UNIV

Patent Information

Application Number
CN202511690946.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-18
Publication Date
2025-12-16
Estimated Expiration
2045-11-18

AI Technical Summary

Technical Problem

In large-scale deep learning model training, existing static parallel schemes cannot adapt to dynamic changes in model freezing and heterogeneous hardware resources, resulting in low computational efficiency and insufficient resource utilization.

Method used

A heterogeneous cluster hybrid parallel training method based on dynamic programming is adopted. By identifying the resources released by the frozen layer, the pipeline segmentation and data parallel configuration are dynamically optimized. Combined with the freezing mechanism and heterogeneous hardware characteristics, the training strategy is adjusted in real time to improve efficiency.

Benefits of technology

It enables efficient training in dynamic model freezing and heterogeneous hardware environments, improving computational efficiency and resource utilization, and shortening training time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145971A_ABST
    Figure CN121145971A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous cluster hybrid parallel training method and system combined with a freezing mechanism, and belongs to the technical field of neural networks. According to the method, in combination with the characteristics of a target heterogeneous hardware cluster, the calculation time and storage requirements of different hardware of different types of hardware of each layer in a model are predicted; solving the optimal assembly line segmentation and data parallel configuration by adopting dynamic programming to obtain a hybrid parallel training scheme; and when the frozen state changes in the training process, the system reallocates the released video memory and computing power resources, and dynamically adjusts the parallel scheme to realize load balancing and improve the training efficiency. According to the method, resource occupation is effectively reduced, and the large-scale model training performance in a heterogeneous environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of neural network technology, and relates to deep learning technology, and in particular to a heterogeneous cluster hybrid parallel training method and system that combines a freezing mechanism. Background Technology

[0002] With the rapid development of deep learning, Transformer-based models (such as GPT-5, DeepSeek-R1, and Tongyi Qianwen 3) have achieved remarkable results in natural language processing, multimodal understanding, and intelligent decision-making. However, the parameter scale of deep learning models is growing exponentially, often reaching tens or even trillions, placing extremely high demands on computing and communication resources. The computing power and memory resources of a single computing device (such as a single GPU) can no longer support the training of such massive models. Therefore, distributed parallel training methods based on large-scale heterogeneous GPU clusters have become the mainstream choice for training large models.

[0003] Existing parallel training strategies mainly include data parallelism, model parallelism, and pipeline parallelism. Data parallelism accelerates training by copying the model across multiple devices and processing different data shards, but as model parameters expand, parameter synchronization communication and GPU memory usage increase dramatically. Model parallelism distributes memory pressure by splitting model parameters or computational layers across multiple devices, but frequent cross-device communication or uneven layer partitioning can lead to idle GPUs and decreased overall throughput. Pipeline parallelism divides the model into multiple stages, with different stages processed on different devices, and improves device utilization through micro-batch overlapping computation and communication. However, its performance depends on reasonable stage partitioning and scheduling strategies and is still constrained by activation value storage and load imbalance.

[0004] To balance performance and resource efficiency across different dimensions, researchers have proposed hybrid parallelism. Hybrid parallelism combines the advantages of data parallelism and pipelined parallelism, integrating data parallelism, model parallelism, and pipelined parallelism in a multi-dimensional manner. For example, pipelined parallelism can be used between stages while data parallelism is used within stages to balance communication and computational load, thus achieving a balance between model scaling and training efficiency to some extent. Typical training frameworks supporting hybrid parallelism include Megatron-LM and DeepSpeed.

[0005] While the aforementioned approaches have achieved some success in large-scale model training, they generally suffer from limitations due to static partitioning and fixed scheduling strategies. This means that before training begins, a parallel strategy (including pipeline stage division and the number of devices for data parallelism within each stage) is determined based on the model structure and GPU cluster hardware environment, and remains unchanged throughout the training process. As the layer parameter update rate and gradient distribution change during model training, static partitioning often leads to some devices being overloaded while others remain idle, resulting in a decrease in overall computational efficiency.

[0006] On the other hand, model freezing has been introduced to further reduce training overhead. This technique identifies model layers that have largely converged or contribute little to the task by monitoring changes in gradient magnitude or activation features, and skips their backpropagation and parameter update operations, thus saving computational and memory resources. However, introducing freezing mechanisms into hybrid parallel training brings new challenges:

[0007] 1. Dynamic load imbalance: The freezing mechanism causes the computational load of each part of the model to change dynamically during training. Static hybrid parallel solutions cannot adapt to this change. The originally balanced pipeline stages may become unevenly loaded due to the freezing of some layers, creating new computational bottlenecks and reducing overall efficiency.

[0008] 2. Complexity of Heterogeneous Environments: Modern GPU clusters often contain GPUs of different models and performance from the same manufacturer (heterogeneous computing resources), and there may be connections with different bandwidths between and within nodes (heterogeneous network interconnection). Static or simple dynamic adjustment strategies are difficult to fully utilize the characteristics of heterogeneous resources. For example, they may fail to allocate computationally intensive tasks to GPUs with high computing power, or fail to prioritize the use of high-bandwidth connections for parallel data communication.

[0009] 3. Difficulty in Optimal Solution Planning: In dynamically changing frozen states and heterogeneous hardware environments, how to evaluate the performance of different hybrid parallel solutions (stage partitioning, data parallelism allocation) in real time and accurately, and quickly find or approximate the optimal solution, is a complex optimization problem. Traditional methods based on static models or simple heuristic rules are inadequate for this task.

[0010] Existing technologies (including methods built into frameworks such as Megatron-LM and DeepSpeed) struggle to achieve a dynamic optimal balance between training efficiency and resource utilization in environments with dynamically frozen models and heterogeneous computing resources. Therefore, there is an urgent need for a training system and method that can adapt to dynamically frozen models, fully utilize heterogeneous computing resources, and optimize hybrid parallel strategies in real time to achieve efficient training of large-scale deep learning models. Summary of the Invention

[0011] To address the issues that existing static parallel schemes cannot adapt to dynamic changes in model freezing and are difficult to efficiently utilize heterogeneous hardware resources when training large-scale models, this invention provides a heterogeneous cluster hybrid parallel training method and system that combines freezing mechanisms. By identifying and utilizing the GPU computing and memory resources released by the frozen layers in the model, the method dynamically optimizes the allocation of heterogeneous cluster resources and the hybrid parallel strategy, thereby improving training efficiency and reducing resource consumption.

[0012] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A heterogeneous cluster hybrid parallel training method incorporating a freezing mechanism includes the following steps: Step 1, Model Analysis: Receive the deep learning model, analyze the model's network structure, and combine the characteristics of the target heterogeneous hardware cluster to predict the computation time and storage requirements of each layer in the model on different types of hardware, and consider the impact of the frozen state on storage requirements and computation time. Step 2, Scheme Planning: Based on the model layer freezing status, model parsing results, and heterogeneous hardware information, dynamic programming is used to solve for the optimal pipeline partitioning and data parallel configuration to obtain a hybrid parallel training scheme. During the dynamic programming process, it is necessary to simultaneously satisfy the constraint that the memory usage of all model layers allocated on a specific GPU device does not exceed the device's upper limit. Step 3, Model Training: The training process of the hybrid parallel training scheme is executed. During the training process, when the frozen state changes, the system reallocates the released video memory and computing resources, dynamically adjusts the parallel scheme, and achieves load balancing and improved training efficiency.

[0013] Furthermore, step 1 specifically includes: (1) Use model parsing tools to obtain the type, dimensions and connection methods of each layer of the model; (2) Use hardware detection tools to obtain information on the GPU model, computing power, memory size, and interconnect bandwidth between GPUs in the heterogeneous cluster; (3) Computation time prediction: For each layer, the prediction is based on its operation type, input / output dimensions, number of parameters, and the computing power of the GPU model on which it will be run. Predict computation time using performance models or empirical formulas. : , Storage requirement forecasting: Calculate the memory required for each layer, including the memory occupied by parameters. Memory occupied by optimizer state Activation of occupied memory and the memory occupied by gradients : , , , , The final result shows the actual memory occupied by one layer: , in, Number of input / output channels The kernel size is the convolution kernel size. To output the height and width of the feature map, and The height and width of the input feature map.

[0014] Furthermore, step 2 specifically includes the following process: The goal of the optimization scheme is to minimize the expected total time for a single training batch. ,in From the start-up phase time Stable phase time and end phase time constitute: , , , , Conditional constraints are imposed on the actual amount of GPU memory used by the model during training: , in, This is the bottleneck stage. Stages Forward and backward propagation times on the allocated heterogeneous GPUs, , Each represents a stage Forward and backward propagation times, Representative stage The backpropagation time, This refers to the quantity in a micro-batch. For the stage The set of parameters involved In the equipment group All-Reduce overhead, This refers to the maximum video memory capacity of the GPU device. This refers to the total number of layers in the model, which is the sum of all convolutional layers, residual blocks, and fully connected layers in the network. For the model number layer, For layer The frozen state The first step in the model analysis process, after considering the frozen state, is the calculation of the... Layer memory requirements, It is a heterogeneous GPU collection; First, identify the bottleneck stage of the production line. Then, dynamic programming is used to solve for the optimal hybrid parallel configuration; when the frozen state of the model changes, replanning will be performed in conjunction with the resources released from the frozen layer; the objective function of the problem is defined as follows. : , , , in, The total number of layers in the model, i.e., the sum of all convolutional layers, residual blocks, and fully connected layers in the network, is the frozen state list. This will affect the computation and communication time at each stage. This represents the total number of GPUs. It is a heterogeneous GPU collection. As the dividing point of the layer, The number of GPUs allocated, For a specific subset of GPUs, the optimization process needs to consider The heterogeneity of GPUs and the interconnect bandwidth between them, Indicates from all GPU collections Selected All possible combinations of GPUs; When predicting the computation time of a hierarchical level, the impact of the frozen state on the computation should be considered: if the level If the process is frozen, its forward propagation still needs to be executed but it does not participate in the backward propagation. The backward propagation time is... =0; In terms of computing resources, the memory requirements of the gradients and optimizer states in the frozen layer are reduced to zero, and the memory and computing resources released by the freeze are reallocated. The optimal solution can be obtained by using the concept of dynamic programming. During training, monitor changes in the frozen state of model layers; when a layer is frozen or thawed, reassess the computational load and memory usage of each pipeline stage; based on the updated load and available resources, recalculate and apply the optimal hybrid parallel scheme.

[0015] Furthermore, the process of solving the problem using the idea of ​​dynamic programming includes: Frozen status list Define the state based on the known information. Before processing Layers that utilize heterogeneous GPU subsets Minimum pipeline time for heterogeneous GPU subset Include One GPU; by iteratively expanding the state, consider moving to the next layer ( Assign a new subset of GPUs Calculate the resulting new stage time and total time, update the DP table, and obtain the final result. This is the optimal solution.

[0016] Furthermore, step 3 uses a computation graph to represent and schedule operations, constructing forward and backward propagation subgraphs for each pipeline stage, with the core scheduling strategy being 1F1B.

[0017] Furthermore, the 1F1B strategy refers to scheduling the corresponding backpropagation as early as possible after the forward propagation of a micro-batch is completed within a pipeline stage, alternating with the forward propagation of subsequent micro-batches.

[0018] Furthermore, in step 3, the following communication mechanism is used for data transmission when handling cross-stage communication: During forward propagation, when the data parallelism of downstream stages increases or the receiving GPU type / bandwidth differs, a separation layer is used to copy or distribute data; when the data parallelism of downstream stages decreases or needs to be converged to a specific type / bandwidth GPU, an aggregation layer is used to collect and integrate data. During backpropagation, gradient copying is used when the data parallelism of the downstream stage increases or the GPU type / bandwidth differs; gradient aggregation is used when the data parallelism of the downstream stage decreases or gradients need to be aggregated to a specific type / bandwidth of GPU, where the overhead of All-Reduce is affected by the bandwidth between the heterogeneous GPUs involved.

[0019] This invention also provides a heterogeneous cluster hybrid parallel training system incorporating a freezing mechanism, including a model parsing module, a scheme planning module, and a model training module, for implementing the aforementioned heterogeneous cluster hybrid parallel training method incorporating a freezing mechanism.

[0020] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. High adaptability: Through dynamic scheme planning, it can adapt to changes in computational load caused by model layer freezing during training, and always maintain high training efficiency.

[0021] 2. Heterogeneous awareness: Clearly consider the differences in GPU performance (computing speed, memory) and network topology heterogeneity (bandwidth), and carry out targeted task allocation and communication optimization to make full use of heterogeneous cluster resources.

[0022] 3. Efficiency Improvement: Combining the scalability of hybrid parallelism, the memory efficiency of 1F1B scheduling, the computational savings of the freezing mechanism, and the optimization for heterogeneous hardware, it significantly shortens the training time of large-scale models.

[0023] 4. Resource optimization: By dynamically adjusting the parallel strategy, computing, memory and bandwidth resources are allocated and utilized more rationally, thereby improving the overall resource utilization rate. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the internal modules of the heterogeneous hybrid parallel training method combined with the freezing mechanism provided by the present invention.

[0025] Figure 2 This is a typical Gantt chart example of hybrid parallel training, demonstrating the 1F1B scheduling strategy.

[0026] Figure 3 This is a schematic diagram of the design of the separation layer and aggregation layer for handling cross-stage communication in hybrid parallel processing.

[0027] Figure 4 It is an algorithm for identifying bottleneck stages.

[0028] Figure 5 It is a dynamic programming algorithm. Detailed Implementation

[0029] The technical solutions provided by the present invention will be described in detail below with reference to specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0030] Reference Figure 1 The overall implementation framework of the heterogeneous hybrid parallel training method combined with a freezing mechanism proposed in this invention comprises three main modules: a model parsing module, a scheme planning module, and a model training module. This framework is equipped with a dynamic adjustment design oriented towards the freezing mechanism. All three modules are software modules, implementing corresponding functions. Therefore, this invention actually includes three steps: a model parsing step, a scheme planning step, and a model training step, each corresponding to the specific functions implemented by the three modules. The functions implemented by the model parsing module, the scheme planning module, and the model training module constitute a system for implementing the heterogeneous hybrid parallel training method combined with a freezing mechanism. For the sake of illustration, the following description is in modular form.

[0031] The framework first calls the model parsing module to obtain information about the model to be trained and the heterogeneous devices, and measures the training performance metrics. These performance metrics are then passed as known information to the solution planning module. The model parsing module is only called once.

[0032] The scheme planning module is responsible for formulating and dynamically adjusting the hybrid parallel strategy; the model training module executes the training process based on the configuration output by the scheme planning module. To accommodate the freeze mechanism during training, the system allows the scheme planning module and the model training module to call each other multiple times. When a freeze occurs during model training, the system pauses the training process, updates the model layer freeze information, and calls the scheme planning module again to construct the latest heterogeneous hybrid parallel training scheme. Then, the model training module is restarted for training.

[0033] The functions of each module (step) are described in detail below.

[0034] Model parsing module:

[0035] The goal of this module is to acquire model information, model layer freeze status, and heterogeneous hardware information, and based on this, predict performance metrics. When predicting layer computation time, if the layer... If the state is frozen, forward propagation still needs to be executed but it does not participate in backward propagation. Backpropagation time is 0, and the memory requirements of gradient and optimizer state are reduced to zero, thereby releasing the corresponding GPU memory and computing resources. These resources can increase the data parallelism of certain stages or adjust the stage boundaries to better balance the load on heterogeneous devices.

[0036] Specifically, the first step is to parse the model structure. This involves receiving the deep learning model file (e.g., PyTorch, TensorFlow, or ONNX format) provided by the user. Using libraries like ONNX, the model is parsed, traversing the computation graph to obtain detailed information for each layer, including layer type (e.g., convolutional, fully connected, Transformer layers), the dimensions of the input and output tensors, the number of parameters, and the connection relationships between layers.

[0037] Secondly, a computational resource and hardware condition analysis is required. The system needs to be deployed on a heterogeneous GPU cluster, which may contain different models (at least two models) of GPUs (such as NVIDIA A100, V100, etc., with different computing capabilities and memory sizes) and different interconnection methods (such as high-speed NVLink, PCIe bus, Ethernet; the heterogeneous cluster should contain at least two different GPU interconnection methods with different transmission rates). Libraries such as nvidia-smi or lower-level libraries are used to query the specific model, computing power (FLOPS), and available memory of each GPU in the cluster, as well as to probe the interconnection topology and point-to-point bandwidth between GPUs. This heterogeneous information forms the basis for subsequent planning.

[0038] Finally, performance metrics are predicted. Based on the model layer freeze status, the parsed model structure, and the collected heterogeneous hardware information, the execution performance of each part of the model on specific hardware is predicted.

[0039] 1. Computation Time Prediction: For each layer, based on its operation type, input / output dimensions, number of parameters, and the computing power of the GPU model to be run on it. Predict computation time using performance models or empirical formulas: .

[0040] 2. Space Cost Prediction: Calculate the memory required for each layer, including the memory occupied by parameters. Memory occupied by optimizer state Activation of occupied memory and the memory occupied by gradients This ultimately yields the actual memory occupied by one layer: , , , , Combining the above formulas, we can obtain the total memory usage of one unfrozen layer. for: , Crucially, this module considers the impact of a frozen state on resource requirements. This can be further described as: , in, For layer The frozen state This indicates that the layer is frozen and does not consume memory related to computation and gradient / optimizer. This indicates that the layer is not frozen and is using memory normally. Number of input / output channels The kernel size is the convolution kernel size. To output the height and width of the feature map, and The height and width of the input feature map.

[0041] These results (time and memory usage of each layer on different heterogeneous GPUs) will be passed as input to the scheme planning module.

[0042] Solution planning module: This module is the core of the invention, responsible for formulating and dynamically adjusting the hybrid parallel strategy for the current model layer freezing situation.

[0043] In terms of computation time, since the frozen layer no longer undergoes backpropagation, therefore, when the layer... When frozen ( =0), Therefore, the total computation time for the stage containing the frozen layer decreases; however, the forward propagation of the frozen layer still needs to be performed.

[0044] In terms of memory usage, the gradient and optimizer state memory requirements of the frozen layer are reduced to zero, thereby releasing the corresponding GPU memory and computing resources. This can be used to increase the data parallelism of certain stages or adjust the stage boundaries to better balance the load on heterogeneous devices.

[0045] The core objective of the scheme planning is to find a way to integrate the model Layers (including frozen layers) are divided into A collection of heterogeneous GPUs ( The hybrid parallel scheme (including pipeline stages and data parallel configurations within each stage) on the pipeline reduces the training time of a single batch. Minimum, training time Startup time Stable time and end time The composition and specific calculation process are as follows: , in, This is the bottleneck stage, which is the longest-running stage in hybrid parallelism. For the stage Forward propagation time, This is the bottleneck stage. The sum of all previous forward propagation times.

[0046] , in, Represents the number of micro-batches. , Each represents a stage The forward and backward propagation times, i.e., the stages The sum of the time for all forward and backward propagations except for the first forward propagation and the last backward propagation.

[0047] , in, Representative stage The All-Reduce overhead (measured by the program) has the following parameter set: GPU devices are ; Representative stage The backpropagation time; That is, the stage The sum of the backpropagation time of all subsequent stages and all All-Reduce overhead.

[0048] , , Where T is the computation time for one batch. This represents the maximum video memory capacity of the device. Considering the frozen result obtained for model analysis Layer memory requirements, This refers to the total number of layers in the model, which is the sum of all convolutional layers, residual blocks, and fully connected layers in the network. For the model number Layers. During training, the hybrid parallel scheme is dynamically adjusted based on changes in the frozen state of the model layers: when the frozen state changes, the system dynamically reassesses memory usage and computational load. By utilizing the memory and computing resources released by the frozen layer to adjust the hybrid parallel configuration, load balancing and training efficiency can be improved while maintaining memory constraints.

[0049] When calculating these times, the effects of heterogeneity are considered as follows: 1. and (stage Forward / backward propagation time depends on the stage The freezing status and the specific model and quantity of GPUs allocated to that stage are measured by the program to obtain the actual values. If the stage... If data parallelism is involved, then it is the time required to perform one data parallel operation (computation + communication) on this group of heterogeneous GPUs. (Stage) The frozen layer in the process does not require backpropagation.

[0050] 2. (Parameter set) The overhead of All-Reduce depends on the number of parameters and the subset of GPUs participating in data parallelism. Internal interconnect bandwidth. The bandwidth between different GPU pairs may vary, affecting the efficiency of All-Reduce.

[0051] 3. The data transfer time across stages also depends on the connection bandwidth between the sending and receiving GPUs.

[0052] 4. The size of the model that a GPU can accommodate is limited by its video memory capacity.

[0053] The specific implementation process of the plan is as follows: (I) Bottleneck Identification Stage

[0054] Due to the stable phase Typically dominates the total time , find The largest stage Crucially, as shown in Algorithm 1, this algorithm dynamically determines the bottleneck stage Z of the pipeline parallelism by scanning the time cost of each training stage in reverse, in order to minimize the training time of a batch. Its core idea is: first, initialize the bottleneck stage Z and the scan pointer s; then, backtrack from the end of the pipeline, comparing the fixed parallel cost of the current stage stage by stage. ) and the cumulative cost of candidate bottleneck stage Z ( If the current stage has a higher cost, it is updated to a new bottleneck stage, and finally the stage with the highest cumulative time cost is selected as the bottleneck optimization target. Through this reverse greedy strategy, the algorithm can... Adaptive adjustment of the parallel scheme within time complexity, considering That is, the number of stages, usually has Therefore, this algorithm can calculate the bottleneck stage in a short time. This effectively balances the computational load across different stages of the pipeline. Specific algorithm 1 is as follows: Figure 4 As shown.

[0055] (ii) Determine the hybrid parallel scheme under the current freezing condition and solve for the optimal configuration. Determine the bottleneck stage using Algorithm 1. Next, a hybrid parallel scheme needs to be determined, with the goal of minimizing the overall pipeline latency by recursively dividing the pipeline stages and enumerating device allocation strategies. This problem can be formalized as follows: , , ,

[0056] in The total number of layers in the model, i.e., the sum of all convolutional layers, residual blocks, and fully connected layers in the network, is the frozen state list. This will affect the computation and communication time at each stage. This represents the total number of GPUs. It is a heterogeneous GPU collection. As the dividing point of the layer, The number of GPUs allocated, For a specific subset of GPUs, the optimization process needs to consider The heterogeneity of GPUs and the interconnect bandwidth between them, Indicates from all GPU collections Selected All possible combinations of GPUs.

[0057] This problem can be solved using the concept of dynamic programming. (Frozen state list) Define the state based on the known information. Before processing Layers that utilize heterogeneous GPU subsets (Include Minimum pipeline time for (multiple GPUs). By iteratively expanding the state, consider the next layer (multiple GPUs). Assign a new subset of GPUs (Given its heterogeneous characteristics), calculate the resulting new stage time and total time, and update the DP table. Finally... That is, the optimal solution, such as Figure 5 As shown in Algorithm 2.

[0058] Model training module:

[0059] This module executes the training process based on the configuration output by the scheme planning module. It uses a computation graph to represent and schedule computations. For each pipeline stage, it constructs a subgraph for forward and backward propagation. The core scheduling strategy is 1F1B, meaning one forward propagation is followed by one backward propagation, such as... Figure 2 As shown. The scheduler ensures micro-batch processing. Backpropagation It must propagate forward. It can only begin after completion, and it alternates with the forward propagation of subsequent micro-batches, which effectively reduces peak memory usage.

[0060] To handle the complex communication requirements arising from heterogeneous environments and hybrid parallelism (e.g., one stage using one GPU, the next stage using two different GPUs via a slower PCIe connection), this module includes specialized communication layers: a separation layer for distributing data from a small number of devices to multiple devices during forward propagation, and an aggregation layer for pooling results from multiple devices to a small number of devices; similar situations are handled in backpropagation using gradient copying and gradient aggregation (such as All-Reduce). The design of these communication operations must consider the differences in capabilities between the source and target devices and the connection bandwidth between them.

[0061] The key to this module lies in handling cross-stage communication, especially in heterogeneous environments and dynamically changing parallelism, referencing Figure 3 :

[0062] 1. Forward propagation: If the stage The output needs to be passed to the downstream stage. , and stage Data parallelism is greater than stage (Or, different GPU types / bandwidths require specific processing), then in the stage A split layer is inserted at the end. This layer is responsible for copying or splitting the activation tensors on a single (or a few) GPUs and sending them to the stage. On all relevant GPUs, the transfer path takes into account the actual bandwidth between them. Conversely, if parallelism decreases or convergence to a specific type / bandwidth of GPU is required, then at the stage... An aggregation layer is inserted at the beginning, which collects data from multiple source GPUs, integrates it, and passes it to subsequent computations.

[0063] 2. Backpropagation: Gradients need to be backpropagated. If the stage... The gradient needs to be propagated back to the downstream stage (the preceding stage on the backpropagation path). , and stage Data parallelism is greater than stage (Or different GPU types / bandwidths), then in the stage Towards the stage After passing the gradient, the stage Internally, gradient replication is required to ensure that all parallel copies of the data receive the gradient. Conversely, if the stage... The parallelism is less than that of the stage Or, when it is necessary to aggregate gradients to a specific type / bandwidth of GPU, then in the stage After calculating the gradient, a gradient aggregation operation (usually All-Reduce) needs to be performed to sum the gradients of all replicas and then pass them to the stage. GPU. All-Reduce performance ( Heavily dependent on GPU participation The interconnection bandwidth between them has been considered in the solution planning section.

[0064] Furthermore, this module features a dynamic adjustment function, which is crucial for handling the freezing mechanism. During training, when the system detects a change in the frozen state of certain layers, the computational cost of these layers ( (This could change to 0 or less) and memory requirements would also change. This could disrupt the existing load balance, leading to bottlenecks. Possible transfer. At this point, reassess the computational load of each pipeline stage s ( Due to factors such as memory usage, the solution planning module needs to be invoked again. It uses the updated layer state (the new...). The system uses the performance predictions provided by the model parsing module (value) and the model analysis module to rerun the optimization process (step (i) identifying the bottleneck stage and step (ii) determining the hybrid parallel scheme under the current freeze condition and solving for the optimal configuration), finds the optimal hybrid parallel scheme under the current state, and applies the new scheme (stage partitioning, GPU allocation, etc.) to the model training module. The memory and computing resources released by the freeze can be reallocated, for example, to increase the data parallelism of certain stages or adjust the stage boundaries to better balance the load on heterogeneous devices.

[0065] Through the collaborative work of these three modules, this invention realizes an efficient hybrid parallel training method on heterogeneous GPU clusters that can dynamically adapt to changes in model freezing. It not only solves the capacity and efficiency problems of training large models, but also fully utilizes the characteristics of heterogeneous hardware and can flexibly respond to dynamic changes during the training process.

[0066] Based on the above-described method, this invention also provides a heterogeneous hybrid parallel training system that incorporates a freezing mechanism, implemented through computer software, including a model parsing module, a scheme planning module, and a model training module. Their specific functions have been described in detail in the preceding paragraphs.

[0067] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.

Claims

1. A heterogeneous cluster hybrid parallel training method incorporating a freezing mechanism, characterized in that, Includes the following steps: Step 1, Model Analysis: Receive the deep learning model, analyze the model's network structure, and combine the characteristics of the target heterogeneous hardware cluster to predict the computation time and storage requirements of each layer in the model on different types of hardware, and consider the impact of the frozen state on storage requirements and computation time. Step 2, Scheme Planning: Based on the model layer freezing status, model parsing results, and heterogeneous hardware information, dynamic programming is used to solve for the optimal pipeline partitioning and data parallel configuration to obtain a hybrid parallel training scheme. During the dynamic programming process, it is necessary to simultaneously satisfy the constraint that the memory usage of all model layers allocated on a specific GPU device does not exceed the device's upper limit. Step 3, Model Training: The training process of the hybrid parallel training scheme is executed. During the training process, when the frozen state changes, the system reallocates the released video memory and computing resources, dynamically adjusts the parallel scheme, and achieves load balancing and improved training efficiency.

2. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 1, characterized in that, Step 1 specifically includes: (1) Use model parsing tools to obtain the type, dimensions and connection methods of each layer of the model; (2) Use hardware detection tools to obtain information on the GPU model, computing power, memory size, and interconnect bandwidth between GPUs in the heterogeneous cluster; (3) Computation time prediction: For each layer, the prediction is based on its operation type, input / output dimensions, number of parameters, and the computing power of the GPU model on which it will be run. Predict computation time using performance models or empirical formulas. : , Storage requirement forecasting: Calculate the memory required for each layer, including the memory occupied by parameters. Memory occupied by optimizer state Activation of occupied memory and the memory occupied by gradients : , , , , The final result shows the actual memory occupied by one layer: , in, Number of input / output channels The kernel size is the convolution kernel size. To output the height and width of the feature map, and The height and width of the input feature map.

3. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 1, characterized in that, Step 2 specifically includes the following process: The goal of the optimization scheme is to minimize the expected total time for a single training batch. ,in From the start-up phase time Stable phase time and end phase time constitute: , , , , Conditional constraints are imposed on the actual amount of GPU memory used by the model during training: , in, This is the bottleneck stage. Stages Forward and backward propagation times on the allocated heterogeneous GPUs, , Each represents a stage Forward and backward propagation times, Representative stage The backpropagation time, This refers to the quantity in a micro-batch. For the stage The set of parameters involved In the equipment group All-Reduce overhead, This refers to the maximum video memory capacity of the GPU device. This refers to the total number of layers in the model, which is the sum of all convolutional layers, residual blocks, and fully connected layers in the network. For the model number layer, For layer The frozen state The first step in the model analysis process, after considering the frozen state, is the calculation of the... Layer memory requirements, It is a heterogeneous GPU collection; First, identify the bottleneck stage of the production line. Then, dynamic programming is used to solve for the optimal hybrid parallel configuration; when the frozen state of the model changes, replanning will be performed in conjunction with the resources released from the frozen layer; the objective function of the problem is defined as follows. : , , , in, The total number of layers in the model, i.e., the sum of all convolutional layers, residual blocks, and fully connected layers in the network, is the frozen state list. This will affect the computation and communication time at each stage. This represents the total number of GPUs. It is a heterogeneous GPU collection. As the dividing point of the layer, The number of GPUs allocated, For a specific subset of GPUs, the optimization process needs to consider The heterogeneity of GPUs and the interconnect bandwidth between them, Indicates from all GPU collections Selected All combinations of GPUs; When predicting the computation time of a hierarchical level, the impact of the frozen state on the computation should be considered: if the level If the process is frozen, its forward propagation still needs to be executed but it does not participate in the backward propagation. The backward propagation time is... =0; In terms of computing resources, the memory requirements of the gradients and optimizer states in the frozen layer are reduced to zero, and the memory and computing resources released by the freeze are reallocated. The optimal solution can be obtained by using the concept of dynamic programming. During training, monitor changes in the frozen state of model layers; when a layer is frozen or thawed, reassess the computational load and memory usage of each pipeline stage; based on the updated load and available resources, recalculate and apply the optimal hybrid parallel scheme.

4. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 3, characterized in that, The process of solving problems using the concept of dynamic programming includes: Frozen status list Define the state based on the known information. Before processing Layers that utilize heterogeneous GPU subsets Minimum pipeline time for heterogeneous GPU subset Include One GPU; by iteratively expanding the state, consider moving to the next layer ( Assign a new subset of GPUs Calculate the resulting new stage time and total time, update the DP table, and obtain the final result. This is the optimal solution.

5. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 1, characterized in that, Step 3 uses a computation graph to represent and schedule operations, constructing forward and backward propagation subgraphs for each pipeline stage. The core scheduling strategy is 1F1B.

6. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 5, characterized in that, The 1F1B strategy refers to scheduling the backpropagation of a microbatch as early as possible after the forward propagation of a microbatch is completed within a pipeline stage, alternating with the forward propagation of subsequent microbatches.

7. The heterogeneous cluster hybrid parallel training method combining a freezing mechanism according to claim 1, characterized in that, Step 3 uses the following communication mechanism for data transmission when handling cross-stage communication: During forward propagation, when the data parallelism of downstream stages increases or the receiving GPU type / bandwidth differs, a separation layer is used to copy or distribute data; when the data parallelism of downstream stages decreases or needs to be converged to a specific type / bandwidth GPU, an aggregation layer is used to collect and integrate data. During backpropagation, gradient copying is used when the data parallelism of the downstream stage increases or the GPU type / bandwidth differs; gradient aggregation is used when the data parallelism of the downstream stage decreases or gradients need to be aggregated to a specific type / bandwidth of GPU, where the overhead of All-Reduce is affected by the bandwidth between the heterogeneous GPUs involved.

8. A heterogeneous cluster hybrid parallel training system incorporating a freezing mechanism, characterized in that, It includes a model parsing module, a scheme planning module, and a model training module, used to implement the heterogeneous cluster hybrid parallel training method with freezing mechanism as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Large electric power model assembly line freezing training optimization method based on reinforcement learning

    CN118674003A

  • Single GPU model training method and device, electronic equipment and storage medium

    CN118798282A

  • Heterogeneous cluster-oriented resource allocation method and device and storage medium

    CN120723469A

  • Deep neural network model freezing training optimization method based on tensor similarity

    CN120851103A

Cited By

  • Automatic optimization method, system and device for model distribution parallelization training process

    CN121390212A