A large model adaptive parallel training method applicable to heterogeneous clusters

By measuring hardware performance in real time in heterogeneous clusters and dynamically adjusting model layer and hardware allocation, and optimizing matching between devices, the uneven resource utilization and video memory limitation of large model training in heterogeneous clusters is solved, the flexibility and efficiency of training are improved, and the stability and adaptability of the training process are ensured.

CN119938327BActive Publication Date: 2025-07-22WEIYE ZHISUAN (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510021206.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-07-22
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

In heterogeneous clusters, large model training has problems such as uneven resource utilization, limited memory space and low training configuration efficiency. Especially in heterogeneous clusters, high-performance devices may be idle due to waiting for poor performance devices, making resource matching and optimization difficult, and it is difficult to quickly adjust during training.

Method used

By measuring the hardware performance of heterogeneous clusters in real time, dynamically compute data parallel values and tensor parallel configurations, automatically adjust the model layer and hardware allocation, optimize performance matching between devices, and adjust the calculation delay through pipeline parallel technology, monitor and adjust the data transmission speed and calculation task allocation in real time, ensure that the training configuration adapts to the current cluster status.

Benefits of technology

It improves resource utilization, eliminates performance bottlenecks, optimizes the utilization of video memory space, enhances the flexibility and efficiency of training, reduces the dependence on expert manpower, and ensures the stability and adaptability of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938327B_ABST
    Figure CN119938327B_ABST
Patent Text Reader

Abstract

The present invention discloses a large model adaptive parallel training method applicable to heterogeneous clusters, which relates to the field of parallel computing technology. This technical solution solves the problems of uneven resource utilization and low efficiency; by dynamically calculating the data parallel value and tensor parallel configuration, it automatically adjusts the allocation of model layers and hardware, optimizes the performance matching and collaborative work between devices, improves resource utilization rate and eliminates performance bottlenecks; at the same time, this method optimizes the use of the video memory and computing power of each device by real-time monitoring and dynamically adjusting the storage demand matching index Cpp and processing duration Csc, and solves the problem of video memory space limitation; in addition, this solution also includes continuously monitoring real-time performance data and dynamically adjusting parameters in pipeline parallelism, such as data transmission speed and computing task reassignment, improving the flexibility and efficiency of training, and ensuring that the parallel training configuration always adapts to the actual operating conditions of the current cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of parallel computing, and particularly to a large model adaptive parallel training method applicable to heterogeneous clusters. Background Art

[0002] Large models generally refer to models with a parameter count of over one million. This type of model has high generality and strong generalization performance and is currently widely used, becoming the mainstream technology in the field of artificial intelligence. However, the main technical defects of large model training in heterogeneous computing clusters can be summarized into the following three points:

[0003] 1. Resource matching and optimization issues: In heterogeneous clusters, different computing devices have different computing capabilities and video memory capacities. If the large model to be trained is evenly sliced and assigned to various computing devices without differential processing, it will lead to uneven resource utilization and low efficiency. High-performance devices may be idle waiting for low-performance devices to complete calculations, resulting in the so-called "barrel short board" phenomenon, that is, the overall performance is limited by the performance of the weakest device;

[0004] 2. Video memory space limitations: Large model training usually requires about 16 times the GPU video memory or the memory of computing chips of the model's parameter count. In heterogeneous clusters, especially in environments with limited resources or uneven performance, a single type of high-performance device is not sufficient to support all the requirements of large models;

[0005] 3. Low training configuration and optimization efficiency: Currently, performing large model training tasks in heterogeneous clusters often requires a large amount of expert manpower for pre-training model layer slicing and configuration. This method not only consumes a large amount of manpower and computing resources, but also, due to being preset before training, it is difficult to quickly and accurately adjust to changes during training (such as model adjustments, data changes, device performance changes, etc.). Each change may require re-designing and validating the training plan, reducing the flexibility and efficiency of training. Summary of the Invention

[0006] In view of the deficiencies of the prior art, the present invention provides a large model adaptive parallel training method applicable to heterogeneous clusters, which solves the problems mentioned in the background art.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A large model adaptive parallel training method applicable to heterogeneous clusters, including the following steps:

[0008] Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices within the heterogeneous cluster, for real-time measurement of the computing capabilities and network bandwidths of different hardware, and update the data-driven training configuration in real time based on the measurement results;

[0009] Step 2: Based on the device video memory and computing power measured in Step 1, dynamically calculate the data parallel value Dps and the tensor parallel configuration, and adjust the device performance matching by automatically adjusting the allocation of model layers and hardware. The specific calculation formula for the data parallel value Dps is as follows:

[0010]

[0011] where N is the total number of devices, GPUm i is the video memory capacity of the i-th device, α is the conversion coefficient of the required video memory for training accuracy, S is the total number of model layers, and Layer k is the number of parameters of the k-th layer;

[0012] Step 3: According to the adjusted device performance matching, preset the parallel configuration scheme for heterogeneous model training, and adjust the calculation delay and the overall training duration by adopting the pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rate and response time of each device;

[0013] Step 4: Before the training starts, estimate the video memory and time cost of the preset parallel configuration scheme, calculate the expected storage demand matching index Cpp and the expected processing duration Csc for each parallel configuration scheme, then fit the expected storage demand matching index Cpp and the expected processing duration Csc to obtain the scheme selection coefficient Fkxs for each parallel configuration scheme and evaluate it. Finally, make the selected parallel configuration scheme run under the current hardware configuration and also be the optimal choice in terms of time efficiency;

[0014] Step 5: Verify the parallel configuration scheme in Step 4 during actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance problems during the training process, and at the same time verifying the accuracy of the time efficiency estimation by calculating and evaluating the actual calculation duration T of the batch data.

[0015] Preferably, Step 1 specifically includes:

[0016] Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing power and inference computing power of each device in real time. At the same time, adopt open-source base models with different architectures to measure the performance under different data input and output lengths; then use network test tools suitable for different hardware platforms to evaluate the communication bandwidth of each device; the measurement results include the floating-point operations per second of the computing device and the megabits per second of the communication performance, which will be recorded in real time and used to update the data-driven training configuration.

[0017] Preferably, Step 2 specifically includes:

[0018] Based on Step 1, by analyzing the device video memory and computing power, and determining the data load and computing tasks borne by each device during training according to the measurement results of the computing power and network bandwidth of different hardware; through an automated algorithm, adjusting the mapping relationship between the layers of the model and the specific hardware, optimizing the performance matching and collaborative work between devices; setting the video memory capacity GPUm and the number of parameters Layer based on the ratio of the total number of model parameters to the total video memory capacity of all participating devices, and finally obtaining the data parallel value Dps; by continuously monitoring real-time performance data, automatically adjusting the parallel configuration in response to the device.

[0019] Preferably, Step 3 specifically includes:

[0020] According to the data parallel value Dps and the tensor parallel configuration determined in Step 2, preset multiple parallel configuration schemes for heterogeneous model training, including the initial configuration of pipelining parallelism for each hardware device, which includes allocating each model layer to different computing nodes, and setting the order and rate of data processing for each node, so that each computing node receives and processes data according to the real-time computing power and network conditions.

[0021] Preferably, Step 3 specifically further includes:

[0022] During the model training process, continuously monitor the real-time computing performance and response time of each hardware device, and dynamically adjust the parameters in the pipelining parallelism according to the actual operating conditions, which includes adjusting the data transmission speed and reallocating the computing tasks, specifically:

[0023] Adjusting the data transmission speed: According to the network communication latency and the data processing capabilities of each device, dynamically adjust the data transmission speed between devices according to the receiving and processing capabilities of each device;

[0024] Reallocating the computing tasks: Based on the performance of the devices, reallocate the computing tasks in the model to each computing node.

[0025] Preferably, Step 4 specifically includes:

[0026] For each preselected parallel configuration scheme, based on the model parameters, the expected computing load, and the hardware performance parameters, including the total number of model parameters Z Layer 、the average video memory capacity of the device Z GPUm 、the expected occupied video memory Zzy, the device computing speed Sjs, the network latency duration Syc, and the model computing complexity value Sfz. Then, after dimensionless processing of the model parameters, the expected computing load, and the hardware performance parameters, calculate the expected storage requirement matching index Cpp and the expected processing duration Csc for each parallel configuration scheme. The specific calculation formulas are as follows:

[0027]

[0028] Csc = Sjs × Syc + Sfz.

[0029] Preferably, step four specifically further includes:

[0030] Fit the expected storage requirement matching index Cpp and the expected processing duration Csc, and obtain the scheme selection coefficient Fkxs of each parallel configuration scheme. The specific calculation formula is as follows:

[0031]

[0032] Preferably, step four specifically further includes:

[0033] Rank and compare the scheme selection coefficients Fkxs of each parallel configuration scheme, and preferably select the parallel configuration scheme with the highest value of the scheme selection coefficient Fkxs. At the same time, monitor the performance during actual operation in real time, including the actual storage usage and processing time; update the hardware performance data and network status information in real time, and preset the deviation threshold between the actual situation and the estimated value. When the deviation between the actual situation and the estimated value exceeds the deviation threshold, readjust the parallel configuration scheme at this time.

[0034] Preferably, step five specifically includes:

[0035] Start the selected parallel configuration scheme and perform actual model training operation. During this period, the system automatically monitors and records the operation data of the model, including the video memory usage and computing efficiency; at the same time, monitor potential performance bottleneck problems in real time during the process, including video memory overflow. When encountering performance bottleneck problems, automatically adjust the operation parameters or pause the training and perform configuration adjustment.

[0036] Preferably, step five specifically further includes:

[0037] During the model training process, calculate and record the actual computing duration T of each batch of data in real time. Among them, the specific formula for the actual computing duration T of each batch of data is as follows:

[0038]

[0039] In the formula, S is the total number of layers of the model, t p,j is the computing time of the device responsible for calculating the jth slice of the pth layer of the model, PP_time p represents the delay in transferring the calculation result to the next layer, that is, the pth layer, after the kth layer of the model is calculated in pipeline parallelism; and DP_time p is the time overhead of tensor parallelism of the pth layer of the model;

[0040] Next, preset the standard processing duration threshold Z, and compare and evaluate the actual calculation duration T, the expected processing duration Csc, and the standard processing duration threshold Z of each batch of data. Since the standard processing duration threshold Z > the expected processing duration Csc, the accuracy of the time efficiency prediction is verified. The specific evaluation content is as follows:

[0041] If the standard processing duration threshold Z > the expected processing duration Csc ≥ the actual calculation duration T, in this case, it indicates that the actual performance of the parallel configuration is normal, and the predetermined efficiency target has been achieved or exceeded, and the expected processing duration Csc is reasonably set;

[0042] If the standard processing duration threshold Z ≥ the actual calculation duration T > the expected processing duration Csc, in this case, it indicates that the actual performance of the parallel configuration is normal. Although the actual calculation duration T reaches and is better than the standard processing duration threshold Z, it fails to reach the expected processing duration Csc; this indicates that the expected processing duration does not fully consider some obstacles in actual operation. At this time, it is necessary to adjust the expected model or consider adjusting the current configuration;

[0043] If the actual calculation duration T > the standard processing duration threshold Z > the expected processing duration Csc, in this case, it indicates that the actual performance of the parallel configuration is abnormal. The actual calculation duration T not only fails to reach the expected processing duration Csc but also exceeds the standard processing duration threshold Z, which indicates that the current parallel configuration scheme has abnormal efficiency and there are performance bottlenecks or configuration problems; at this time, focus on checking and solving the factors affecting performance, including insufficient hardware performance, improper software optimization, low data management efficiency, and network latency.

[0044] The present invention provides a large model adaptive parallel training method applicable to heterogeneous clusters. It has the following beneficial effects:

[0045] (1) For the large model adaptive parallel training method applicable to heterogeneous clusters, in a heterogeneous cluster, each computing device has different computing capabilities and video memory capacities; if the large model is evenly sliced and allocated to various computing devices without differential processing, it will lead to uneven resource utilization and low efficiency; in addition, high-performance devices may be idle waiting for performance-poorer devices to complete calculations, resulting in the "barrel short board" phenomenon, that is, the overall performance is limited by the performance of the weakest device; this solution automatically adjusts the allocation of model layers and hardware by dynamically calculating the data parallel value and tensor parallel configuration, effectively optimizing the performance matching and collaborative work between devices, improving resource utilization and eliminating performance bottlenecks;

[0046] (2) The large model adaptive parallel training method for heterogeneous clusters. Generally, the number of model parameters required for large model training is about 16 times that of the GPU video memory or the memory of the computing chip. In a heterogeneous cluster environment with limited resources or uneven performance, a single type of high-performance device is often insufficient to support all the requirements of the large model. This solution solves the problem of video memory space limitation by real-time monitoring and dynamically adjusting the storage requirement matching index Cpp and the processing duration Csc to ensure that the video memory and computing power of each device are optimally utilized.

[0047] (3) The large model adaptive parallel training method for heterogeneous clusters. Currently, when heterogeneous clusters execute large model training tasks, a large amount of expert manpower is often required for pre-model layer slicing and configuration. This method not only consumes a large amount of manpower and computing resources, but also, due to being preset before training, it is difficult to quickly and accurately adjust to changes during the training process, including model adjustments, data changes, and device performance changes. Each change may require re-designing and validating the training plan, reducing the flexibility and efficiency of training. This solution continuously monitors real-time performance data and dynamically adjusts various parameters in pipeline parallelism, including data transfer speed and computing task reallocation, to ensure that the parallel training configuration always adapts to the actual operating conditions of the current cluster, improving the flexibility and efficiency of training. Brief Description of the Drawings

[0048] Figure 1 It is a schematic diagram of the step flow of a large model adaptive parallel training method for heterogeneous clusters according to the present invention. Detailed Embodiments

[0049] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] Embodiment 1

[0051] Please refer to Figure 1 , a large model adaptive parallel training method for heterogeneous clusters, including the following steps:

[0052] Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices in the heterogeneous cluster for real-time measurement of the computing power and network bandwidth of different hardware, and update the data-driven training configuration in real time according to the measurement results.

[0053] Step 2: Based on the device video memory and computing power measured in Step 1, dynamically calculate the data parallelism value Dps and tensor parallel configuration, and adjust the device performance matching by automatically adjusting the allocation of model layers and hardware. The specific calculation formula for the data parallelism value Dps is as follows:

[0054]

[0055] where N is the total number of devices, GPUm i is the video memory capacity of the i-th device, α is the conversion coefficient of the required video memory for training accuracy, S is the total number of model layers, and Layer k is the number of parameters of the k-th layer;

[0056] Step 3: According to the adjusted device performance matching, preset the parallel configuration scheme for heterogeneous model training, and adjust the calculation delay and overall training duration by adopting the pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rate and response time of each device;

[0057] Step 4: Before the training starts, estimate the video memory and time cost of the preset parallel configuration scheme, calculate the expected storage demand matching index Cpp and the expected processing duration Csc for each parallel configuration scheme, then fit the expected storage demand matching index Cpp and the expected processing duration Csc, obtain the scheme selection coefficient Fkxs for each parallel configuration scheme and evaluate it. Finally, make the selected parallel configuration scheme operate under the current hardware configuration and also be the optimal choice in terms of time efficiency;

[0058] Step 5: Verify the parallel configuration scheme in Step 4 during actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance problems during the training process, and verifying the accuracy of the time efficiency estimation by calculating and evaluating the actual calculation duration T of the batch data.

[0059] In this embodiment, in Step 1, by using a general computing framework to automatically measure the computing performance and network performance of various hardware devices in the heterogeneous cluster, the computing capabilities and network bandwidths of different hardware are measured in real time, and the data-driven training configuration is updated in real time according to the measurement results, ensuring that the training configuration is always based on the latest performance data, improving the resource utilization efficiency and the accuracy of the configuration;

[0060] Based on the device video memory and computing power measured in Step 1, Step 2 dynamically calculates the data parallelism value Dps and tensor parallel configuration, and adjusts the device performance matching by automatically adjusting the allocation of model layers and hardware, optimizing the allocation and utilization of hardware resources and reducing the performance loss caused by resource mismatch;

[0061] Step 3: According to the adjusted device performance matching situation, execute the heterogeneous parallel strategy for model training, and use pipelining parallel technology to adjust the calculation delay and the overall training duration. This strategy minimizes the delay caused by different device performance differences by adapting to the computing rates and response times of each device, improving the overall speed and efficiency of training.

[0062] Step 4: Before training starts, estimate the video memory and time costs of the preselected parallel configuration schemes, calculate the expected storage demand matching index Cpp and the expected processing duration Csc for each parallel configuration scheme, then fit the expected storage demand matching index Cpp and the expected processing duration Csc, obtain the scheme selection coefficient Fkxs for each parallel configuration scheme and evaluate it, so that the selected parallel configuration scheme runs under the current hardware configuration and is also the optimal choice in terms of time efficiency, ensuring the optimization of the training scheme and improving the efficiency and scientificity of decision-making.

[0063] Step 5: Verify the parallel configuration scheme in Step 4 during actual operation, run the training process and calibrate the estimated model parameters, monitor and prevent video memory overflow and other performance problems during the training process, and at the same time verify the accuracy of the time efficiency estimation by calculating and evaluating the actual calculation duration T of the batch data, ensuring the stability of the training process and the accuracy of prediction, and optimizing and adjusting the parallel configuration through actual operation feedback to achieve the best operating state.

[0064] Example 2

[0065] Step 1 specifically includes:

[0066] Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing ability and inference computing ability of each device in real time, and at the same time use open-source base models with different architectures to measure the performance under different data input and output lengths; then use network test tools suitable for different hardware platforms to evaluate the communication bandwidth of each device; the measurement results include the floating-point operations per second of the computing device and the megabits per second of the communication performance, which will be recorded in real time and used to update the data-driven training configuration.

[0067] Step 2 specifically includes:

[0068] Based on Step 1, by analyzing the device video memory and computing power, and determining the data load and computing tasks borne by each device during training according to the measurement results of the computing power and network bandwidth of different hardware; through an automated algorithm, adjusting the mapping relationship between the layers of the model and the specific hardware, optimizing the performance matching and collaborative work between devices; setting the video memory capacity GPUm and the number of parameters Layer based on the ratio of the total number of model parameters to the total video memory capacity of all participating devices, and finally obtaining the data parallel value Dps; by continuously monitoring the real-time performance data, automatically adjusting the parallel configuration to respond to the device.

[0069] In this embodiment, Step 1 measures the fixed matrix computing power and inference computing power of each device in real time by running predefined computing tasks and network bandwidth tests, and measures the performance under different data input and output lengths using open-source base models with different architectures, which enables the system to accurately configure resources according to the real-time computing power and network bandwidth of each device, ensuring that the data-driven training configuration is always updated based on the latest performance data.

[0070] Then, based on these measurement results, Step 2 analyzes the device video memory and computing power, adjusts the mapping relationship between the model layers and the hardware using an automated algorithm, optimizes the performance matching and collaborative work between devices, and accurately calculates the data parallel value Dps by setting the video memory capacity GPUmi and the number of parameters Layeri of the model layer. This not only enhances the flexibility of the system but also improves the training efficiency; this method allows the system to dynamically adjust the parallel configuration to respond to changes in device performance, ensuring that the parallel training can achieve the best performance state under various operating conditions; in summary, this training method effectively improves the utilization rate of resources, reduces the performance loss caused by resource mismatch, and at the same time enhances the stability and predictability of the training process, providing an efficient and reliable solution for the training of large models in a heterogeneous cluster environment.

[0071] Embodiment 3

[0072] Step 3 specifically includes:

[0073] According to the data parallel value Dps and the tensor parallel configuration determined in Step 2, preset multiple parallel configuration schemes for heterogeneous model training, including the initial configuration of pipeline parallelism for each hardware device, which includes allocating each model layer to different computing nodes, and setting the order and rate of data processing for each node, so that each computing node receives and processes data according to the real-time computing power and network conditions.

[0074] Step 3 specifically further includes:

[0075] During the model training process, continuously monitor the real-time computing performance and response time of each hardware device, and dynamically adjust various parameters in pipeline parallelism according to the actual operating conditions, including adjusting the data transfer speed and reallocating computing tasks. Specifically:

[0076] Adjust the data transfer speed: Dynamically adjust the data transfer speed between devices according to the network communication latency and the data processing capabilities of each device, based on the receiving and processing capabilities of each device.

[0077] Reallocate computing tasks: Reallocate the computing tasks in the model to each computing node according to the performance of the devices.

[0078] Step four specifically includes:

[0079] Step four specifically includes:

[0080] For each preselected parallel configuration scheme, based on the model parameters, expected computing load, and hardware performance parameters, including the total number of model parameters Z Layer 、the average video memory capacity of the device Z GPUm 、the expected occupied video memory Zzy, the device computing speed Sjs, the network latency duration Syc, and the model computing complexity value Sfz. Then, after dimensionless processing of the model parameters, expected computing load, and hardware performance parameters, calculate the expected storage demand matching index Cpp and the expected processing duration Csc for each parallel configuration scheme. The specific calculation formulas are as follows:

[0081]

[0082] Csc = Sjs × Syc + Sfz.

[0083] Step four specifically also includes:

[0084] Fit the expected storage demand matching index Cpp and the expected processing duration Csc to obtain the scheme selection coefficient Fkxs for each parallel configuration scheme. The specific calculation formula is as follows:

[0085]

[0086] Step four specifically also includes:

[0087] Rank and compare the scheme selection coefficients Fkxs of each parallel configuration scheme, and preferentially select the parallel configuration scheme with the highest value of the scheme selection coefficient Fkxs. At the same time, monitor the performance during actual operation in real time, including the actual storage usage and processing time; update the hardware performance data and network status information in real time, and preset the deviation threshold between the actual situation and the predicted value. When the deviation between the actual situation and the predicted value exceeds the deviation threshold, readjust the parallel configuration scheme at this time.

[0088] In this embodiment, step three performs an initial configuration of pipeline parallelism on the hardware device based on the data parallelism value Dps and the tensor parallel configuration, ensuring that each model layer is effectively allocated to a computing node with appropriate computing power and network conditions. This step enables each computing node to receive and process data according to its real-time computing power and network conditions, optimizing the utilization of computing resources and the order of data processing. At the same time, step three also includes continuously monitoring the real-time computing performance and response time of each hardware device, and dynamically adjusting the data transmission speed and reallocating computing tasks according to the actual situation. These adjustments enable the system to flexibly respond to performance changes and maintain the operating efficiency.

[0089] Step four conducts an in-depth performance evaluation for each preselected parallel configuration scheme. Based on parameters such as the total amount of model parameters Zmc, the average video memory capacity of the device Zcl, the expected occupied video memory Zzy, the computing speed of the device Sjs, the network latency duration Syc, and the model computing complexity value Sfz, after these parameters are dimensionless processed, they are used to calculate the storage requirement matching index Cpp and the expected processing duration Csc of each configuration scheme. This not only helps to accurately evaluate the storage and processing capabilities of each scheme, but also enables obtaining the scheme selection coefficient Fkxs by fitting these indexes, which is used to compare and select the parallel configuration scheme with the optimal performance. In addition, step four also includes real-time monitoring of system performance and updating of hardware and network status information to ensure that the selected scheme can continuously perform optimally in the actual environment. Overall, this training method greatly improves the flexibility, efficiency, and reliability of model training by comprehensively utilizing real-time performance monitoring and automatic adjustment strategies, ensuring the optimal training results of large models in heterogeneous cluster environments.

[0090] Example 4

[0091] Step five specifically includes:

[0092] Start the selected parallel configuration scheme and perform actual model training runs. During this period, the system automatically monitors and records the running data of the model, including video memory usage and computing efficiency. At the same time, potential performance bottleneck problems are monitored in real time during the process, including video memory overflow. When encountering performance bottleneck problems, the running parameters are automatically adjusted or the training is paused and configuration adjustments are made.

[0093] Step five specifically also includes:

[0094] During the model training process, the actual computing duration T of each batch of data is calculated in real time and recorded. Among them, the specific formula for the actual computing duration T of each batch of data is as follows:

[0095]

[0096] In the formula, S is the total number of layers of the model, t p,jIt is the computing time of the device responsible for the j-th slice of the p-th layer of the computing model, PP_time p It represents the delay in transmitting the calculation result to the next layer, that is, the p-th layer, after the calculation of the k-th layer of the model is completed in pipeline parallelism; while DP_time p is the time overhead of tensor parallelism for the p-th layer of the model;

[0097] Next, a preset standard processing duration threshold Z is set, and the actual calculation duration T, the expected processing duration Csc, and the standard processing duration threshold Z of each batch of data are compared and evaluated. And the standard processing duration threshold Z > the expected processing duration Csc to verify the accuracy of the estimated time efficiency. The specific evaluation content is as follows:

[0098] If the standard processing duration threshold Z > the expected processing duration Csc ≥ the actual calculation duration T, in this case, it shows that the actual performance of the parallel configuration is normal, and the predetermined efficiency target has been reached or exceeded, and the expected processing duration Csc is reasonably set;

[0099] If the standard processing duration threshold Z ≥ the actual calculation duration T > the expected processing duration Csc, in this case, it shows that the actual performance of the parallel configuration is normal. Although the actual calculation duration T reaches and is better than the standard processing duration threshold Z, it fails to reach the expected processing duration Csc; it indicates that the expected processing duration does not fully consider some obstacles in actual operation. At this time, it is necessary to adjust the expected model or consider adjusting the current configuration;

[0100] If the actual calculation duration T > the standard processing duration threshold Z > the expected processing duration Csc, in this case, it shows that the actual performance of the parallel configuration is abnormal. The actual calculation duration T not only fails to reach the expected processing duration Csc, but also exceeds the standard processing duration threshold Z. This indicates that the current parallel configuration scheme has abnormal efficiency and there are performance bottlenecks or configuration problems; at this time, focus on checking and solving the factors affecting performance, including insufficient hardware performance, improper software optimization, low data management efficiency, and network latency.

[0101] In this embodiment, Step Five conducts crucial actual operation and monitoring on the large model adaptive parallel training method in the heterogeneous cluster, ensuring the efficiency and stability of the training process. By starting the selected parallel configuration scheme and conducting actual model training, the system can monitor and record key operation data in real time, including video memory usage and computing efficiency. This enables timely detection of any potential performance bottleneck issues, such as video memory overflow, and necessary configuration adjustments can be made by automatically adjusting operation parameters or pausing training, thus avoiding possible system crashes or performance degradation. In addition, by calculating the actual computing duration T of each batch of data in real time and comparing it with the preset standard processing duration threshold Z and the expected processing duration Csc, not only the accuracy of the time efficiency estimation is verified, but also a real-time feedback mechanism is provided to evaluate whether the parallel configuration scheme meets the predetermined efficiency target. During this process, the actual computing duration T of each batch of data is calculated based on the total number of layers S of the model, the tensor parallel time overhead of each layer, and the latency in pipeline parallelism. These calculations reflect the actual performance of the current configuration scheme in processing complex data structures, enabling the system to adjust the model training strategy according to the actual situation, optimize the hardware and software configuration, and improve the overall training efficiency. Through this comprehensive evaluation and adjustment, Step Five not only enhances the transparency and controllability of the training process, but also significantly improves the system's adaptability and efficiency in processing complex data, providing an efficient and reliable training solution for large-scale machine learning tasks.

[0102] Among them, t p,j The value depends on the batch data volume, input and output lengths, data precision, and the computing performance of the device (embodiment.

[0103] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An adaptive parallel training method for large models applicable to heterogeneous clusters, characterized in that: It includes the following steps: Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices in a heterogeneous cluster, for real-time measurement of the computing capabilities and network bandwidths of different hardware, and update the data-driven training configuration in real time according to the measurement results; Step 2: Based on the device video memory and computing capabilities measured in Step 1, dynamically calculate the data parallelism value Dps and the tensor parallel configuration, and adjust the device performance matching situation by automatically adjusting the allocation of model layers and hardware. The specific calculation formula for the data parallelism value Dps is as follows: Among them, N is the total number of devices, and GPUm i is the video memory capacity of the i-th device, α is the conversion coefficient of the video memory required for the training accuracy, S is the total number of model layers, and Layer k is the number of parameters of the k-th layer; Step 3: According to the adjusted device performance matching situation, preset the parallel configuration scheme for heterogeneous model training, and adjust the computing latency and the overall training duration by adopting the pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rates and response times of each device; Step 4: Before the training starts, estimate the video memory and time costs of the preset parallel configuration scheme, calculate the expected storage requirement matching index Cpp and the expected processing duration Csc for each parallel configuration scheme, then fit the expected storage requirement matching index Cpp and the expected processing duration Csc, obtain the scheme selection coefficient Fkxs for each parallel configuration scheme and evaluate it. Finally, make the selected parallel configuration scheme run under the current hardware configuration and also be the optimal choice in terms of time efficiency; Step 5: Verify the parallel configuration scheme in Step 4 during actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance problems during the training process, and at the same time verifying the accuracy of the time efficiency estimation by calculating and evaluating the actual computing duration T of the batch data.

2. The large model adaptive parallel training method for heterogeneous clusters according to claim 1, wherein: Specifically, Step 1 includes: Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing capabilities and inference computing capabilities of each device in real time. At the same time, adopt open-source base models with different architectures to measure the performance under different data input and output lengths; then use network test tools suitable for different hardware platforms to evaluate the communication bandwidths of each device; the measurement results include the floating-point operations per second of the computing devices and the megabits per second of the communication performance, which will be recorded in real time and used to update the data-driven training configuration.

3. The large model adaptive parallel training method for heterogeneous clusters according to claim 1, wherein: Specifically, Step 2 includes: Based on Step 1, by analyzing the device video memory and computing capabilities, and determining the data load and computing tasks borne by each device during training according to the measurement results of the computing capabilities and network bandwidths of different hardware; through an automated algorithm, adjust the mapping relationship between the layers of the model and the specific hardware to optimize the performance matching and collaborative work among devices; set the video memory capacity GPUm and the number of parameters Layer based on the ratio of the total number of model parameters to the total video memory capacity of all participating devices, and finally obtain the data parallelism value Dps; continuously monitor the real-time performance data and automatically adjust the parallel configuration in response to the devices.

4. The large model adaptive parallel training method for heterogeneous clusters according to claim 1, wherein: Specifically, Step 3 includes: According to the data parallel value Dps and tensor parallel configuration determined in step two, preset multiple parallel configuration schemes for heterogeneous model training, including the initial configuration of pipeline parallelism for each hardware device, which includes allocating each model layer to different computing nodes, and setting the order and rate of data processing for each node, so that each computing node receives and processes data according to its real-time computing ability and network conditions.

5. The large model adaptive parallel training method for heterogeneous clusters according to claim 4, characterized in that: Step three specifically further includes: During the model training process, continuously monitor the real-time computing performance and response time of each hardware device, and dynamically adjust various parameters in the pipeline parallelism according to the actual running situation, which includes adjusting the data transmission speed and reallocating computing tasks, specifically: Adjust the data transmission speed: According to the network communication delay and the data processing ability of each device, dynamically adjust the data transmission speed between devices according to the receiving and processing abilities of each device; Reallocate computing tasks: According to the performance level of the devices, reallocate the computing tasks in the model to each computing node.

6. The large model adaptive parallel training method for heterogeneous clusters according to claim 1, characterized in that: Step four specifically includes: For each preselected parallel configuration scheme, based on the model parameters, the expected computational load, and the hardware performance parameters, including the total number of model parameters Z Layer 、the average video memory capacity Z of the device GPUm 、the expected video memory occupancy Zzy, the device computing speed Sjs, the network latency duration Syc, and the model computational complexity value Sfz. Then, after dimensionless processing of the model parameters, the expected computational load, and the hardware performance parameters, calculate the expected storage requirement matching index Cpp and the expected processing duration Csc for each parallel configuration scheme. The specific calculation formulas are as follows: Csc = Sjs × Syc + Sfz.

7. A large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step four specifically further includes: Fit the expected storage requirement matching index Cpp and the expected processing duration Csc, and obtain the scheme selection coefficient Fkxs for each parallel configuration scheme. The specific calculation formula is as follows:

8. A large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step four specifically further includes: Rank and compare the scheme selection coefficients Fkxs of each parallel configuration scheme, and preferentially select the parallel configuration scheme with the highest value of the scheme selection coefficient Fkxs. At the same time, monitor the actual performance during real-time operation, including the actual storage usage and processing time; update the hardware performance data and network condition information in real time, and preset the deviation threshold between the actual situation and the estimated value. When the deviation between the actual situation and the estimated value exceeds the deviation threshold, readjust the parallel configuration scheme at this time.

9. The large model adaptive parallel training method for heterogeneous clusters according to claim 1, wherein: Step five specifically includes: Start the selected parallel configuration scheme and perform actual model training operation. During this period, the system automatically monitors and records the running data of the model, including the video memory usage and computing efficiency; at the same time, monitor potential performance bottleneck problems in real time during the process, including video memory overflow. When encountering performance bottleneck problems, automatically adjust the running parameters or pause the training and perform configuration adjustment.

10. A large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step five specifically further includes: During the model training process, calculate and record the actual computing duration T of each batch of data in real time. Among them, the specific formula for the actual computing duration T of each batch of data is as follows: where S is the total number of layers of the model, and t p,j is the computing time of the device responsible for calculating the j-th slice of the p-th layer of the model, and PP_time p represents the delay of passing the calculation result to the next layer, i.e., the p-th layer, after the calculation of the k-th layer of the model is completed in pipeline parallelism; while DP_time p is the time overhead of tensor parallelism of the p-th layer of the model; Next, preset the standard processing duration threshold Z, and compare and evaluate the actual computing duration T of each batch of data, the expected processing duration Csc, and the standard processing duration threshold Z. And the standard processing duration threshold Z > the expected processing duration Csc, to verify the accuracy of the estimated time efficiency. The specific evaluation content is as follows: If the standard processing duration threshold Z > the expected processing duration Csc ≥ the actual computing duration T, in this case, it shows that the actual performance of the parallel configuration is normal, and the predetermined efficiency target has been achieved or exceeded, and the expected processing duration Csc is reasonably set; If the standard processing duration threshold Z ≥ the actual calculation duration T > the expected processing duration Csc, in this case, it shows that the actual performance of the parallel configuration is normal. Although the actual calculation duration T reaches and is better than the standard processing duration threshold Z, it fails to reach the expected processing duration Csc; this indicates that the expected processing duration does not fully consider some obstacles in actual operation. At this time, it is necessary to adjust the expected model or consider adjusting the current configuration; If the actual calculation duration T > the standard processing duration threshold Z > the expected processing duration Csc, in this case, it shows that the actual performance of the parallel configuration is abnormal. The actual calculation duration T not only fails to reach the expected processing duration Csc but also exceeds the standard processing duration threshold Z. This indicates that the current parallel configuration scheme has abnormal efficiency and there are performance bottlenecks or configuration problems; at this time, focus on checking and solving the factors affecting performance, including insufficient hardware performance, improper software optimization, low data management efficiency, and network latency.

Citation Information

Patent Citations

  • Data co-processing method and system based on large model

    CN117608866A

  • Model parallel training method and device

    CN118410859A