Large model adaptive parallel training method capable of being used for heterogeneous cluster

By automatically measuring equipment performance in heterogeneous clusters and dynamically adjusting parallel configuration and resource allocation, the problem of low resource matching, video memory space and training configuration efficiency of large model training in heterogeneous clusters is solved, and an efficient and flexible training process is achieved.

CN119938327AActive Publication Date: 2025-05-06WEIYE ZHISUAN (BEIJING) TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510021206.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06
Estimated Expiration
2045-01-07

AI Technical Summary

Technical Problem

In heterogeneous clusters, large-model training faces the problems of resource matching and optimization, memory space limitations, and low training configuration and optimization efficiency.

Method used

By automatically measuring the computing performance and network performance of devices in heterogeneous clusters, dynamically compute data parallel values ​​and tensor parallel configurations, automatically adjust the allocation of model layers and hardware, use pipeline parallel technology to adjust the calculation delay, and monitor and adjust the parallel configuration in real time to adapt to changes in equipment performance.

Benefits of technology

It effectively optimizes performance matching between devices, improves resource utilization, solves the limitation of video memory space, improves training flexibility and efficiency, and avoids performance losses caused by resource mismatch and configuration problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938327A_ABST
    Figure CN119938327A_ABST
Patent Text Reader

Abstract

The invention discloses a large model adaptive parallel training method capable of being used for a heterogeneous cluster, relates to the technical field of parallel computing, and solves the problems of non-uniform resource utilization and low efficiency. Through dynamic calculation of data parallel values and tensor parallel configuration, distribution of a model layer and hardware is automatically adjusted, performance matching and cooperative work between devices are optimized, the resource utilization rate is increased, and performance bottlenecks are eliminated; meanwhile, by monitoring and dynamically adjusting a storage demand matching index Cpp and a processing duration Csc in real time, the video memory and the computing power of each device are optimally utilized, and the problem of space limitation of the video memory is solved; besides, the scheme also comprises the steps of continuously monitoring real-time performance data and dynamically adjusting parameters in pipeline parallelism, such as data transmission speed and calculation task reallocation, so that the flexibility and efficiency of training are improved, and the parallel training configuration is ensured to always adapt to the actual operation condition of the current cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of parallel computing, and in particular to a large model adaptive parallel training method that can be used for heterogeneous clusters. Background Art

[0002] Large models generally refer to models with more than one million parameters. This type of model has high versatility and strong generalization performance. It is currently widely used and has become a mainstream technology in the field of artificial intelligence. However, the main technical defects of large model training in heterogeneous computing clusters can be summarized as follows:

[0003] 1. Resource matching and optimization issues: In a heterogeneous cluster, different computing devices have different computing capabilities and memory capacities. If the large model to be trained is evenly divided and assigned to various computing devices without differentiated processing, it will lead to uneven resource utilization and low efficiency. High-performance devices may be idle while waiting for devices with poor performance to complete calculations, resulting in the so-called "barrel short board" phenomenon, that is, the overall performance is limited by the performance of the weakest device;

[0004] 2. Memory space limitation: Large model training usually requires GPU memory or computing chip memory that is about 16 times the model parameters. In a heterogeneous cluster, especially in an environment with limited resources or uneven performance, a single model of high-performance equipment is not enough to support all the needs of a large model;

[0005] 3. Low efficiency of training configuration and optimization: Current heterogeneous clusters often require a large number of experts to perform pre-slicing and configuration of model layers when performing large model training tasks. This method not only consumes a lot of manpower and computing resources, but also, because it is preset before training, it is difficult to quickly and accurately adjust changes that occur during the training process (such as model adjustments, data changes, equipment performance changes, etc.). Each change may require redesigning and verifying the training plan, reducing the flexibility and efficiency of training. Summary of the invention

[0006] In view of the deficiencies in the prior art, the present invention provides a large model adaptive parallel training method that can be used for heterogeneous clusters, which solves the problems mentioned in the background technology.

[0007] To achieve the above objectives, the present invention is implemented by the following technical solution: a large model adaptive parallel training method that can be used for heterogeneous clusters, comprising the following steps:

[0008] Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices in the heterogeneous cluster. This is used to measure the computing power and network bandwidth of different hardware in real time, and update the data-driven training configuration in real time based on the measurement results.

[0009] Step 2: Based on the device memory and computing power measured in step 1, dynamically calculate the data parallel value Dps and tensor parallel configuration, and adjust the device performance matching by automatically adjusting the allocation of model layers and hardware. The specific calculation formula of the data parallel value Dps is as follows:

[0010]

[0011] Where N is the total number of devices, GPUm i is the video memory capacity of the i-th device, α is the video memory conversion coefficient required for training accuracy, S is the total number of model layers, Layer i is the parameter quantity of the i-th layer;

[0012] Step 3: Based on the performance matching of the adjusted devices, preset the parallel configuration scheme for heterogeneous model training, and adjust the computing delay and the duration of the overall training by adopting pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rate and response time of each device;

[0013] Step 4: Before training begins, estimate the video memory and time costs of the preset parallel configuration scheme, calculate the expected storage requirement matching index Cpp and the expected processing time Csc of each parallel configuration scheme, then fit the expected storage requirement matching index Cpp and the expected processing time Csc, obtain the scheme optional coefficient Fkxs of each parallel configuration scheme and evaluate it, and finally make the selected parallel configuration scheme the optimal choice in terms of time efficiency while running under the current hardware configuration;

[0014] Step 5: Verify the parallel configuration scheme in step 4 in actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance issues during the training process, and verifying the accuracy of the estimated time efficiency by calculating the actual calculation time T of the batch data.

[0015] Preferably, step one specifically includes:

[0016] Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing power and inference computing power of each device in real time. Use open source base models with different architectures to measure performance under different data input and output lengths. Then use network testing tools suitable for different hardware platforms to evaluate the communication bandwidth of each device. The measurement results include the number of floating-point operations per second of the computing device and the megabits per second of communication performance, which will be recorded in real time and used to update the data-driven training configuration.

[0017] Preferably, step 2 specifically includes:

[0018] Based on step one, by analyzing the device video memory and computing power, and based on the measurement results of the computing power and network bandwidth of different hardware, determine the data load and computing tasks undertaken by each device in training; through automated algorithms, adjust the mapping relationship between the model layers and specific hardware to optimize the performance matching and collaborative work between devices; set the video memory capacity GPUm and parameter quantity Layer based on the ratio of the entire model parameter volume to the total video memory capacity of all participating devices, and finally obtain the data parallel value Dps; automatically adjust the parallel configuration to respond to the device by continuously monitoring real-time performance data.

[0019] Preferably, step three specifically includes:

[0020] According to the data parallel value Dps and tensor parallel configuration determined in step 2, preset parallel configuration schemes for multiple heterogeneous model training, including the initial configuration of pipeline parallelism for each hardware device, including allocating each model layer to different computing nodes, and setting the order and rate of data processing by each node, so that each computing node receives and processes data according to the real-time computing capability and network conditions.

[0021] Preferably, step three specifically further includes:

[0022] During the model training process, the real-time computing performance and response time of each hardware device are continuously monitored, and various parameters in the pipeline parallelization are dynamically adjusted according to the actual operation situation, including adjusting the data transmission speed and reallocating computing tasks, specifically:

[0023] Adjust data transmission speed: According to network communication delay and the data processing capacity of each device, dynamically adjust the data transmission speed between devices according to the receiving and processing capabilities of each device;

[0024] Computing task reallocation: Redistribute computing tasks in the model to various computing nodes based on the performance of the device.

[0025] Preferably, step 4 specifically includes:

[0026] For each pre-selected parallel configuration, based on model parameters, expected computational load, and hardware performance parameters, including the total model parameter quantity Z Layer 、The average video memory capacity of the device Z GPUm , the expected amount of video memory occupied Zzy, the device computing speed Sjs, the network delay time Syc, and the model computing complexity value Sfz. Then, after dimensionless processing of the model parameters, the expected computing load, and the hardware performance parameters, the expected storage demand matching index Cpp and the expected processing time Csc of each parallel configuration scheme are calculated. The specific calculation formula is as follows:

[0027]

[0028] Csc=Sjs×Syc+Sfz.

[0029] Preferably, step 4 specifically further includes:

[0030] Fit the expected storage demand matching index Cpp and the expected processing time Csc to obtain the optional coefficient Fkxs of each parallel configuration solution. The specific calculation formula is as follows:

[0031]

[0032] Preferably, step 4 specifically further includes:

[0033] The optional coefficient Fkxs of each parallel configuration scheme is ranked and compared, and the parallel configuration scheme with the highest optional coefficient Fkxs value is given priority. At the same time, the performance during actual runtime is monitored in real time, including actual storage usage and processing time; hardware performance data and network status information are updated in real time, and the deviation threshold between the actual situation and the estimated value is preset. When the deviation between the actual situation and the estimated value exceeds the deviation threshold, the parallel configuration scheme is readjusted.

[0034] Preferably, step five specifically includes:

[0035] Start the selected parallel configuration scheme and perform actual model training. During this period, the system automatically monitors and records the model's operating data, including video memory usage and computing efficiency. At the same time, it monitors potential performance bottlenecks in real time, including video memory overflow. When encountering performance bottlenecks, it automatically adjusts operating parameters or pauses training and makes configuration adjustments.

[0036] Preferably, step five specifically further includes:

[0037] During the model training process, the actual calculation time T of each batch of data is calculated in real time and recorded. The specific formula for the actual calculation time T of each batch of data is as follows:

[0038]

[0039] Where S is the total number of layers in the model, t p,j PP_time is the computation time of the device responsible for computing the jth slice of the pth layer of the model. p It indicates the delay of transferring the calculation result to the next layer after the i-th layer of the model is calculated in the pipeline parallel; while DP_time p is the time cost of tensor parallelism in the pth layer of the model;

[0040] Next, a standard processing time threshold Z is preset, and the actual calculation time T, expected processing time Csc and standard processing time threshold Z of each batch of data are compared and evaluated. The standard processing time threshold Z> expected processing time Csc to verify the accuracy of time efficiency estimation. The specific evaluation contents are as follows:

[0041] If the standard processing time threshold Z> expected processing time Csc≥ actual calculation time T, in this case, the actual performance of the parallel configuration is normal, has reached or exceeded the predetermined efficiency target, and the expected processing time Csc is reasonably set;

[0042] If the standard processing time threshold Z ≥ actual calculation time T > expected processing time Csc, the actual performance of the parallel configuration is normal. Although the actual calculation time T reaches and exceeds the standard processing time threshold Z, it fails to reach the expected processing time Csc. This indicates that the expected processing time does not fully consider some obstacles in actual operation. In this case, it is necessary to adjust the expected model or consider adjusting the current configuration.

[0043] If the actual calculation time T> the standard processing time threshold Z> the expected processing time Csc, the actual performance of the parallel configuration is abnormal. The actual calculation time T not only fails to reach the expected processing time Csc, but also exceeds the standard processing time threshold Z. This indicates that the current parallel configuration solution is abnormally efficient and there are performance bottlenecks or configuration problems. At this time, focus on checking and solving factors that affect performance, including insufficient hardware performance, improper software optimization, inefficient data management, and network delays.

[0044] The present invention provides a large model adaptive parallel training method that can be used for heterogeneous clusters. It has the following beneficial effects:

[0045] (1) This is a large model adaptive parallel training method that can be used in heterogeneous clusters. In a heterogeneous cluster, each computing device has different computing power and memory capacity. If the large model is evenly divided and allocated to various computing devices without differentiated processing, it will lead to uneven resource utilization and low efficiency. In addition, high-performance devices may be idle while waiting for devices with poor performance to complete calculations, resulting in the "short board of the barrel" phenomenon, that is, the overall performance is limited by the performance of the weakest device. This solution dynamically calculates data parallel values ​​and tensor parallel configurations, automatically adjusts the allocation of model layers and hardware, effectively optimizes the performance matching and collaborative work between devices, improves resource utilization and eliminates performance bottlenecks;

[0046] (2) This is a large model adaptive parallel training method that can be used in heterogeneous clusters. The model parameters required for large model training are usually about 16 times the GPU memory or computing chip memory. In a heterogeneous cluster environment with limited resources or uneven performance, a single model of high-performance device is often insufficient to support all the needs of a large model. This solution ensures that the memory and computing power of each device are optimally utilized by real-time monitoring and dynamic adjustment of the storage demand matching index Cpp and processing time Csc, thereby solving the problem of memory space limitation.

[0047] (3) This is a large model adaptive parallel training method that can be used for heterogeneous clusters. Current heterogeneous clusters often require a large number of experts to perform pre-slicing and configuration of model layers when performing large model training tasks. This method not only consumes a lot of manpower and computing resources, but also because it is preset before training, it is difficult to quickly and accurately adjust to changes that occur during the training process, including model adjustments, data changes, and equipment performance changes. Each change may require redesigning and verifying the training plan, reducing the flexibility and efficiency of training. This solution continuously monitors real-time performance data and dynamically adjusts various parameters in the pipeline parallelism, including data transmission speed and computing task reallocation, to ensure that the parallel training configuration always adapts to the actual operating conditions of the current cluster, thereby improving the flexibility and efficiency of training. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic diagram of the steps of a large model adaptive parallel training method applicable to heterogeneous clusters according to the present invention; DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] Example 1

[0051] See also Figure 1 , a large model adaptive parallel training method that can be used for heterogeneous clusters, comprising the following steps:

[0052] Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices in the heterogeneous cluster. This is used to measure the computing power and network bandwidth of different hardware in real time, and update the data-driven training configuration in real time based on the measurement results.

[0053] Step 2: Based on the device memory and computing power measured in step 1, dynamically calculate the data parallel value Dps and tensor parallel configuration, and adjust the device performance matching by automatically adjusting the allocation of model layers and hardware. The specific calculation formula of the data parallel value Dps is as follows:

[0054]

[0055] Where N is the total number of devices, GPUm i is the video memory capacity of the i-th device, α is the video memory conversion coefficient required for training accuracy, S is the total number of model layers, Layer i is the parameter quantity of the i-th layer;

[0056] Step 3: Based on the performance matching of the adjusted devices, preset the parallel configuration scheme for heterogeneous model training, and adjust the computing delay and the duration of the overall training by adopting pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rate and response time of each device;

[0057] Step 4: Before training begins, estimate the video memory and time costs of the preset parallel configuration scheme, calculate the expected storage requirement matching index Cpp and the expected processing time Csc of each parallel configuration scheme, then fit the expected storage requirement matching index Cpp and the expected processing time Csc, obtain the scheme optional coefficient Fkxs of each parallel configuration scheme and evaluate it, and finally make the selected parallel configuration scheme the optimal choice in terms of time efficiency while running under the current hardware configuration;

[0058] Step 5: Verify the parallel configuration scheme in step 4 in actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance issues during the training process, and verifying the accuracy of the estimated time efficiency by calculating the actual calculation time T of the batch data.

[0059] In this embodiment, step 1 automatically measures the computing performance and network performance of various hardware devices in the heterogeneous cluster by using a general computing framework, measures the computing power and network bandwidth of different hardware in real time, and updates the data-driven training configuration in real time based on the measurement results, ensuring that the training configuration is always based on the latest performance data, thereby improving resource utilization efficiency and configuration accuracy;

[0060] Step 2 dynamically calculates the data parallel value Dps and tensor parallel configuration based on the device video memory and computing power measured in step 1, and automatically adjusts the allocation of model layers and hardware to adjust the device performance matching, thereby optimizing the allocation and utilization of hardware resources and reducing performance losses caused by resource mismatch;

[0061] Step 3: Based on the performance matching of the adjusted devices, the heterogeneous parallel strategy of model training is executed, and the pipeline parallel technology is used to adjust the computing delay and the duration of the overall training. This strategy minimizes the delay caused by the performance differences of different devices by adapting to the computing rate and response time of each device, thereby improving the overall speed and efficiency of training.

[0062] Step 4: Before training begins, estimate the memory and time costs of the pre-selected parallel configuration schemes, calculate the expected storage requirement matching index Cpp and the expected processing time Csc for each parallel configuration scheme, then fit the expected storage requirement matching index Cpp and the expected processing time Csc, obtain the scheme optional coefficient Fkxs of each parallel configuration scheme and evaluate it, so that the selected parallel configuration scheme can run under the current hardware configuration while also being the best choice in terms of time efficiency, ensuring the optimization of the training plan and improving the efficiency and scientificity of decision-making.

[0063] Step 5: Verify the parallel configuration scheme in step 4 in actual operation, run the training process and calibrate the estimated model parameters, monitor and prevent video memory overflow and other performance problems during the training process, and verify the estimated accuracy of time efficiency by calculating the actual calculation time T of the batch data, ensuring the stability of the training process and the accuracy of the prediction. Optimize and adjust the parallel configuration through actual operation feedback to achieve the best operation status.

[0064] Example 2

[0065] Step 1 specifically includes:

[0066] Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing power and inference computing power of each device in real time. Use open source base models with different architectures to measure performance under different data input and output lengths. Then use network testing tools suitable for different hardware platforms to evaluate the communication bandwidth of each device. The measurement results include the number of floating-point operations per second of the computing device and the megabits per second of communication performance, which will be recorded in real time and used to update the data-driven training configuration.

[0067] Step 2 specifically includes:

[0068] Based on step one, by analyzing the device video memory and computing power, and based on the measurement results of the computing power and network bandwidth of different hardware, determine the data load and computing tasks undertaken by each device in training; through automated algorithms, adjust the mapping relationship between the model layers and specific hardware to optimize the performance matching and collaborative work between devices; set the video memory capacity GPUm and parameter quantity Layer based on the ratio of the entire model parameter volume to the total video memory capacity of all participating devices, and finally obtain the data parallel value Dps; automatically adjust the parallel configuration to respond to the device by continuously monitoring real-time performance data.

[0069] In this embodiment, step 1 measures the fixed matrix computing power and reasoning computing power of each device in real time by running predefined computing tasks and network bandwidth tests, and uses open source base models of different architectures to measure performance under different data input and output lengths. This enables the system to accurately configure resources according to the real-time computing power and network bandwidth of each device, ensuring that the data-driven training configuration is always updated based on the latest performance data;

[0070] Then, in step 2, based on these measurement results, by analyzing the device memory and computing power, an automated algorithm is used to adjust the mapping relationship between the model layer and the hardware, optimize the performance matching and collaboration between devices, and accurately calculate the data parallel value Dps by setting the memory capacity GPUmi and the model layer parameter Layeri, which not only enhances the flexibility of the system, but also improves the efficiency of training; this method allows the system to dynamically adjust the parallel configuration in response to changes in device performance, ensuring that parallel training can achieve optimal performance under various operating conditions; in summary, this training method effectively improves resource utilization, reduces performance losses caused by resource mismatch, and enhances the stability and predictability of the training process, providing an efficient and reliable solution for the training of large models in a heterogeneous cluster environment.

[0071] Example 3

[0072] Step three specifically includes:

[0073] According to the data parallel value Dps and tensor parallel configuration determined in step 2, preset parallel configuration schemes for multiple heterogeneous model training, including the initial configuration of pipeline parallelism for each hardware device, including allocating each model layer to different computing nodes, and setting the order and rate of data processing by each node, so that each computing node receives and processes data according to the real-time computing capability and network conditions.

[0074] Step three specifically includes:

[0075] During the model training process, the real-time computing performance and response time of each hardware device are continuously monitored, and various parameters in the pipeline parallelization are dynamically adjusted according to the actual operation situation, including adjusting the data transmission speed and reallocating computing tasks, specifically:

[0076] Adjust data transmission speed: According to network communication delay and the data processing capacity of each device, dynamically adjust the data transmission speed between devices according to the receiving and processing capabilities of each device;

[0077] Computing task reallocation: Redistribute computing tasks in the model to various computing nodes based on the performance of the device.

[0078] Step 4 specifically includes:

[0079] Step 4 specifically includes:

[0080] For each pre-selected parallel configuration, based on model parameters, expected computational load, and hardware performance parameters, including the total model parameter quantity Z Layer 、The average video memory capacity of the device Z GPUm , the expected amount of video memory occupied Zzy, the device computing speed Sjs, the network delay time Syc, and the model computing complexity value Sfz. Then, after dimensionless processing of the model parameters, the expected computing load, and the hardware performance parameters, the expected storage demand matching index Cpp and the expected processing time Csc of each parallel configuration scheme are calculated. The specific calculation formula is as follows:

[0081]

[0082] Csc=Sjs×Syc+Sfz.

[0083] Step 4 specifically includes:

[0084] Fit the expected storage demand matching index Cpp and the expected processing time Csc to obtain the optional coefficient Fkxs of each parallel configuration solution. The specific calculation formula is as follows:

[0085]

[0086] Step 4 specifically includes:

[0087] The optional coefficient Fkxs of each parallel configuration scheme is ranked and compared, and the parallel configuration scheme with the highest optional coefficient Fkxs value is given priority. At the same time, the performance during actual runtime is monitored in real time, including actual storage usage and processing time; hardware performance data and network status information are updated in real time, and the deviation threshold between the actual situation and the estimated value is preset. When the deviation between the actual situation and the estimated value exceeds the deviation threshold, the parallel configuration scheme is readjusted.

[0088] In this embodiment, step three ensures that each model layer is effectively allocated to a computing node with appropriate computing power and network conditions by performing initial configuration of pipeline parallelism on the hardware device based on the data parallel value Dps and tensor parallel configuration. This step enables each computing node to receive and process data according to its real-time computing power and network conditions, thereby optimizing the utilization of computing resources and the order of data processing. At the same time, step three also includes continuously monitoring the real-time computing performance and response time of each hardware device, and dynamically adjusting the data transmission speed and reallocating computing tasks according to actual conditions. These adjustments enable the system to flexibly respond to performance changes and maintain operating efficiency.

[0089] Step 4 conducts an in-depth performance evaluation for each pre-selected parallel configuration scheme, based on parameters such as the total model parameter Zmc, the average device memory capacity Zcl, the expected memory usage Zzy, the device computing speed Sjs, the network delay time Syc, and the model computing complexity value Sfz. These parameters are dimensionlessly processed and used to calculate the storage requirement matching index Cpp and the expected processing time Csc for each configuration scheme. This not only helps to accurately evaluate the storage and processing capabilities of each scheme, but also enables the scheme optional coefficient Fkxs to be obtained by fitting these indices. The coefficient is used to compare and select the parallel configuration scheme with the best performance. In addition, step 4 also includes real-time monitoring of system performance and updating of hardware and network status information to ensure that the selected scheme can continue to perform best in the actual environment. Overall, this training method greatly improves the flexibility, efficiency and reliability of model training by comprehensively utilizing real-time performance monitoring and automatic adjustment strategies, ensuring the optimal training results of large models in heterogeneous cluster environments.

[0090] Example 4

[0091] Step 5 specifically includes:

[0092] Start the selected parallel configuration scheme and perform actual model training. During this period, the system automatically monitors and records the model's operating data, including video memory usage and computing efficiency. At the same time, it monitors potential performance bottlenecks in real time, including video memory overflow. When encountering performance bottlenecks, it automatically adjusts operating parameters or pauses training and makes configuration adjustments.

[0093] Step 5 specifically includes:

[0094] During the model training process, the actual calculation time T of each batch of data is calculated in real time and recorded. The specific formula for the actual calculation time T of each batch of data is as follows:

[0095]

[0096] Where S is the total number of layers in the model, t p,jPP_time is the computation time of the device responsible for computing the jth slice of the pth layer of the model. p It indicates the delay of transferring the calculation result to the next layer after the i-th layer of the model is calculated in the pipeline parallel; while DP_time p is the time cost of tensor parallelism in the pth layer of the model;

[0097] Next, a standard processing time threshold Z is preset, and the actual calculation time T, expected processing time Csc and standard processing time threshold Z of each batch of data are compared and evaluated. The standard processing time threshold Z> expected processing time Csc to verify the accuracy of time efficiency estimation. The specific evaluation contents are as follows:

[0098] If the standard processing time threshold Z> expected processing time Csc≥ actual calculation time T, in this case, the actual performance of the parallel configuration is normal, has reached or exceeded the predetermined efficiency target, and the expected processing time Csc is reasonably set;

[0099] If the standard processing time threshold Z ≥ actual calculation time T > expected processing time Csc, the actual performance of the parallel configuration is normal. Although the actual calculation time T reaches and exceeds the standard processing time threshold Z, it fails to reach the expected processing time Csc. This indicates that the expected processing time does not fully consider some obstacles in actual operation. In this case, it is necessary to adjust the expected model or consider adjusting the current configuration.

[0100] If the actual calculation time T> the standard processing time threshold Z> the expected processing time Csc, the actual performance of the parallel configuration is abnormal. The actual calculation time T not only fails to reach the expected processing time Csc, but also exceeds the standard processing time threshold Z. This indicates that the current parallel configuration solution is abnormally efficient and there are performance bottlenecks or configuration problems. At this time, focus on checking and solving factors that affect performance, including insufficient hardware performance, improper software optimization, inefficient data management, and network delays.

[0101] In this embodiment, step five performs key actual operation and monitoring of the large model adaptive parallel training method in the heterogeneous cluster, ensuring the efficiency and stability of the training process; by starting the selected parallel configuration scheme and performing actual model training, the system can monitor and record key operating data in real time, including video memory usage and computing efficiency, so that any potential performance bottleneck problems, such as video memory overflow, can be discovered in time and necessary configuration adjustments can be made by automatically adjusting operating parameters or pausing training, thereby avoiding possible system crashes or performance degradation; in addition, the actual calculation time T of each batch of data is calculated in real time and compared with the preset standard processing time threshold Z and the expected processing time Csc, which not only verifies the accuracy of the estimated time efficiency It not only provides a real-time feedback mechanism to evaluate whether the parallel configuration scheme has achieved the predetermined efficiency target; in this process, the actual calculation time T of each batch of data is calculated based on the total number of layers S of the model, the tensor parallel time overhead of each layer, and the delay in pipeline parallelism. These calculations reflect the actual performance of the current configuration scheme when processing complex data structures, so that the system can adjust the model training strategy according to actual conditions, optimize hardware and software configuration, and improve the overall training efficiency; through this comprehensive evaluation and adjustment, step five not only enhances the transparency and controllability of the training process, but also significantly improves the system's adaptability and efficiency to complex data processing, providing an efficient and reliable training solution for large-scale machine learning tasks.

[0102] Among them, t p,j The value depends on the batch data size, input and output length, data accuracy, and the computing performance of the device (Example.

[0103] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A large model adaptive parallel training method that can be used for heterogeneous clusters, characterized by: The following steps are involved: Step 1: Use a general computing framework to automatically measure the computing performance and network performance of various hardware devices in the heterogeneous cluster. This is used to measure the computing power and network bandwidth of different hardware in real time, and update the data-driven training configuration in real time based on the measurement results. Step 2: Based on the device memory and computing power measured in step 1, dynamically calculate the data parallel value Dps and tensor parallel configuration, and adjust the device performance matching by automatically adjusting the allocation of model layers and hardware. The specific calculation formula of the data parallel value Dps is as follows: Where N is the total number of devices, GPUm i is the video memory capacity of the i-th device, α is the video memory conversion coefficient required for training accuracy, S is the total number of model layers, Layer i is the parameter quantity of the i-th layer; Step 3: Based on the performance matching of the adjusted devices, preset the parallel configuration scheme for heterogeneous model training, and adjust the computing delay and the duration of the overall training by adopting pipeline parallel technology, where the dynamic configuration of the pipeline adapts to the computing rate and response time of each device; Step 4: Before training begins, estimate the video memory and time costs of the preset parallel configuration scheme, calculate the expected storage requirement matching index Cpp and the expected processing time Csc of each parallel configuration scheme, then fit the expected storage requirement matching index Cpp and the expected processing time Csc, obtain the scheme optional coefficient Fkxs of each parallel configuration scheme and evaluate it, and finally make the selected parallel configuration scheme the optimal choice in terms of time efficiency while running under the current hardware configuration; Step 5: Verify the parallel configuration scheme in step 4 in actual operation, including running the training process and calibrating the estimated model parameters, monitoring and preventing video memory overflow and other performance issues during the training process, and verifying the estimated accuracy of time efficiency by calculating the actual calculation time T of the batch data.

2. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 1 specifically includes: Run predefined computing tasks and network bandwidth tests to measure the fixed matrix computing power and inference computing power of each device in real time. Use open source base models with different architectures to measure performance under different data input and output lengths. Then use network testing tools suitable for different hardware platforms to evaluate the communication bandwidth of each device. The measurement results include the number of floating-point operations per second of the computing device and the megabits per second of communication performance, which will be recorded in real time and used to update the data-driven training configuration.

3. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 2 specifically includes: Based on step one, by analyzing the device video memory and computing power, and based on the measurement results of the computing power and network bandwidth of different hardware, determine the data load and computing tasks undertaken by each device in training; through automated algorithms, adjust the mapping relationship between the model layers and specific hardware to optimize the performance matching and collaborative work between devices; set the video memory capacity GPUm and parameter quantity Layer based on the ratio of the entire model parameter volume to the total video memory capacity of all participating devices, and finally obtain the data parallel value Dps; automatically adjust the parallel configuration to respond to the device by continuously monitoring real-time performance data.

4. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step three specifically includes: According to the data parallel value Dps and tensor parallel configuration determined in step 2, preset parallel configuration schemes for multiple heterogeneous model training, including the initial configuration of pipeline parallelism for each hardware device, including allocating each model layer to different computing nodes, and setting the order and rate of data processing by each node, so that each computing node receives and processes data according to the real-time computing capability and network conditions.

5. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 4, characterized in that: Step three specifically includes: During the model training process, the real-time computing performance and response time of each hardware device are continuously monitored, and various parameters in the pipeline parallelization are dynamically adjusted according to the actual operation situation, including adjusting the data transmission speed and reallocating computing tasks, specifically: Adjust data transmission speed: According to network communication delay and the data processing capacity of each device, dynamically adjust the data transmission speed between devices according to the receiving and processing capabilities of each device; Computing task reallocation: Redistribute computing tasks in the model to various computing nodes based on the performance of the device.

6. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 4 specifically includes: For each pre-selected parallel configuration, based on model parameters, expected computational load, and hardware performance parameters, including the total model parameter quantity Z Layer 、The average video memory capacity of the device Z GPUm , the expected amount of video memory occupied Zzy, the device computing speed Sjs, the network delay time Syc, and the model computing complexity value Sfz. Then, after dimensionless processing of the model parameters, the expected computing load, and the hardware performance parameters, the expected storage demand matching index Cpp and the expected processing time Csc of each parallel configuration scheme are calculated. The specific calculation formula is as follows: Csc=Sjs×Syc+Sfz.

7. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 4 specifically includes: Fit the expected storage demand matching index Cpp and the expected processing time Csc to obtain the optional coefficient Fkxs of each parallel configuration solution. The specific calculation formula is as follows:

8. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 4 specifically includes: The optional coefficient Fkxs of each parallel configuration scheme is ranked and compared, and the parallel configuration scheme with the highest optional coefficient Fkxs value is given priority. At the same time, the performance during actual runtime is monitored in real time, including actual storage usage and processing time; hardware performance data and network status information are updated in real time, and the deviation threshold between the actual situation and the estimated value is preset. When the deviation between the actual situation and the estimated value exceeds the deviation threshold, the parallel configuration scheme is readjusted.

9. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 5 specifically includes: Start the selected parallel configuration scheme and perform actual model training. During this period, the system automatically monitors and records the model's operating data, including video memory usage and computing efficiency. At the same time, it monitors potential performance bottlenecks in real time, including video memory overflow. When encountering performance bottlenecks, it automatically adjusts operating parameters or pauses training and makes configuration adjustments.

10. The large model adaptive parallel training method applicable to heterogeneous clusters according to claim 1, characterized in that: Step 5 specifically includes: During the model training process, the actual calculation time T of each batch of data is calculated in real time and recorded. The specific formula for the actual calculation time T of each batch of data is as follows: Where S is the total number of layers in the model, t p,j PP_time is the computation time of the device responsible for computing the jth slice of the pth layer of the model. p It indicates the delay of transferring the calculation result to the next layer after the i-th layer of the model is calculated in the pipeline parallel; DP_time p is the time cost of tensor parallelism in the pth layer of the model; Next, a standard processing time threshold Z is preset, and the actual calculation time T, expected processing time Csc and standard processing time threshold Z of each batch of data are compared and evaluated. The standard processing time threshold Z> expected processing time Csc to verify the accuracy of time efficiency estimation. The specific evaluation contents are as follows: If the standard processing time threshold Z> expected processing time Csc≥ actual calculation time T, in this case, the actual performance of the parallel configuration is normal, has reached or exceeded the predetermined efficiency target, and the expected processing time Csc is reasonably set; If the standard processing time threshold Z ≥ actual calculation time T > expected processing time Csc, the actual performance of the parallel configuration is normal. Although the actual calculation time T reaches and exceeds the standard processing time threshold Z, it fails to reach the expected processing time Csc. This indicates that the expected processing time does not fully consider some obstacles in actual operation. In this case, it is necessary to adjust the expected model or consider adjusting the current configuration. If the actual calculation time T> the standard processing time threshold Z> the expected processing time Csc, the actual performance of the parallel configuration is abnormal. The actual calculation time T not only fails to reach the expected processing time Csc, but also exceeds the standard processing time threshold Z. This indicates that the current parallel configuration solution is abnormally efficient and there are performance bottlenecks or configuration problems. At this time, focus on checking and solving factors that affect performance, including insufficient hardware performance, improper software optimization, inefficient data management, and network delays.

Citation Information

Patent Citations

  • Data co-processing method and system based on large model

    CN117608866A

  • Model parallel training method and device

    CN118410859A

  • Heterogeneous cluster model training method and system, electronic equipment and medium

    CN118428450A

  • Large-model heterogeneous cluster scheduling system and method based on adaptive parallel co-optimization

    CN118916156A

  • Intelligent scheduling method and system for heterogeneous equipment system

    CN118939438A

Cited By

  • Processor performance hardware optimization method and system, electronic equipment and storage medium

    CN120196527A

  • Dynamic resource optimization method and system based on large model incremental learning

    CN120223553A

  • Dynamic resource optimization method and system based on large model incremental learning

    CN120223553B

  • Model training method and device based on heterogeneous GPU cluster and storage medium

    CN120508395A

  • Server cluster monitoring method, device and system and electronic equipment

    CN120639596A