Method and device for determining model training configuration parameters and storage medium
By dividing the configuration items of model training into two categories, fixed and to be determined, predicting and optimizing the parameters of the configuration items to be determined, the problems of large hardware resource demand and high cost in large model training are solved, and more efficient resource utilization and cost reduction are achieved.
Patent Information
- Application Number
- CN202510273331.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-27
AI Technical Summary
There are problems of large hardware resource demand and high cost in the training stage of large model, which limits the wide application and development of large model technology.
By dividing the configuration items trained by the model into fixed configuration items and pending configuration items, the training performance of the model to be trained under different configuration parameters in the pending configuration items is predicted, and the target configuration parameters are selected based on the predicted performance to optimize training efficiency and reduce costs.
It significantly reduces the cost required for model training, while maximizing resource utilization efficiency, reducing resource waste, and improving training efficiency, solving the problems of large demands and high costs of hardware resources.
Smart Images

Figure CN120218192A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a method, apparatus, and storage medium for determining model training configuration parameters. Background Art
[0002] With the continuous increase in model scale, the problems of computing power and cost are becoming increasingly severe. The effective application of large models depends on powerful computing capabilities, and a large amount of computing resources are required during both training and inference processes. The larger the model structure, the more computing resources are needed.
[0003] Currently, the amount of computation used in the largest-scale artificial intelligence training doubles every 3.4 months, indicating that the demand and cost for hardware resources, especially GPUs (Graphics Processing Units), in large model training have reached unprecedented heights.
[0004] However, there are problems of large computing power requirements and high costs during the training phase of current large models, which limit the wide application and development of large model technologies. Summary of the Invention
[0005] This application provides a method, apparatus, and storage medium for determining model training configuration parameters to at least solve the problems of large hardware resource requirements and high costs during the training phase of large models in related technologies.
[0006] This application provides a method for determining model training configuration parameters, including: obtaining a plurality of configuration items required for a model to be trained during training; the plurality of configuration items include a set of fixed configuration items and a set of undetermined configuration items; when the set of fixed configuration items fixedly adopt specified configurations, predicting the training performance of the model to be trained when the set of undetermined configuration items adopt multiple configuration parameters for training, and obtaining predicted training performances corresponding to the multiple configuration parameters; selecting a target configuration parameter from the multiple configuration parameters according to the predicted training performances corresponding to the multiple configuration parameters; the target configuration parameter is the configuration parameter adopted by the set of undetermined configuration items when training the model to be trained.
[0007] The present application also provides an apparatus for determining configuration parameters for model training, including: a configuration item acquisition module, configured to acquire a plurality of configuration items required for a model to be trained during training; the plurality of configuration items including a set of fixed configuration items and a set of undetermined configuration items; a training performance prediction module, configured to predict the training performance corresponding to the model to be trained when the set of undetermined configuration items adopt multiple configuration parameters during training, given that the set of fixed configuration items are fixedly configured with specified configurations, and obtain predicted training performances corresponding to the multiple configuration parameters; a configuration parameter determination module, configured to select a target configuration parameter from the multiple configuration parameters according to the predicted training performances corresponding to the multiple configuration parameters; the target configuration parameter being the configuration parameter adopted by the set of undetermined configuration items when the model to be trained is trained.
[0008] The present application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above methods for determining configuration parameters for model training when executing the computer program.
[0009] The present application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above methods for determining configuration parameters for model training.
[0010] The present application also provides a computer program product including a computer program, which, when executed by a processor, implements the steps of any of the above methods for determining configuration parameters for model training.
[0011] Through the present application, the configuration items are divided into two categories: fixed configuration items and undetermined configuration items. This design allows, given the known fixed configuration items, to focus on optimizing the selection of undetermined configuration items to improve training efficiency and reduce costs; under the constraint that the fixed configuration items are configured with specified configurations, predict the training performance of the model to be trained when the undetermined configuration items adopt different configuration parameters during training, and based on the predicted training performances corresponding to the multiple configuration parameters, select a target configuration parameter from the multiple configuration parameters of the set of undetermined configuration items. When the model to be trained is trained, selecting the target configuration parameter for the set of undetermined configuration items can significantly reduce the costs required for model training, maximize the resource utilization efficiency, reduce resource waste, improve training efficiency, and solve the problems of high hardware resource requirements and high costs in the large model training stage in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0013] Figure 1 Schematic diagram of the application of a method for determining model training configuration parameters provided by an embodiment of the present application;
[0014] Figure 2 Flowchart of a method for determining model training configuration parameters provided by an embodiment of the present application;
[0015] Figure 3 Structural diagram of a terminal device provided by an embodiment of the present application;
[0016] Figure 4 Schematic diagram of the performance prediction module provided by an embodiment of the present application predicting the predicted training performance corresponding to multiple configuration parameters;
[0017] Figure 5 Schematic diagram of the video memory prediction module provided by an embodiment of the present application predicting the total maximum video memory occupancy corresponding to multiple configuration parameters;
[0018] Figure 6 Structural diagram of a device for determining model training configuration parameters provided by an embodiment of the present application. Detailed implementation manners
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0020] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0021] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following will further elaborate on the present application in conjunction with the drawings and specific implementation manners.
[0022] Currently, the computing power used in the largest-scale AI training is doubling at a rate of 3.4 months per month, indicating that the demand and cost for hardware resources, especially GPUs, in large model training have reached unprecedented heights. For example, the training of GPT-4 (Generative Pre-trained Transformer) requires tens of thousands of GPUs, and the training cost of the word model exceeds $12 million. There are problems of large computing power demand and high cost in the current training stage of large models, which limit the wide application and development of large model technology. To solve the above problems, the embodiments of the present application provide a method for determining model training configuration parameters. Before training the model to be trained, predict the training performance of the model to be trained when using various configuration parameters for multiple configuration items, and select the target configuration parameters most suitable for the model to be trained from various configuration parameters according to the training performance corresponding to different configuration parameters, effectively solving the problem of high computing power cost in large model training, improving the model training performance and efficiency, and promoting the popularization and innovation of large model technology.
[0023] According to one aspect of the embodiments of the present application, a method for determining model training configuration parameters is provided. Optionally, in this embodiment, the above method for determining model training configuration parameters may be but is not limited to being applied to a hardware environment such as Figure 1 shown in Figure 1 including a terminal device 102 and a server 104. The server 104 can be connected to the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal device 102 or a client installed on the terminal device 102. A database can be set on the server 104 or independently of the server 104 to provide data storage services for the server 104.
[0024] The above network may include but is not limited to at least one of the following: wired network, wireless network. The above wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC (Personal Computer), mobile phone, tablet computer, etc. The server 104 may be but is not limited to a cloud server, server cluster or other server types.
[0025] The method for determining model training configuration parameters in the embodiments of the present application may be executed by the server 104, or may be executed by the terminal device 102, or may be jointly executed by the server 104 and the terminal device 102. Among them, the execution of the method for determining model training configuration parameters by the terminal device 102 may also be executed by a client installed thereon.
[0026] Taking the method for determining the model training configuration parameters in this embodiment executed by the terminal device 102 as an example, Figure 2 is a flow chart of an optional method for determining model training configuration parameters according to an embodiment of the present application, such as Figure 2 As shown, the process of the method may include the following steps:
[0027] Step S202, obtaining multiple configuration items required for training the model to be trained; the multiple configuration items include a group of fixed configuration items and a group of pending configuration items.
[0028] Among them, the model to be trained refers to a model whose final architecture and parameter settings have not been determined before training. It can be a GPT-based structure or any other model architecture, such as BERT (Bidirectional Encoder Representations from Transformers), ViT (Vision Transformer), etc.
[0029] Configuration items refer to various parameters and strategies that can affect the utilization of computing resources during model training. For example, configuration items can be hardware conditions (such as graphics card model, network bandwidth, NvLink (NVIDIA connection) bandwidth), model parameters (such as the number of hidden layers, the number of model layers, the number of vocabulary, the number of hidden layers of MLP layers, etc.) and parallel strategies (tensor parallel paths, pipeline parallel paths, data parallel paths, expert parallel paths, sequence parallel paths, etc.). In addition, configuration items can also include but are not limited to the number of micro-batches, learning rate, optimizer type, activation function, model pre-training and fine-tuning strategies, data pre-processing methods, etc. When selecting models and formulating strategies, it is key to reasonably adjust these configuration items to improve training efficiency and model performance.
[0030] When configuring the training of the model to be trained, the configuration items can be divided into two groups: a group of fixed configuration items and a group of pending configuration items. Fixed configuration items refer to configuration items whose configuration parameters are pre-specified among multiple configuration items of the model to be trained. For example, fixed configuration items can be hardware conditions and optimizer types, which are usually determined in the preparation stage of model training and will not be dynamically adjusted during the training process. The pending configuration items are those that need to be determined through prediction and optimization before training, such as model structure parameters, parallel strategy parameters, and the number of micro-batches. These configuration items directly affect the memory usage and computing efficiency, and need to be carefully selected according to the current hardware conditions and expected training goals.
[0031] Optionally, the terminal device determines a training framework of the model to be trained and multiple configuration items required for training the model to be trained, and divides the multiple configuration items into a group of fixed configuration items and a group of pending configuration items according to actual needs.
[0032] Step S204, when a group of fixed configuration items adopts a fixed specified configuration, predict the training performance of the to-be-trained model when a group of to-be-determined configuration items adopts multiple configuration parameters for training, and obtain the predicted training performance corresponding to the multiple configuration parameters.
[0033] Among them, the specified configuration refers to the configuration parameters pre-specified from the configuration parameter range of the fixed configuration items, wherein the configuration parameter range of the fixed configuration items refers to a series of possible values or conditions set for the fixed configuration items during the entire training process before starting prediction and model training. For example, if a set of fixed configuration items includes hardware conditions, the configuration parameter range of the hardware conditions can be all mainstream GPU models currently on the market, and the specified configuration of the hardware conditions can be a specified graphics card model selected from the corresponding configuration parameter range. For example, if a set of fixed configuration items includes model parameters, the configuration parameter range of the model parameters includes a model layer number range of 6 to 48 layers, and a hidden layer size range of 512 to 8192. The specified configuration of the model parameters can be that the model layer number is fixed to 24 layers and the hidden layer is fixed to 1024.
[0034] Each pending configuration item in a set of pending configuration items corresponds to a plurality of configuration parameters, and the plurality of configuration parameters corresponding to each pending configuration item refer to a plurality of candidate values set for each pending configuration item during the prediction process, and the plurality of configuration parameters corresponding to each pending configuration item are selected from the configuration parameter range of each pending configuration item, wherein the configuration parameter range of each pending configuration item refers to a series of possible values or conditions set for the pending configuration item during the entire training process before starting the prediction and model training. For example, if the pending configuration item includes model parameters, the configuration parameter range of the number of model layers includes 12, 24 or 36 layers, and the configuration parameter range of the hidden layer size includes 1024, 2048 or 4096 units.
[0035] The embodiments of this application adopt a strategy of combining a set of fixed configuration items with a set of to-be-determined configuration items to explore the optimal configuration combination under a specific framework. The fixed configuration items adopt the specified configuration during the prediction process, providing a consistent training environment, while the to-be-determined configuration items allow for flexible adjustment of the to-be-determined configuration items by selecting specific configuration parameters to form a configuration parameter combination. It can be understood that: a set of fixed configuration items fixedly adopt the specified configuration, and in combination with a set of to-be-determined configuration items adopting one configuration parameter, a configuration parameter combination is obtained. For example, among multiple configuration items related to hardware conditions, model parameters, and parallel strategies, the hardware conditions are used as fixed configuration items, and their specified configurations such as GPU model, network bandwidth, etc. remain unchanged, while the model parameters (such as the number of layers, the number of hidden units) and the parallel strategy (such as the number of tensor parallel paths) are used as to-be-determined configuration items, each selecting one configuration parameter and combining it with the specified configuration of the hardware conditions to form a complete configuration parameter combination.
[0036] When training the model, the training speed and efficiency of the model under each configuration parameter combination, that is, the training performance, is the core of the evaluation. The training performance can be measured by calculating the total time consumed for model training, the utilization efficiency of computing resources (such as the usage rate of the GPU), and the progress of model training per unit time. By predicting the training performance corresponding to the model to be trained when a set of to-be-determined configuration items adopt multiple configuration parameters for training, the predicted training performance corresponding to multiple configuration parameters can be obtained. It can be understood that the predicted training performance corresponding to multiple configuration parameters refers to, before model training, according to different configuration parameter options of a set of to-be-determined configuration items, estimating the performance and efficiency of the model during training under each specific configuration parameter through a prediction algorithm.
[0037] Optionally, the terminal device selects different multiple configuration parameters from the configuration parameter ranges of each to-be-determined configuration item, takes a set of fixed configuration items that fixedly adopt the specified configuration and a set of to-be-determined configuration items that adopt one configuration parameter as a complete configuration parameter combination, predicts the total time consumed for model training corresponding to the model to be trained when training under each configuration parameter combination, and based on the total time consumed for model training corresponding to each configuration parameter combination, predicts the training performance corresponding to the model to be trained when training under each configuration parameter combination, and obtains the predicted training performance corresponding to multiple configuration parameters.
[0038] Step S206, select the target configuration parameter from the multiple configuration parameters; the target configuration parameter is the configuration parameter adopted by a set of to-be-determined configuration items when training the model to be trained.
[0039] Among them, in the preparation stage of model training, the fixed configuration items are usually determined by existing conditions and remain unchanged throughout the training process. The undetermined configuration items, on the other hand, need to dynamically select the optimal configuration parameters (i.e., the target configuration parameters) through prediction and evaluation according to the conditions of the fixed configuration items. By traversing multiple configuration parameters of the undetermined configuration items and combining them with the fixed configuration items, multiple combinations of configuration parameters are obtained, and the training performance of the model to be trained under each combination of configuration parameters is calculated. Thus, a set of optimal configuration parameters (i.e., the target configuration parameters) of the undetermined configuration items with the highest training performance is selected when a set of fixed configuration items adopts the specified configuration. This dynamic selectivity ensures that, under a given set of fixed configuration items, the computing resources can be utilized to the maximum extent, improving the efficiency and cost-effectiveness of model training.
[0040] For example, among multiple configuration items involving hardware conditions, model parameters, and parallel strategies, the hardware conditions are used as the fixed configuration items, and the model parameters and parallel strategies are used as the undetermined configuration items. By calculating the optimal configuration parameters of the model parameters and parallel strategies with the highest training performance when the hardware conditions adopt the specified configuration, it is possible to identify and select the combination of model parameters and parallel strategies that is most suitable for efficient model training under the given hardware conditions, ensuring that the selected model parameters and parallel strategies can seamlessly match the hardware conditions, thereby achieving the most efficient training effect.
[0041] For example, among multiple configuration items involving hardware conditions, model parameters, and parallel strategies, the model parameters are used as the fixed configuration items, and the hardware conditions and parallel strategies are used as the undetermined configuration items. By calculating the optimal configuration parameters of the hardware conditions and parallel strategies with the highest training performance when the model parameters adopt the specified configuration, it is possible to identify and select the combination of hardware conditions and parallel strategies that is most suitable for efficient model training under the given model parameters, ensuring that the selected hardware conditions and parallel strategies can seamlessly match the model parameters, thereby achieving the most efficient training effect.
[0042] Optionally, the terminal device uses one of the configuration parameters corresponding to the maximum predicted training performance among the predicted training performances corresponding to multiple configuration parameters as the target configuration parameter, and sets the target configuration parameter adopted by a set of undetermined configuration items when performing model training on the model to be trained.
[0043] Through the embodiments of the present application, the configuration items are divided into two categories: fixed configuration items and to-be-determined configuration items. This design allows, when the fixed configuration items are known, to focus on optimizing the selection of the to-be-determined configuration items to improve the training efficiency and reduce the cost. Under the constraint of adopting the specified configuration for the fixed configuration items, predict the prediction training performance of the model to be trained when different configuration parameters are adopted for the to-be-determined configuration items during training. Based on the prediction training performance corresponding to multiple configuration parameters, select the target configuration parameters from multiple configuration parameters of a group of to-be-determined configuration items. When the model to be trained performs model training, select the target configuration parameters for a group of to-be-determined configuration items, which can significantly reduce the cost required for model training, maximize the resource utilization efficiency, reduce resource waste, improve the training efficiency, and solve the problems of large hardware resource requirements and high costs in the large model training stage in the related art.
[0044] In an exemplary embodiment, Figure 3 is a structural diagram of a terminal device provided by the embodiments of the present application. As Figure 3 shown, taking multiple configuration items including the hardware conditions of the hardware device on which the model to be trained runs, the model parameters of the model to be trained, and the parallel strategy adopted during the training of the model to be trained as an example, the terminal device includes a model selection module. Among them, the model selection module includes a performance prediction module. The input of the performance prediction module includes model parameters and a parallel strategy, and is used to predict the training performance corresponding to the model to be trained when multiple configuration parameters are adopted for a group of to-be-determined configuration items during training, and output the prediction training performance corresponding to multiple configuration parameters. The model selection module selects the target configuration parameters from multiple configuration parameters according to the prediction training performance corresponding to multiple configuration parameters.
[0045] Among multiple configuration items related to hardware conditions, model parameters, and parallel strategies, take the configuration parameters among multiple configuration parameters of a group of to-be-determined configuration items as the current configuration parameters and perform the following prediction operations to obtain the prediction training performance corresponding to multiple configuration parameters:
[0046] First, according to the pre-constructed graphics card performance database, the specified configuration, and the current configuration parameters, calculate the graphics card computing performance corresponding to the model to be trained when multiple configuration parameters are adopted for a group of to-be-determined configuration items during training, and obtain the graphics card computing performance corresponding to the current configuration parameters. The graphics card performance database stores the graphics card computing performance corresponding to the model to be trained when different hardware conditions, different model parameters, and different parallel strategies are adopted during training.
[0047] Among them, the graphics card performance database is used to store and manage the actual computing performance data of the graphics card (i.e., the graphics card computing performance) during the training of the model to be trained under different hardware conditions, model parameters, and parallel strategy combinations. For matrices of different shapes and sizes, the single-card computing performance may be somewhat different from the theoretical performance. Therefore, in this embodiment, a large number of real data of different sizes are tested based on various servers, and relevant fittings are performed on the untested sizes, establishing a complete graphics card performance database to ensure the authenticity of the computing efficiency as much as possible. In the untested models, tests can be carried out through test cases, and relevant data can be imported to supplement the relevant data in the graphics card performance database. Thus, through the previous actual measurements and data analysis, the graphics card performance database covers a wide range of hardware configurations and strategy combinations, providing reliable data support for performance prediction during model training.
[0048] In some embodiments, the construction process of the graphics card performance database is as follows: collect and organize the corresponding graphics card computing performance after adopting different model structures and different parallel strategies under different hardware conditions, and perform data fitting to obtain the graphics card performance database.
[0049] Optionally, the terminal device searches in the pre-constructed graphics card performance database using the specified configuration and the current configuration parameters to obtain the graphics card computing performance corresponding to the model to be trained when training with a set of undetermined configuration items using the current configuration parameters.
[0050] Second, predict the current actual training performance of the model to be trained according to the graphics card computing performance corresponding to the current configuration parameters.
[0051] Among them, the current actual training performance refers to the real computing efficiency and resource utilization rate shown by the model to be trained during the training process under specific hardware conditions and parallel strategy configurations.
[0052] Optionally, the terminal device calculates the model training time required for the model to be trained when training with a set of undetermined configuration items using the current configuration parameters according to the graphics card computing performance corresponding to the current configuration parameters, and determines the ratio between the model training time and the total model calculation amount of the model to be trained as the current actual training performance of the model to be trained under the current configuration parameters.
[0053] Among them, the training time of the model refers to the total time from the start of model training until the training is completed, including the pure calculation time, communication time, and idle time during training. The pure calculation time during training refers to the time for the model to perform direct calculation tasks (such as forward propagation, backward propagation, and parameter update) during training. The communication time refers to the time required for data exchange and synchronization between computing nodes during parallel training of the model. The idle time refers to the time during model training when the graphics card or computing node is not fully utilized due to mismatched parallel strategies or improper hardware resource allocation.
[0054] The total model computation is the sum of all computational operations when the model completes a full training cycle (including forward propagation and backward propagation). The method for predicting the total model computation includes the following steps: First, evaluate the inherent parameter scale of the model by analyzing the model architecture, such as the number of layers, the number of hidden units, etc. Second, quantify the computational requirements of each training step in combination with parameters such as the micro-batch size, sequence length, and gradient accumulation times set during training. Then, consider the impact of different parallel strategies (such as tensor, data, or pipeline parallelism) on the computational distribution and the resulting additional communication costs. Finally, based on the above-mentioned inherent parameter scale of the model itself, the computational requirements of each training step, and the additional communication costs, the total model computation can be obtained.
[0055] For example, Figure 4 is a schematic diagram for the performance prediction module provided by the embodiment of the present application to predict the predicted training performance corresponding to multiple configuration parameters. As Figure 4 shown, among multiple configuration items related to hardware conditions, model parameters, and parallel strategies, the hardware condition is used as a fixed configuration item, and the model parameters and parallel strategies are used as undetermined configuration items. Different configuration parameters are selected from the configuration parameter range of the selected model parameters, and different configuration parameters are selected from the configuration parameter range of the selected parallel strategy, and combined with the fixed hardware conditions to generate multiple configuration parameter combinations. Different configuration parameter combinations enter the performance prediction module in sequence, query the graphics card performance database through the performance prediction module to obtain the graphics card computing performance corresponding to the current configuration parameter combination, and calculate the pure calculation time, communication time, and idle time required during training when the model to be trained is trained with a set of current configuration parameters of the undetermined configuration items under the specified configuration. The pure calculation time, communication time, and idle time during training are superimposed to obtain the model training time required when the model to be trained is trained with the current configuration parameters. The ratio between the model training time and the total model computation of the model to be trained is determined as the current actual training performance of the model to be trained under the current configuration parameters.
[0056] Third, determine the predicted training performance corresponding to the current configuration parameters according to the ratio between the current actual training performance and the theoretical training performance corresponding to the current configuration parameters.
[0057] Among them, the theoretical training performance corresponding to the current configuration parameters refers to the maximum computing efficiency that the graphics card can achieve when a set of fixed configuration items adopt the specified configuration and a set of undetermined configuration items adopt the current configuration parameters. In addition to storing the graphics card computing performance corresponding to the model to be trained when training with different model parameters and different parallel strategies under different hardware conditions, the graphics card performance database also stores the graphics card peak performance (i.e., the maximum computing efficiency) that the graphics card can achieve when training with different model parameters and different parallel strategies under different hardware conditions. In this embodiment, the graphics card peak performance is used as the theoretical training performance corresponding to the given hardware conditions, given model parameters, and given parallel strategies. Therefore, the theoretical training performance corresponding to the current configuration parameters can be obtained by looking up in the graphics card performance database using the specified configuration and the current configuration parameters.
[0058] Through this embodiment, by performing single-card actual measurements on matrices of different shapes, collecting and recording the actual computing performance of the graphics card when processing datasets of various sizes and dimensions, a rich and accurate graphics card performance database is established. When evaluating the training performance corresponding to the model to be trained when a set of undetermined configuration items adopt multiple configuration parameters, the graphics card computing performance in the graphics card performance database is used to predict the current actual training performance of the model to be trained, significantly improving the accuracy of the predicted training performance corresponding to multiple configuration parameters. The training performance prediction based on actual measurement data not only ensures the efficient use of hardware resources, reduces resource waste, but also reduces the cost of model training. In addition, according to the characteristics of the hardware environment, the graphics card performance database can be supplemented through test cases, making the data in the graphics card performance database closer to the actual usage scenario, thereby improving the pertinence and accuracy when predicting the model training performance, and enhancing the flexibility and application scenarios of the graphics card performance database.
[0059] In an exemplary embodiment, in large model training, as the complexity and scale of the model increase, the demand for video memory also increases. There is a problem that video memory overflows during the model training process, which seriously affects the efficiency and reliability of the model training, and even leads to training interruption or failure. Therefore, this embodiment predicts the maximum total video memory occupancy corresponding to the model to be trained when a set of undetermined configuration items adopt multiple configuration parameters, and eliminates a set of configuration parameters with a maximum total video memory occupancy greater than the maximum video memory of a single card from the multiple configuration parameters of the set of undetermined configuration items, ensuring that the model training does not exceed the limits of available hardware resources, avoiding the risk of video memory overflow, solving the problem of the impact of model scale and parallel strategy on video memory occupancy, and ensuring that the model training is carried out within reasonable resource limits.
[0060] In some embodiments, according to the predicted training performance corresponding to multiple configuration parameters, selecting target configuration parameters from the multiple configuration parameters includes:
[0061] Predict the total maximum video memory occupancy corresponding to a set of to-be-determined configuration items when the model to be trained is trained with multiple configuration parameters, and obtain the total maximum video memory occupancy corresponding to the multiple configuration parameters; exclude from the multiple configuration parameters a set of configuration parameters whose corresponding total maximum video memory occupancy is greater than the maximum video memory of a single card; determine, as the target configuration parameter, a configuration parameter corresponding to the maximum predicted training performance among the multiple configuration parameters after exclusion.
[0062] Among them, the total video memory occupancy refers to the total space in the video memory used to store all data and states related to model training during model training. It includes, but is not limited to, the storage of model parameters, temporary data required for the execution of the training process (such as intermediate calculation results and gradient information), and the video memory space occupied by the state information of the optimizer when updating model parameters. The total maximum video memory occupancy refers to the peak value of the video memory occupancy in a single training step.
[0063] The maximum video memory of a single card refers to the maximum video memory capacity that the graphics card can provide, which directly determines the upper limit of the video memory resources that can be allocated to the model during model training. When evaluating the model training configuration, if the video memory occupancy of the model exceeds the maximum video memory of a single card, it will cause the training process to be unable to proceed or the efficiency to drop significantly.
[0064] Optionally, as Figure 3 shown, the model selection module further includes a video memory prediction module. The terminal device uses the configuration parameters among the multiple configuration parameters as the current configuration parameters. When the model to be trained is trained with the current configuration parameters for a set of to-be-determined configuration items, the input of the video memory prediction module includes model parameters and a parallel strategy. The video memory prediction module predicts the video memory occupancy of each graphics card under the current hardware conditions, obtains the total video memory occupancy of each graphics card at the current moment, then performs peak analysis on the total video memory occupancy at different moments, obtains the total maximum video memory occupancy corresponding to the current configuration parameters, and outputs the total maximum video memory occupancy of each graphics card under the given model parameters and parallel strategy. The terminal device excludes from the multiple configuration parameters a set of configuration parameters whose corresponding total maximum video memory occupancy is greater than the maximum video memory of a single card through the model selection module, and determines, as the target configuration parameter, a configuration parameter corresponding to the maximum predicted training performance among the multiple configuration parameters after exclusion.
[0065] Through this embodiment, configuration parameters whose total maximum video memory occupancy exceeds the maximum video memory limit of a single card are automatically excluded from multiple configuration parameters under a set of to-be-determined configuration items, thus avoiding training failures caused by video memory overflow. This screening mechanism solves the problem that the model configuration exceeds the existing computing resource limit, ensures that all the remaining configuration parameters are within the hardware capabilities, and improves the feasibility and efficiency of model training.
[0066] In an exemplary embodiment, in a scenario where given hardware resources are available and it is necessary to determine the model parameters and parallel strategy that are most suitable for training on the given hardware resources, as Figure 3 shown, a set of fixed configuration items is set to include the hardware conditions of the hardware device on which the model to be trained runs, and a set of to-be-determined configuration items includes the model parameters of the model to be trained and the parallel strategy adopted during the training of the model to be trained. In this way, the target configuration parameters of the set of to-be-determined configuration items include the first target configuration parameters of the model parameters and the second target configuration parameters of the parallel strategy.
[0067] In some embodiments, determining the configuration parameter corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the target configuration parameter includes:
[0068] Determining the configuration parameter of the model parameters corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the first target configuration parameter; determining the configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the second target configuration parameter.
[0069] Optionally, as Figure 3 shown, the terminal device selects different configuration parameters from the configured parameter ranges of the selected model parameters, selects different configuration parameters from the configured parameter ranges of the selected parallel strategies, and combines the fixed hardware conditions to generate multiple configuration parameter combinations; different configuration parameter combinations sequentially enter the performance prediction module to obtain the predicted training performance corresponding to each configuration parameter combination; different configuration parameter combinations sequentially enter the video memory prediction module to obtain the maximum total video memory occupancy corresponding to each configuration parameter combination, eliminate the configuration parameter combinations whose maximum total video memory occupancy exceeds the single-card maximum video memory limit from the multiple configuration parameter combinations, determine the model parameters corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the first target configuration parameter, that is, the model structure configuration most suitable for the current hardware resources, and determine the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the second target configuration parameter, that is, the parallel strategy configuration that can most optimize the hardware computing efficiency.
[0070] For example, among multiple configuration items related to hardware conditions, model parameters, and parallel strategies, the hardware conditions are fixed configuration items, while the model parameters and parallel strategies are to-be-determined configuration items. Among them, the hardware conditions of the hardware device include 8 graphics cards of model A, NvLink network bandwidth, etc. The model parameters include, but are not limited to, the number of model layers, the number of neurons in the hidden layer, the vocabulary size, etc.; the configuration parameter range of the model parameters includes: the number of model layers: {6, 12, 24, 36}, the number of neurons in the hidden layer: {1024, 2048, 4096, 8192}; the parallel strategies include the number of tensor parallelism, the number of pipeline parallelism, the number of data parallelism, etc.; the configuration parameter range of the parallel strategies includes: the number of tensor parallelism: {1, 2, 4}, the number of pipeline parallelism: {1, 2, 4}, the number of data parallelism: {1, 2, 4}.
[0071] Perform video memory prediction for each of the above configuration parameter combinations (the combination of hardware conditions, model parameters, and parallel strategies) to obtain the corresponding video memory occupancy. For example, when the number of model layers is 36, the number of neurons in the hidden layer is 8192, the number of tensor parallelism is 4, the number of pipeline parallelism is 4, and the number of data parallelism is 1, the predicted video memory occupancy is 32GB / card (assuming the occupancy of all GPUs is the same). Eliminate those configuration parameter combinations whose video memory requirements exceed 40GB to ensure that the training can be carried out on the existing hardware. Perform performance prediction on the remaining combinations, and calculate the model training time consumption, computing efficiency, etc. For example, after video memory screening, the following several configurations are retained:
[0072] a) The number of model layers = 24, the hidden layer = 4096, the tensor = 2, the pipeline = 2, the data = 2;
[0073] b) The number of model layers = 12, the hidden layer = 8192, the tensor = 4, the pipeline = 4, the data = 1;
[0074] c) The number of model layers = 6, the hidden layer = 4096, the tensor = 1, the pipeline = 8, the data = 1.
[0075] Through calculation, obtain the predicted training performance under each configuration, and identify the combination with the highest predicted training performance from the remaining configurations. For example, if the predicted training performance of combination b) is the highest, then the model parameters in combination b) (the number of model layers = 12, the hidden layer = 8192) are determined as the first target configuration parameters, and the parallel strategy (tensor = 4, pipeline = 4, data = 1) is determined as the second target configuration parameters.
[0076] Through this embodiment, the model parameters and parallel strategies are flexibly adjusted under fixed hardware conditions, enhancing the scenario adaptability of model training; the highest computing efficiency is obtained by traversing the video memory evaluation and efficiency calculation, and the optimal model parameters (i.e., the first target configuration parameters) and the optimal parallel strategy (i.e., the second target configuration parameters) that can maximize the computing efficiency of the given hardware conditions are selected, which can select the most cost-effective model structure and parallel strategy under the given hardware conditions, thereby providing a reference solution for model training, accelerating the training speed, improving the model performance, and saving the training cost.
[0077] In an exemplary embodiment, in a scenario where given model parameters, the hardware conditions and parallel strategies suitable for training under the given model parameters need to be determined, such as Figure 3 shown, a set of fixed configuration items is set to include the model parameters of the model to be trained; a set of undetermined configuration items includes the hardware conditions of the hardware device on which the model to be trained runs, and the parallel strategy adopted during the training of the model to be trained. In this way, the target configuration parameters include the second target configuration parameter of the parallel strategy and the third target configuration parameter of the hardware conditions.
[0078] In some embodiments, the configuration parameter corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the target configuration parameter, including:
[0079] The configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the second target configuration parameter; the configuration parameter of the hardware conditions corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the third target configuration parameter.
[0080] Optionally, as Figure 3 shown, the terminal device selects different configuration parameters from the range of configuration parameters of the selected hardware conditions through the model selection module, selects different configuration parameters from the range of configuration parameters of the selected parallel strategy, and combines the fixed model parameters to generate multiple configuration parameter combinations; different configuration parameter combinations sequentially enter the performance prediction module to obtain the predicted training performance corresponding to each configuration parameter combination; different configuration parameter combinations sequentially enter the video memory prediction module to obtain the maximum total video memory occupancy corresponding to each configuration parameter combination, and the configuration parameter combinations with the maximum total video memory occupancy exceeding the single-card maximum video memory limit are eliminated from the multiple configuration parameter combinations. The parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the second target configuration parameter, that is, the parallel strategy configuration that can optimize the hardware computing efficiency the most, and the hardware conditions corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination are determined as the third target configuration parameter, that is, the hardware resource configuration most suitable for the current model structure.
[0081] For example, among multiple configuration items related to hardware conditions, model parameters, and parallel strategies, the model parameters are fixed configuration items, and the hardware conditions and parallel strategies are to-be-determined configuration items. Among them, the model parameters include that the number of model layers is 36, the number of neurons in the hidden layer is 4096, and the vocabulary size is 50,000.
[0082] The hardware conditions include but are not limited to the graphics card model, network bandwidth, etc.; the configuration parameter range of the hardware conditions includes: graphics card model: {Model A, Model B, Model C}, network bandwidth: {100 Gbps, 200 Gbps}. The parallel strategies include but are not limited to the number of tensor parallelism, pipelining parallelism, data parallelism paths, etc.; the configuration parameter range of the parallel strategies includes: number of tensor parallelism: {1, 2, 4}, number of pipelining parallelism: {1, 2, 4}, number of data parallelism: {1, 2, 4}.
[0083] Perform video memory prediction and training performance prediction for each of the above configuration parameter combinations (combinations of hardware conditions, model parameters, and parallel strategies). For example, for the case where the model parameters remain unchanged, using 8 graphics cards of Model A, the network bandwidth is 200 Gbps, the number of tensor parallelism is 4, the number of pipelining parallelism is 2, and the number of data parallelism is 1, it is predicted that the video memory occupancy is within the hardware capabilities, and the training performance prediction shows that it can efficiently complete the training task. Eliminate those hardware conditions that will cause the video memory to exceed the maximum video memory capacity of the hardware even under the given model parameters. For example, if using graphics cards of Model B, even under the smallest parallel strategy, the video memory occupancy may exceed its maximum video memory capacity, so this part of the hardware conditions will be eliminated. From the remaining combinations of hardware conditions and parallel strategies, identify the combination with the highest predicted training performance. For example, in combination a): when using 8 graphics cards of Model A, the network bandwidth is 200 Gbps, the number of tensor parallelism is 4, the number of pipelining parallelism is 2, and the number of data parallelism is 1, the predicted training performance is the best. Therefore, the parallel strategy (tensor = 4, pipelining = 2, data = 1) in combination a) is determined as the second target configuration parameter, and the hardware conditions (8 graphics cards of Model A, network bandwidth 200 Gbps) in combination a) are determined as the third target configuration parameter.
[0084] Through this embodiment, when the model parameters are known, by predicting the model training efficiency under different hardware conditions and parallel strategies, find the combination that can maximize the training performance from multiple hardware conditions and multiple parallel strategies, so as to obtain the hardware conditions and parallel strategies that maximize the model training performance when the model parameters are fixed, improve the training efficiency, reduce the cost, ensure the effective utilization of resources, avoid the problem of low training efficiency caused by insufficient hardware resources or improper parallel strategies, and solve the computing power resource and cost problems in the related technology.
[0085] In an exemplary embodiment, in the related art, the method for evaluating the video memory occupancy during model training is relatively rough. The rough results obtained cannot fully utilize the memory of the computing resources, which affects the judgment of the optimal configuration parameters and causes waste of computing resources. Therefore, to solve the above problems, in this embodiment, the video memory occupancy is subdivided into three parts: model parameter allocation, optimizer data, and temporary activation. Through accurate video memory prediction, training failures or resource waste caused by video memory overflow are avoided, ensuring the stability and efficiency of model training, and solving the problem of rough evaluation of video memory occupancy during the above-mentioned model training.
[0086] In this embodiment, the maximum total video memory occupancy corresponding to a plurality of configuration parameters when the model to be trained adopts various configuration parameters in a set of pending configuration items is predicted, and the maximum total video memory occupancy corresponding to the plurality of configuration parameters is obtained, including: using the configuration parameter in the plurality of configuration parameters as the current configuration parameter, and performing the following prediction operations to obtain the maximum total video memory occupancy corresponding to the plurality of configuration parameters:
[0087] 1. According to the specific number of parameters of the model to be trained, determine the first video memory occupancy required for the storage parameter amounts corresponding to multiple video cards in the hardware device on which the model to be trained runs when the model to be trained adopts the current configuration parameter in a set of pending configuration items.
[0088] Among them, the specific number of parameters of the model to be trained refers to the quantization indexes of each component constituting the deep learning model. For example, the specific number of parameters of the model to be trained includes the number of model layers, the number of neurons in the hidden layer, the vocabulary size, etc. The specific number of parameters of the model to be trained determines the overall scale and complexity of the model. During the training process, the specific number of parameters of the model to be trained needs to be allocated among multiple video cards participating in the training. If the tensor parallel strategy is adopted, the specific number of parameters of the model to be trained will be divided into multiple parts, and each video card stores a part of them; if the data parallel strategy is adopted, a copy of the entire model parameters will be stored on each video card, but the data set will be divided; pipelining parallelism involves the division of model layers. It can be seen that the first video memory occupancy refers to the total video memory space required for multiple video cards to store the parameters of the model to be trained when the current configuration parameter is adopted in a set of pending configuration items. The calculation of the first video memory occupancy takes into account the influence of the parallel strategy on parameter allocation, ensuring that each video card can correctly store the model parameters allocated to it during the parallel training process and avoiding the risk of video memory overflow.
[0089] Optionally, Figure 5 is a schematic diagram of the video memory prediction module provided in the embodiment of the present application for predicting the maximum total video memory occupancy corresponding to a plurality of configuration parameters, as Figure 5As shown, the terminal device directly obtains the specific parameter quantity information of the model to be trained through the model architecture definition, or extracts it through the model structure analysis tool, including but not limited to the number of model layers, the number of neurons in the hidden layer, the vocabulary size, etc. The terminal device allocates each graphics card parameter through the video memory prediction module, specifically including: determining the parallel strategy according to the current configuration parameters, and calculating the model parameter quantity allocated to the storage of each graphics card according to the selected parallel strategy and the specific parameter quantity of the model to be trained. The terminal device calculates the first video memory occupancy of each graphics card through the video memory prediction module, specifically including: according to the model parameter quantity allocated to each graphics card, combined with the hardware characteristics of the graphics card (such as video memory size, computing unit ability, etc.), calculating the first video memory occupancy required for the storage parameter quantity corresponding to multiple graphics cards in the hardware device where the model to be trained runs when the model to be trained adopts the current configuration parameters in a set of undetermined configuration items.
[0090] Second, determine the second video memory occupancy required for multiple graphics cards to store the optimizer state when the model to be trained adopts the current configuration parameters in a set of undetermined configuration items according to the preset optimizer data of the model to be trained;
[0091] Among them, the preset optimizer data of the model to be trained refers to the data and settings used or generated by the optimizer to update the model parameters during the model training process. The optimizer is a crucial component in the training of deep learning models. Its task is to adjust the model parameters according to the gradient of the loss function after each iteration to minimize the loss function and gradually improve the model performance. It can be understood that the preset optimizer data, that is, the optimization method, includes but not limited to whether to store some data on the CPU (Central Processing Unit), and the selection of data format (such as storing with 16-bit or 8-bit precision).
[0092] The second video memory occupancy refers to the total video memory space required for multiple graphics cards to store the optimizer state when the model to be trained adopts the current configuration parameters in a set of undetermined configuration items. The selection of the optimization method has a direct impact on the video memory occupancy. Reasonable optimization can reduce the video memory demand and improve the computing efficiency. The second video memory occupancy will be calculated according to the optimizer state and the optimization method to ensure that the optimization steps during the training process will not cause insufficient video memory.
[0093] Optionally, such as Figure 5As shown in the figure, the terminal device calculates the second video memory occupancy of each graphics card through the video memory prediction module, which specifically includes: The video memory prediction module determines the type of optimizer used (such as Adam (Adaptive Moment Estimation), SGD (Stochastic Gradient Descent), etc.) and its specific parameters according to the preset optimizer data of the model to be trained, and calculates the video memory required for the optimizer state based on the type and parameters of the optimizer. For example, the Adam optimizer needs to store the first moment (average gradient) and the second moment (average of the squared gradients) of the model parameters, and these state quantities are usually twice the number of model parameters. The terminal device allocates the video memory required for the optimizer state according to the parallel strategy in the current configuration parameters, and obtains the second video memory occupancy required for multiple graphics cards to store the optimizer state when the model to be trained adopts the current configuration parameters in a set of undetermined configuration items.
[0094] Third, analyze the peak video memory occupancy corresponding to the model to be trained when adopting the current configuration parameters in a set of undetermined configuration items, and obtain the third video memory occupancy required for multiple graphics cards to store the intermediate results generated by the temporary activation when the model to be trained adopts the current configuration parameters in a set of undetermined configuration items;
[0095] Among them, the third video memory occupancy refers to the peak amount of video memory space required for multiple graphics cards to store the temporary activation (i.e., intermediate calculation results) when adopting the current configuration parameters in a set of undetermined configuration items. The video memory occupancy of the temporary activation changes with the training process, and the peak appears during the logits (unnormalized probability) calculation or the reverse propagation recalculation. By analyzing the video memory occupancy at these peak moments, it can be ensured that the hardware resources can meet the requirements during the most intensive calculation stage. Among them, Logits refers to the output of the last layer of the neural network, and these outputs are the original values without being transformed by the activation function. The Logits calculation is a common step in the classification task, used to obtain the original scores of each category from the model, and these scores will subsequently be used to calculate the probability distribution for prediction or decision-making. For example, in the text generation task, the model will output the Logits values corresponding to each vocabulary, and after being transformed by the activation function, the probability distribution of the vocabulary is obtained, thereby determining the generation of the next vocabulary.
[0096] Optionally, such as Figure 5As shown in the figure, the terminal device calculates the third video memory occupancy of each graphics card through the video memory prediction module, which specifically includes: the terminal device calculates the output data volume of each activation layer of the model to be trained based on the first video memory occupancy required for storing the corresponding storage parameter quantities of multiple graphics cards and the model parameters of the model to be trained, calculates the activation data volume actually required to be stored on each graphics card according to the parallel strategy, analyzes the changes in video memory occupancy during the forward propagation and backward propagation processes, finds the moment of peak video memory occupancy on each graphics card, and uses the peak video memory occupancy corresponding to the moment of peak video memory occupancy on each graphics card as the third video memory occupancy required for storing the intermediate results generated by the temporary activation of each graphics card when the model to be trained adopts the current configuration parameters in a set of to-be-determined configuration items.
[0097] IV. Determine the total video memory occupancy corresponding to multiple graphics cards when the model to be trained adopts the current configuration parameters in a set of to-be-determined configuration items according to the first video memory occupancy, the second video memory occupancy, and the third video memory occupancy;
[0098] Among them, the total video memory occupancy refers to the sum of the first video memory occupancy, the second video memory occupancy, and the third video memory occupancy of each graphics card, which represents the total video memory occupancy of multiple graphics cards during the model training process when the current configuration parameters are adopted in a set of to-be-determined configuration items. The calculation of the total video memory occupancy provides a comprehensive perspective for the resource planning of model training, ensuring that the resource allocation of all graphics cards not only meets the storage requirements of model parameters and optimizer states but also covers the maximum activation video memory occupancy that may occur during the calculation process.
[0099] Optionally, for each graphics card, the terminal device adds the first video memory occupancy, the second video memory occupancy, and the third video memory occupancy of each graphics card to obtain the total video memory occupancy of each graphics card.
[0100] V. Determine the maximum total video memory occupancy among the total video memory occupancies of multiple graphics cards as the maximum total video memory occupancy corresponding to the current configuration parameters.
[0101] Among them, the maximum total video memory occupancy refers to the video memory space occupied by the graphics card with the highest total video memory occupancy among multiple graphics cards under a set of given configuration parameters. The maximum total video memory occupancy is the strictest requirement for video memory prediction, ensuring that even on the graphics card with the largest video memory occupancy, the training requirements can be met, avoiding the overall training failure caused by the video memory limit of a single card. By calculating the maximum total video memory occupancy, it can be ensured that the model training does not exceed the limit of the hardware video memory, and at the same time, it provides a basis for the upgrade or reasonable allocation of hardware resources.
[0102] Optionally, as Figure 5As shown, the terminal device analyzes the maximum video memory of each graphics card through the video memory prediction module, specifically including: comparing the total video memory occupation of multiple graphics cards, and determining the largest total video memory occupation as the maximum total video memory occupation corresponding to the current configuration parameters.
[0103] For example, the model parameters of the model to be trained include: the number of hidden layers is 128, and the number of model layers is 24. The parallel strategy of the model to be trained includes: the number of tensor parallel paths is 2, the number of pipeline parallel paths is 4, and the number of data parallel paths is 8. The preset optimizer data includes: using the Adam optimizer, and each parameter requires additional storage space to save the optimization state.
[0104] The specific number of parameters of the model to be trained is 10GB (this is a rough assumption, and the actual number of parameters is determined by the specific architecture of the model). Under 2-way tensor parallelism, each graphics card will store 5GB of model parameters. Using the Adam optimizer, the preset optimizer data requires additional storage of first-order and second-order moment estimates, which means that the storage space of the preset optimizer data will actually increase to 15GB (5GB * 3). In addition, assume that the temporary activation of each layer occupies 1GB of video memory. Considering 4-way pipeline parallelism, this means that each graphics card will generate 4GB of temporary activation video memory occupation when processing data. At a certain moment during the training process, assume that the parameter storage on graphics card M reaches the peak, that is, 15GB (because of data parallelism, each GPU needs to store the complete parameters and optimizer state), plus the 4GB of temporary activation video memory processed by graphics card M, the total video memory occupation reaches 19GB. For other graphics cards, due to different activation data or parameters processed, assume that their maximum video memory occupations are 17GB, 16GB, 18GB, 17GB, 16GB, 18GB, 17GB respectively. Then the final output maximum video memory occupation is 19GB, and the corresponding graphics card number is graphics card M.
[0105] Through this embodiment, based on the specific parameter amount of the model and the current configuration parameters, the storage parameter amount allocated on each graphics card is calculated to obtain the first video memory occupancy of each graphics card, and the video memory occupancy prediction of the parameters is refined; according to the preset optimizer data, combined with the first video memory occupancy under the current configuration parameters, the video memory space required by the optimizer state on each graphics card, that is, the second video memory occupancy, is calculated to ensure accurate prediction of the video memory occupancy of the optimizer state; the video memory occupancy peak value of the model during the training process under the current configuration parameters is refined and analyzed, and the video memory occupancy of the temporary activation data on each graphics card is calculated in detail, and determined as the third video memory occupancy, providing a more refined video memory occupancy prediction; The first, second and third video memory occupancy are aggregated to obtain the total video memory occupancy of each graphics card, providing a comprehensive video memory occupancy estimate. By subdividing the video memory occupancy into three parts: model parameter allocation, optimizer data and temporary activation, the problem of rough assessment of video memory occupancy during model training is solved. Furthermore, the maximum video memory occupancy is identified in the total video memory occupancy of each graphics card, and it is determined as the corresponding maximum total video memory occupancy under the current configuration parameters. By identifying the highest peak value of video memory occupancy, it can ensure that the memory configuration of the hardware device can meet the most demanding video memory requirements, avoid training failure or resource waste caused by video memory overflow, and ensure the stability of training and full utilization of resources.
[0106] In an exemplary embodiment, the activation data (i.e., the output of the network layer, the intermediate result in the forward propagation and backward propagation process) has a great impact on the video memory occupancy. However, the related art lacks a detailed analysis of the activation during the training process. Since the demand for video memory of the activation data is not accurately evaluated, improper hardware configuration selection may be caused. For example, if the video memory capacity of the hardware is less than the peak video memory occupancy (i.e., the third video memory occupancy), the model training process may fail due to video memory overflow, or frequent video memory paging operations may be required, which seriously affects the training efficiency. On the other hand, if the hardware configuration is much larger than the actual requirement, although the video memory overflow problem can be avoided, it will cause unnecessary waste of computing power and cost. Therefore, to solve this problem, this embodiment calculates the first video memory occupancy peak of multiple graphics cards during the back propagation recalculation process, and calculates the second video memory occupancy peak of multiple graphics cards during the fully connected calculation process, so as to refine the video memory analysis, more accurately evaluate the video memory demand during the model training process, avoid the video memory overflow problem, and avoid resource waste.
[0107] In this embodiment, the peak value of the video memory usage corresponding to the model to be trained when the current configuration parameters are adopted in a set of pending configuration items is analyzed to obtain the third video memory usage required for storing the intermediate results generated by temporary activation of multiple graphics cards when the model to be trained adopts the current configuration parameters, including:
[0108] When the model to be trained adopts the current configuration parameters for a set of undetermined configuration items, calculate the first peak video memory occupancy of multiple graphics cards during the reverse propagation recomputation process, and calculate the second peak video memory occupancy of multiple graphics cards during the fully connected calculation process; determine the maximum value among the first peak video memory occupancy and the second peak video memory occupancy of multiple graphics cards as the third video memory occupancy required for multiple graphics cards to store the intermediate results generated by the temporary activation.
[0109] Among them, according to the analysis, the peak of the activated video memory may appear during logits calculation or reverse propagation recomputation. Therefore, it is necessary to discuss the two situations separately. Among them, logits calculation usually occurs in the fully connected layer (also known as the dense layer or linear layer) of the neural network, especially in the model of the classification task. In the architecture of the deep learning network, the fully connected layer is usually located at the end of the network, and its function is to map the feature maps generated by the previous layer to the classification space, generating the original scores corresponding to each category, that is, the so-called logits.
[0110] Optionally, when the terminal device adopts the current configuration parameters for the model to be trained for a set of undetermined configuration items, the calculation of the first peak video memory occupancy of each graphics card during the reverse propagation recomputation process can be calculated using the following formula (1), and the calculation of the second peak video memory occupancy of each graphics card during the fully connected calculation process can be calculated using the following formula (2):
[0111]
[0112] Among them, M 3-1 is the first peak video memory occupancy corresponding to the reverse propagation recomputation; M 3-2 is the second peak video memory occupancy corresponding to the logits calculation, B is the micro-batch size, S is the sequence length, H is the number of hidden layers, V is the vocabulary size, A is the number of gradient accumulation times, K is the KV chanel (key-value channel) number, L is the total number of model layers, T p is the number of tensor parallel paths, P p is the number of pipeline parallel paths.
[0113] The terminal device determines the maximum value among the first peak video memory occupancy and the second peak video memory occupancy of each graphics card as the third video memory occupancy required for each graphics card to store the intermediate results generated by the temporary activation. The corresponding mathematical expression is shown in formula (3):
[0114] M3 = max (M 3-1 , M 3-2 ) (3)
[0115] Among them, M3 is the third video memory occupancy required for each graphics card to store the intermediate results generated by the temporary activation.
[0116] Through this embodiment, the first peak video memory occupancy during the refined prediction backpropagation recomputation process and the second peak video memory occupancy during the fully connected computation process are predicted. The maximum value of the first peak video memory occupancy and the second peak video memory occupancy is determined as the third video memory occupancy required for multiple graphics cards to store temporary activation data (i.e., intermediate results). This means that the highest demand for video memory during the model training process will be used as the standard for hardware configuration selection, which can more accurately evaluate the actual demand for video memory during model training, avoid problems of insufficient or excessive hardware configuration, as well as avoid problems of resource waste and high costs.
[0117] From the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation manner.
[0118] The embodiments of the present application also provide a device for determining model training configuration parameters, as Figure 6 shown. The device includes:
[0119] A configuration item acquisition module 602, configured to acquire a plurality of configuration items required for a model to be trained during training; the plurality of configuration items include a set of fixed configuration items and a set of undetermined configuration items;
[0120] A training performance prediction module 604, configured to predict the training performance corresponding to the model to be trained when a set of undetermined configuration items adopt multiple configuration parameters during training, when a set of fixed configuration items are fixedly adopted with specified configurations, and obtain predicted training performances corresponding to the multiple configuration parameters;
[0121] A configuration parameter determination module 606, configured to select a target configuration parameter from the multiple configuration parameters according to the predicted training performances corresponding to the multiple configuration parameters; the target configuration parameter is the configuration parameter adopted by a set of undetermined configuration items when performing model training on the model to be trained.
[0122] In an exemplary embodiment, the multiple configuration items include the hardware conditions of the hardware device on which the model to be trained runs, the model parameters of the model to be trained, and the parallel strategy adopted during the training of the model to be trained; the training performance prediction module 604 is further configured to use the configuration parameters among the multiple configuration parameters as the current configuration parameters to perform the following prediction operations to obtain the predicted training performance corresponding to the multiple configuration parameters: calculate the graphics card computing performance corresponding to the model to be trained when trained with the current configuration parameters in a set of to-be-determined configuration items according to the pre-constructed graphics card performance database, the specified configuration, and the current configuration parameters, so as to obtain the graphics card computing performance corresponding to the current configuration parameters; the graphics card performance database stores the graphics card computing performance corresponding to the model to be trained when trained with different model parameters and different parallel strategies under different hardware conditions; predict the current actual training performance of the model to be trained according to the graphics card computing performance corresponding to the current configuration parameters; determine the predicted training performance corresponding to the current configuration parameters according to the ratio between the current actual training performance and the theoretical training performance corresponding to the current configuration parameters.
[0123] In an exemplary embodiment, the configuration parameter determination module 606 is further configured to predict the total maximum video memory occupancy corresponding to the model to be trained when trained with multiple configuration parameters in a set of to-be-determined configuration items, so as to obtain the total maximum video memory occupancy corresponding to the multiple configuration parameters; eliminate a set of configuration parameters among the multiple configuration parameters whose corresponding total maximum video memory occupancy is greater than the maximum video memory of a single card; determine the configuration parameter corresponding to the maximum predicted training performance among the remaining multiple configuration parameters as the target configuration parameter.
[0124] In an exemplary embodiment, a set of fixed configuration items includes the hardware conditions of the hardware device on which the model to be trained runs; a set of to-be-determined configuration items includes the model parameters of the model to be trained and the parallel strategy adopted during the training of the model to be trained; the target configuration parameters include the first target configuration parameter of the model parameters and the second target configuration parameter of the parallel strategy; the configuration parameter determination module 606 is further configured to determine the configuration parameter of the model parameters corresponding to the maximum predicted training performance among the remaining multiple configuration parameters as the first target configuration parameter; determine the configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the remaining multiple configuration parameters as the second target configuration parameter.
[0125] In an exemplary embodiment, a set of fixed configuration items includes model parameters of the model to be trained; a set of undetermined configuration items includes hardware conditions of the hardware device on which the model to be trained runs, and a parallel strategy adopted during the training of the model to be trained; the target configuration parameters include a second target configuration parameter of the parallel strategy and a third target configuration parameter of the hardware conditions; the configuration parameter determination module 606 is further configured to determine, as the second target configuration parameter, the configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination; and determine, as the third target configuration parameter, the configuration parameter of the hardware conditions corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination.
[0126] In an exemplary embodiment, the configuration parameter determination module 606 is further configured to use the configuration parameter among the multiple configuration parameters as the current configuration parameter, perform the following prediction operations to obtain the maximum total video memory occupancy corresponding to the multiple configuration parameters: determine, according to the specific number of parameters of the model to be trained, a first video memory occupancy required for storing the storage parameter amounts corresponding to multiple video cards in the hardware device on which the model to be trained runs when the model to be trained adopts the current configuration parameter in a set of undetermined configuration items; determine, according to the preset optimizer data of the model to be trained, a second video memory occupancy required for optimizing the states of multiple video cards when the model to be trained adopts the current configuration parameter in a set of undetermined configuration items; analyze the peak video memory occupancy corresponding to the model to be trained when adopting the current configuration parameter in a set of undetermined configuration items to obtain a third video memory occupancy required for storing intermediate results generated by the temporary activation of multiple video cards when the model to be trained adopts the current configuration parameter in a set of undetermined configuration items; determine the total video memory occupancy corresponding to multiple video cards when the model to be trained adopts the current configuration parameter in a set of undetermined configuration items according to the first video memory occupancy, the second video memory occupancy, and the third video memory occupancy; and determine the maximum total video memory occupancy among the total video memory occupancies of multiple video cards as the maximum total video memory occupancy corresponding to the current configuration parameter.
[0127] In an exemplary embodiment, the configuration parameter determination module 606 is further configured to calculate, when the model to be trained adopts the current configuration parameter in a set of undetermined configuration items, a first peak video memory occupancy of multiple video cards during the reverse propagation recalculation process, and calculate a second peak video memory occupancy of multiple video cards during the fully connected calculation process; and determine the maximum value among the first peak video memory occupancy and the second peak video memory occupancy of multiple video cards as the third video memory occupancy required for storing intermediate results generated by the temporary activation of multiple video cards.
[0128] For the description of the features in the corresponding embodiment of the device for determining model training configuration parameters, reference can be made to the relevant description in the corresponding embodiment of the method for determining model training configuration parameters, which will not be elaborated here one by one.
[0129] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above-described method embodiments for determining model training configuration parameters.
[0130] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above-described method embodiments for determining model training configuration parameters when running.
[0131] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drives, read-only memories (ROMs), random access memories (RAMs), external hard drives, magnetic disks, or optical discs, and other media that can store computer programs.
[0132] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for determining model training configuration parameters.
[0133] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for determining model training configuration parameters.
[0134] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0135] The above has introduced in detail a method for determining model training configuration parameters provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method of this application and its core idea. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for determining model training configuration parameters, characterized in that: include: Get multiple configuration items required for training the model to be trained; The multiple configuration items include a group of fixed configuration items and a group of pending configuration items; In the case where the set of fixed configuration items adopts a fixed specified configuration, predicting the training performance of the to-be-trained model when the set of undetermined configuration items adopts a plurality of configuration parameters for training, and obtaining predicted training performance corresponding to the plurality of configuration parameters; According to the predicted training performance corresponding to the multiple configuration parameters, target configuration parameters are selected from the multiple configuration parameters; the target configuration parameters are configuration parameters adopted by the set of pending configuration items when model training is performed on the model to be trained.
2. The method according to claim 1, characterized in that The multiple configuration items include hardware conditions of the hardware device on which the model to be trained runs, model parameters of the model to be trained, and a parallel strategy adopted when training the model to be trained; The predicting the training performance of the to-be-trained model when the set of pending configuration items is trained using a plurality of configuration parameters, and obtaining the predicted training performance corresponding to the plurality of configuration parameters, includes: The following prediction operation is performed using a configuration parameter among the multiple configuration parameters as the current configuration parameter to obtain the prediction training performance corresponding to the multiple configuration parameters: According to the pre-built graphics card performance database, the specified configuration and the current configuration parameters, the graphics card computing performance corresponding to the model to be trained when the current configuration parameters are used for training in the set of pending configuration items is calculated to obtain the graphics card computing performance corresponding to the current configuration parameters; the graphics card performance database stores the graphics card computing performance corresponding to the model to be trained when different model parameters and different parallel strategies are used for training under different hardware conditions; Predicting the current actual training performance of the model to be trained according to the computing performance of the graphics card corresponding to the current configuration parameters; The predicted training performance corresponding to the current configuration parameters is determined according to a ratio between the current actual training performance and the theoretical training performance corresponding to the current configuration parameters.
3. The method according to claim 1, characterized in that The selecting a target configuration parameter from the multiple configuration parameters according to the predicted training performance corresponding to the multiple configuration parameters includes: Predicting the maximum total memory usage corresponding to the model to be trained when the set of pending configuration items is trained using the multiple configuration parameters, and obtaining the maximum total memory usage corresponding to the multiple configuration parameters; Eliminate a group of configuration parameters whose corresponding maximum total video memory usage is greater than the maximum video memory of a single card from the multiple configuration parameters; A configuration parameter having the maximum predicted training performance among the eliminated configuration parameters corresponds to the configuration parameter, and the configuration parameter is determined as the target configuration parameter.
4. The method according to claim 3, characterized in that The set of fixed configuration items includes the hardware conditions of the hardware device on which the model to be trained runs; the set of undetermined configuration items includes the model parameters of the model to be trained and the parallel strategy adopted when training the model to be trained; the target configuration parameters include the first target configuration parameters of the model parameters and the second target configuration parameters of the parallel strategy; The step of determining a configuration parameter corresponding to the maximum predicted training performance among the plurality of configuration parameters after elimination as the target configuration parameter includes: Determine the configuration parameter of the model parameter corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the first target configuration parameter; The configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the second target configuration parameter.
5. The method according to claim 3, characterized in that: The set of fixed configuration items includes model parameters of the model to be trained; the set of undetermined configuration items includes hardware conditions of the hardware device on which the model to be trained runs, and the parallel strategy adopted when training the model to be trained; the target configuration parameters include the second target configuration parameters of the parallel strategy and the third target configuration parameters of the hardware conditions; The step of determining a configuration parameter corresponding to the maximum predicted training performance among the plurality of configuration parameters after elimination as the target configuration parameter includes: Determine the configuration parameter of the parallel strategy corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination as the second target configuration parameter; The configuration parameter of the hardware condition corresponding to the maximum predicted training performance among the multiple configuration parameters after elimination is determined as the third target configuration parameter.
6. The method according to claim 3, characterized in that The predicting the maximum total memory usage corresponding to the to-be-trained model when the set of pending configuration items adopts the multiple configuration parameters, and obtaining the maximum total memory usage corresponding to the multiple configuration parameters, includes: A configuration parameter among the multiple configuration parameters is used as a current configuration parameter, and the following prediction operation is performed to obtain the maximum total video memory occupancy corresponding to the multiple configuration parameters: According to the specific parameter quantity of the model to be trained, determining a first video memory occupancy amount required for multiple graphics cards in a hardware device running the model to be trained to store corresponding storage parameter quantities when the model to be trained adopts the current configuration parameters in the set of pending configuration items; Determine, according to the preset optimizer data of the model to be trained, the second video memory occupancy required by the plurality of graphics card storage optimizer states when the model to be trained adopts the current configuration parameters of the set of pending configuration items; Analyze the peak value of video memory occupancy corresponding to the model to be trained when the set of pending configuration items adopts the current configuration parameters, and obtain the third video memory occupancy required for the multiple graphics cards to store intermediate results generated by temporary activation when the model to be trained adopts the current configuration parameters; Determine, according to the first video memory occupancy, the second video memory occupancy, and the third video memory occupancy, a total video memory occupancy corresponding to the multiple graphics cards when the to-be-trained model adopts the current configuration parameters in the set of pending configuration items; The maximum total memory occupancy among the total memory occupancy of the multiple graphics cards is determined as the maximum total memory occupancy corresponding to the current configuration parameter.
7. The method according to claim 6, characterized in that The step of analyzing the peak value of the video memory occupancy corresponding to the model to be trained when the set of pending configuration items adopts the current configuration parameters to obtain the third video memory occupancy required for the multiple graphics cards to store the intermediate results generated by temporary activation when the model to be trained adopts the current configuration parameters includes: When the model to be trained adopts the current configuration parameters in the set of pending configuration items, calculating the first peak values of the video memory usage of the multiple graphics cards in the back propagation recalculation process, and calculating the second peak values of the video memory usage of the multiple graphics cards in the full connection calculation process; The maximum of the first video memory occupancy peak value and the second video memory occupancy peak value of the multiple graphics cards is determined as the third video memory occupancy required for the multiple graphics cards to store the intermediate results generated by the temporary activation.
8. A device for determining model training configuration parameters, characterized in that: include: A configuration item acquisition module is used to obtain multiple configuration items required for training the model to be trained; The multiple configuration items include a group of fixed configuration items and a group of pending configuration items; A training performance prediction module, used to predict the training performance of the to-be-trained model when the to-be-determined configuration items are trained with multiple configuration parameters, and obtain the predicted training performance corresponding to the multiple configuration parameters, when the to-be-trained model is trained with multiple configuration parameters, when the to-be-determined configuration items are fixedly configured with a specified configuration; A configuration parameter determination module is used to select target configuration parameters from the multiple configuration parameters according to the predicted training performance corresponding to the multiple configuration parameters; the target configuration parameters are the configuration parameters used by the set of pending configuration items when the model to be trained is trained.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for determining the model training configuration parameters as described in any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for determining the model training configuration parameters as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Method for adjusting resources in processor and related product
CN121116647A
Configuration adjustment method of model training system
CN121189524A
A model training system configuration adjustment method
CN121189524B
Deep learning model training method and device and electronic equipment
CN121599025A
Parallel strategy generation method, model training method and equipment
CN121599162A