Fine adjustment method of deep learning model, computer equipment and medium
By setting the low-rank matrix in the target encoding and decoding layer of the deep learning model and updating the matrix parameters, the graphics card memory usage and computing overhead problems during fine-tuning of large pretrained models are solved, and the model fine-tuning efficiency is improved.
Patent Information
- Application Number
- CN202510346168.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-08
AI Technical Summary
When the prior art fine-tuning of large pretrained models, the graphics card memory usage and computing overhead are too large, resulting in low model fine-tuning efficiency.
By setting the low-rank matrix in the target encoding and decoding layer of the deep learning model, only the matrix parameters in the low-rank matrix are updated, other model parameters are frozen, and combined with the adjustment of the number of bits of floating-point operation, the model training process is optimized.
It reduces the memory usage and computing overhead of graphics card, improves the efficiency of model fine-tuning, and optimizes the model training process.
Smart Images

Figure CN120278214A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of model fine-tuning, and specifically relates to a method for fine-tuning a deep learning model, a computer device, and a medium. Background Art
[0002] With the rapid development of deep learning and natural language processing technologies, large pre-trained models (big models) occupy an important position in the field of artificial intelligence due to their excellent generalization ability and multi-task processing ability. However, with the continuous expansion of the model scale, significant technical challenges have been encountered in their actual deployment and application. These challenges not only limit the application scope of big models but also pose higher requirements for building an efficient and scalable big model development platform.
[0003] When fine-tuning a big model, all model parameters in the big model are often adjusted, which will bring huge graphics card memory occupation and computing overhead, and also reduce the efficiency of model fine-tuning. Summary of the Invention
[0004] Embodiments of this application provide a method for fine-tuning a deep learning model, a computer device, and a medium, aiming to reduce the graphics card memory occupation and computing overhead during model fine-tuning and improve the efficiency of model fine-tuning.
[0005] In a first aspect, embodiments of this application provide a method for fine-tuning a deep learning model, and the method for fine-tuning the deep learning model includes:
[0006] Obtain a deep learning model and a model training data set, where the deep learning model includes multiple encoder-decoder layers, and each encoder-decoder layer includes linear transformation modules for queries, keys, and values;
[0007] Select at least one target encoder-decoder layer from the multiple encoder-decoder layers, and set a low-rank matrix in the target linear transformation module of the target encoder-decoder layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the target linear transformation module;
[0008] During the process of performing forward propagation processing on the model training data set using the deep learning model, determine a first output result of the target encoder-decoder layer and a second output result of the low-rank matrix, and determine an output result of the deep learning model according to the first output result and the second output result;
[0009] During the process of performing backward propagation processing on the model training data set using the deep learning model, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters;
[0010] Determine the fine-tuned deep learning model by using the target matrix parameters and the weight matrix of the target linear transformation module.
[0011] In some embodiments, during the process of determining the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, calculations are performed according to floating-point operations with a first preset number of digits.
[0012] During the process of updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model, calculations are performed according to floating-point operations with a second preset number of digits, and the second preset number of digits is greater than the first preset number of digits.
[0013] In some embodiments, the determining the output result of the deep learning model according to the first output result and the second output result includes:
[0014] Merge the first output result and the second output result to obtain the output result of the target encoding and decoding layer;
[0015] Perform forward propagation processing on the output result of the target encoding and decoding layer by using the deep learning model to obtain the output result of the deep learning model.
[0016] In some embodiments, the updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model includes:
[0017] Use the output result of the deep learning model to determine the loss function of the low-rank matrix;
[0018] Update the matrix parameters in the low-rank matrix based on the minimization strategy of the loss function of the low-rank matrix.
[0019] In some embodiments, the determining the fine-tuned deep learning model by using the target matrix parameters and the weight matrix of the target linear transformation module includes:
[0020] Merge the target matrix parameters and the weight matrix of the target linear transformation module to obtain an updated weight matrix;
[0021] Determine the fine-tuned deep learning model according to the updated weight matrix.
[0022] In some embodiments, before selecting at least one target encoding and decoding layer from multiple encoding and decoding layers, it further includes:
[0023] Obtain the set of model parameters in the deep learning model;
[0024] According to a preset allocation strategy, allocate the model parameters in the set of model parameters to the video memories of multiple graphics cards, where the model parameters allocated to the video memories of different graphics cards are different.
[0025] In some embodiments, the step of allocating the model parameters in the set of model parameters to the video memories of multiple graphics cards according to a preset allocation strategy includes:
[0026] Obtain the number of model parameters and the input size in the deep learning model;
[0027] Obtain the expected maximum batch size of the deep learning model;
[0028] Based on the number of model parameters, the input size, and the expected maximum batch size, determine the expected memory capacity of the deep learning model;
[0029] According to the expected memory capacity, determine the expected occupancy value of the video memory of the deep learning model for each graphics card;
[0030] Allocate the model parameters in the set of model parameters to the video memories of multiple graphics cards, so that the data volume of the model parameters allocated to the video memory of each graphics card matches the corresponding expected occupancy value of the video memory.
[0031] In some embodiments, after determining the expected occupancy value of the video memory of the deep learning model for each graphics card according to the expected memory capacity, it further includes:
[0032] If the difference between the remaining capacity of the video memory of the corresponding graphics card and the expected occupancy value of the video memory is less than a preset difference value, then reduce the expected maximum batch size according to the remaining capacity of the video memory of the corresponding graphics card to adjust the expected occupancy value of the video memory;
[0033] The step of allocating the model parameters in the set of model parameters to the video memories of multiple graphics cards, so that the data volume of the model parameters allocated to the video memory of each graphics card matches the corresponding expected occupancy value of the video memory, includes:
[0034] Allocate the model parameters in the set of model parameters to the video memories of multiple graphics cards, so that the data volume of the model parameters allocated to the video memory of each graphics card matches the adjusted corresponding expected occupancy value of the video memory.
[0035] In a second aspect, an embodiment of the present application provides a fine-tuning device for a deep learning model, and the fine-tuning device for the deep learning model includes:
[0036] A first acquisition module, configured to acquire a deep learning model and a model training data set, where the deep learning model includes a plurality of encoding and decoding layers, and each of the encoding and decoding layers includes a linear transformation module for queries, keys, and values;
[0037] A second acquisition module, configured to select at least one target encoding and decoding layer from the multiple encoding and decoding layers, and set a low-rank matrix in a target linear transformation module of the target encoding and decoding layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the target linear transformation module;
[0038] A first determination module, configured to determine a first output result of the target encoding and decoding layer and a second output result of the low-rank matrix during the forward propagation process of the model training dataset using the deep learning model, and determine an output result of the deep learning model according to the first output result and the second output result;
[0039] A second determination module, configured to freeze all model parameters in the deep learning model that do not include the low-rank matrix during the backward propagation process of the model training dataset using the deep learning model, and update matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters;
[0040] A third determination module, configured to determine a fine-tuned deep learning model using the target matrix parameters and the weight matrix of the target linear transformation module.
[0041] In a third aspect, an embodiment of the present application provides a computer device, where the computer device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is configured to be executed by the processor to implement the fine-tuning method of the deep learning model as described in any one of the above.
[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the fine-tuning method of the deep learning model as described in any one of the above.
[0043] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction, and the computer program or instruction is executed by a processor to implement the fine-tuning method of the deep learning model as described in any one of the above.
[0044] Beneficial effects of the embodiments of the present application:
[0045] In the embodiments of the present application, by setting a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, and using the model training dataset of the deep learning model to update the matrix parameters in the low-rank matrix, so as to implement the fine-tuning of the deep learning model, avoiding updating all model parameters in the deep learning model, thereby reducing the graphics card memory occupation and calculation overhead, and improving the efficiency of model fine-tuning. Brief Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0047] Figure 1 is a schematic flowchart of an embodiment of the fine-tuning method for a deep learning model provided by an embodiment of the present application;
[0048] Figure 2 is another schematic flowchart of an embodiment of the fine-tuning method for a deep learning model provided by an embodiment of the present application;
[0049] Figure 3 is still another schematic flowchart of an embodiment of the fine-tuning method for a deep learning model provided by an embodiment of the present application;
[0050] Figure 4 is yet another schematic flowchart of an embodiment of the fine-tuning method for a deep learning model provided by an embodiment of the present application;
[0051] Figure 5 is a schematic structural diagram of an embodiment of a computer device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.
[0053] In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined. In addition, in the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features.
[0054] To reduce the GPU memory occupancy and computational overhead during model fine-tuning and improve the efficiency of model fine-tuning, the embodiments of the present application provide a method for fine-tuning a deep learning model, a computer device, and a medium. By setting a low-rank matrix in the target linear transformation module of the target encoder-decoder layer and using the model training dataset of the deep learning model to update the matrix parameters in the low-rank matrix, the fine-tuning of the deep learning model is achieved, avoiding updating all model parameters in the deep learning model to reduce GPU memory occupancy and computational overhead and improve the efficiency of model fine-tuning. For specific solutions, please refer to the following specific descriptions.
[0055] In a first aspect, embodiments of the present application provide a method for fine-tuning a deep learning model. Specifically, referring to Figure 1 , Figure 1 is a schematic flowchart of an embodiment of the method for fine-tuning a deep learning model. In Figure 1 , the method for fine-tuning a deep learning model may include:
[0056] 101. Obtain a deep learning model and a model training dataset, where the deep learning model includes multiple encoder-decoder layers, and each encoder-decoder layer includes linear transformation modules for queries, keys, and values.
[0057] In the embodiments of the present application, the deep learning model may be a model that needs post-training reinforcement inference. Post-training is relative to pre-training. By performing post-training on the deep learning model, the inference ability of the deep learning model can be strengthened. Deep learning models that need post-training reinforcement inference may, for example, be large pre-trained models (large models). The model training dataset is a dataset used for post-training the deep learning model, and this dataset can be set based on actual needs.
[0058] The deep learning model includes multiple encoder-decoder (encode-decode) layers, and each encoder-decoder layer includes linear transformation modules for queries (query), keys (key), and values (value), so that the encoder-decoder layer forms a Transformer layer.
[0059] A weight matrix for linear transformation processing is set in the linear transformation module, and the weight matrix may include query vectors, keys, values, etc. Through the linear transformation processing of this weight matrix, the input data of the encoder-decoder layer is converted into corresponding output data, realizing the data transformation function of the encoder-decoder layer.
[0060] 102. Select at least one target encoder-decoder layer from the multiple encoder-decoder layers, and set a low-rank matrix in the target linear transformation module of the target encoder-decoder layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module.
[0061] In an embodiment of the present application, the target encoding and decoding layer may be an encoding and decoding layer in a deep learning model that has a greater impact on performance. The target encoding and decoding layer can be set based on actual requirements, or all the encoding and decoding layers in the deep learning model can be used as the target encoding and decoding layer respectively to perform subsequent steps.
[0062] In some embodiments, the low-rank matrix can be obtained by decomposing the weight matrix of the linear transformation module, and specifically, it can be implemented by using Low-Rank Adaptation (LoRA). Low-rank adaptation is a fine-tuning technique for large language models. Its core idea is to approximate the full-rank matrix of the original model by introducing a low-rank matrix, thereby reducing the number of parameters and computational complexity. Therefore, by adding a low-rank adapter to the target encoding and decoding layer, a corresponding low-rank matrix can be generated based on the weight matrix of the linear transformation module.
[0063] 103. During the process of performing forward propagation processing on the model training dataset using the deep learning model, determine the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model based on the first output result and the second output result.
[0064] In an embodiment of the present application, the model training of the deep learning model may specifically include a forward propagation processing stage, and the forward propagation processing stage refers to the stage of determining the output result of the deep learning model by using the model training dataset of the deep learning model. In the forward propagation processing stage, using the model training dataset, the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix can be determined respectively. That is, in the forward propagation processing stage, for the target encoding and decoding layer, in addition to performing the original standard calculation process, it is also necessary to additionally calculate the incremental output caused by the low-rank matrix. And the output result of the deep learning model can be comprehensively determined based on the first output result and the second output result, so that the influence of the low-rank matrix is considered in the process of forward propagation processing.
[0065] 104. During the process of performing backward propagation processing on the model training dataset using the deep learning model, freeze all the model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain the target matrix parameters.
[0066] In an embodiment of the present application, the model training of the deep learning model may further include a backpropagation processing stage, which refers to the stage of updating the target low-rank matrix using the output result of the deep learning model. In the backpropagation processing stage, only the matrix parameters in the low-rank matrix are updated, and all other model parameters in the deep learning model except the low-rank matrix are frozen (for example, the weight matrix of the target linear transformation module is not updated) to reduce the graphics card memory occupancy and computational overhead and improve the efficiency of model fine-tuning.
[0067] 105. Determine the fine-tuned deep learning model using the target matrix parameters and the weight matrix of the target linear transformation module.
[0068] In an embodiment of the present application, after updating the matrix parameters in the low-rank matrix, the target matrix parameters that meet the requirements can be obtained. At this time, the training process of the deep learning model ends, and the fine-tuned deep learning model can be determined based on the target matrix parameters and the weight matrix of the target linear transformation module to achieve the fine-tuning of the deep learning model. Among them, the model training can be implemented using training models such as DeepSpeed. Parameters such as ZeRO Stage3 and the number of activated checkpoints and the number of gradient accumulation steps can be specified in the configuration file of DeepSpeed.
[0069] It can be seen that in the above embodiment of the present application, by setting a low-rank matrix in the target linear transformation module of the target encoding and decoding layer and using the model training dataset of the deep learning model to update the matrix parameters in the low-rank matrix, the fine-tuning of the deep learning model is achieved, avoiding updating all model parameters in the deep learning model to reduce the graphics card memory occupancy and computational overhead and improve the efficiency of model fine-tuning.
[0070] In some embodiments of the present application, as Figure 2 shown, on the basis of the embodiment shown in Figure 1 the fine-tuning method of the deep learning model may further include:
[0071] 201. During the process of determining the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, perform calculations according to the floating-point operation of the first preset number of digits.
[0072] In an embodiment of the present application, floating-point operation (floating point, FP) refers to real number operation. The floating-point operation of the first preset number of digits may be, for example, 16-bit floating-point operation, that is, FP16. Since the model training of the deep learning model may specifically include a forward propagation processing stage, which refers to the stage of determining the output result of the deep learning model using the model training dataset of the deep learning model, the model calculation rules in the forward propagation processing stage can be restricted according to the floating-point operation of the first preset number of digits.
[0073] 202. In the process of updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model, calculations are performed according to floating-point operations with a second preset number of digits, and the second preset number of digits is greater than the first preset number of digits.
[0074] In an embodiment of the present application, since the model training of the deep learning model may further include a backpropagation processing stage, and the backpropagation processing stage refers to the stage of updating the matrix parameters in the low-rank matrix by using the output result of the deep learning model, the model calculation rules in the backpropagation processing stage can be restricted according to floating-point operations with a second preset number of digits.
[0075] Among them, the second preset number of digits is greater than the first preset number of digits. The floating-point operation with the second preset number of digits can be, for example, 32-bit floating-point operation, that is, FP32, and the floating-point operation with the first preset number of digits can be, for example, 16-bit floating-point operation. It can be seen that by setting a floating-point operation with a smaller number of digits in the forward propagation processing stage of the deep learning model, the occupancy of video memory in the forward propagation processing stage can be reduced and the training speed can be accelerated. In the backpropagation processing stage, a floating-point operation with a larger number of digits is set to ensure the stability and accuracy when updating the matrix parameters in the low-rank matrix.
[0076] In some embodiments of the present application, determining the output result of the deep learning model according to the first output result and the second output result may include: combining the first output result and the second output result to obtain the output result of the target encoding and decoding layer, that is, adding the incremental output caused by the low-rank matrix to the output of the original standard calculation process to obtain the output result of the target encoding and decoding layer; using the deep learning model to continue the forward propagation processing on the output result of the target encoding and decoding layer to obtain the output result of the deep learning model.
[0077] In some embodiments of the present application, updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model may include: using the output result of the deep learning model to determine the loss function of the low-rank matrix. The type of the loss function can be preset based on actual needs, such as mean square error, mean absolute error, cross-entropy loss, etc., which are not limited herein; based on the minimization strategy of the loss function of the low-rank matrix, updating the matrix parameters in the low-rank matrix. For example, with the minimization of the loss function of the low-rank matrix as the goal, updating the matrix parameters in the low-rank matrix until the loss function of the low-rank matrix meets the expected requirements to obtain the target matrix parameters, without using the loss function of the weight matrix of the target linear transformation module to reduce the computational overhead and accelerate the convergence of the loss function.
[0078] In some embodiments of the present application, as Figure 3 shown, in Figure 1 or Figure 2Based on the illustrated embodiments, determining a fine-tuned deep learning model using the target matrix parameters and the weight matrix of the linear transformation module may include:
[0079] 301. Combine the target matrix parameters and the weight matrix of the target linear transformation module to obtain an updated weight matrix.
[0080] In the embodiments of the present application, after obtaining the target matrix parameters, it indicates that the model training of the deep learning model is completed. At this time, the target matrix parameters and the weight matrix of the target linear transformation module can be combined to obtain an updated weight matrix, so as to avoid updating the weight matrix of the target linear transformation module in each iteration process during the model training of the deep learning model, thereby reducing the video card memory occupancy and calculation overhead. Among them, the combination process of the target matrix parameters and the weight matrix of the target linear transformation module can be, for example, matrix addition processing.
[0081] 302. Determine the fine-tuned deep learning model according to the updated weight matrix.
[0082] In the embodiments of the present application, replacing the weight matrix of the original target linear transformation module in the deep learning model with the updated weight matrix can obtain the fine-tuned deep learning model. It can be seen that during the fine-tuning process of the deep learning model, the weight matrix of the target linear transformation module is updated less frequently. For example, the weight matrix of the target linear transformation module can be updated only once, so as to reduce the video card memory occupancy and calculation overhead and improve the efficiency of model fine-tuning.
[0083] In some embodiments of the present application, as Figure 4 shown, before selecting at least one target encoding and decoding layer from multiple said encoding and decoding layers based on any one of the embodiments shown in Figures 1 to 3 it may further include:
[0084] 401. Obtain the model parameter set in the deep learning model.
[0085] In the embodiments of the present application, the model parameter set refers to the set of all model parameters in the deep learning model.
[0086] 402. Allocate the model parameters in the model parameter set to the video memories of multiple video cards according to a preset allocation strategy, and the model parameters allocated to the video memories of different video cards are different.
[0087] In an embodiment of the present application, multiple graphics cards are simultaneously set in a host for fine-tuning a deep learning model, and the fine-tuning of the deep learning model can be implemented distributively on the multiple graphics cards. Therefore, according to a preset allocation strategy, the model parameters in the model parameter set can be allocated to the video memories of the multiple graphics cards, and the model parameters allocated to the video memories of different graphics cards are different. For example, a part of the model parameters in the model parameter set can be allocated to the first graphics card among the multiple graphics cards, and another part of the model parameters in the model parameter set can be allocated to the second graphics card among the multiple graphics cards, so as to implement the distributed fine-tuning of the deep learning model. The preset allocation strategy can be determined based on the automatic partitioning function of DeepSpeed and follow the rules of Zero Stage 3.
[0088] In some embodiments of the present application, when configuring the video memory for a deep learning model, relevant technical personnel often manually allocate the corresponding graphics card memory size based on past experience. However, the graphics card memory allocated in this way is often not reasonable enough, and when the graphics card memory consumed by the deep learning model is too large, problems such as out-of-memory or a decrease in the model inference speed are likely to occur. In order to make the graphics card memory allocation of the deep learning model more reasonable, the setting parameters such as the expected maximum batch size can also be adjusted according to hardware resources such as the graphics card video memory to improve the performance when fine-tuning the deep learning model. Specifically, step 402 may include:
[0089] Obtain the number of model parameters and the input size in the deep learning model; obtain the expected maximum batch size of the deep learning model; determine the expected memory capacity of the deep learning model based on the number of model parameters, the input size, and the expected maximum batch size; determine the expected video memory occupancy value of the deep learning model for each graphics card according to the expected memory capacity; allocate the model parameters in the model parameter set to the video memories of the multiple graphics cards, so that the data volume of the model parameters allocated to the video memory of each graphics card matches the corresponding expected video memory occupancy value. For example, the data volume of the model parameters allocated to the video memory of each graphics card can be equal to the corresponding expected video memory occupancy value.
[0090] Among them, the number of model parameters refers to the number of model parameters. The input size refers to the size of the input data of the deep learning model. For example, it may include: the batch size and the number of input features of the fully connected layer in the deep learning model, the batch size, the number of input channels, height, and width of the convolutional layer, the batch size, the number of input channels, height, and width of the pooling layer, the batch size and the number of input features of the normalization layer, and the batch size and the number of input features of the activation layer. The expected maximum batch size refers to the expected maximum batch size (batchsize) of the deep learning model, which can be set within a preset numerical range based on actual requirements.
[0091] In the step of determining the expected memory capacity of a deep learning model based on the number of model parameters, the input size, and the expected maximum batch size, since the larger the number of model parameters, the input size, and the expected maximum batch size, the more complex the deep learning model itself is, and usually more video card memory is required for model calculation. Therefore, based on previous experiments, a better correlation relationship between the number of model parameters, the input size, the expected maximum batch size, and the expected memory capacity can be generated. Then, based on this correlation relationship, the expected memory capacity of the deep learning model can be determined.
[0092] In the step of determining the expected video memory occupancy value of the deep learning model for each video card according to the expected memory capacity, the expected memory capacity can be divided according to the equal distribution strategy to obtain the expected video memory occupancy value of the deep learning model for each video card, that is, the expected video memory occupancy value of the deep learning model for each video card can be equal.
[0093] It can be seen that in the above embodiments of the present application, by determining the expected memory capacity of the deep learning model based on the number of model parameters, the input size, and the expected memory capacity in the deep learning model, and then allocating it to the video memories of multiple video cards, the video card memory allocation of the deep learning model is made more reasonable to reduce the situation of memory overflow or the decrease in model inference speed.
[0094] In some embodiments of the present application, after determining the expected video memory occupancy value of the deep learning model for each video card according to the expected memory capacity, it may further include: for each video card, if the difference between the remaining video memory capacity of the corresponding video card and the expected video memory occupancy value is less than the preset difference, it indicates that if the video card memory is directly allocated to the deep learning model according to the expected memory capacity, phenomena such as memory overflow or the decrease in model inference speed may occur. Therefore, the expected maximum batch size can be reduced according to the remaining video memory capacity of the corresponding video card to adjust the expected video memory occupancy value. For example, the expected maximum batch size can be reduced according to a preset amplitude so that the expected video memory occupancy value is also reduced until the difference between the remaining video memory capacity of the corresponding video card and the adjusted expected video memory occupancy value (i.e., the reduced expected video memory occupancy value) is greater than or equal to the preset difference.
[0095] When the difference between the remaining video memory capacity of the corresponding video card and the adjusted expected video memory occupancy value is greater than or equal to the preset difference, it indicates that the current expected video memory occupancy value and the expected maximum batch size (i.e., the reduced expected video memory occupancy value and the expected maximum batch size) are appropriate. Adaptively, allocating the model parameters in the model parameter set to the video memories of multiple video cards so that the data volume of the model parameters allocated to the video memory of each video card matches the corresponding expected video memory occupancy value may include: allocating the model parameters in the model parameter set to the video memories of multiple video cards so that the data volume of the model parameters allocated to the video memory of each video card matches the adjusted corresponding expected video memory occupancy value.
[0096] Second aspect, based on the fine-tuning method of the deep learning model in the above embodiments, an embodiment of the present application provides a deep learning model fine-tuning device, and the deep learning model fine-tuning device is used to execute the steps in any one of the embodiments of the above deep learning model fine-tuning method. Specifically, the deep learning model fine-tuning device may include:
[0097] A first acquisition module, configured to acquire a deep learning model and a model training data set, wherein the deep learning model includes a plurality of encoding and decoding layers, and each encoding and decoding layer includes a linear transformation module for queries, keys, and values;
[0098] A second acquisition module, configured to select at least one target encoding and decoding layer from the plurality of encoding and decoding layers, and set a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, wherein the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module;
[0099] A first determination module, configured to determine a first output result of the target encoding and decoding layer and a second output result of the low-rank matrix during the forward propagation process of using the deep learning model to process the model training data set, and determine the output result of the deep learning model according to the first output result and the second output result;
[0100] A second determination module, configured to freeze all model parameters in the deep learning model that do not include the low-rank matrix during the backward propagation process of using the deep learning model to process the model training data set, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters;
[0101] A third determination module, configured to determine the fine-tuned deep learning model by using the target matrix parameters and the weight matrix of the linear transformation module.
[0102] Third aspect, an embodiment of the present application provides a computer device, which integrates any one of the deep learning model fine-tuning devices provided in the embodiments of the present application. The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is configured to be executed by the processor to implement the deep learning model fine-tuning method described in any one of the above embodiments, for example:
[0103] Obtain a deep learning model and a model training dataset, where the deep learning model includes multiple encoding and decoding layers, and each encoding and decoding layer includes a linear transformation module for queries, keys, and values; select at least one target encoding and decoding layer from the multiple encoding and decoding layers, and set a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module; during the process of forward propagation processing of the model training dataset using the deep learning model, determine the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model based on the first output result and the second output result; during the process of backpropagation processing of the model training dataset using the deep learning model, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix based on the output result of the deep learning model to obtain target matrix parameters; use the target matrix parameters and the weight matrix of the linear transformation module to determine the fine-tuned deep learning model.
[0104] In a fourth aspect, an embodiment of the present application provides a computer device that integrates any fine-tuning device of the deep learning model provided by the embodiments of the present application. As Figure 5 shown, it shows a schematic structural diagram of the computer device involved in the embodiments of the present application. Specifically:
[0105] The computer device may include a processor 501 with one or more processing cores, a storage unit 502 with one or more computer-readable storage media, a power supply 503, an input unit 504, and other components. Those skilled in the art can understand that Figure 5 the structural diagram of the computer device shown in does not constitute a limitation on the computer device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:
[0106] The processor 501 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines. By running or executing software programs and / or modules stored in the storage unit 502, and calling the data stored in the storage unit 502, it executes various functions of the computer device and processes data, thereby monitoring the computer device as a whole. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, where the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communication. It can be understood that the above modem processor may not be integrated into the processor 501.
[0107] The storage unit 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the storage unit 502. The storage unit 502 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the storage unit 502 can include a high-speed random access memory and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the storage unit 502 can also include a memory controller to provide the processor 501 with access to the storage unit 502.
[0108] The computer device further includes a power supply 503 for supplying power to each component. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 503 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0109] The computer device may further include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0110] Although not shown, the computer device may further include a display unit and the like, which will not be elaborated here. Specifically, in the embodiment of the present application, the processor 501 in the computer device will load the executable files corresponding to the processes of one or more application programs into the storage unit 502 according to the following instructions, and the processor 501 will run the application programs stored in the storage unit 502 to implement various functions, such as:
[0111] Obtain a deep learning model and a model training dataset. The deep learning model includes multiple encoding and decoding layers, and each encoding and decoding layer includes linear transformation modules for queries, keys, and values. Select at least one target encoding and decoding layer from the multiple encoding and decoding layers, and set a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module. During the forward propagation process of using the deep learning model to process the model training dataset, determine the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model based on the first output result and the second output result. During the backward propagation process of using the deep learning model to process the model training dataset, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters. Use the target matrix parameters and the weight matrix of the linear transformation module to determine the fine-tuned deep learning model.
[0112] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disc, etc. The computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the fine-tuning method of the deep learning model as described in any one of the above, for example:
[0113] Obtain a deep learning model and a model training dataset. The deep learning model includes multiple encoding and decoding layers, and each encoding and decoding layer includes linear transformation modules for queries, keys, and values. Select at least one target encoding and decoding layer from the multiple encoding and decoding layers, and set a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module. During the forward propagation process of using the deep learning model to process the model training dataset, determine the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model based on the first output result and the second output result. During the backward propagation process of using the deep learning model to process the model training dataset, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters. Use the target matrix parameters and the weight matrix of the linear transformation module to determine the fine-tuned deep learning model.
[0114] In a sixth aspect, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute and implement the fine-tuning method of the deep learning model as described in any one of the above, for example:
[0115] Obtain a deep learning model and a model training data set. Among them, the deep learning model includes multiple encoder-decoder layers, and each encoder-decoder layer includes a linear transformation module for query, key, and value; select at least one target encoder-decoder layer from the multiple encoder-decoder layers, and set a low-rank matrix in the target linear transformation module of the target encoder-decoder layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the linear transformation module; during the forward propagation process of using the deep learning model to process the model training data set, determine the first output result of the target encoder-decoder layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model according to the first output result and the second output result; during the backward propagation process of using the deep learning model to process the model training data set, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters; use the target matrix parameters and the weight matrix of the linear transformation module to determine the fine-tuned deep learning model.
[0116] The embodiments of the present application have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A fine-tuning method for a deep learning model, characterized in that, The fine-tuning method of the deep learning model includes: Obtain a deep learning model and a model training dataset, where the deep learning model includes multiple encoding and decoding layers, and each encoding and decoding layer includes a linear transformation module for queries, keys, and values; Select at least one target encoding and decoding layer from the multiple encoding and decoding layers, and set a low-rank matrix in the target linear transformation module of the target encoding and decoding layer, where the rank of the low-rank matrix is less than the rank of the weight matrix of the target linear transformation module; During the process of forward propagation processing of the model training dataset using the deep learning model, determine the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, and determine the output result of the deep learning model according to the first output result and the second output result; During the process of backpropagation processing of the model training dataset using the deep learning model, freeze all model parameters in the deep learning model that do not include the low-rank matrix, and update the matrix parameters in the low-rank matrix according to the output result of the deep learning model to obtain target matrix parameters; Determine the fine-tuned deep learning model using the target matrix parameters and the weight matrix of the target linear transformation module.
2. The fine-tuning method of the deep learning model according to claim 1, characterized in that During the process of determining the first output result of the target encoding and decoding layer and the second output result of the low-rank matrix, perform calculations according to floating-point operations with a first preset number of digits; During the process of updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model, perform calculations according to floating-point operations with a second preset number of digits, and the second preset number of digits is greater than the first preset number of digits.
3. The fine-tuning method of the deep learning model according to claim 1, characterized in that The determining the output result of the deep learning model according to the first output result and the second output result includes: Merge the first output result and the second output result to obtain the output result of the target encoding and decoding layer; Perform forward propagation processing on the output result of the target encoding and decoding layer using the deep learning model to obtain the output result of the deep learning model.
4. The fine-tuning method of the deep learning model according to claim 1, characterized in that The updating the matrix parameters in the low-rank matrix according to the output result of the deep learning model includes: Determine the loss function of the low-rank matrix using the output result of the deep learning model; Update the matrix parameters in the low-rank matrix based on the minimization strategy of the loss function of the low-rank matrix.
5. The fine-tuning method of the deep learning model according to claim 1, characterized in that, The determining the fine-tuned deep learning model using the target matrix parameters and the weight matrix of the target linear transformation module includes: Merge the target matrix parameters and the weight matrix of the target linear transformation module to obtain an updated weight matrix; Determine the fine-tuned deep learning model according to the updated weight matrix.
6. The fine-tuning method of the deep learning model according to claim 1, characterized in that, Before selecting at least one target encoding and decoding layer from the multiple encoding and decoding layers, it further includes: Obtain the model parameter set in the deep learning model; According to a preset allocation strategy, allocate the model parameters in the model parameter set to the video memories of multiple graphics cards, and the allocated model parameters of different graphics cards are different.
7. The fine-tuning method of the deep learning model according to claim 6, wherein Allocating the model parameters in the model parameter set to the video memories of multiple graphics cards according to a preset allocation strategy includes: Obtaining the number of model parameters and the input size in the deep learning model; Obtaining the expected maximum batch size of the deep learning model; Determining the expected memory capacity of the deep learning model based on the number of model parameters, the input size, and the expected maximum batch size; Determining the expected video memory occupancy value of the deep learning model for each graphics card according to the expected memory capacity; Allocating the model parameters in the model parameter set to the video memories of multiple graphics cards so that the data volume of the model parameters allocated to the video memory of each graphics card matches the corresponding expected video memory occupancy value.
8. The fine-tuning method of the deep learning model according to claim 7, wherein After determining the expected video memory occupancy value of the deep learning model for each graphics card according to the expected memory capacity, it further includes: If the difference between the remaining video memory capacity of the corresponding graphics card and the expected video memory occupancy value is less than a preset difference, reducing the expected maximum batch size according to the remaining video memory capacity of the corresponding graphics card to adjust the expected video memory occupancy value; The step of allocating the model parameters in the model parameter set to the video memories of multiple graphics cards so that the data volume of the model parameters allocated to the video memory of each graphics card matches the corresponding expected video memory occupancy value includes: Allocating the model parameters in the model parameter set to the video memories of multiple graphics cards so that the data volume of the model parameters allocated to the video memory of each graphics card matches the adjusted corresponding expected video memory occupancy value.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory. The computer program is configured to be executed by the processor to implement the fine-tuning method of the deep learning model according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is configured to be executed by a processor to implement the fine-tuning method of the deep learning model according to any one of claims 1 to 8.