Large model fine tuning training method and device, electronic equipment and computer readable medium

By chunking quantization and downsampling of large model weights to generate companion networks, combined with asynchronous heterogeneous training, the problem of resource waste in existing large model fine-tuning training methods is solved, and more efficient resource utilization and model training performance is achieved.

CN120218164APending Publication Date: 2025-06-27SHANGHAI QI ZHI INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285349.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing large-model fine-tuning training methods have insufficient efficiency and resource utilization, resulting in waste of computing resources and video memory resources.

Method used

The weights of the trained large model are quantized in blocks and loaded the quantized weights and models into the central processor. The large model is downsampled to generate a companion network and loaded into the graphics processor. Based on the preset number of iterations, perform asynchronous heterogeneous training steps to optimize the use of computing and video memory resources.

Benefits of technology

This method effectively reduces the waste of computing resources and video memory resources, improves the efficiency of fine-tuning training of large models, and improves the performance of overall model training by optimizing resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218164A_ABST
    Figure CN120218164A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large model fine tuning training method and device, electronic equipment and a computer readable medium. A specific embodiment of the method comprises the steps of performing block quantization on a weight of a to-be-trained large model to obtain a quantized weight, and loading the to-be-trained large model and the quantized weight into a central processing unit to serve as a quantized large model; performing down-sampling on the to-be-trained large model to obtain an associated network, and loading the associated network into a graphics processor; executing a training step on the quantized large model and the associated network; inputting the original data into the quantized large model to obtain a forward propagation vector set; performing asynchronous heterogeneous training on the associated network to obtain an updated associated network; determining the updated associated network as a trained associated network in response to determining that the current number of iterations is equal to the preset number of iterations; and determining the quantized large model and the trained associated network as a trained large model. According to the embodiment, consumption of calculation and video memory resources in model training can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and more particularly, to large model fine-tuning training methods, devices, electronic devices, and computer-readable media. Background Art

[0002] During the fine-tuning process of large models, a large amount of computing and storage resources are usually required, which limits the application and wide deployment of large models. Currently, when fine-tuning and training large models, the commonly used methods are: co-training method, parameter-efficient fine-tuning method (PEFT), and model quantization method. Among them, the co-training method transfers part of the storage tasks to the central processing unit (CPU) to reduce the burden on the graphics processing unit (GPU), thereby efficiently completing the fine-tuning training of large models; the parameter-efficient fine-tuning method reduces the memory occupation during the fine-tuning of large models by introducing lightweight modules or selectively freezing some parameters; the model quantization method reduces the memory requirement during the fine-tuning of large models by quantifying the model weights.

[0003] However, when using the above methods to fine-tune and train large models, the following technical problems often exist:

[0004] First, existing methods only regard the CPU as an auxiliary device, making it difficult to improve the efficiency of large model fine-tuning training, resulting in waste of computing resources and video memory resources.

[0005] Second, the parameter-efficient fine-tuning method and the model quantization method still need to cache a large number of intermediate activation values. Especially when dealing with long sequences or large batches of data, it may cause a large consumption of video memory resources.

[0006] The above information disclosed in this background art section is only used to enhance the understanding of the background of the concept of the present disclosure. Therefore, it may include information that does not constitute the prior art known to ordinary technicians in the art of this country. Summary of the Invention

[0007] The content part of the present disclosure is used to introduce the concepts in a brief form, and these concepts will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0008] Some embodiments of the present disclosure propose large model fine-tuning training methods, devices, electronic devices, and computer-readable media to solve one or more of the technical problems mentioned in the above background art section.

[0009] In a first aspect, some embodiments of the present disclosure provide a method for fine-tuning and training a large model. The method includes: in response to detecting that the video memory occupancy of a server performing fine-tuning and training of the large model exceeds a preset occupancy condition, performing block quantization on the weights of the large model to be trained to obtain quantized weights, and loading the large model to be trained and the quantized weights into a central processing unit as a quantized large model; performing downsampling on the large model to be trained to obtain an associated network, and loading the associated network into a graphics processing unit; based on a preset number of iterations, performing the following training steps on the quantized large model and the associated network: inputting original data into the quantized large model to obtain a set of forward propagation vectors, where the set of forward propagation vectors is a set composed of forward propagation vectors output by each network layer of the quantized large model; performing asynchronous heterogeneous training on the associated network according to the set of forward propagation vectors to obtain an updated associated network; in response to determining that the current number of iterations is not equal to the preset number of iterations, using the updated associated network as the associated network to continue performing the above training steps; in response to determining that the current number of iterations is equal to the preset number of iterations, determining the updated associated network as the trained associated network; and determining the quantized large model and the trained associated network as the trained large model.

[0010] In a second aspect, some embodiments of the present disclosure provide a device for fine-tuning and training a large model. The device includes: a block quantization unit configured to, in response to detecting that the video memory occupancy of a server performing fine-tuning and training of the large model exceeds a preset occupancy condition, perform block quantization on the weights of the large model to be trained to obtain quantized weights, and load the large model to be trained and the quantized weights into a central processing unit as a quantized large model; a downsampling unit configured to perform downsampling on the large model to be trained to obtain an associated network, and load the associated network into a graphics processing unit; a training unit configured to, based on a preset number of iterations, perform the following training steps on the quantized large model and the associated network: inputting original data into the quantized large model to obtain a set of forward propagation vectors, where the set of forward propagation vectors is a set composed of forward propagation vectors output by each layer of the quantized large model; performing asynchronous heterogeneous training on the associated network according to the set of forward propagation vectors to obtain an updated associated network; in response to determining that the current number of iterations is not equal to the preset number of iterations, using the updated associated network as the associated network to continue performing the above training steps; in response to determining that the current number of iterations is equal to the preset number of iterations, determining the updated associated network as the trained associated network; and a determination unit configured to determine the quantized large model and the trained associated network as the trained large model.

[0011] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the above first aspect.

[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the above first aspect is implemented.

[0013] The above embodiments of the present disclosure have the following beneficial effects: Through the large model fine-tuning training method of some embodiments of the present disclosure, the waste of computing resources and video memory resources can be reduced. Specifically, the reasons for the waste of computing resources and video memory resources are as follows: The existing methods only regard the CPU as an auxiliary device and it is difficult to improve the efficiency of large model fine-tuning training, resulting in the waste of computing resources and video memory resources. Based on this, in the large model fine-tuning training method of some embodiments of the present disclosure, first, in response to detecting that the video memory occupancy of the server executing the large model fine-tuning training exceeds the preset occupancy condition, the weights of the large model to be trained are block-quantized to obtain quantized weights, and the above-mentioned large model to be trained and the above-mentioned quantized weights are loaded into the central processing unit as a quantized large model. By block-quantizing the weights, the memory occupancy during large model training can be reduced, and by loading the large model and the quantized weights into the CPU, data can be efficiently stored and processed, reducing the usage pressure on the GPU and avoiding video memory overflow. Next, the above-mentioned large model to be trained is downsampled to obtain an associated network, and the above-mentioned associated network is loaded into the graphics processing unit. Thus, a lightweight network model can be obtained, reducing the amount of computation during training. After that, based on a preset number of iterations, the following training steps are performed on the above-mentioned quantized large model and the above-mentioned associated network: The original data is input into the above-mentioned quantized large model to obtain a forward propagation vector set. Among them, the above-mentioned forward propagation vector set is a set composed of forward propagation vectors output by each network layer of the above-mentioned quantized large model. By performing forward propagation on the large model in the CPU, the waiting and idle time during training can be reduced, the utilization of hardware resources can be maximized, and the overall efficiency of model training can be improved. According to the above-mentioned forward propagation vector set, asynchronous heterogeneous training is performed on the above-mentioned associated network to obtain an updated associated network. By using the forward propagation result of the large model to guide the training of the associated network, parallel work between the CPU and the GPU is realized, resource waste during training is reduced, and the resource utilization efficiency is improved. In response to determining that the current number of iterations is not equal to the preset number of iterations, the above-mentioned updated associated network is used as the associated network to continue performing the above-mentioned training steps. In response to determining that the current number of iterations is equal to the preset number of iterations, the above-mentioned updated associated network is determined as the trained associated network. Finally, the above-mentioned quantized large model and the above-mentioned trained associated network are determined as the trained large model. Thus, the usage efficiency of computing resources and video memory resources during the large model fine-tuning training process can be improved, and the waste of computing and video memory resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more obvious. Throughout the accompanying drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.

[0015] Figure 1 is a flowchart of some embodiments of the large model fine-tuning training method according to the present disclosure;

[0016] Figure 2 is a structural diagram of asynchronous heterogeneous training according to some embodiments of the large model fine-tuning training method of the present disclosure;

[0017] Figure 3 is a schematic structural diagram of some embodiments of the large model fine-tuning training device according to the present disclosure;

[0018] Figure 4 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners

[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0020] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0021] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.

[0022] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0023] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.

[0024] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.

[0025] Figure 1 Flow 100 of some embodiments of the large model fine-tuning training method according to the present disclosure is shown. The large model fine-tuning training method includes the following steps:

[0026] Step 101, in response to detecting that the video memory occupancy of the server performing large model fine-tuning training exceeds a preset occupancy condition, perform block quantization on the weights of the large model to be trained to obtain quantized weights, and load the large model to be trained and the quantized weights into the central processing unit as a quantized large model.

[0027] In some embodiments, the execution subject of the large model fine-tuning training method may, in response to detecting that the video memory occupancy of the server performing large model fine-tuning training exceeds a preset occupancy condition, perform block quantization on the weights of the large model to be trained to obtain quantized weights, and load the above-mentioned large model to be trained and the above-mentioned quantized weights into the central processing unit as a quantized large model. Among them, the above-mentioned preset occupancy condition may be that the video memory occupancy of the server performing large model fine-tuning training exceeds 90% of the video memory quota.

[0028] In some optional implementation manners of some embodiments, the above-mentioned execution subject, in response to detecting that the video memory occupancy of the server performing large model fine-tuning training exceeds a preset occupancy condition, performs block quantization on the weights of the large model to be trained to obtain quantized weights, and loads the above-mentioned large model to be trained and the above-mentioned quantized weights into the central processing unit as a quantized large model, may include the following steps:

[0029] First step, divide the weight tensor corresponding to the above-mentioned weights into blocks to obtain a set of weight blocks. Among them, first, the weight tensor corresponding to the above-mentioned weights can be converted into a one-dimensional array. After that, the above-mentioned one-dimensional array can be divided into blocks to obtain a set of weight blocks.

[0030] As an example, assume that there is a weight tensor of size b×h. The above-mentioned weight tensor can be converted into a one-dimensional data, and the size of the above-mentioned one-dimensional data is 1×(b×h). After that, the above-mentioned one-dimensional data can be divided into n weight blocks of size 1×d. Among them, b is the first dimension of the tensor, which usually refers to the batch size in model training. h is the second dimension of the tensor, which usually refers to the hidden dimension in model training. n is the number of weight blocks after division, n = (b×h) / d. d is the length of the weight block.

[0031] Second step, for each weight block in the above-mentioned set of weight blocks, perform the following quantization steps to generate a quantized weight block to obtain a set of quantized weight blocks:

[0032] First sub-step, determine the maximum value among the absolute values of each element in the above-mentioned weight block as the quantization maximum value. Among them, the above-mentioned quantization maximum value may be the largest numerical value among the absolute values of all elements in the above-mentioned weight block.

[0033] Second sub-step: Determine the ratio between the above quantization maximum value and a preset target quantization bit number as the quantization constant. Specifically, the ratio between the above quantization maximum value and the preset target quantization bit number can be determined as the quantization constant. The above target quantization bit number can be the bit number of elements in the quantized weight block, which can be 8 or 16.

[0034] Third sub-step: Determine the ratio between each element in the above weight block and the above quantization constant as the quantized element value, obtaining a set of quantized element values.

[0035] Fourth sub-step: Merge the various quantized element values in the above set of quantized element values to obtain a quantized weight block. Specifically, the various quantized element values in the above set of quantized element values can be merged according to the positions of the elements corresponding to each quantized element value in the weight block, obtaining a quantized weight block.

[0036] Third step: Determine the above set of quantized weight blocks as the quantized weight. Specifically, the various quantized weight blocks in the above set of quantized weight blocks can be merged to obtain the quantized weight.

[0037] In practice, by quantizing the weights of the large model to be trained, the precision of the weights can be reduced, thereby reducing the memory occupied by the weights.

[0038] Step 102: Downsample the large model to be trained to obtain an associated network, and load the associated network into a graphics processing unit.

[0039] In some embodiments, the above execution entity can downsample the above large model to be trained to obtain an associated network, and load the above associated network into a graphics processing unit.

[0040] In some optional implementation manners of some embodiments, the above execution entity downsampling the above large model to be trained to obtain an associated network, and loading the above associated network into a graphics processing unit may include the following steps:

[0041] First step: Based on a preset sampling rate, sample each neuron in each network layer of the above large model to be trained to generate lightweight network layers, obtaining a set of lightweight network layers. The above sampling rate can be the percentage of the number of neurons in the lightweight network layer to the number of neurons in the original network layer. Optionally, each neuron in each network layer of the above large model to be trained can be sampled through a preset downsampling algorithm to generate lightweight network layers, obtaining a set of lightweight network layers.

[0042] As an example, the above preset downsampling algorithm may include, but is not limited to, at least one of the following: Adapter algorithm, AdapterLinear algorithm.

[0043] In the second step, each lightweight network layer in the above lightweight network layer set is concatenated to obtain an associated network. Among them, the number of layers of the above associated network is the same as the number of layers of the large model to be trained.

[0044] In practice, the above associated network is a lightweight version of the large model to be trained, which is obtained by downsampling each layer of the large model. The number of layers (structure) of the associated network is the same as that of the large model, but the feature dimension of each layer is lower.

[0045] In the third step, the above associated network is loaded into the graphics processing unit, and the above associated network is initialized according to the weights of the large model to be trained.

[0046] As an example, assume there is a 3-layer neural network with 5 neurons in the first layer. First, the 5 neurons in the first layer can be reduced to 3 neurons through the above downsampling algorithm. Each layer of the above neural network is downsampled and then concatenated to obtain the corresponding associated network. Here, the number of layers (structure) of the above associated network is the same as that of the neural network, and the feature dimension of each layer is lower than that of the original neural network. Finally, the weight parameters corresponding to each neuron in the above associated network can be determined according to the weights of the neural network, and the initialization of the above associated network is completed.

[0047] Step 103, based on a preset number of iterations, perform the following training steps on the quantized large model and the associated network:

[0048] Step 1031, input the original data into the quantized large model to obtain a forward propagation vector set.

[0049] In some embodiments, the above execution entity can input the original data into the above quantized large model to obtain a forward propagation vector set. Among them, the above forward propagation vector set is a set composed of vectors output by each network layer of the above quantized large model.

[0050] In practice, the above forward propagation process is performed in the central processing unit, which can reduce the computational burden of the graphics processing unit and improve the efficiency of model fine-tuning training.

[0051] Step 1032, perform asynchronous heterogeneous training on the associated network according to the forward propagation vector set to obtain an updated associated network.

[0052] In some embodiments, the above execution entity can perform asynchronous heterogeneous training on the above associated network according to the above forward propagation vector set to obtain an updated associated network.

[0053] In some alternative implementations of some embodiments, the execution subject performs asynchronous heterogeneous training on the associated network according to the foregoing forward propagation vector set to obtain an updated associated network, which may include the following steps:

[0054] First step, according to the foregoing forward propagation vector set, output an associated network output vector through the associated network.

[0055] Optionally, the execution subject outputs an associated network output vector through the associated network according to the foregoing forward propagation vector set, which may include the following steps:

[0056] First sub-step, in response to determining that the current network layer is the first network layer of the associated network, use the forward propagation vector corresponding to the first one in the foregoing forward propagation vector set as the input of the first network layer of the associated network, and output an associated network forward propagation vector through the first network layer of the associated network.

[0057] Second sub-step, in response to determining that the current network layer is not the first network layer of the associated network, perform the following forward propagation steps:

[0058] Sub-step one, determine the forward propagation vector corresponding to the current layer in the foregoing forward propagation vector set as the first input vector of the current network layer.

[0059] Sub-step two, determine the associated network forward propagation vector output by the previous network layer of the current network layer as the second input vector of the current network layer.

[0060] Sub-step three, add the first input vector and the second input vector to obtain a fused input vector. Among them, the first input vector and the second input vector can be added through matrix addition to obtain a fused input vector.

[0061] Sub-step four, input the fused input vector into the current network layer to obtain an associated network current output vector.

[0062] Sub-step five, in response to determining that the current network layer is the last network layer of the associated network, determine the associated network current output vector as the associated network output vector.

[0063] Sub-step six, in response to determining that the current network layer is not the last network layer of the associated network, use the associated network current output vector as the associated network forward propagation vector, and perform the above forward propagation steps again.

[0064] In the second step, the output vector of the large model and the output vector of the above-mentioned associated network are fused to obtain a fused output vector. Among them, the output vector of the large model is the forward propagation vector corresponding to the last one in the above-mentioned forward propagation vector set. Here, the output vector of the large model and the output vector of the above-mentioned associated network can be fused through a preset fusion operation to obtain a fused output vector.

[0065] As an example, the above-mentioned fusion operation may include but is not limited to at least one of the following: element-wise addition (add), feature concatenation (concat), etc.

[0066] In the third step, the above-mentioned fused output vector is input into a preset output layer to obtain a training result. Among them, the above-mentioned output layer may be composed of a linear transformation layer (Linear Layer) and a Softmax layer. First, the above-mentioned fused output vector can be input into the above-mentioned linear transformation layer to obtain a linearly transformed vector. After that, the above-mentioned linearly transformed vector can be input into the above-mentioned Softmax layer to obtain a training result.

[0067] In practice, when the purpose of fine-tuning training is to generate a question-and-answer model, the above-mentioned original data may be question text sample data, and the above-mentioned training result may be the probability distribution of answer texts. The above-mentioned probability distribution of answer texts includes multiple answer texts and the probability values corresponding to the answer texts. When the purpose of fine-tuning training is to obtain a graphic drawing model, the above-mentioned original data may be description text sample data or image sample data, and the above-mentioned training result may be the probability distribution of images. The above-mentioned probability distribution of images includes multiple images and the probability values corresponding to the images.

[0068] In the fourth step, according to the loss value between the above-mentioned training result and the true value corresponding to the above-mentioned original data, backpropagation is performed on the above-mentioned associated network to obtain an updated associated network.

[0069] In practice, since the forward propagation of the large model to be trained is performed in the central processing unit, and the backpropagation step of the associated network is performed in the graphics processing unit, the large model to be trained and the associated network can be trained in parallel, improving the speed and efficiency of model training and reducing the consumption of computing resources and video memory resources during the training process.

[0070] Optionally, the above-mentioned execution subject performs backpropagation on the above-mentioned associated network according to the loss value between the above-mentioned training result and the true value corresponding to the above-mentioned original data to obtain an updated associated network, which may include the following steps:

[0071] In the first step, according to a preset loss function, the loss value between the above-mentioned training result and the true value corresponding to the above-mentioned original data is determined. Among them, the above-mentioned loss function may be a cross-entropy loss function (Cross-Entropy Loss). The above-mentioned true value may be the true value of the answer text sample or the true value of the image sample.

[0072] In practice, the purposes of fine-tuning training are different, and the data forms of the original data, training results, and true values are different, which are not specifically limited herein.

[0073] In the second step, according to the above loss value, perform backpropagation on the above-mentioned associated network to obtain an updated associated network. Among them, gradient descent can be performed on the above loss value to realize the update of the backpropagation weights of the above associated network, and an updated associated network can be obtained.

[0074] In practice, the structural diagram of the above asynchronous heterogeneous training is as Figure 2 shown. Among them, the solid arrows represent forward propagation, and the dashed arrows represent backpropagation.

[0075] The above step 1032 and its related content are an inventive point of the embodiment of the present disclosure, which solves the above-mentioned second technical problem of "causing a large consumption of video memory resources". The factors that cause the above technical problems are often as follows: The parameter-efficient fine-tuning method and the model quantization method still need to cache a large number of intermediate activation values, especially when processing long sequences or large batches of data. If the above factors are solved, the consumption of video memory resources can be reduced. To achieve this effect, first, the output of the large model to be trained is used as the input of the associated network, which reduces the computational amount of the associated network on the graphics processor, avoids the storage of unnecessary intermediate data, and thus reduces the burden on the graphics processor. Secondly, by fusing the forward propagation vector output by the large model to be trained with the forward propagation vector of the associated network itself, complex features can be better learned. Then, by fusing the output vectors of the large model and the associated network, the accuracy and generalization ability of the model can be improved. Finally, by only performing backpropagation in the graphics processor to update the associated network, parallel processing between the forward propagation of the large model and the forward and backpropagation of the associated network is realized, reducing the memory occupancy of the intermediate activation values and avoiding a large consumption of video memory resources. At the same time, when the user uses the large model fine-tuning training method of the present disclosure to obtain a trained large model for question-and-answer operations, the above-mentioned trained large model can process the question text submitted by the user in a timely manner, reducing the consumption of video memory resources and network traffic.

[0076] Step 1033, in response to determining that the current iteration number is not equal to the preset iteration number, use the updated associated network as the associated network to continue executing the above training steps;

[0077] In some embodiments, the above execution subject may, in response to determining that the current iteration number is not equal to the preset iteration number, use the above updated associated network as the associated network to continue executing the above training steps. Among them, the above preset iteration number may be 1000 times or 2000 times, which is not specifically limited herein.

[0078] Step 1034, in response to determining that the current iteration number is equal to the preset iteration number, determine the updated associated network as the trained associated network.

[0079] In some embodiments, the above-mentioned execution entity may, in response to determining that the current iteration number is equal to the preset iteration number, determine the above-mentioned updated associated network as the trained associated network.

[0080] Step 104, determine the quantized large model and the trained associated network as the trained large model.

[0081] In some embodiments, the above-mentioned execution entity may determine the above-mentioned quantized large model and the above-mentioned trained associated network as the trained large model. Among them, the above-mentioned quantized large model, the above-mentioned trained associated network, and the above-mentioned output layer may jointly serve as the trained large model. Here, the set of forward propagation vectors output by the above-mentioned quantized large model may serve as the first input to the corresponding layer of the above-mentioned trained associated network, and the output of each layer of the above-mentioned trained associated network serves as the second input to the next layer. After adding the first input and the second input, input them into the corresponding layer of the associated network. Then, fuse the outputs of the last layer of the above-mentioned quantized large model and the above-mentioned trained associated network and input them into the above-mentioned output layer. Finally, output the result through the above-mentioned output layer.

[0082] In practice, when a user uses a trained large model for question-and-answer operations or other operations, the technical problem often faced is that the trained large model often only uses a graphics processing unit to process user questions, and the response speed is slow, resulting in the consumption of network traffic. To avoid excessive consumption of network traffic, the following solution is proposed.

[0083] Optionally, the above-mentioned original data includes question text sample data, the above-mentioned true value includes answer text sample data, and the above-mentioned execution entity may further perform the following steps:

[0084] The first step is to obtain, from a terminal communicatively connected to the server for performing fine-tuning training of the large model, the question input by the user to be analyzed and answered by the server, and parse the question to generate a question text sequence. Among them, the above-mentioned question text sequence may be a sequence composed of at least one character or word.

[0085] The second step is to input the above-mentioned question text sequence into the above-mentioned trained large model and perform the following steps to obtain an answer text sequence:

[0086] The first sub-step is to input the above-mentioned question text sequence into the above-mentioned quantized large model to obtain a set of text vectors. Among them, the text vector corresponding to the last one in the above-mentioned set of text vectors is the text output vector. The above-mentioned quantized large model may be deployed in the central processing unit of the above-mentioned server.

[0087] The second sub-step is to output the associated network text output vector through the above-mentioned trained associated network according to the above-mentioned text vector set. Among them, the associated network text output vector can be output through the above-mentioned trained associated network according to the above-mentioned text vector set in the manner of the above-mentioned step 1032 or the above-mentioned step 104. This will not be elaborated here. The above-mentioned trained associated network can be deployed in the graphics processor of the above-mentioned server.

[0088] The third sub-step is to fuse the above-mentioned text output vector and the above-mentioned associated network text output vector to obtain a text fusion vector. Among them, the above-mentioned text output vector and the above-mentioned associated network text output vector can be fused through the fusion operation in the second step of the above-mentioned step 1032 to obtain a text fusion vector. This will not be elaborated here.

[0089] The fourth sub-step is to input the above-mentioned text fusion vector into the above-mentioned output layer to obtain an answer text sequence. Among them, the above-mentioned answer text sequence is a text sequence composed of at least one character or word. The above-mentioned answer text sequence can be a complete sentence or a paragraph.

[0090] The third step is to send the above-mentioned answer text sequence to the above-mentioned terminal so that the terminal can present the above-mentioned answer text sequence in the form of voice or text. Among them, the above-mentioned terminal can convert the above-mentioned answer text sequence into an answer voice or an answer text, and play or display the above-mentioned answer voice or answer text through a voice player or a display device communicatively connected to the terminal.

[0091] The above-mentioned optional steps and their related content of the present disclosure are an inventive point of the embodiments of the present disclosure, which solve the above-mentioned technical problem of "causing the consumption of network traffic". The factors leading to the above-mentioned technical problem are often as follows: the trained large model often only uses the graphics processor to process user questions, and the response speed is slow, resulting in the consumption of network traffic. If the above-mentioned factors are solved, the consumption of network traffic can be reduced. To achieve this effect, first, by deploying the quantized large model in the central server, the pressure on the graphics processor is reduced, so that the time for the model to analyze and calculate user questions can be reduced. Then, by fusing the text output vector obtained by the central processor and the associated network text output vector obtained in the graphics processor, a text fusion vector is obtained. Secondly, the answer text sequence corresponding to the text fusion vector is determined. Thus, while improving the accuracy of the answer text, the response speed of the answer can be guaranteed, and the consumption of network traffic can be reduced.

[0092] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the large model fine-tuning training method of some embodiments of the present disclosure, the waste of computing resources and video memory resources can be reduced. Specifically, the reasons for the waste of computing resources and video memory resources are as follows: The existing methods only regard the CPU as an auxiliary device and it is difficult to improve the efficiency of large model fine-tuning training, resulting in the waste of computing resources and video memory resources. Based on this, in the large model fine-tuning training method of some embodiments of the present disclosure, first, in response to detecting that the video memory occupancy of the server executing the large model fine-tuning training exceeds the preset occupancy condition, the weights of the large model to be trained are block-quantized to obtain quantized weights, and the above-mentioned large model to be trained and the above-mentioned quantized weights are loaded into the central processing unit as a quantized large model. Block quantization of the weights can reduce the memory occupancy during large model training, and loading the large model and the quantized weights into the CPU can efficiently store and process data, reduce the usage pressure of the GPU, and avoid video memory overflow. Next, the above-mentioned large model to be trained is downsampled to obtain an associated network, and the above-mentioned associated network is loaded into the graphics processing unit. Thus, a lightweight network model can be obtained, reducing the amount of calculation during the training process. After that, based on a preset number of iterations, the following training steps are performed on the above-mentioned quantized large model and the above-mentioned associated network: The original data is input into the above-mentioned quantized large model to obtain a forward propagation vector set. Wherein, the above-mentioned forward propagation vector set is a set composed of forward propagation vectors output by each network layer of the above-mentioned quantized large model. By performing forward propagation on the large model in the CPU, the waiting and idle time during the training process can be reduced, the utilization of hardware resources can be maximized, and the overall efficiency of model training can be improved. According to the above-mentioned forward propagation vector set, asynchronous heterogeneous training is performed on the above-mentioned associated network to obtain an updated associated network. The training of the associated network is guided by the forward propagation result of the large model, realizing parallel work between the CPU and the GPU, reducing resource waste during the training process, and improving resource utilization efficiency. In response to determining that the current number of iterations is not equal to the preset number of iterations, the above-mentioned updated associated network is used as the associated network to continue performing the above-mentioned training steps. In response to determining that the current number of iterations is equal to the preset number of iterations, the above-mentioned updated associated network is determined as the trained associated network. Finally, the above-mentioned quantized large model and the above-mentioned trained associated network are determined as the trained large model. Thus, the usage efficiency of computing resources and video memory resources during the large model fine-tuning training process can be improved, and the waste of computing and video memory resources can be reduced.

[0093] Further referring to Figure 3 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a large model fine-tuning training device. These device embodiments correspond to Figure 1 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0094] As shown inFigure 3 As shown in Figure 3 , the large model fine-tuning training device 300 of some embodiments includes: a block quantization unit 301, a downsampling unit 302, a training unit 303, and a determination unit 304. Among them, the block quantization unit 301 is configured to, in response to detecting that the video memory occupancy of the server executing the large model fine-tuning training exceeds a preset occupancy condition, perform block quantization on the weights of the large model to be trained to obtain quantized weights, and load the large model to be trained and the quantized weights into the central processing unit as a quantized large model; the downsampling unit 302 is configured to perform downsampling on the large model to be trained to obtain an associated network, and load the associated network into the graphics processing unit; the training unit 303 is configured to perform the following training steps on the quantized large model and the associated network based on a preset number of iterations: input the original data into the quantized large model to obtain a forward propagation vector set, where the forward propagation vector set is a set composed of forward propagation vectors output by each layer of the quantized large model; perform asynchronous heterogeneous training on the associated network according to the forward propagation vector set to obtain an updated associated network; in response to determining that the current iteration number is not equal to the preset iteration number, use the updated associated network as the associated network to continue performing the above training steps; in response to determining that the current iteration number is equal to the preset iteration number, determine the updated associated network as the trained associated network; the determination unit 304 is configured to determine the quantized large model and the trained associated network as the trained large model.

[0095] It can be understood that the various units described in the device 300 correspond to the respective steps in the method described with reference to Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 300 and the units included therein, and will not be repeated here.

[0096] Next, refer to Figure 4 , which shows a schematic structural diagram of an electronic device (for example, a computing device) 400 suitable for implementing some embodiments of the present disclosure. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.

[0097] As Figure 4As shown, the electronic device 400 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 401, which may perform various appropriate actions and processes according to a program stored in the read-only memory 402 or a program loaded from the storage device 408 into the random access memory 403. In the random access memory 403, various programs and data required for the operation of the electronic device 400 are also stored. The processing device 401, the read-only memory 402, and the random access memory 403 are connected to each other through a bus 404. The input / output interface 405 is also connected to the bus 404.

[0098] Generally, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 an electronic device 400 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 4 Each block shown in may represent one device or, as needed, multiple devices.

[0099] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through the communication device 409, or installed from the storage device 408, or installed from the read-only memory 402. When the computer program is executed by the processing device 401, the above functions defined in the methods of some embodiments of the present disclosure are executed.

[0100] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0101] In some embodiments, the client and the server may communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and may be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0102] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to detecting that the video memory occupancy of the server performing large model fine-tuning training exceeds a preset occupancy condition, perform block quantization on the weights of the large model to be trained to obtain quantized weights, and load the above large model to be trained and the above quantized weights into the central processing unit as a quantized large model; perform downsampling on the above large model to be trained to obtain an associated network, and load the above associated network into the graphics processing unit; based on a preset number of iterations, perform the following training steps on the above quantized large model and the above associated network: input the original data into the above quantized large model to obtain a forward propagation vector set, where the above forward propagation vector set is a set composed of forward propagation vectors output by each network layer of the above quantized large model; according to the above forward propagation vector set, perform asynchronous heterogeneous training on the above associated network to obtain an updated associated network; in response to determining that the current number of iterations is not equal to the preset number of iterations, use the above updated associated network as the associated network to continue performing the above training steps; in response to determining that the current number of iterations is equal to the preset number of iterations, determine the above updated associated network as the trained associated network; determine the above quantized large model and the above trained associated network as the trained large model.

[0103] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0104] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0105] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a block quantization unit, a downsampling unit, a training unit, and a determination unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the block quantization unit can also be described as "a unit for performing block quantization on the weights of the large model to be trained".

[0106] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0107] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with technical features having similar functions (but not limited to) disclosed in the embodiments of the present disclosure.

Claims

1. A large model fine-tuning training method, comprising: In response to detecting that the video memory occupancy of the server performing the large model fine-tuning training exceeds a preset occupancy condition, the weight of the large model to be trained is quantized in blocks to obtain quantized weights, and the large model to be trained and the quantized weights are loaded into a central processing unit as a quantized large model; Downsampling the large model to be trained to obtain a companion network, and loading the companion network into a graphics processor; Based on a preset number of iterations, the following training steps are performed on the quantized large model and the companion network: Inputting the original data into the quantized large model to obtain a forward propagation vector set, wherein the forward propagation vector set is a set consisting of forward propagation vectors output by each network layer of the quantized large model; According to the forward propagation vector set, asynchronous heterogeneous training is performed on the companion network to obtain an updated companion network; In response to determining that the current number of iterations is not equal to a preset number of iterations, continuing to perform the training step using the updated companion network as the companion network, and updating the number of iterations based on a preset value; In response to determining that the current number of iterations is equal to a preset number of iterations, determining the updated companion network as a trained companion network; The quantized large model and the trained companion network are determined as the trained large model.

2. The method according to claim 1, wherein: The original data includes question text sample data, the true value includes answer text sample data; and the method further includes: Obtaining, from a terminal in communication with a server that performs large model fine-tuning training, questions input by a user to be analyzed and answered by the server, and parsing the questions to generate a question text sequence; The question text sequence is input into the trained large model, and the following steps are performed to obtain the answer text sequence: Input the question text sequence into the quantitative large model to obtain a text vector set, wherein the text vector corresponding to the last one in the text vector set is a text output vector; Outputting a companion network text output vector through the trained companion network according to the text vector set; Fusing the text output vector with the associated network text output vector to obtain a text fusion vector; Input the text fusion vector into the output layer to obtain an answer text sequence; The answer text sequence is sent to the terminal so that the terminal presents the answer text sequence in the form of voice or text.

3. The method according to claim 1, wherein: The weights of the large model to be trained are quantized in blocks to obtain quantized weights, including: Dividing the weight tensor corresponding to the weight into blocks to obtain a weight block set; For each weight block in the weight block set, the following quantization steps are performed to generate a quantized weight block, thereby obtaining a quantized weight block set: Determine the maximum value among the absolute values ​​of the elements in the weight block as the quantization maximum value; Determine the ratio between the maximum quantization value and a preset target quantization bit number as a quantization constant; Determine the ratio between each element in the weight block and the quantization constant as a quantized element value to obtain a quantized element value set; Merging each quantized element value in the quantized element value set to obtain a quantized weight block; The quantized weight block set is determined as the quantized weight.

4. The method according to claim 1, wherein: The step of downsampling the large model to be trained to obtain a companion network, and loading the companion network into a graphics processor, includes: Based on a preset sampling rate, each neuron in each network layer in the large model to be trained is sampled to generate a lightweight network layer, thereby obtaining a lightweight network layer set; splicing the lightweight network layers in the lightweight network layer set to obtain a companion network; The companion network is loaded into a graphics processor, and the companion network is initialized according to the weights of the large model to be trained.

5. The method according to claim 1, wherein: The step of performing asynchronous heterogeneous training on the companion network according to the forward propagation vector set to obtain an updated companion network includes: Outputting a companion network output vector through the companion network according to the set of forward propagation vectors; Fusing the large model output vector and the companion network output vector to obtain a fused output vector, wherein the large model output vector is the last forward propagation vector corresponding to the forward propagation vector set; Inputting the fused output vector into a preset output layer to obtain a training result; According to the loss value between the training result and the true value corresponding to the original data, the companion network is back-propagated to obtain an updated companion network.

6. The method according to claim 5, wherein: The step of back-propagating the companion network according to the loss value between the training result and the true value corresponding to the original data to obtain an updated companion network includes: Determine the loss value between the training result and the true value corresponding to the original data according to a preset loss function; According to the loss value, the companion network is back-propagated to obtain an updated companion network.

7. A large model fine-tuning training device, comprising: A block quantization unit is configured to, in response to detecting that the video memory occupancy of the server performing the large model fine-tuning training exceeds a preset occupancy condition, block quantize the weight of the large model to be trained to obtain the quantized weight, and load the large model to be trained and the quantized weight into a central processing unit as a quantized large model; A downsampling unit is configured to downsample the large model to be trained to obtain a companion network, and load the companion network into a graphics processor; The training unit is configured to perform the following training steps on the quantized large model and the companion network based on a preset number of iterations: Inputting the original data into the quantized large model to obtain a forward propagation vector set, wherein the forward propagation vector set is a set consisting of forward propagation vectors output by each layer of the quantized large model; According to the forward propagation vector set, asynchronous heterogeneous training is performed on the companion network to obtain an updated companion network; In response to determining that the current number of iterations is not equal to the preset number of iterations, continuing to perform the training step using the updated companion network as the companion network; In response to determining that the current number of iterations is equal to a preset number of iterations, determining the updated companion network as a trained companion network; A determination unit is configured to determine the quantized large model and the trained companion network as the trained large model.

8. An electronic device, comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.