Model fine tuning method and device, electronic equipment and storage medium
By performing parameter decomposition and isometric block training on the multi-head attention layer of large language models, the problem of insufficient model stability and robustness after low-cost fine-tuning is solved, and the stability and robustness of the model are improved while reducing the computational complexity.
Patent Information
- Application Number
- CN202411716066.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-22
AI Technical Summary
The prior art is difficult to ensure the stability and robustness of large language models after fine-tuning on a low-cost basis.
By decomposing the multi-head attention layer of large language models, freezing the original weight matrix, and training the sample data is equidistantly blocked, and fine-tuning is used to use the learning weight matrix to reduce the computational complexity and training costs.
While reducing training costs, it ensures stability and robustness after fine-tuning of large language models.
Smart Images

Figure CN120354899A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of neural networks, and more specifically, to a method, apparatus, electronic device, and storage medium for fine-tuning a model. Background Art
[0002] With the rapid development of technology, neural networks are increasingly widely used in people's daily lives. Among them, large language models can be used to perform different tasks on text. In order to achieve different tasks, the large language model can be fine-tuned to be more suitable for the corresponding tasks. However, current methods for fine-tuning large language models cannot guarantee the stability and robustness of the large language model after fine-tuning on a low-cost basis. Therefore, how to ensure the stability and robustness of the large language model after fine-tuning on a low-cost basis has become an urgent problem to be solved. Summary of the Invention
[0003] In view of the above problems, embodiments of this application propose a method, apparatus, electronic device, and storage medium for fine-tuning a model to improve the above problems.
[0004] According to one aspect of the embodiments of this application, a method for fine-tuning a model is provided, including: determining a network to be adjusted in the multi-head attention layer of a large language model to be adjusted, and obtaining a data set for fine-tuning the large language model, where the data set includes a plurality of sample data; determining a first weight matrix corresponding to the multi-head attention layer and a second weight matrix corresponding to the network to be adjusted, and performing parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix; equally spacing and dividing the strings corresponding to the plurality of sample data to obtain a plurality of character blocks; freezing the original weight matrix, and training the large language model to be adjusted through the plurality of character blocks to obtain a training result; and updating the learning weight matrix according to the training result to obtain an adjusted large language model.
[0005] In some embodiments, the plurality of character blocks include a plurality of first character blocks and a plurality of second character blocks. The equally spacing and dividing the strings corresponding to the plurality of sample data to obtain a plurality of character blocks includes: determining a block window length, and equally spacing and dividing the strings corresponding to the plurality of sample data according to the block window length to obtain the plurality of first character blocks; determining a target length according to the block window length, and dividing the plurality of first character blocks according to the target length to obtain a plurality of sub-character blocks; and combining the plurality of sub-character blocks to obtain the plurality of second character blocks.
[0006] In some embodiments, the combining of the plurality of sub-character blocks to obtain the plurality of second character blocks includes: combining adjacent sub-character blocks among the plurality of sub-character blocks and combining the first and last character blocks among the plurality of sub-character blocks according to the order of the plurality of first character blocks, to obtain the plurality of second character blocks.
[0007] In some embodiments, the training of the large language model to be adjusted by the plurality of character blocks to obtain a training result includes: respectively performing attention calculation on the plurality of first character blocks to obtain first attention scores corresponding to the plurality of first character blocks respectively; respectively performing attention calculation on the plurality of second character blocks to obtain second attention scores corresponding to the plurality of second character blocks respectively; determining a target attention score based on the first attention scores and the second attention scores, and obtaining the training result according to the target attention score.
[0008] In some embodiments, the determining of the chunk window length includes: determining the character lengths corresponding to the plurality of sample data respectively; determining the maximum character length among the character lengths corresponding to the plurality of sample data respectively as the context length; and determining the chunk window length according to the context length.
[0009] In some embodiments, the determining of the network to be adjusted in the multi-head attention layer of the large language model to be adjusted includes: determining the fine-tuning task of the large language model to be adjusted; and determining the network to be adjusted according to the fine-tuning task.
[0010] In some embodiments, the updating of the learning weight matrix according to the training result to obtain the adjusted large language model includes: determining the sample output labels corresponding to the plurality of sample data respectively, and determining the target output labels corresponding to the plurality of sample data respectively according to the training result; determining a loss value according to the target output labels and the sample output labels; performing parameter update on the learning weight matrix according to the loss value, and obtaining the adjusted large language model according to the learning weight matrix after parameter update.
[0011] In some embodiments, the determining of the loss value according to the target output labels and the sample output labels includes: determining the fine-tuning task of the large language model to be adjusted; determining a loss function according to the fine-tuning task; and determining the loss value of the loss function based on the loss function according to the target output labels and the sample output labels.
[0012] According to one aspect of the embodiments of the present application, there is provided a fine-tuning device for a model, including: a network to be adjusted determination module, configured to determine a network to be adjusted in the multi-head attention layer of a large language model to be adjusted, and obtain a data set for fine-tuning the large language model, where the data set includes a plurality of sample data; a parameter decomposition module, configured to determine a first weight matrix corresponding to the multi-head attention layer and a second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learned weight matrix and an original weight matrix; a chunking module, configured to perform equidistant chunking on the strings corresponding to the plurality of sample data respectively to obtain a plurality of character chunks; a training module, configured to freeze the original weight matrix, and train the large language model to be adjusted through the plurality of character chunks to obtain a training result; an adjustment module, configured to update the learned weight matrix according to the training result to obtain an adjusted large language model.
[0013] According to one aspect of the embodiments of the present application, there is provided an electronic device, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above-mentioned fine-tuning method of the model is implemented.
[0014] According to one aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the above-mentioned fine-tuning method of the model is implemented.
[0015] In the solution of the present application, first, the network to be adjusted in the multi-head attention layer of the large language model to be adjusted is determined, so as to determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain the original weight matrix and the learned weight matrix. Then, equidistant chunking is performed on the strings corresponding to the data set for fine-tuning the large language model, which includes a plurality of sample data, to obtain a plurality of character chunks. In this way, the large language model to be adjusted can be trained through the plurality of character chunks while freezing the original weight matrix to obtain a training result. Finally, the learned weight matrix is updated according to the training result to obtain an adjusted large language model. By chunking a plurality of strings, the computational complexity corresponding to the long context of the input sample data is reduced, and on the basis of reducing the training cost of the large language model to be adjusted, the stability and robustness of the large language model after fine-tuning are ensured.
[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings
[0017] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other accompanying drawings based on these drawings without creative efforts.
[0018] Figure 1 is a flowchart of a method for fine-tuning a model shown according to an embodiment of this application;
[0019] Figure 2 is a schematic flowchart of a method for fine-tuning a model shown according to another embodiment of this application;
[0020] Figure 3 is a specific step flowchart of step S240 shown according to an embodiment of this application;
[0021] Figure 4 is a specific step flowchart of step S260 shown according to an embodiment of this application;
[0022] Figure 5 is a schematic diagram of calculating attention scores for a first character block and a second character block shown according to an embodiment of this application;
[0023] Figure 6 is a schematic flowchart of a method for fine-tuning a model shown according to still another embodiment of this application;
[0024] Figure 7 is a schematic flowchart of a method for fine-tuning a model shown according to yet another embodiment of this application;
[0025] Figure 8 is a schematic flowchart of a method for fine-tuning a model shown according to yet another embodiment of this application;
[0026] Figure 9 is a block diagram of a device for fine-tuning a model shown according to an embodiment of this application;
[0027] Figure 10 shows a block diagram of an electronic device for executing a scrambling method of a model according to an embodiment of this application;
[0028] Figure 11 shows a storage unit for storing or carrying program code for implementing a scrambling method of a model according to an embodiment of this application. Detailed implementation manners
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this application will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0030] In addition, the described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of this application. However, those skilled in the art will realize that the technical solutions of this application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of this application.
[0031] The block diagrams shown in the drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0032] The flowcharts shown in the drawings are merely illustrative and do not necessarily include all the content and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0033] Currently, large language models (LLMs) are usually trained with a predefined context length. For example, LLaMA supports 2048 tokens, and LLaMA-2 supports 4096 tokens. In some increasingly common scenarios, LLMs need to summarize long documents or combine multiple rounds of historical conversations for interactive responses, etc., which poses new challenges to the context length supported by LLMs. To enable LLMs to adapt to longer context lengths, some existing methods retrain LLMs with longer sequence inputs or fine-tune existing LLMs, but this requires a large amount of computing resources and results in high training costs. Another approach is to fine-tune pre-trained LLMs by modifying the linear projection layer in the self-attention module with low-rank matrices using LoRA. Although this reduces the number of trainable parameters, simple low-rank adaptation fine-tuning leads to high perplexity in the model after long context expansion. In addition, since LoRA still uses the standard self-attention calculation mechanism with an operation complexity of O(n^2), the computational cost and training time will still increase as the context scale expands. Other solutions use memory mechanisms to compress past inputs to find tokens relevant to the current input. However, since the attention calculation obtained by compression is very different from the standard attention mechanism, this approach limits the possibility of fine-tuning pre-trained LLMs.
[0034] To address the above problems, the inventors have found through long-term research and proposed a model fine-tuning method, device, electronic device, and storage medium provided in the embodiments of this application. By chunking multiple strings, the computational complexity corresponding to the long context of the input sample data is reduced, and on the basis of reducing the training cost of the large language model to be adjusted, the stability and robustness of the large language model after fine-tuning are ensured.
[0035] To better understand the solutions of the embodiments of this application, the following first explains the technical terms used in the embodiments of this application.
[0036] LLaMA (Large Language Model Meta AI): Large Language Model Meta AI, a series of large language models released by Meta AI starting from February 2023.
[0037] Transformer: A deep learning architecture that relies on parallel multi-head attention mechanisms, proposed by Ashish Vaswani et al. of the Google Brain team in 2017. It requires less training time than previous recurrent neural network architectures, and its variants have been widely used to train large language models.
[0038] Large Language Model (LLM): A language model composed of artificial neural networks with billions or more parameters, trained on a large amount of unlabeled text using self-supervised learning or semi-supervised learning.
[0039] Low-Rank Adaptation (LoRA): Fine-tune the model by freezing the weights of the pre-trained model and injecting trainable layers into each Transformer block, greatly reducing the computational cost of the fine-tuned model.
[0040] The embodiments related to the present application will be described below in conjunction with the accompanying drawings.
[0041] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a method for fine-tuning a model provided by an embodiment of the present application. By chunking multiple strings, the method for fine-tuning the model reduces the computational complexity corresponding to the long context of the input sample data, and on the basis of reducing the training cost of the large language model to be adjusted, ensures the stability and robustness of the large language model after fine-tuning. In a specific embodiment, the method for fine-tuning the model can be applied to a fine-tuning device 200 of a model as shown in Figure 9 and an electronic device 100 configured with the fine-tuning device 200 of the model ( Figure 10 ). The following will elaborate in detail on the Figure 1 shown process. The method for fine-tuning the model may specifically include the following steps:
[0042] Step S110, determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, and obtain a data set for fine-tuning the large language model, where the data set includes multiple sample data.
[0043] As a way, since large language models can perform a wide range of tasks, including text summarization, translation, sentiment analysis, etc., by training on a large amount of text data. Usually, the training process of large language models includes two main steps: pre-training and fine-tuning. Among them, in the pre-training stage, the model learns from a huge and diverse dataset, usually containing billions of words from different sources, such as websites, books, and articles. This stage allows the model to learn general language patterns and representations; in the fine-tuning stage, the model is further trained on a more specific and smaller dataset related to the target task or domain. This helps the model fine-tune its understanding and adapt to the special requirements of the task. Therefore, the corresponding task can be achieved by fine-tuning the pre-trained large language model. Optionally, the user can customize the network to be adjusted corresponding to the large language model to be adjusted, and generate an identification to be adjusted for the set network to be adjusted, so as to determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted by obtaining the identification to be adjusted.
[0044] Optionally, the dataset corresponding to the sample data for fine-tuning can be set in advance according to the fine-tuning task of the large language model to be adjusted, so as to facilitate the adjustment of the large language model based on the dataset corresponding to the specific fine-tuning task.
[0045] Step S120, determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain the learning weight matrix and the original weight matrix.
[0046] As a way, each head attention layer in the multi-head attention layer can use different linear transformations to learn different dimensions of features in the input data, so as to increase the learning ability of the model, capture the diversity of data from different angles, and enhance the model's understanding and generalization ability for complex sequence tasks. There are multiple neural networks in the multi-head attention layer, and each neural network has its own weight parameters. The weight parameters of each head attention layer are composed of the weight parameters of the neural network corresponding to this layer. Therefore, the first weight matrix can be obtained by obtaining the weight parameters corresponding to each neural network of each head attention layer.
[0047] Optionally, the neural network to be adjusted in the multi-head attention layer can be set in advance, and the weight parameters corresponding to the neural network to be adjusted are marked, and the weight corresponding to the network to be adjusted also has this mark accordingly. Therefore, after the first weight matrix is determined, the weight parameters with marks in the first weight matrix are determined to determine the second weight matrix.
[0048] Optionally, to avoid overfitting of the corresponding parameters caused by the participation of weight parameters unrelated to the adjustment task in the training process during the adjustment of the large language model to be adjusted, resulting in low performance of the trained model. Therefore, when the first weight matrix and the second weight matrix are determined, parameter decomposition can be performed according to the first weight matrix and the second weight matrix to obtain the original weight matrix that does not need to participate in the training and the learning weight matrix that needs to participate in the training.
[0049] Optionally, to avoid the problem of low training efficiency caused by a large number of parameters in the learning weight matrix resulting in a parameter matrix with a large rank, the learning weight matrix can be split into multiple matrices with lower ranks, and then the weight matrices with lower ranks can be adjusted during the training process. Optionally, the learning weight matrix can be disassembled according to the method of matrix multiplication. For example, if the first weight matrix is W ∈ R d×k , where d and k are the rows and columns corresponding to the first weight matrix respectively, and the second weight matrix corresponding to the network to be adjusted is ΔW. Thus, it can be determined that the first weight matrix is decomposed into W1 + ΔW, where W1 is the original weight matrix and ΔW is the learning weight matrix. Then, a target value is determined and used as the rank of the decomposition of the learning weight matrix. Then, based on this rank, the learning weight matrix is decomposed into low rank, that is, W = W1 + ΔW = W1 + BA, where B ∈ R d×r , A ∈ R r×k , r is the determined rank, and r is much smaller than d and k. Optionally, the value of r can be 2 or other values, as long as it is ensured that r is much smaller than d and k. In this way, the large language model to be adjusted can be trained and fine-tuned with less resource consumption.
[0050] Step S130: Perform equidistant partitioning on the strings corresponding to the multiple sample data to obtain multiple character blocks.
[0051] As a method, during the process of fine-tuning the large language model to be adjusted through the multi-head attention layer, it is necessary to calculate the attention for each input token and all the remaining tokens, and its complexity is O(n 2) Thus, when training based on sample data, the memory cost is high and the speed is slow in the application of long sequences. Therefore, multiple sample data can be chunked by using a fixed-length window for their respective corresponding strings to obtain multiple character chunks of the same length, where the toke lengths corresponding to the multiple character chunks are consistent, or the length of the last character chunk is different from that of the other character chunks. For example, for the case where the length of the input string is 6144 tokens, a window length of 2048 is selected as the chunking length, and it can be divided into three character chunks of length 2048. In this way, the large language model to be adjusted is trained with multiple shorter strings to improve the training efficiency and reduce the storage pressure of the model.
[0052] Step S140: Freeze the original weight matrix and train the large language model to be adjusted with the multiple character chunks to obtain a training result.
[0053] As a method, to avoid the reduction in the accuracy of the large language model to be adjusted caused by overfitting of parameters other than the weight parameters corresponding to the network to be adjusted due to all weight parameters participating in training, the weight parameters corresponding to the network that do not need to be adjusted can be frozen. In this way, without sacrificing the model performance, the parameter update rate is reduced, the training speed is increased, and the robustness and stability of the model are improved.
[0054] Optionally, the large language model to be adjusted can be used to calculate the attention scores corresponding to the multiple character chunks respectively, then calculate the weights corresponding to the input string according to the attention scores corresponding to the multiple character chunks, and then multiply the weights corresponding to the input string by the recognition value corresponding to the input string to obtain an output result. Among them, the attention scores (Attention Scores) are indicators used to measure the importance degree of different elements in the sequence to the current element. In natural language processing, when the model needs to decide how much "attention" to give to a certain token in the sequence, it will calculate the attention scores between this token and other tokens, and these scores indicate the influence degree of each token in the sequence on the output result when generating the output.
[0055] Optionally, the training result can be the recognition result corresponding to the language processing of the input long string. For example, the translation text for translating long characters, the semantic recognition result for semantic recognition of the string, and the answer result for intelligent answering according to the string, etc. The specific training result is related to the specific execution task of the large language model to be adjusted, and no specific limitation is made here. It can be limited according to actual needs.
[0056] Step S150: Update the learning weight matrix according to the training result to obtain an adjusted large language model.
[0057] As a way, after determining the training result corresponding to the large language model to be adjusted, first determine the loss function related to the execution task corresponding to the large language model to be adjusted, so as to determine the loss value of the loss function based on the training result, and then determine whether the loss function converges according to the loss value. Furthermore, when it is determined that the loss function converges, end the training of the large language model to be adjusted to obtain an adjusted large language model; if it is determined that the loss function does not converge, adjust the learning weight parameters according to the loss value, update the learning weight matrix accordingly, and perform corresponding recognition or processing tasks based on the updated learning weight matrix and the frozen original weight matrix to obtain a training result, and then determine whether the loss function converges based on this training result until the loss function converges and stop training the large language model to be adjusted to obtain an adjusted large language model.
[0058] In the embodiment of the present application, first determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, so as to determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain the original weight matrix and the learning weight matrix. Then, equally divide the corresponding strings in the dataset for fine-tuning the large language model, which includes multiple sample data, into multiple character blocks, so that the large language model to be adjusted can be trained through multiple character blocks while freezing the original weight matrix to obtain a training result. Finally, update the learning weight matrix according to the training result to obtain an adjusted large language model. This solution divides multiple strings into blocks, reduces the computational complexity corresponding to the long context of the input sample data, and ensures the stability and robustness of the large language model after fine-tuning while reducing the training cost of the large language model to be adjusted.
[0059] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of a method for fine-tuning a model provided by an embodiment of the present application. The following will elaborate in detail on the Figure 2 flow shown. The method for fine-tuning the model may specifically include the following steps:
[0060] Step S210: Determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, and obtain a dataset for fine-tuning the large language model, where the dataset includes multiple sample data.
[0061] Step S220: Determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition based on the first weight matrix and the second weight matrix to obtain a learned weight matrix and an original weight matrix.
[0062] Step S230: Determine the block window length, and perform equidistant blocking on the strings corresponding to the multiple sample data according to the block window length to obtain the multiple first character blocks.
[0063] As a way, the user can pre-customize and set the block window length in advance and store the block window length in the electronic device, so that the block window length can be obtained from the storage device of the electronic device. Optionally, it can also be determined according to the string lengths of the multiple sample data in the dataset. For example, it can be determined based on the length of the longest string among the multiple sample data.
[0064] Optionally, after determining the block window length, perform equidistant blocking on the strings corresponding to each sample data in the dataset according to the block window length, so as to obtain multiple first character blocks.
[0065] Step S240: Determine the target length according to the block window length, and perform blocking on the multiple first character blocks according to the target length to obtain multiple sub-character blocks.
[0066] As a way, the target length is used to perform further blocking on the first character blocks. The half of the block window length can be determined as the target length, or other block coefficient values can be determined first, and then the block coefficient is multiplied by the block window length to obtain the target length.
[0067] Optionally, since the target length is less than the block window length, the multiple first character blocks can be directly blocked according to the target length to obtain multiple sub-character blocks.
[0068] As another way, the multiple first character blocks can also be blocked by the strategy of shifting the block window forward, where the moving length can be half of the block window length. For example, taking a context length of 6144 tokens, the starting point of the first block window is shifted to 1024. Then the first attention sub-block starts from the 1025th token and ends at the 3072nd token, and so on. The first 1024 tokens and the last remaining 1024 tokens are divided into the same group.
[0069] In some embodiments, as Figure 3 shown, step S240 includes:
[0070] Step S241, determine the character length corresponding to each of the multiple sample data.
[0071] As a way, to avoid the inconvenient calculation of attention scores due to inconsistent lengths of the chunks of multiple sample data, which may lead to a decrease in the training efficiency of the large language model to be adjusted, the character length corresponding to each of the multiple sample data can be determined in advance, and based on the character length corresponding to each sample data, the chunk window length can be determined. Optionally, the character length corresponding to each of the multiple sample data can be the number of tokens corresponding to each of them.
[0072] Step S242, determine the maximum character length among the character lengths corresponding to each of the multiple sample data as the context length.
[0073] As a way, to avoid the reduction in the accuracy of the large language model to be adjusted caused by the inability of some inputs in the sample data to participate in training due to the chunk window length, the character length corresponding to each of the multiple sample data can be determined first, and the maximum character length can be determined as the context length. For example, among the multiple sample data, there are character lengths of 6114, 3378, 2048, etc., and 6114 is determined as the context length.
[0074] Step S243, determine the chunk window length according to the context length.
[0075] As a way, to avoid the high memory cost and slow task execution speed caused by the large language model to be adjusted processing long strings, the context length can be evenly divided to determine the chunk window length. Optionally, evenly dividing the context length can be done by first determining the number of chunks and then performing the division. Among them, the number of chunks can be set by the user, or can be intelligently determined based on the context length, which is not specifically limited here and can be set according to actual needs.
[0076] Please continue to refer to Figure 2 , step S250, combine the multiple sub-character chunks to obtain the multiple second character chunks.
[0077] As a way, to introduce information interaction between different chunk windows and reduce model perplexity, multiple sub-strings can be combined into character chunks with the same length as the first character chunk, thereby obtaining multiple second character chunks. Optionally, the sub-character chunks that make up the second character chunks can be from different first character chunks, which ensures information interaction between different sub-character chunks and improves the accuracy of the large language model to be adjusted.
[0078] In some embodiments, step S250 includes: combining adjacent sub-character blocks among the multiple sub-character blocks and combining the first and last character blocks among the multiple sub-character blocks according to the order of the multiple first character blocks to obtain the multiple second character blocks.
[0079] As a way, to avoid the problem of chaotic information interaction caused by randomly combining multiple sub-character blocks, the order of the multiple sub-character blocks can be determined in sequence according to the order of the first character blocks for chunking, and then the adjacent sub-character blocks that are the first character blocks among the multiple sub-character blocks are combined to obtain the second character blocks. Since there are uncombined sub-character blocks in the first character block at the head of the order and the first character block at the tail of the order, the first and last character blocks can be combined to ensure the integrity of the second character blocks.
[0080] Step S260: Freeze the original weight matrix and train the large language model to be adjusted with the multiple character blocks to obtain a training result.
[0081] In some embodiments, as Figure 4 shown, step S260 includes:
[0082] Step S261: Calculate the attention for each of the multiple first character blocks respectively to obtain the first attention scores corresponding to the multiple first character blocks.
[0083] As a way, the attention score is an important concept in the attention mechanism for measuring the similarity between a query and a key. In deep learning and natural language processing, the attention score is usually used to calculate the attention weights, and then determine the importance of each value in the attention pooling process. The attention scores corresponding to the multiple first character blocks can be calculated respectively according to the attention scoring function to obtain the first attention scores corresponding to the multiple first character blocks.
[0084] Step S262: Calculate the attention for each of the multiple second character blocks respectively to obtain the second attention scores corresponding to the multiple second character blocks.
[0085] As a way, the attention scores corresponding to the multiple second character blocks can be calculated respectively according to the attention scoring function to obtain the second attention scores corresponding to the multiple second character blocks, where the attention scoring function for calculating the second attention scores can be the same function as the attention scoring function for calculating the first attention scores.
[0086] Step S263: Determine a target attention score based on the first attention score and the second attention score, and obtain the training result according to the target attention score.
[0087] As a way, in order to achieve information flow interaction between different attention sub-blocks without increasing additional computational costs, add up the attention scores corresponding to all character blocks to obtain the total attention score corresponding to each sample data, and then adjust the large language model to be adjusted based on the total attention score corresponding to each sample data. As Figure 5 shown, obtain the total attention score by adding up the first attention scores corresponding to all first character blocks and the second attention scores corresponding to all second character blocks.
[0088] Step S270: Update the learning weight matrix according to the training result to obtain the adjusted large language model.
[0089] Among them, for the specific step descriptions of steps S210 - S220 and step S270, refer to steps S110 - step S120 and step S150, which will not be elaborated here.
[0090] In this embodiment, by chunking the strings corresponding to multiple sample data according to the determined chunk window length to obtain multiple first character blocks, then chunking the multiple first strings again according to the target length determined by the chunk window length to obtain multiple sub-character blocks, and then determining multiple second character blocks based on the multiple sub-character blocks, it is possible to train the large language model to be adjusted based on the first attention score of the first character blocks and the second attention score of the second character blocks, thereby ensuring the information flow between attention blocks, fully retaining the attention interaction results, ensuring the model inference performance, and greatly reducing the training optimization cost.
[0091] Please refer to Figure 6 , Figure 6 which shows a schematic flowchart of the fine-tuning method of the model provided in an embodiment of the present application. The following will elaborate in detail on the Figure 6 flow shown. The fine-tuning method of the model may specifically include the following steps:
[0092] Step 310: Determine the fine-tuning task of the large language model to be adjusted.
[0093] As a way, the user can customize the fine-tuning task of the large language model to be adjusted, and then the electronic device generates a task identifier based on this fine-tuning task, and then determines the fine-tuning task of the large language model to be adjusted based on this task identifier, where the mapping relationship between different task identifier words and the corresponding fine-tuning tasks is preset in advance.
[0094] Optionally, it can also be determined by a dataset for a fine-tuning task to determine the fine-tuning task of the large language model to be adjusted, where the dataset includes multiple sample data and the results of the multiple sample data based on the fine-tuning task, and thus the fine-tuning task of the large language model to be adjusted is determined based on the results of the fine-tuning task in the dataset.
[0095] Step S320: Determine the network to be adjusted according to the fine-tuning task, and obtain a dataset for fine-tuning the large language model, where the dataset includes multiple sample data.
[0096] As a way, different fine-tuning tasks correspond to different networks to be adjusted. Different networks to be adjusted corresponding to different fine-tuning tasks can be preset in advance, so that after the fine-tuning task of the large language model to be adjusted is determined, the network to be adjusted can be determined based on the fine-tuning task.
[0097] Step S330: Determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix.
[0098] Step S340: Perform equidistant chunking on the strings corresponding to the multiple sample data respectively to obtain multiple character chunks.
[0099] Step S350: Freeze the original weight matrix, and train the large language model to be adjusted through the multiple character chunks to obtain a training result.
[0100] Step S360: Update the learning weight matrix according to the training result to obtain the adjusted large language model.
[0101] Among them, for the specific step descriptions of steps S330 - S360, please refer to steps S120 - S150, and will not be elaborated here.
[0102] In this embodiment, by determining the fine-tuning task of the large language model to be adjusted to determine the network to be adjusted, the original weight matrix and the learning weight matrix can be determined based on the network to be adjusted, so that the learning weight matrix is adjusted in the case of freezing the original weight matrix, realizing the fine-tuning of the large language model to be adjusted, improving the efficiency of fine-tuning the large language model to be adjusted, and avoiding overfitting caused by training the original weight matrix.
[0103] Please refer to Figure 7 , Figure 7 shows a schematic flowchart of a method for fine-tuning a model provided in an embodiment of the present application. Next, it will be directed to Figure 7The following is a detailed description of the process shown. The method for fine-tuning the model can specifically include the following steps:
[0104] Step S410: Determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, and obtain a dataset for fine-tuning the large language model, where the dataset includes multiple sample data.
[0105] Step S420: Determine the first weight matrix corresponding to the multi-head attention layer and the second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix.
[0106] Step S430: Perform equidistant chunking on the strings corresponding to the multiple sample data to obtain multiple character chunks.
[0107] Step S440: Freeze the original weight matrix, and train the large language model to be adjusted with the multiple character chunks to obtain a training result.
[0108] Among them, the specific step descriptions of steps S410 - S440 can refer to steps S110 - step S140, and will not be elaborated here.
[0109] Step S450: Determine the sample output labels corresponding to the multiple sample data, and determine the target output labels corresponding to the multiple sample data according to the training result.
[0110] As a way, in order to better train the large language model to be adjusted, it is necessary to determine the loss value for training the sample data, so that the weight parameters of the large language model to be adjusted can be adjusted based on the loss value, and the large language model to be adjusted can be adjusted. The loss value can be determined by the sample outputs corresponding to the multiple sample data and the training outputs after training the large language model to be adjusted with the multiple sample data. Optionally, in order to reduce the storage pressure of the sample data on the memory among the multiple sample data, the sample output labels indicating the sample data for the fine-tuning task can be associated and stored with the sample data, so that the output results of the sample data can be determined by obtaining the sample output labels corresponding to the multiple sample data.
[0111] Optionally, in order to better determine the loss value, the target output labels for indicating the target output results can be generated based on the target output results corresponding to the multiple sample data in the training result.
[0112] Step S460: Determine the loss value according to the target output label and the sample output label.
[0113] As a way, after determining the target output label and the sample output label, the loss value can be determined according to the loss function corresponding to the large language model to be adjusted. This loss function is related to the fine-tuning task of the large language model to be adjusted, and the loss function can be set according to the fine-tuning task of the large language model to be adjusted.
[0114] In some embodiments, the step S460 includes: determining the fine-tuning task of the large language model to be adjusted; determining the loss function according to the fine-tuning task; and determining the loss value of the loss function based on the loss function according to the target output label and the sample output label.
[0115] As a way, the corresponding relationship between the loss functions corresponding to different fine-tuning tasks can be preset in advance. In this way, after determining the fine-tuning task of the large language model to be adjusted, the loss function can be determined through the corresponding relationship and the fine-tuning task.
[0116] Optionally, after determining the loss function, the loss value can be calculated by substituting the vector or value corresponding to the sample output label and the vector or value corresponding to the target output label into the value loss function.
[0117] Step S470, update the parameters of the learning weight matrix according to the loss value, and obtain the adjusted large language model according to the learning weight matrix after parameter update.
[0118] As a way, it can be determined whether to end the training of the large language model to be adjusted to obtain the adjusted large language model by determining whether the loss function converges according to the loss value. Optionally, if it is determined that the loss function converges, end the training of the large language model to be adjusted to obtain the adjusted large language model; if it is determined that the loss function does not converge, adjust the learning weight parameters according to the loss value, update the learning weight matrix accordingly, and perform corresponding recognition or processing tasks based on the updated learning weight matrix and the frozen original weight matrix to obtain the training result, and then determine whether the loss function converges based on the training result until the loss function converges and stop training the large language model to be adjusted to obtain the adjusted large language model.
[0119] In this embodiment, the loss value of the loss function corresponding to the fine-tuning task of the large language model to be adjusted is determined according to the sample output labels corresponding to multiple sample data and the target output labels of the multiple sample data corresponding to the training result. In this way, the learning weight matrix can be updated according to the loss value, so as to ensure the accuracy of the adjusted large language model.
[0120] Figure 8 It is a schematic flowchart of a method for fine-tuning a model shown in an embodiment of the present application, as Figure 8As shown, first select an instruction fine-tuning dataset. The instruction fine-tuning dataset includes multiple sample data. Each sample (x i , y i ) contains input text x i and the corresponding output label y i . Then determine the context length based on the multiple sample data, and determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted that needs to be fine-tuned. And based on the weight parameter matrix corresponding to the network to be adjusted and the weight parameter matrix corresponding to the multi-head attention layer, the original weight matrix and the learning weight matrix are determined respectively. Among them, the learning weight matrix can be a matrix obtained by multiplying multiple low-rank matrices. Then freeze the original weight matrix. Among them, the learning weight matrix can be AB, where B ∈ R d×r , A ∈ R r×k , r is the rank of the learning weight matrix, d and r are the rows and columns of B, and r and k are the rows and columns of A.
[0121] Then, the input sample data is tokenized by a pre-trained LLaMa tokenizer, processed into tokens of length d, and used as the input of the large language model to be adjusted as shown Figure 8 . In the multi-head attention mechanism, feature extraction is performed according to the tokens of length d of the input. During feature extraction, the multiple sample data are divided into blocks to obtain multiple first character blocks and multiple second character blocks, and the first attention scores of the multiple first character blocks and the second attention scores of the multiple second character blocks are calculated respectively. Then, feature extraction is performed according to the first attention score and the second attention score, and then through the feed-forward layer calculation of the Transformer, the output of the model is obtained.
[0122] Finally, based on the output of the model, the loss value is determined according to the loss function corresponding to the fine-tuning task of the large language model to be adjusted, and then gradient backpropagation is performed according to the loss value to optimize the weight parameters in the learning weight matrices A and B. Among them, different fine-tuning tasks correspond to different loss functions. For example, for a classification task, the corresponding loss function can be a cross-entropy loss function.
[0123] Figure 9 is a model fine-tuning device shown according to an embodiment of the present application. As shown Figure 9 , the model fine-tuning device 200 includes: a network to be adjusted determination module 210, a parameter decomposition module 220, a block division module 230, a training module 240, and an adjustment module 250.
[0124] The network to be adjusted determination module 210 is configured to determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, and obtain a data set for fine-tuning the large language model, where the data set includes a plurality of sample data; the parameter decomposition module 220 is configured to determine a first weight matrix corresponding to the multi-head attention layer and a second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix; the chunking module 230 is configured to equally distance chunk the strings corresponding to the plurality of sample data to obtain a plurality of character chunks; the training module 240 is configured to freeze the original weight matrix, and train the large language model to be adjusted through the plurality of character chunks to obtain a training result; the adjustment module 250 is configured to update the learning weight matrix according to the training result to obtain an adjusted large language model.
[0125] In some embodiments, the chunking module 230 includes: a first character chunk determination sub-module, configured to determine a chunking window length, and equally distance chunk the strings corresponding to the plurality of sample data according to the chunking window length to obtain the plurality of first character chunks; a sub-character chunk determination sub-module, configured to determine a target length according to the chunking window length, and chunk the plurality of first character chunks according to the target length to obtain a plurality of sub-character chunks; a second character chunk determination sub-module, configured to combine the plurality of sub-character chunks to obtain the plurality of second character chunks.
[0126] In some embodiments, the second character chunk determination includes: a second character chunk determination unit, configured to combine adjacent sub-character chunks and combine the head and tail character chunks among the plurality of sub-character chunks according to the order of the plurality of first character chunks to obtain the plurality of second character chunks.
[0127] In some embodiments, the training module 240 includes: a first attention score determination sub-module, configured to perform attention calculation on the plurality of first character chunks respectively to obtain first attention scores corresponding to the plurality of first character chunks; a second attention score determination sub-module, configured to perform attention calculation on the plurality of second character chunks respectively to obtain second attention scores corresponding to the plurality of second character chunks; a training result determination sub-module, configured to determine a target attention score based on the first attention scores and the second attention scores, and obtain the training result according to the target attention score.
[0128] In some embodiments, the first character block determination sub-module includes: a character length determination unit configured to determine the character length corresponding to each of the plurality of sample data; a context length determination unit configured to determine the maximum character length among the character lengths corresponding to each of the plurality of sample data as the context length; and a chunk window length determination unit configured to determine the chunk window length according to the context length.
[0129] In some embodiments, the network to be adjusted determination module 210 includes: a fine-tuning task determination sub-module configured to determine the fine-tuning task of the large language model to be adjusted; and a network to be adjusted determination sub-module configured to determine the network to be adjusted according to the fine-tuning task.
[0130] In some embodiments, the adjustment module 250 includes: an output label determination sub-module configured to determine the sample output labels corresponding to each of the plurality of sample data and determine the target output labels corresponding to each of the plurality of sample data according to the training result; a loss value determination sub-module configured to determine a loss value according to the target output labels and the sample output labels; and an adjustment determination sub-module configured to update the parameters of the learning weight matrix according to the loss value and obtain the adjusted large language model according to the learning weight matrix after parameter update.
[0131] In some embodiments, the loss value determination sub-module includes: a fine-tuning task determination unit configured to determine the fine-tuning task of the large language model to be adjusted; a loss function determination unit configured to determine a loss function according to the fine-tuning task; and a loss value determination unit configured to determine the loss value of the loss function based on the loss function according to the target output labels and the sample output labels.
[0132] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0133] In several embodiments provided in the present application, the coupling between modules may be electrical, mechanical, or other forms of coupling.
[0134] In addition, in each embodiment of the present application, the various functional modules may be integrated into one processing module, or each module may exist physically alone, or two or more modules may be integrated into one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0135] Please refer to Figure 10, which shows a structural block diagram of an electronic device provided by an embodiment of the present application. The electronic device 300 may be an electronic device such as a smart phone, a tablet computer, an e-book, etc. that can run application programs. The electronic device 300 in the present application may include one or more of the following components: a processor 310, a processor 320, and one or more application programs, where one or more application programs may be stored in the processor 320 and configured to be executed by one or more processors 310, and one or more programs are configured to execute the methods described in the foregoing method embodiments.
[0136] The processor 310 may include one or more processing cores. The processor 310 connects various parts within the entire electronic device 300 through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the processor 320, and by calling data stored in the processor 320, it executes various functions of the electronic device 300 and processes data. Optionally, the processor 310 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 310 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 310 and may be implemented separately through a communication chip.
[0137] The processor 320 may include a random access memory (RAM) and may also include a read-only memory. The processor 320 can be used to store instructions, programs, code, code sets, or instruction sets. The processor 320 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device 300 (such as phone books, audio and video data, chat record data, etc.).
[0138] Please refer to Figure 11, which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 400, and the program code can be called by a processor to execute the method described in the above method embodiment.
[0139] The computer-readable storage medium 400 can be an electronic memory such as a flash memory, EEPROM (electrically erasable programmable read-only memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has a storage space for the program code 410 that executes any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code 410 can be compressed in a suitable form, for example.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for fine-tuning a model, characterized in that, The method includes: Determine the network to be adjusted in the multi-head attention layer of the large language model to be adjusted, and obtain a data set for fine-tuning the large language model, where the data set includes multiple sample data; Determine a first weight matrix corresponding to the multi-head attention layer and a second weight matrix corresponding to the network to be adjusted, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix; Perform equidistant chunking on the strings corresponding to the multiple sample data respectively to obtain multiple character chunks; Freeze the original weight matrix, and train the large language model to be adjusted through the multiple character chunks to obtain a training result; Update the learning weight matrix according to the training result to obtain an adjusted large language model.
2. The method according to claim 1, characterized in that, The multiple character chunks include multiple first character chunks and multiple second character chunks. The step of performing equidistant chunking on the strings corresponding to the multiple sample data respectively to obtain multiple character chunks includes: Determine the chunking window length, and perform equidistant chunking on the strings corresponding to the multiple sample data respectively according to the chunking window length to obtain the multiple first character chunks; Determine a target length according to the chunking window length, and perform chunking on the multiple first character chunks according to the target length to obtain multiple sub-character chunks; Combine the multiple sub-character chunks to obtain the multiple second character chunks.
3. The method according to claim 2, characterized in that, The step of combining the multiple sub-character chunks to obtain the multiple second character chunks includes: According to the order of the multiple first character chunks, combine adjacent sub-character chunks and combine the first and last character chunks among the multiple sub-character chunks to obtain the multiple second character chunks.
4. The method according to claim 2, wherein The step of training the large language model to be adjusted through the multiple character chunks to obtain a training result includes: Perform attention calculation on the multiple first character chunks respectively to obtain first attention scores corresponding to the multiple first character chunks; Perform attention calculation on the multiple second character chunks respectively to obtain second attention scores corresponding to the multiple second character chunks; Determine a target attention score based on the first attention score and the second attention score, and obtain the training result according to the target attention score.
5. The method according to claim 2, characterized in that The step of determining the chunking window length includes: Determine the character lengths corresponding to the multiple sample data respectively; Determine the maximum character length among the character lengths corresponding to the multiple sample data as the context length; Determine the chunking window length according to the context length.
6. The method according to claim 1, characterized in that The step of determining the network to be adjusted in the multi-head attention layer of the large language model to be adjusted includes: Determine the fine-tuning task of the large language model to be adjusted; Determine the network to be adjusted according to the fine-tuning task.
7. The method according to any one of claims 1-6, characterized in that, The step of updating the learning weight matrix according to the training result to obtain an adjusted large language model includes: Determine the sample output labels corresponding to the multiple sample data respectively, and determine the target output labels corresponding to the multiple sample data respectively according to the training result; Determine a loss value based on the target output label and the sample output label; Update the parameters of the learning weight matrix according to the loss value, and obtain the adjusted large language model based on the learning weight matrix after parameter update.
8. The method according to claim 7, wherein The determining the loss value according to the target output label and the sample output label includes: Determine the fine-tuning task of the large language model to be adjusted; Determine a loss function according to the fine-tuning task; Based on the loss function, determine the loss value of the loss function according to the target output label and the sample output label.
9. A fine-tuning device for a model, characterized in that, The apparatus includes: An unadjusted network determination module, configured to determine an unadjusted network in the multi-head attention layer of the large language model to be adjusted, and obtain a data set for fine-tuning the large language model, where the data set includes a plurality of sample data; A parameter decomposition module, configured to determine a first weight matrix corresponding to the multi-head attention layer and a second weight matrix corresponding to the unadjusted network, and perform parameter decomposition according to the first weight matrix and the second weight matrix to obtain a learning weight matrix and an original weight matrix; A chunking module, configured to perform equidistant chunking on the strings corresponding to the plurality of sample data respectively to obtain a plurality of character chunks; A training module, configured to freeze the original weight matrix, and train the large language model to be adjusted through the plurality of character chunks to obtain a training result; An adjustment module, configured to update the learning weight matrix according to the training result to obtain an adjusted large language model.
10. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory electrically connected to the one or more processors; One or more applications, where the one or more applications are stored in the memory and configured to be executed by the one or more processors, and the one or more applications are configured to execute the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, Program code is stored in the computer-readable storage medium, and the program code can be called by a processor to execute the method according to any one of claims 1 to 8.
12. A computer program product comprising computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the method according to any one of claims 1 to 8.