LoRA-based large model fine tuning method and device, computer equipment and storage medium
By accumulating fine-tuning parameters in LoRA and adjusting the matrix according to the weight matrix, the limitations of LoRA when processing complex tasks are solved, the flexibility and prediction effect of the model are improved, and the adaptation to different scenarios and tasks are adapted to different scenarios and tasks.
Patent Information
- Application Number
- CN202510117847.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
LoRA has limitations when dealing with complex tasks, and the fixed structure of the low-rank matrix limits the processing flexibility and prediction effect of the model.
By accumulating fine-tuning parameters on fixed parameters, the fused parameters are obtained, and the first and second matrices are weighted to calculate the first and second matrices according to the weight matrix, and the fine-tuning parameters are adjusted to improve the flexibility and adaptability of the model.
It improves the processing flexibility and prediction effect of the model, can better adapt to different scenarios and tasks, quickly realize the processing functions of specific scenarios and tasks, and improve processing speed.
Smart Images

Figure CN120046733A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a method, device, computer device, and storage medium for fine-tuning large models based on LoRA. Background Art
[0002] Fine-tuning large models can improve the performance and adaptability of large models in specific tasks or fields, as well as save costs.
[0003] Low-Rank Adaptation of Large Language Models (LoRA) inserts low-rank matrices into the parameters. By fine-tuning the low-rank matrices, the number of parameters to be fine-tuned can be reduced, achieving improved fine-tuning efficiency.
[0004] However, the fixed structure of such low-rank matrices makes LoRA have limitations when dealing with complex tasks. Summary of the Invention
[0005] Based on this, to address the above technical problems, it is necessary to provide a method, device, computer device, computer-readable storage medium, and computer program product for fine-tuning large models based on LoRA, which can improve the processing flexibility and prediction effect of the model.
[0006] In a first aspect, this application provides a method for fine-tuning a large model based on LoRA, including:
[0007] Fix the parameters of the pre-trained initial model;
[0008] Accumulate fine-tuning parameters on the fixed parameters to obtain the fused parameters; wherein, the fine-tuning parameters are obtained by weighted calculation of the first matrix and the second matrix according to the weight matrix;
[0009] Process the sample data using the fused parameters, and adjust the first matrix, weight matrix, and second matrix to obtain the trained first matrix, weight matrix, and second matrix;
[0010] Calculate the adjustment result of the fine-tuning parameters according to the trained first matrix, weight matrix, and second matrix;
[0011] Accumulate the adjustment result of the fine-tuning parameters on the parameters of the initial model to obtain the trained target deep learning model.
[0012] In a second aspect, this application provides a device for fine-tuning a large model based on LoRA, including:
[0013] A fixing module, configured to fix the parameters of the pre-trained initial model;
[0014] A fusion module, configured to accumulate fine-tuning parameters on fixed parameters to obtain fused parameters; wherein, the fine-tuning parameters are obtained by performing weighted calculation on a first matrix and a second matrix according to a weight matrix;
[0015] A parameter adjustment module, configured to process sample data by using the fused parameters, and adjust the first matrix, the weight matrix, and the second matrix to obtain a trained first matrix, weight matrix, and second matrix;
[0016] The parameter adjustment module is further configured to calculate an adjustment result of the fine-tuning parameters according to the trained first matrix, weight matrix, and second matrix;
[0017] The parameter adjustment module is further configured to add the adjustment result of the fine-tuning parameters to the parameters of the initial model to obtain a trained target deep learning model.
[0018] In a third aspect, the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps in the above method are implemented.
[0019] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are implemented.
[0020] In a fifth aspect, the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in the above method are implemented.
[0021] In the above LoRA-based large model fine-tuning method, device, computer device, computer-readable storage medium, and computer program product, by configuring the fine-tuning parameters to be obtained by performing weighted calculation on a first matrix and a second matrix according to a weight matrix, the first matrix and the second matrix can be fused by using flexible weights, improving the flexibility of the model. At the same time, for different scenarios and different tasks, the structure of the corresponding fine-tuning parameters can be determined, thereby improving the adaptability of the model to scenarios and tasks. In specific scenarios and specific tasks, the prediction effect of the model can be improved, and through the fine-tuning method, the processing function of specific scenarios and specific tasks can be quickly realized, and the processing speed of the model for specific scenarios and specific tasks can be improved. Description of the Drawings
[0022] Figure 1 It is an application environment diagram of a LoRA-based large model fine-tuning method provided by an embodiment of the present application;
[0023] Figure 2 It is a flowchart of a LoRA-based large model fine-tuning method provided by an embodiment of the present application;
[0024] Figure 3 It is a schematic flowchart of another LoRA-based large model fine-tuning method provided by an embodiment of the present application;
[0025] Figure 4 It is a structural block diagram of a LoRA-based large model fine-tuning device provided by an embodiment of the present application;
[0026] Figure 5 It is an internal structure diagram of a computer device provided by an embodiment of the present application;
[0027] Figure 6 It is another internal structure diagram of a computer device provided by an embodiment of the present application;
[0028] Figure 7 It is an internal structure diagram of a computer-readable storage medium provided by an embodiment of the present application. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0030] The LoRA-based large model fine-tuning method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 where the terminal 102 communicates with the server 104 through a communication network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or can be placed in the cloud or other network servers. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. The portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0031] As shown in Figure 2 the embodiments of the present application provide a LoRA-based large model fine-tuning method, which is described by taking the method applied to the terminal 102 or the server 104 in Figure 1 as an example. It can be understood that the computer device can include at least one of a terminal and a server. The method includes the following steps:
[0032] S201. Fix the parameters of the pre-trained initial model.
[0033] Among them, the functions of the initial model can be set as needed. In some embodiments, the initial model is a large language model, which is used to process natural language texts to obtain response texts. For another example, the initial model is a text processing model, which is used to translate texts to obtain translation texts. For another example, the initial model is an image processing model, which can be used to perform object detection on images to obtain object detection results. For another example, the initial model is an audio processing model, which can be used to perform speech recognition processing on speech to obtain speech recognition results. For another example, the initial model can be a video processing model, which can be used to perform object motion detection on videos to obtain object motion data. Fixing the parameters of the initial model may mean that the parameters of the initial model will not be adjusted or changed.
[0034] S202. Accumulate the fine-tuning parameters on the fixed parameters to obtain the fused parameters; among them, the fine-tuning parameters are obtained by performing weighted calculations on the first matrix and the second matrix according to the weight matrix.
[0035] Among them, the fine-tuning parameters can be adjusted and changed. The fused parameters can be the sum of the fixed parameters and the fine-tuning parameters. The fine-tuning parameters can be the product of the first matrix, the weight matrix, and the second matrix. The weight matrix can be adjusted, and the weight matrix is used to fuse the first matrix and the second matrix with adjustable weights to flexibly adjust the fixed structure of the fine-tuning parameters.
[0036] In some embodiments, the fused parameter M satisfies the formula: M = M0 + ΔM = M0 + AWB. Where M0 is the fixed parameter, M0 is a matrix, and ΔM is the fine-tuning parameter. A is the first matrix, W is the weight matrix, and B is the second matrix.
[0037] S203. Process the sample data using the fused parameters, and adjust the first matrix, the weight matrix, and the second matrix to obtain the trained first matrix, weight matrix, and second matrix.
[0038] Among them, the initial model encodes or decodes the sample data using the fused parameters to obtain the processing result. According to the processing result, the parameters of the initial model are adjusted. Among them, the fixed parameters remain unchanged, and only the first matrix, the weight matrix, and the second matrix in the fine-tuning parameters are adjusted.
[0039] S204. Calculate the adjustment result of the fine-tuning parameters according to the trained first matrix, weight matrix, and second matrix.
[0040] Among them, calculate the product of the trained first matrix, weight matrix, and second matrix to obtain the adjustment result of the fine-tuning parameters.
[0041] S205. Add the adjustment result of the fine-tuning parameter to the parameter of the initial model to obtain the trained target deep learning model.
[0042] Among them, the parameter of the target deep learning model is the sum of the parameter fixed by the initial model and the adjustment result.
[0043] It can be seen that in the embodiment of the present application, by configuring the fine-tuning parameter to be obtained by weighted calculation of the first matrix and the second matrix according to the weight matrix, the first matrix and the second matrix can be fused using flexible weights, improving the flexibility of the model. At the same time, for different scenarios and different tasks, the structure of the corresponding fine-tuning parameter can be determined, thereby improving the adaptability of the model to scenarios and tasks. In terms of specific scenarios and specific tasks, the prediction effect of the model can be improved, and the processing function of specific scenarios and specific tasks can be quickly realized through the fine-tuning method, improving the processing speed of the model for specific scenarios and specific tasks.
[0044] In some embodiments, adding the fine-tuning parameter to the fixed parameter to obtain the fused parameter includes:
[0045] Obtain multiple first sub-matrices obtained by decomposing the first matrix;
[0046] Obtain multiple second sub-matrices obtained by decomposing the second matrix;
[0047] According to the weight matrix, perform weighted calculation on each first sub-matrix and each second sub-matrix to obtain multiple fine-tuning sub-matrices;
[0048] Determine each fine-tuning sub-matrix as the fine-tuning parameter;
[0049] Add the corresponding elements in the fine-tuning parameter to the fixed parameter to obtain the fused parameter.
[0050] Among them, the number of the first sub-matrices is at least two, and the number of the second sub-matrices is at least two. The number of the first sub-matrices is the same as the number of the second sub-matrices. Different sub-matrices are responsible for processing different parts of the input data. Decomposing the first matrix and the second matrix into corresponding sub-matrices is actually performing multi-subspace decomposition on the first matrix and the second matrix. The output of the sub-matrix is used as the representation of a subspace, which can enhance the model's ability to capture data diversity. Performing weighted calculation on each first sub-matrix and each second sub-matrix by the weight matrix is actually performing weighted summation on the output of the subspaces, which can realize dynamically adjusting the output weight of each subspace to adapt to the requirements of different tasks and scenarios. The fixed parameter can be a matrix. Add the elements in the matrix to the elements at the corresponding positions in the fine-tuning parameter to obtain a new matrix, and determine the new matrix as the fused parameter. Through the learnable weight matrix, the model can adaptively adjust the weights of each subspace according to different tasks and data, greatly enhancing the flexibility.
[0051] In some embodiments, the first matrix A is decomposed into a plurality of smaller sub - matrices, the first sub - matrix: A 1 , A 2 ……A h . The second matrix B is decomposed into a plurality of smaller sub - matrices, the second sub - matrix: B 1 , B 2 ……B h . The weight matrix: W ∈ R r×r , and the weight matrix corresponds to the sub - spaces of two low - rank matrices respectively. For example, for the i - th A matrix (the i - th first sub - matrix) and the j - th B matrix (the j - th second sub - matrix), the corresponding weight is w ij . Correspondingly, after AB is split, wherein, and after adding the weight matrix, wherein, and The fine - tuning parameter is calculated based on the following formula:
[0052]
[0053] It can be seen that in this embodiment, by decomposing the first matrix and the second matrix respectively to obtain a plurality of sub - matrices, and performing weighted summation on each sub - matrix based on the weight matrix, the subspace decomposition of the low - rank matrix is realized, and the weighted fusion of the outputs of the sub - controls is carried out, so as to process different parts of the input data. Moreover, the weight matrix is also tuned to automatically adjust the weights of the outputs of each sub - control. Thus, the fusion output is based on the flexibly adjusted weights, which can be flexibly adapted to different tasks and different scenarios. At the same time, it can increase the diversity of the extracted data, and perform weighted summation on the outputs of each subspace, which is equivalent to fusing the outputs of all subspaces, so that the output of each subspace participates in the calculation of the final result. It can avoid selecting the output of some subspaces as the final result, resulting in waste of some computing resources, thereby reducing the latency problem of model calculation brought by this selection mechanism and improving the processing flexibility of the model.
[0054] In some embodiments, both the first matrix and the second matrix are low - rank matrices; the number of rows of the first matrix is the same as that of the matrix corresponding to the fixed parameter, and the number of columns of the second matrix is the same as that of the matrix corresponding to the fixed parameter; the number of columns of the first matrix, the number of rows and columns of the weight matrix, and the number of rows of the second matrix are all ranks; the sum of the ranks of a plurality of first sub - matrices is equal to the rank of the first matrix; the sum of the ranks of a plurality of second sub - matrices is equal to the rank of the second matrix.
[0055] Among them, the fine-tuning parameters are obtained by calculating the product of the first matrix, the weight matrix, and the second matrix. Since the dimensions between the fixed parameters and the fine-tuning parameters are the same, the number of rows of the first matrix is the same as the number of rows of the fixed parameters. The number of columns of the second matrix is the same as the number of columns of the fixed parameters. The number of rows and columns of the weight matrix are both ranks. Correspondingly, the weight matrix is a low-rank matrix.
[0056] In some embodiments, the number of rows of the matrix corresponding to the fixed parameters is d 1 , and the number of columns is d 2 . The first matrix is The second matrix is where r represents the size of the rank, and usually r << min{d1, d2}. The first matrix can be decomposed into two first sub-matrices, A 1 and A 2 , r 1 +r 2 = r.
[0057] It can be seen that by configuring the first matrix, the second matrix, and the weight matrix as low-rank matrices, the number of parameters to be adjusted can be reduced, the fine-tuning efficiency can be improved, and the accuracy of the model after fine-tuning can be taken into account; by determining the relationship between the ranks of the sub-matrices and the rank of the matrix, accurate decomposition can be achieved and the accuracy of matrix decomposition can be improved.
[0058] In some embodiments, the initial model is a large language model;
[0059] Using the fused parameters to process the sample data, and adjusting the first matrix, the weight matrix, and the second matrix to obtain the trained first matrix, weight matrix, and second matrix, including:
[0060] Using the fused parameters to process the question text in the sample data to obtain a response text;
[0061] According to the difference between the response text and the expected content in the sample data, adjust the first matrix, the weight matrix, and the second matrix respectively to obtain the trained first matrix, weight matrix, and second matrix.
[0062] Among them, the problem text can refer to the input data of the large language model. The response text can refer to the output data of the large language model. The parameters in the large language model are the fused parameters. The fine-tuned large language model encodes the problem text to obtain a feature vector, and decodes the feature vector to obtain the response text. Among them, the fused parameters can be used to encode the problem text to obtain a feature vector, and / or decode the feature vector to obtain the response text. The expected content can refer to the true value of the problem text. The difference between the response text and the expected content is used to adjust the first matrix, the weight matrix, and the second matrix. The backpropagation algorithm can be used to adjust the parameters of the three matrices according to the loss function.
[0063] In some embodiments, the problem text can be multimodal content, and the large language model can be a multimodal large language model. The problem text can be text asking about the method of using an API (Application Programming Interface). The response text can be text about the usage method of the API. The large language model uses the API documentation as a knowledge base and queries the usage of relevant functions, instances, and parameters in the knowledge base.
[0064] It can be seen that by adjusting the three matrices according to the difference between the response text output by the large language model and the expected content for the large language model, fine-tuning of the large language model can be achieved, and the output combination method of the matrices can be adaptively adjusted to enable the large language model to better adapt to complex tasks and scenarios.
[0065] In some embodiments, the initial model is an image processing model;
[0066] Processing the sample data with the fused parameters and adjusting the first matrix, the weight matrix, and the second matrix to obtain the trained first matrix, weight matrix, and second matrix includes:
[0067] Processing the sample image in the sample data with the fused parameters to obtain a defect detection result;
[0068] Adjusting the first matrix, the weight matrix, and the second matrix respectively according to the difference between the defect detection result and the standard detection result in the sample data to obtain the trained first matrix, weight matrix, and second matrix.
[0069] Among them, the image processing model can be used to detect the defect detection result on the surface of a product based on the sample image of the product surface. The labeled detection result can be the real defect in the sample image. The backpropagation algorithm can be used to adjust the parameters of the three matrices according to the loss function.
[0070] It can be seen that by adjusting the three matrices according to the difference between the response text output by the image processing model and the expected content for the image processing model, so as to achieve fine-tuning of the image processing model, the output combination mode of the matrices can be adaptively adjusted, enabling the image processing model to better adapt to complex tasks and scenarios.
[0071] In some embodiments, such as Figure 3 shown, the embodiment of the present application provides another LoRA-based large model fine-tuning method. This method includes the following steps:
[0072] S301. Fix the parameters of the pre-trained initial model.
[0073] S302. Accumulate the fine-tuning parameters on the fixed parameters to obtain the fused parameters; wherein, the fine-tuning parameters are obtained by weighted calculation of the first matrix and the second matrix according to the weight matrix.
[0074] S303. Process the sample data using the fused parameters, and adjust the first matrix, the weight matrix, and the second matrix to obtain the trained first matrix, weight matrix, and second matrix.
[0075] S304. Calculate the adjustment result of the fine-tuning parameters according to the trained first matrix, weight matrix, and second matrix.
[0076] S305. Add the adjustment result of the fine-tuning parameters to the parameters of the initial model to obtain the trained target deep learning model.
[0077] S306. Input the target problem into the trained target deep learning model for processing, and output the target response to the target problem.
[0078] Among them, the target deep learning model encodes and decodes the target problem using the fused parameters to obtain the target response. Specifically, the fused parameters are obtained by accumulating the trained fine-tuning parameters on the fixed pre-trained parameters, and the trained fine-tuning parameters are obtained by weighted fusion of the first sub-matrices obtained by decomposing the first matrix and the second sub-matrices obtained by decomposing the second matrix according to the weight matrix.
[0079] It can be seen that by using the target deep learning model obtained by adding the fine-tuned parameters after fine-tuning to the original trained parameters to process the target problem and output the target response, no additional calculation operations between the fine-tuned parameters and the input data are required during the inference process, ensuring that the inference time of the model will not increase, thus ensuring that the inference efficiency is not affected. While maintaining the high-efficiency inference of the model, it is still possible to make full use of the flexibility and expressiveness brought by the fine-tuned parameters. Moreover, during the model processing and calculation process, the weight matrix performs weighted summation on the outputs of each subspace, realizing the utilization of the outputs of all subspaces, avoiding the waste of computing resources caused by the outputs of subspaces not being applied, and avoiding the inference delay problem caused by selecting the outputs of subspaces, improving the model processing efficiency. And using the trained weight matrix to fuse the outputs of subspaces can improve the representation ability of the model, and further improve the response accuracy of the model.
[0080] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0081] Based on the same inventive concept, an embodiment of the present application also provides a large model fine-tuning device based on LoRA. The implementation solution provided by this device to solve problems is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the large model fine-tuning device based on LoRA provided below can refer to the limitations on the large model fine-tuning method based on LoRA in the above text, and will not be repeated here.
[0082] As Figure 4 shown, an embodiment of the present application provides a large model fine-tuning device 400 based on LoRA, including:
[0083] A fixing module 401, configured to fix the parameters of the pre-trained initial model;
[0084] A fusion module 402, configured to accumulate fine-tuned parameters on the fixed parameters to obtain fused parameters; wherein, the fine-tuned parameters are obtained by performing weighted calculation on the first matrix and the second matrix according to the weight matrix;
[0085] The parameter tuning module 403 is used to process the sample data with the fused parameters, adjust the first matrix, the weight matrix, and the second matrix, and obtain the trained first matrix, weight matrix, and second matrix;
[0086] The parameter tuning module 403 is further used to calculate the adjustment result of the fine-tuning parameters according to the trained first matrix, weight matrix, and second matrix;
[0087] The parameter tuning module 403 is further used to add the adjustment result of the fine-tuning parameters to the parameters of the initial model to obtain the trained target deep learning model.
[0088] In some embodiments, in terms of adding the fine-tuning parameters to the fixed parameters to obtain the fused parameters, the fusion module 402 is specifically used for:
[0089] Obtain multiple first sub-matrices obtained by decomposing the first matrix;
[0090] Obtain multiple second sub-matrices obtained by decomposing the second matrix;
[0091] According to the weight matrix, perform weighted calculation on each first sub-matrix and each second sub-matrix to obtain multiple fine-tuning sub-matrices;
[0092] Determine each fine-tuning sub-matrix as a fine-tuning parameter;
[0093] Add the corresponding elements in the fine-tuning parameters to the fixed parameters to obtain the fused parameters.
[0094] In some embodiments, both the first matrix and the second matrix are low-rank matrices; the number of rows of the first matrix is the same as the number of rows of the matrix corresponding to the fixed parameters, and the number of columns of the second matrix is the same as the number of columns of the matrix corresponding to the fixed parameters; the number of columns of the first matrix, the number of rows and columns of the weight matrix, and the number of rows of the second matrix are all ranks; the sum of the ranks of the multiple first sub-matrices is equal to the rank of the first matrix; the sum of the ranks of the multiple second sub-matrices is equal to the rank of the second matrix.
[0095] In some embodiments, the initial model is a large language model; in terms of processing the sample data with the fused parameters, adjusting the first matrix, the weight matrix, and the second matrix, and obtaining the trained first matrix, weight matrix, and second matrix, the parameter tuning module 403 is specifically used for:
[0096] Process the problem text in the sample data with the fused parameters to obtain a response text;
[0097] According to the difference between the response text and the expected content in the sample data, adjust the first matrix, the weight matrix, and the second matrix respectively to obtain the trained first matrix, weight matrix, and second matrix.
[0098] In some embodiments, the initial model is an image processing model; in terms of processing the sample data with the fused parameters and adjusting the first matrix, the weight matrix, and the second matrix to obtain the trained first matrix, weight matrix, and second matrix, the parameter tuning module 403 is specifically configured to:
[0099] Process the sample images in the sample data with the fused parameters to obtain defect detection results;
[0100] Adjust the first matrix, the weight matrix, and the second matrix respectively according to the difference between the defect detection results and the standard detection results in the sample data to obtain the trained first matrix, weight matrix, and second matrix.
[0101] In some embodiments, the device further includes a question-answering module;
[0102] The question-answering module is configured to input the target question into the trained target deep learning model for processing and output the target answer to the target question.
[0103] Each module in the above LoRA-based large model fine-tuning device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0104] In some embodiments, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store relevant data of the LoRA-based large model fine-tuning method. The input / output interface of the computer device is used for the processor to exchange information with external devices. The communication interface of the computer device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it implements the steps in the above LoRA-based large model fine-tuning method.
[0105] In some embodiments, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be asFigure 6 As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements the steps in the above-mentioned LoRA-based large model fine-tuning method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen; the input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0106] Those skilled in the art can understand that Figure 5 or Figure 6 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0107] In some embodiments, a computer device is provided. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps in the above-mentioned method embodiments.
[0108] In some embodiments, as Figure 7 shown, an internal structure diagram of a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0109] In some embodiments, a computer program product is provided. The computer program product includes a computer program, and when the computer program is executed by the processor, it implements the steps in the above-mentioned method embodiments.
[0110] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0111] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include Read-Only Memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0112] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0113] The above embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A large model fine-tuning method based on LoRA, characterized in that: include: Fix the parameters of the trained initial model; Accumulating fine-tuning parameters on fixed parameters to obtain fused parameters; wherein the fine-tuning parameters are obtained by weighted calculation of the first matrix and the second matrix according to the weight matrix; Processing the sample data using the fused parameters, adjusting the first matrix, the weight matrix, and the second matrix to obtain a trained first matrix, a weight matrix, and a second matrix; Calculating the adjustment result of the fine-tuning parameter according to the trained first matrix, weight matrix and second matrix; The adjustment results of the fine-tuning parameters are added to the parameters of the initial model to obtain a trained target deep learning model.
2. The method according to claim 1, characterized in that The accumulating fine-tuning parameters on the fixed parameters to obtain the fused parameters includes: Obtain multiple first sub-matrices obtained by decomposing the first matrix; Obtain a plurality of second sub-matrices obtained by decomposing the second matrix; According to the weight matrix, weighted calculation is performed on each of the first sub-matrices and each of the second sub-matrices to obtain a plurality of fine-tuning sub-matrices; Determining each of the fine-tuning sub-matrices as a fine-tuning parameter; The corresponding elements in the fine-tuning parameters are accumulated on the fixed parameters to obtain the fused parameters.
3. The method according to claim 2, characterized in that The first matrix and the second matrix are both low-rank matrices; the number of rows of the first matrix is the same as the number of rows of the matrix corresponding to the fixed parameters, and the number of columns of the second matrix is the same as the number of columns of the matrix corresponding to the fixed parameters; the number of columns of the first matrix, the number of rows and columns of the weight matrix, and the number of rows of the second matrix are all ranks; the sum of the ranks of the multiple first sub-matrices is equal to the rank of the first matrix; the sum of the ranks of the multiple second sub-matrices is equal to the rank of the second matrix.
4. The method according to claim 1, characterized in that: The initial model is a large language model; The method of using the fused parameters to process the sample data, adjusting the first matrix, the weight matrix and the second matrix to obtain the trained first matrix, the weight matrix and the second matrix includes: Using the fused parameters to process the question text in the sample data to obtain a reply text; According to the difference between the reply text and the expected content in the sample data, the first matrix, the weight matrix and the second matrix are adjusted respectively to obtain the trained first matrix, weight matrix and second matrix.
5. The method according to claim 1, characterized in that The initial model is an image processing model; The method of using the fused parameters to process the sample data, adjusting the first matrix, the weight matrix and the second matrix to obtain the trained first matrix, the weight matrix and the second matrix includes: Using the fused parameters to process the sample image in the sample data to obtain a defect detection result; According to the difference between the defect detection result and the standard detection result in the sample data, the first matrix, the weight matrix and the second matrix are adjusted respectively to obtain the trained first matrix, weight matrix and second matrix.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: The target question is input into the trained target deep learning model for processing, and the target response to the target question is output.
7. A large model fine-tuning device based on LoRA, characterized in that: include: The fixed module is used to fix the parameters of the trained initial model; A fusion module, used for accumulating fine-tuning parameters on fixed parameters to obtain fused parameters; wherein the fine-tuning parameters are obtained by weighted calculation of the first matrix and the second matrix according to the weight matrix; A parameter adjustment module, used to process the sample data using the fused parameters, adjust the first matrix, the weight matrix and the second matrix, and obtain the trained first matrix, weight matrix and second matrix; The parameter adjustment module is further used to calculate the adjustment result of the fine-tuning parameter according to the trained first matrix, weight matrix and second matrix; The parameter adjustment module is also used to accumulate the adjustment results of the fine-tuning parameters to the parameters of the initial model to obtain a trained target deep learning model.
8. The device according to claim 7, characterized in that In terms of accumulating fine-tuning parameters on the fixed parameters to obtain fused parameters, the fusion module is specifically used for: Obtain multiple first sub-matrices obtained by decomposing the first matrix; Obtain a plurality of second sub-matrices obtained by decomposing the second matrix; According to the weight matrix, weighted calculation is performed on each of the first sub-matrices and each of the second sub-matrices to obtain a plurality of fine-tuning sub-matrices; Determining each of the fine-tuning sub-matrices as a fine-tuning parameter; The corresponding elements in the fine-tuning parameters are accumulated on the fixed parameters to obtain the fused parameters.
9. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.