Large model parameter fine-tuning method and device based on incremental learning, equipment and medium
Patent Information
- Application Number
- CN202310809667.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-06-30
AI Technical Summary
[0006](1)每个任务的特征和难度不同,如果直接将所有任务一起训练,很容易出现“一哄而上”的情况,导致模型无法兼顾各种任务的性能
[0023]本申请提供一种基于增量学习的大模型参数微调方法、装置、设备及存储介质,本申请方法包括获取目标任务的任务数据;基于数据大模型中各层任务子模型,对所述任务数据进行数据预测处理,获得所述各层任务子模型输出的模型预测结果;基于权重计算公式和所述任务数据的数据特征,计算所述各层任务子模型对于所述目标任务的注意力权重;基于所述各层任务子模型对应的所述注意力权重,对所述各层任务子模型对应的所述模型预测结果进行加权计算,获得目标预测结果。通过上述方式,通过数据大模型中各层任务子模型对任务数据进行数据预测处理,获得各层任务子模型的模型预测结果,实现对目标任务的多角度分析,从而完整提取目标任务的任务数据的特征,以提高医疗文本的识别准确性;根据任务数据的数据特征和权重计算公式,计算各层任务子模型对于目标任务的注意力权重,进一步地根据注意力权重,将各层任务子模型输出的模型预测结果进行加权计算,可以提取任务数据中的重要特征,降低非重要特征对于文本分类识别结果的影响,提高目标预测结果的准确性。
Smart Images

Figure CN116822651B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of text classification and digital medical technology, and in particular to a method, apparatus, device and medium for fine-tuning large model parameters based on incremental learning. Background Technology
[0002] Text classification refers to the process by which a computer maps a text containing information to one or more pre-defined categories. While current text classification technologies perform well in some scenarios, they still have some shortcomings in the medical field.
[0003] On the one hand, the medical field has a vast and complex terminology and knowledge system, making it difficult to accurately model using traditional machine learning methods. On the other hand, medical texts contain a large amount of noise and class imbalance, causing trained models to often perform poorly on a few classes.
[0004] Furthermore, most existing text classification techniques use single-task learning for model training, lacking the ability to share and transfer knowledge across different tasks, resulting in high model complexity and poor generalization ability. At the same time, traditional incremental learning methods are sensitive to data changes and cannot consistently improve model performance.
[0005] Traditional methods for fine-tuning large-scale medical models train all tasks together in the model, which leads to the following problems in practical applications:
[0006] (1) Each task has different characteristics and difficulty. If all tasks are trained together, it is easy to cause a "rush to train", which will make the model unable to take into account the performance of various tasks.
[0007] (2) During the training process, if some tasks are overfitted or underfitted, it will affect the overall performance of the model.
[0008] (3) Due to the complexity of medical data, traditional fine-tuning methods may have problems in recognizing certain medical texts, leading to serious consequences such as misdiagnosis and missed diagnosis.
[0009] Therefore, how to solve the low accuracy of current large medical models in recognizing multi-task medical texts has become an urgent technical problem. Summary of the Invention
[0010] This application provides a method, apparatus, device, and storage medium for fine-tuning parameters of a large medical model based on incremental learning, aiming to improve the recognition accuracy of a large medical model for multi-task medical text.
[0011] Firstly, this application provides a method for fine-tuning large model parameters based on incremental learning, the method comprising:
[0012] Obtain the task data for the target task;
[0013] Based on the task sub-models at each layer in the large data model, data prediction processing is performed on the task data to obtain the model prediction results output by each task sub-model.
[0014] Based on the weight calculation formula and the data characteristics of the task data, the attention weights of each layer of task sub-models for the target task are calculated.
[0015] Based on the attention weights corresponding to each task sub-model, the model prediction results corresponding to each task sub-model are weighted and calculated to obtain the target prediction result.
[0016] Secondly, this application also provides an apparatus for a large model parameter fine-tuning method based on incremental learning, the apparatus comprising:
[0017] The task data acquisition module is used to acquire task data for the target task.
[0018] The data processing module is used to perform data prediction processing on the task data based on the task sub-models at each layer in the large data model, and to obtain the model prediction results output by the task sub-models at each layer.
[0019] The attention weight calculation module is used to calculate the attention weight of each task sub-model for the target task based on the weight calculation formula and the data characteristics of the task data.
[0020] The target prediction result acquisition module is used to perform weighted calculation on the model prediction results corresponding to each task sub-model based on the attention weights corresponding to each task sub-model to obtain the target prediction result.
[0021] Thirdly, this application also provides a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the large model parameter fine-tuning method based on incremental learning as described above.
[0022] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the above-described method for fine-tuning large model parameters based on incremental learning.
[0023] This application provides a method, apparatus, device, and storage medium for fine-tuning large model parameters based on incremental learning. The method includes acquiring task data for a target task; performing data prediction processing on the task data based on task sub-models at each layer in the large data model to obtain model prediction results output by each task sub-model; calculating the attention weights of each task sub-model for the target task based on the weight calculation formula and the data characteristics of the task data; and performing a weighted calculation on the model prediction results corresponding to each task sub-model based on the attention weights to obtain the target prediction result. By employing the above method, task data is processed through task sub-models at each layer of the large data model to obtain the model prediction results of each task sub-model. This enables multi-angle analysis of the target task, thereby fully extracting the features of the target task data and improving the recognition accuracy of medical text. Based on the data features and weight calculation formula of the task data, the attention weight of each task sub-model for the target task is calculated. Furthermore, based on the attention weight, the model prediction results output by each task sub-model are weighted and calculated to extract important features from the task data, reduce the impact of unimportant features on the text classification and recognition results, and improve the accuracy of the target prediction results. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 An embodiment of this application provides a large model parameter fine-tuning system based on incremental learning;
[0026] Figure 2 A flowchart illustrating the first embodiment of a method for fine-tuning large model parameters based on incremental learning provided in this application;
[0027] Figure 3 A flowchart illustrating a second embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application;
[0028] Figure 4 A flowchart illustrating a third embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application;
[0029] Figure 5 A schematic diagram of a low-rank matrix structure for a LoRA model provided in this application;
[0030] Figure 6A flowchart illustrating the fourth embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application;
[0031] Figure 7 This is a schematic diagram of the structure of a first embodiment of a large model parameter fine-tuning device based on incremental learning provided in this application;
[0032] Figure 8 This is a schematic block diagram of the structure of a computer device provided in an embodiment of this application.
[0033] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] The flowchart shown in the attached diagram is for illustrative purposes only and does not necessarily include all content and operations / steps, nor does it necessarily have to be performed in the order described. For example, some operations / steps can be broken down, combined, or partially merged, so the actual execution order may change depending on the actual situation.
[0036] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0037] Embodiments of this application provide a method, apparatus, device, and storage medium for fine-tuning large model parameters based on incremental learning. This method is used to perform weighted fusion of model prediction results of multi-layer task sub-models according to the data characteristics of the target task, thereby improving the accuracy of the large data model's data prediction results for the target task.
[0038] like Figure 1 As shown, Figure 1 An embodiment of this application provides a large model parameter fine-tuning system based on incremental learning. The system includes a terminal and a server, the terminal and the server being communicatively connected, and the server being communicatively connected to a medical text database.
[0039] The terminals include electronic devices such as mobile phones, tablets, laptops, desktop computers, personal digital assistants, and wearable devices.
[0040] The server may be a single independent server or a server cluster.
[0041] The following section will provide a detailed description of the large model parameter fine-tuning method based on incremental learning provided in the embodiments of this application, based on the large model parameter fine-tuning system based on incremental learning.
[0042] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating a first embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application. This large model parameter fine-tuning method based on incremental learning can be used in the server of a large model parameter fine-tuning system based on incremental learning.
[0043] like Figure 2 As shown, the large model parameter fine-tuning method based on incremental learning includes steps S101 to S104.
[0044] Step S101: Obtain the task data for the target task;
[0045] In one embodiment, when a medical data processing task is required, such as a medical text classification task, the text data is acquired and input into a large medical model, which then automatically performs text classification and prediction on the text data.
[0046] Step S102: Based on the task sub-models at each layer in the large data model, perform data prediction processing on the task data to obtain the model prediction results output by each task sub-model.
[0047] In one embodiment, when medical text data is input into the big data model, each task sub-model in the big data model needs to perform a data prediction process on the medical text data in sequence, and then output the model prediction results of each task sub-model for the medical text data.
[0048] In one embodiment, because the task sub-models at each layer of the large data model have different focuses on text data processing, for example, the first-layer task sub-model is suitable for extracting and classifying medical professional terms, while the second-layer task sub-model is suitable for recognizing and predicting medical consultation information, the focus of each task sub-model on the same medical text data will be different, resulting in different model prediction results. This can also adapt to a variety of different medical text classification tasks.
[0049] Step S103: Based on the weight calculation formula and the data characteristics of the task data, calculate the attention weight of each layer of task sub-model for the target task;
[0050] In one embodiment, attention weights for the task data of the target task are calculated for each task sub-model layer based on an attention mechanism.
[0051] In one embodiment, the weight calculation formula includes:
[0052]
[0053] Where X is the task data, T is the number of tasks, and ω t L is the attention weight for task t. t It is a task sub-model with task t.
[0054] In one embodiment, the attention mechanism is a special structure embedded in machine learning models to automatically learn and calculate the contribution of input data to output data. When performing text classification tasks, the data model first performs feature engineering, transforming the raw text into numerical vectors.
[0055] In one embodiment, feature engineering is an application of attention mechanisms in data science. It helps models select effective and appropriately sized features, enabling them to complete tasks efficiently and effectively. For example, by using stepwise regression analysis to filter the original feature set and obtain a high-quality subset of features, downstream models can focus on signals most closely related to the task.
[0056] In one embodiment, the attention mechanism scores the dimensions of each task sub-model and then weights the features according to the scores to highlight the impact of important features on downstream models or modules. For example, in the process of extracting consultation information from medical texts, the consultation extraction task sub-model scores the highest, while other task sub-models are assigned different scores based on relevance or influence. In this case, the consultation extraction task sub-model has the greatest influence on the text information extraction results of the medical text, and its output best meets the actual needs of the task, resulting in the most accurate output.
[0057] Step S104: Based on the attention weights corresponding to each task sub-model, perform weighted calculation on the model prediction results corresponding to each task sub-model to obtain the target prediction result.
[0058] In one embodiment, the model prediction results corresponding to each task sub-model are weighted and summed according to the attention weights of each task sub-model to obtain the target prediction result.
[0059] In one embodiment, the attention mechanism calculates the relevance between each task sub-model and the task data based on the historical output of each task sub-model and the task data of the target task. This allows for the calculation of the attention weight of each task sub-model for the target task, thereby highlighting the model prediction results of the task sub-model with the highest relevance and reducing the impact of the output results of other task sub-models. This adapts to the needs of different medical text classification tasks and improves the prediction accuracy of the task data for the target task.
[0060] In one embodiment, after calculating the attention weights of each layer of task sub-models for the target task based on the weight calculation formula and the data characteristics of the task data, the method further includes: verifying, calculating, and adjusting the attention weights corresponding to each layer of task sub-models based on cross-validation.
[0061] In one embodiment, after adjusting the attention weights of each task sub-model according to the characteristics of the target task, the attention weights of each task sub-model can be verified and adjusted through methods such as cross-validation to make the model prediction results more accurate.
[0062] In one embodiment, cross-validation is a method used to observe the stability of a model. The dataset is divided multiple times, with one part used as the training set to train the model and the other part used as the test set. This process is repeated multiple times to evaluate the model's performance.
[0063] In one embodiment, cross-validation divides the data into n parts, using one part as the test set and the remaining n-1 parts as the training set, and repeatedly calculating the model's accuracy to evaluate its average performance. The splitting of the training and test sets can interfere with the model's results; therefore, the average of the results from n cross-validations is a better measure of the model's performance. For example, if a split coefficient of 0.3 is specified, then 7 / 10 of the data is used as the training set, and the remaining 3 / 10 is used as the test set.
[0064] This embodiment provides a method for fine-tuning the parameters of a large model based on incremental learning. The method uses task sub-models at each layer of the large data model to perform data prediction processing on the task data, obtaining the model prediction results of each task sub-model. This enables multi-angle analysis of the target task, thereby fully extracting the features of the target task's data to improve the accuracy of medical text recognition. Based on the data features and weight calculation formulas of the task data, the attention weights of each task sub-model for the target task are calculated. Furthermore, based on these attention weights, the model prediction results output by each task sub-model are weighted, which can extract important features from the task data, reduce the impact of unimportant features on the text classification and recognition results, and improve the accuracy of the target prediction results.
[0065] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating a second embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application.
[0066] like Figure 3 As shown, based on the above Figure 2 In the illustrated embodiment, after step S104, the method further includes:
[0067] Step S201: Upon receiving a model prediction result query request sent by the user terminal, obtain the target task label in the model prediction result query request;
[0068] Step S202: Based on the target task label, query the task sub-model corresponding to the target task label, and output the model prediction result of the task sub-model.
[0069] In one embodiment, when querying the data prediction results for a target task, the user can directly query the target prediction result obtained after weighted summation based on attention weights. Alternatively, based on task requirements or the task tag of the target task, the user can query the task sub-model corresponding to that task tag and output the model prediction result corresponding to that task sub-model.
[0070] In one embodiment, the task label can be the focus of the task content corresponding to each task sub-model, such as medical consultation task label, medical extraction task label, medical diagnosis task label, disease treatment task label, etc.
[0071] Please refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of a large model parameter fine-tuning method based on incremental learning provided in this application.
[0072] like Figure 4 As shown, based on the above Figure 2 In the illustrated embodiment, prior to step S101, the method further includes:
[0073] Step S301: Obtain the basic large model and the task labels and task data corresponding to at least one model task, wherein the basic large model includes at least one basic sub-model.
[0074] In one embodiment, the base large model can be a GPT large model. GPT refers to Generative Pre-Training, which enables unsupervised pre-training and supervised fine-tuning on downstream tasks.
[0075] In one embodiment, GPT uses the decoder structure of the Transformer machine learning model, and removes the multi-head attention used to introduce the encoder output in the decoder, and combines it with the downstream model of the features to form the model structure of GPT.
[0076] Step S302: Based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task, perform data training on the basic sub-model to obtain the task sub-model corresponding to each task label;
[0077] In one embodiment, for each medical task, such as diagnosis or treatment, we will train a LoRA model for each task. This model focuses only on the specific task, and when using this model, only the corresponding task label needs to be input. This can be represented by the following formula:
[0078] L t =f(X) t W t )
[0079] Where t is the task label, X t It's task data, W t These are the model parameters of the task sub-model.
[0080] Understandably, the models are overparameterized, and they have a smaller intrinsic dimension. The models mainly rely on this low intrinsic dimension to adapt to the task.
[0081] In one embodiment, it is assumed that the changes in weights during task adaptation are low-rank, thus allowing the use of the Low-Rank Adaptation (LoRA) method. LoRA allows for the indirect training of some dense layers in the neural network by optimizing the rank decomposition matrix of the dense layers during adaptation, while keeping the pre-trained weights unchanged.
[0082] like Figure 5 As shown, LoRA adds a bypass to the original pre-trained language model (PLM) to perform a dimensionality reduction and then increase operation to simulate intrinsic rank. During training, the parameters of PLM are fixed, and only the dimensionality reduction matrix A and the dimensionality increase matrix B are trained. The input and output dimensions of the model remain unchanged, and the output is the superposition of the parameters of BA and PLM. A is initialized with a random Gaussian distribution, and B is initialized with a zero matrix to ensure that the bypass matrix is still a zero matrix at the beginning of training.
[0083] Furthermore, such as Figure 6 As shown, step S302 specifically further includes:
[0084] Step S401: Obtain the multidimensional weight matrix based on the matrix product of the fundamental matrix of the basic sub-model and the low-rank matrix;
[0085] In one embodiment, assuming that a pre-trained language model needs to be fine-tuned, the pre-trained model parameters need to be updated, as expressed by the following formula:
[0086] W0+W
[0087] Where W0 is the parameter initialized in the pre-trained model, and ΔW is the parameter that needs to be updated.
[0088] Step S402: Based on the data characteristics of the current task, adjust the weight parameters of the low-rank matrix in the multidimensional weight matrix to obtain the weight change.
[0089] In one embodiment, for LoRA, only fine-tuning of ΔW is required. During LoRA training, W0 remains constant, and only A and B are training parameters.
[0090] Step S403: Based on the weighted summation formula, perform a weighted summation calculation on the weight change and the basic weights of the basic matrix to obtain the target weight;
[0091] In one embodiment, in each fine-tuning training, multi-dimensional weights are added based on the characteristics and difficulties of the target task to help the model better adapt to new data. In order to make the model better adapt to new data, multi-dimensional incremental learning techniques can be used to weight the model based on the characteristics and difficulties of the target task.
[0092] In one embodiment, incremental learning is a method based on repeated learning. Its basic principle is to segment the data during model training, using only the data from the previous segment for each segment, until the desired accuracy is achieved. This allows for rapid improvement of model performance with limited data. In natural language processing, incremental learning can be used for tasks such as text classification, sentiment analysis, and named entity recognition.
[0093] In one embodiment, incremental learning is the ability to continuously process a continuous stream of information in the real world, absorbing new knowledge while retaining or even integrating and optimizing old knowledge. Incremental learning is multi-tasking, but it allows data from the current task to be processed multiple times before moving on to the next task.
[0094] In one embodiment, incremental learning techniques retain most of the previously learned knowledge while learning new knowledge, meaning the model performs well on both old and new tasks. The computational power and memory requirements of incremental learning should remain constant or grow slowly with the number of categories. Ideally, once learning a task is complete, all observations for that task are discarded. The incremental learning model can continuously learn new knowledge from new tasks and new data; it remains trainable even when new tasks appear at different times.
[0095] In one embodiment, the weighted summation formula includes:
[0096]
[0097] Where t is the current task label, and N is the current iteration number. These are the base weights for the current task t. It's the learning rate. It is the change in the weight.
[0098] Step S404: Based on the target weights, perform iterative parameter adjustment on the multidimensional weight matrix of the basic sub-model to obtain the task sub-model.
[0099] In one embodiment, the weights of each task can be decomposed into multi-dimensional weights, such as feature weights and category weights. In each round of training, the corresponding weights are iteratively calculated and adjusted according to the different features of the target task, so that the model is more adapted to the current task.
[0100] For example, in a medical text classification task, a base model is first trained. Then, in the first stage, a medical extraction task is selected for fine-tuning, training a LoRa model to obtain a medical extraction model. Next, in the second stage, a medical consultation task is selected for fine-tuning, training another LoRa model to obtain a medical consultation model. Finally, these two models are merged into a large model based on multi-task learning. Cross-validation is used to determine the weights of each task, and the corresponding weights are adjusted according to the characteristics of the target task in each training round to achieve better performance and generalization ability.
[0101] Step S303: Based on the task sub-model, replace the basic sub-model in the basic large model to obtain the data large model.
[0102] In one embodiment, after training the task sub-models corresponding to different task labels, it is necessary to merge the task sub-models at each layer with the basic large model and replace the basic sub-models in the basic large model, thereby obtaining a large data model that can meet the needs of various different text classification tasks.
[0103] In one embodiment, multidimensional incremental learning is used to fine-tune the parameters of a large data model, ensuring sufficient training for each task and avoiding a "rush to learn" situation, thus giving appropriate attention to different tasks. By training the LORA model separately and then weighting and fusing them, multi-task learning is achieved, effectively improving the model's performance and generalization ability. By introducing multidimensional incremental learning and weighting the model according to the characteristics and difficulties of the target task, the model is further optimized, improving its ability to recognize medical text, thereby improving the accuracy and efficiency of medical diagnosis.
[0104] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a first embodiment of a large model parameter fine-tuning device based on incremental learning provided in this application. This device is used to execute the aforementioned large model parameter fine-tuning method based on incremental learning. The device can be configured in a server.
[0105] like Figure 7 As shown, the large model parameter fine-tuning device 300 based on incremental learning includes: a task data acquisition module 301, a data processing module 302, an attention weight calculation module 303, and a target prediction result acquisition module 304.
[0106] The task data acquisition module 301 is used to acquire task data of the target task.
[0107] Data processing module 302 is used to perform data prediction processing on the task data based on the task sub-models at each layer in the large data model, and obtain the model prediction results output by the task sub-models at each layer.
[0108] Attention weight calculation module 303 is used to calculate the attention weight of each layer of task sub-model for the target task based on the weight calculation formula and the data characteristics of the task data;
[0109] The target prediction result acquisition module 304 is used to perform weighted calculation on the model prediction results corresponding to each task sub-model based on the attention weights corresponding to each task sub-model to obtain the target prediction result.
[0110] In one embodiment, the large model parameter fine-tuning device 300 based on incremental learning further includes a model training module, which includes:
[0111] A data acquisition unit is used to acquire a basic large model and task tags and task data corresponding to at least one model task, wherein the basic large model includes at least one basic sub-model.
[0112] The task sub-model acquisition unit is used to perform data training on the basic sub-model based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task, so as to obtain the task sub-model corresponding to each task label.
[0113] The data large model acquisition unit is used to replace the basic sub-model in the basic large model based on the task sub-model to obtain the data large model.
[0114] In one embodiment, the task sub-model acquisition unit includes:
[0115] The multidimensional weight matrix acquisition sub-unit is used to obtain the multidimensional weight matrix based on the matrix product of the basic matrix of the basic sub-model and the low-rank matrix;
[0116] The weight change acquisition subunit is used to adjust the weight parameters of the low-rank matrix in the multidimensional weight matrix based on the data characteristics of the current task, so as to obtain the weight change.
[0117] The target weight acquisition sub-unit is used to calculate the target weight by performing a weighted summation of the weight change and the basic weights of the basic matrix based on the weighted summation formula.
[0118] The task sub-model obtains a sub-unit, which is used to perform parameter iterative adjustment on the multidimensional weight matrix of the basic sub-model based on the target weight, to obtain the task sub-model.
[0119] In one embodiment, the weighted summation formula includes:
[0120]
[0121] Where t is the current task label, and N is the current iteration number. These are the base weights for the current task t. It's the learning rate. It is the change in the weight.
[0122] In one embodiment, the weight calculation formula includes:
[0123]
[0124] Where X is the task data, T is the number of tasks, and ω t L is the attention weight for task t. t It is a task sub-model with task t.
[0125] In one embodiment, the large model parameter fine-tuning device 300 based on incremental learning further includes a task sub-model data query module, which includes:
[0126] The target task tag acquisition unit is used to acquire the target task tag in the model prediction result query request when a model prediction result query request is received from the user terminal.
[0127] The model prediction result query unit is used to query the task sub-model corresponding to the target task label based on the target task label, and output the model prediction result of the task sub-model.
[0128] In one embodiment, the large model parameter fine-tuning device 300 based on incremental learning further includes an attention weight verification module, which is used to verify, calculate and adjust the attention weights corresponding to the task sub-models of each layer based on cross-validation.
[0129] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the device and each module described above can be referred to the corresponding process in the aforementioned embodiment of the large model parameter fine-tuning method based on incremental learning, and will not be repeated here.
[0130] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 8 It runs on the computer device shown.
[0131] Please see Figure 8 , Figure 8 This is a schematic block diagram illustrating the structure of a computer device according to an embodiment of this application. The computer device may be a server.
[0132] See Figure 8 The computer device includes a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.
[0133] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any method for fine-tuning large model parameters based on incremental learning.
[0134] The processor provides computing and control capabilities, supporting the operation of the entire computer device.
[0135] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When these computer programs are executed by the processor, the processor can perform any method for fine-tuning the parameters of a large model based on incremental learning.
[0136] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0137] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.
[0138] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps:
[0139] Obtain the task data for the target task;
[0140] Based on the task sub-models at each layer in the large data model, data prediction processing is performed on the task data to obtain the model prediction results output by each task sub-model.
[0141] Based on the weight calculation formula and the data characteristics of the task data, the attention weights of each layer of task sub-models for the target task are calculated.
[0142] Based on the attention weights corresponding to each task sub-model, the model prediction results corresponding to each task sub-model are weighted and calculated to obtain the target prediction result.
[0143] In one embodiment, before acquiring the task data for the target task, the processor is further configured to:
[0144] Obtain a basic large model and task labels and task data corresponding to at least one model task, wherein the basic large model includes at least one basic sub-model;
[0145] Based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task, the basic sub-model is trained to obtain the task sub-model corresponding to each task label.
[0146] Based on the task sub-model, the basic sub-model in the basic large model is replaced to obtain the data large model.
[0147] In one embodiment, when the processor performs data training on the basic sub-model based on the low-rank matrix corresponding to each of the task labels and the task data corresponding to each of the model tasks to obtain the task sub-model corresponding to each of the task labels, it is configured to:
[0148] A multidimensional weight matrix is obtained based on the matrix product of the fundamental matrix of the basic sub-model and the low-rank matrix.
[0149] Based on the data characteristics of the current task, the weight parameters of the low-rank matrix in the multidimensional weight matrix are adjusted to obtain the weight change.
[0150] Based on the weighted summation formula, the weight change and the basic weights of the basic matrix are weighted and summed to obtain the target weight;
[0151] Based on the target weights, the multidimensional weight matrix of the basic sub-model is iteratively adjusted to obtain the task sub-model.
[0152] In one embodiment, the weighted summation formula includes:
[0153]
[0154] Where t is the current task label, and N is the current iteration number. These are the base weights for the current task t. It's the learning rate. It is the change in the weight.
[0155] In one embodiment, the weight calculation formula includes:
[0156]
[0157] Where X is the task data, T is the number of tasks, and ω t L is the attention weight for task t. t It is a task sub-model with task t.
[0158] In one embodiment, after implementing the task sub-models at each layer of the large data model, performing data prediction processing on the task data, and obtaining the model prediction results output by each task sub-model, the processor is further configured to implement:
[0159] Upon receiving a model prediction result query request from a user, the target task label in the model prediction result query request is obtained.
[0160] Based on the target task label, query the task sub-model corresponding to the target task label, and output the model prediction result of the task sub-model.
[0161] In one embodiment, after implementing the calculation of the attention weights of each layer of task sub-models for the target task based on the weight calculation formula and the data features of the task data, the processor is further configured to implement:
[0162] Based on cross-validation, the attention weights corresponding to the task sub-models of each layer are verified, calculated, and adjusted.
[0163] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the large model parameter fine-tuning methods based on incremental learning provided in the embodiments of this application.
[0164] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiments, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, SmartMedia Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the computer device.
[0165] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for fine-tuning large model parameters based on incremental learning, characterized in that, The method includes: Obtain a basic large model and task labels and task data corresponding to at least one model task, wherein the basic large model includes at least one basic sub-model; Based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task, the basic sub-model is trained to obtain the task sub-model corresponding to each task label. Based on the task sub-model, the basic sub-model in the basic large model is replaced to obtain the data large model; Obtain task data for the target task, wherein the task data is medical text data; Based on the task sub-models at each layer in the large data model, data prediction processing is performed on the task data to obtain the model prediction results output by each task sub-model. Based on the weight calculation formula and the data characteristics of the task data, the attention weights of each layer of task sub-models for the target task are calculated. Based on the attention weights corresponding to each task sub-model, the model prediction results corresponding to each task sub-model are weighted and calculated to obtain the target prediction result. The step of training the basic sub-model based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task to obtain the task sub-model corresponding to each task label includes: A multidimensional weight matrix is obtained based on the matrix product of the fundamental matrix of the basic sub-model and the low-rank matrix. Based on the data characteristics of the current task, the weight parameters of the low-rank matrix in the multidimensional weight matrix are adjusted to obtain the weight change. Based on the weighted summation formula, the weight change and the basic weights of the basic matrix are weighted and summed to obtain the target weight; Based on the target weights, the multidimensional weight matrix of the basic sub-model is iteratively adjusted to obtain the task sub-model.
2. The method for fine-tuning large model parameters based on incremental learning according to claim 1, characterized in that, The weighted summation formula include: in, t The current task label, N This represents the current iteration number. These are the base weights for the current task t. It's the learning rate. It is the change in the weight.
3. The method for fine-tuning large model parameters based on incremental learning according to claim 1, characterized in that, The weight calculation formula includes: Where X is the task data, and T is the number of tasks. These are the attention weights for task t. It is a task sub-model with task t.
4. The method for fine-tuning large model parameters based on incremental learning according to claim 1, characterized in that, After the task data is processed by the task sub-models at each layer of the large data model to obtain the model prediction results output by each task sub-model, the method further includes: Upon receiving a model prediction result query request from a user, the target task label in the model prediction result query request is obtained. Based on the target task label, query the task sub-model corresponding to the target task label, and output the model prediction result of the task sub-model.
5. The method for fine-tuning large model parameters based on incremental learning according to any one of claims 1-4, characterized in that, After calculating the attention weights of each task sub-model for the target task based on the weight calculation formula and the data features of the task data, the method further includes: Based on cross-validation, the attention weights corresponding to the task sub-models of each layer are verified, calculated, and adjusted.
6. A device for fine-tuning large model parameters based on incremental learning, characterized in that, The apparatus for fine-tuning large model parameters based on incremental learning includes: A model training module is used to obtain a basic large model and task labels and task data corresponding to at least one model task, wherein the basic large model includes at least one basic sub-model; based on the low-rank matrix corresponding to each task label and the task data corresponding to each model task, the basic sub-model is trained to obtain task sub-models corresponding to each task label; based on the task sub-models, the basic sub-models in the basic large model are replaced to obtain a data large model; The task data acquisition module is used to acquire task data of the target task, wherein the task data is medical text data; The data processing module is used to perform data prediction processing on the task data based on the task sub-models at each layer in the large data model, and to obtain the model prediction results output by the task sub-models at each layer. The attention weight calculation module is used to calculate the attention weight of each layer of task sub-model for the target task based on the weight calculation formula and the data characteristics of the task data. The target prediction result acquisition module is used to perform weighted calculation on the model prediction results corresponding to each task sub-model based on the attention weights corresponding to each task sub-model to obtain the target prediction result. The model training module is further configured to: obtain a multidimensional weight matrix based on the matrix product of the base matrix and the low-rank matrix of the base sub-model; adjust the weight parameters of the low-rank matrix in the multidimensional weight matrix based on the data features of the current task to obtain the weight change; calculate the weight change and the base weights of the base matrix using a weighted summation formula to obtain the target weight; and iteratively adjust the parameters of the multidimensional weight matrix of the base sub-model based on the target weight to obtain the task sub-model.
7. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the large model parameter fine-tuning method based on incremental learning as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the large model parameter fine-tuning method based on incremental learning as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Training method and device of multi-task pre-training model, electronic equipment and medium
CN113704388A
System and Method for Low Rank Training of Neural Networks
US20230057387A1