Medical analysis-oriented large language model fine tuning method and related equipment

The large language model is trained through sparse masks and adaptive low-rank adaptation algorithms, which reduces the scale of training parameters, solves the problem of high computing resource consumption in medical analysis, and realizes efficient medical analysis on resource-limited devices.

CN120373388APending Publication Date: 2025-07-25HUNAN UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510325831.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing large language model consumes a lot of computing resources in medical analysis, resulting in limited applications on devices with limited resources. The traditional fine-tuning methods consume huge computing resources, excessive video memory usage, and low fine-tuning efficiency.

Method used

By obtaining the sparse mask to describe the training parameter position of the large language model, combining the adaptive low-rank adaptation algorithm and sparse mask training, a joint optimization goal is built to obtain the adaptive low-rank adapter and sparse adapter, and merge it into the final large language model to reduce the training parameter scale, realize low-rank decomposition, and reduce computing resource consumption.

Benefits of technology

While improving the accuracy of fine-tuning of large language models, it effectively reduces computing resource consumption and is suitable for medical analysis using equipment with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373388A_ABST
    Figure CN120373388A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language model fine tuning, and provides a medical analysis-oriented large language model fine tuning method and related equipment. The method comprises the steps that sparse masks used for describing the positions of multiple training parameters of the large language model are acquired; training the large language model according to a self-adaptive low-rank adaptation algorithm to obtain a self-adaptive low-rank adapter; training the large language model according to the sparse mask to obtain a sparse adapter; constructing a joint optimization target of the self-adaptive low-rank adapter and the sparse adapter, and performing minimization solution on the joint optimization target to obtain a final self-adaptive low-rank adapter and a final sparse adapter; and combining the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain a final large language model. According to the method, the computing resource consumption of medical analysis can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of large language model fine-tuning, and particularly to a large language model fine-tuning method and related devices for medical analysis. Background Art

[0002] With the rapid growth of medical data and the progress of computing technology, how to efficiently extract useful information from massive medical data has become an important issue in the current medical field. Although traditional medical image analysis and medical record information processing methods have made progress, they still face challenges such as data diversity and high processing complexity. Especially for doctors, quickly and accurately diagnosing diseases and proposing treatment plans are the keys to improving medical efficiency and patient treatment outcomes.

[0003] In recent years, with the rapid development of large-scale pre-trained models (LLMs, Large Language Models) in the field of natural language processing, models based on the Transformer architecture (such as the bidirectional encoder representations from Transformers, Generative Pretrained Transformer, Text-to-Text Transfer Transformer, etc.) have achieved remarkable success in multiple tasks. These models can capture rich language structures and context information through pre-training on massive text data and possess strong generalization capabilities.

[0004] The application of these advanced large language models (LLMs, Large Language Model) in the field of medical diagnosis is accompanied by huge computational and memory costs, especially when training these models from scratch. In this context, fine-tuning the LLM using a limited amount of data has become an effective and popular method, which can significantly improve the performance of the model in specific medical tasks (such as disease prediction, auxiliary diagnosis of image analysis, etc.) or make the LLM better adapt to the user needs in the medical scenario. However, traditional fine-tuning methods usually require updating all parameters of the entire model, which often leads to problems such as huge consumption of computing resources, excessive video memory occupation, and low fine-tuning efficiency in practical applications. These challenges limit the wide application of large language models on resource-limited devices (such as edge devices in hospitals, low-configuration servers, etc.). Although purchasing high-performance servers can alleviate this problem to a certain extent, it will greatly increase the computing power cost, which is unacceptable for medical-centered institutions. Thus, there is a problem of large consumption of computing resources for medical analysis in the current large language model fine-tuning method for medical analysis. Summary of the Invention

[0005] This application provides a large language model fine-tuning method for medical analysis, which can solve the problem of large consumption of computing resources for medical analysis.

[0006] In a first aspect, an embodiment of the present application provides a method for fine-tuning a large language model for medical analysis, and the method for fine-tuning the large language model includes:

[0007] Obtain a sparse mask for describing the positions of multiple training parameters of the large language model; the large language model data is used to analyze medical data, and the importance value of the training parameters is greater than each other parameter in the large language model;

[0008] Train the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter; the adaptive low-rank adapter is the large language model after training the parameters according to the adaptive low-rank adaptation algorithm;

[0009] Train the large language model according to the sparse mask to obtain a sparse adapter; the sparse adapter is the large language model after training the parameters according to the sparse mask;

[0010] Construct a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and perform a minimization solution on the joint optimization objective to obtain a final adaptive low-rank adapter and a final sparse adapter; the joint optimization objective is used to describe the accuracy of the adaptive low-rank adapter and the sparse adapter in performing medical data analysis;

[0011] Merge the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain a final large language model.

[0012] Optionally, obtaining a sparse mask for describing the positions of multiple training parameters of the large language model includes:

[0013] Calculate the importance values of all parameters in the large language model;

[0014] Sort all the importance values in descending order, and use the first multiple importance values as the target importance values, and use the parameters corresponding to each target importance value as the training parameters;

[0015] Generate a sparse mask for describing the positions of all training parameters.

[0016] Optionally, calculating the importance values of all parameters in the large language model includes:

[0017] Through the formula:

[0018]

[0019] Calculate the importance value matrix of all parameters

[0020] Where N represents the number of samples in the training dataset, y represents the output of the large language model, Represents the gradient operation on the parameter vector, xr denotes the r-th sample, p θ (y|x r ) represents the output distribution generated on y after inputting x r into the large language model, and θ represents the parameter vector composed of all parameters of the large language model.

[0021] Optionally, train the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter, including:

[0022] For each weight matrix in the large language model, update the weight matrix according to the adaptive low-rank adaptation algorithm to obtain the updated weight matrix;

[0023] Substitute all the updated weight matrices into the large language model to obtain the adaptive low-rank adapter.

[0024] Optionally, update the weight matrix according to the adaptive low-rank adaptation algorithm to obtain the updated weight matrix, including:

[0025] Perform singular value decomposition on the incremental matrix of the weight matrix to obtain the left singular vector, right singular vector, and singular value matrix of the incremental matrix;

[0026] For each element in the left singular vector, take the elements corresponding to the element in the right singular vector and singular value matrix as target elements, and combine the element with the two target elements to form a triple, and take the target element in the singular value matrix as the importance score of the triple;

[0027] Increment the iteration count by one, prune all triples according to the importance scores of all triples to obtain multiple new triples;

[0028] Calculate the budget values of all new triples and determine whether the budget values of all new triples are less than or equal to the preset budget value;

[0029] If so, calculate the final incremental matrix based on all new triples, and update the weight matrix according to the final incremental matrix to obtain the updated weight matrix;

[0030] Otherwise, take the multiple new triples as all triples, and return the step of incrementing the iteration count by one and pruning all triples according to the importance scores of all triples to obtain multiple new triples.

[0031] Optionally, prune all triples according to the importance scores of all triples to obtain multiple new triples, including:

[0032] Through the formula:

[0033]

[0034] Calculate the singular value matrix of the k-th weight matrix at the (t + 1)-th iteration

[0035] where represents the intermediate singular value matrix of the k-th weight matrix at the t-th iteration, represents all importance scores, k = 1, 2,..., K, where K represents the total number of weight matrices, represents the singular value pruning result, represents the intermediate singular value corresponding to the i-th triple of the k-th weight matrix at the t-th iteration, represents the importance score corresponding to the i-th triple of the k-th weight matrix at the t-th iteration, b (t) represents the budget of the remaining singular values at the t-th iteration, S (t) represents the set of all importance scores at the t-th iteration, represents the singular value matrix of the k-th weight matrix at the t-th iteration, η represents the learning rate, represents the training objective at the t-th iteration:

[0036]

[0037] where represents the training cost at the t-th iteration, represents the left singular vectors of all weight matrices at the t-th iteration, ε t represents the singular value matrices of all weight matrices at the t-th iteration, represents the right singular vectors of all weight matrices at the t-th iteration, γ represents the regularization coefficient, n represents the number of weight matrices, represents the regularizer:

[0038]

[0039] where represents the left singular vector of the k-th weight matrix at the t-th iteration, represents the right singular vector of the k-th weight matrix at the t-th iteration, I represents the regularization parameter,

[0040] Take the triples corresponding to the non-zero elements in the singular value matrix of the k-th weight matrix at the (t + 1)-th iteration as new triples.

[0041] Optionally, train the large language model according to the sparse mask to obtain a sparse adapter, including:

[0042] Update only all the training parameters according to the positions of all the training parameters in the large language model described by the sparse mask to obtain a sparse adapter.

[0043] Optionally, the joint optimization objective is:

[0044]

[0045] where, Δ L represents all the parameter perturbations between the adaptive low-rank adapter and the large language model, and Δ S represents all the parameter perturbations between the sparse adapter and the large language model. represents the data set, represents the remaining parameters in the large language model that are not updated, represents the loss function, s.t. represents the constraint, F represents the number of parameters in the large language model, represents the f-th parameter perturbation in Δ L , represents the f-th parameter perturbation in Δ S , d represents the perturbation density, r represents a constant, m f represents the number of rows of the sparse matrix, and n f represents the number of columns of the sparse matrix.

[0046] In a second aspect, an embodiment of the present application provides a large language model fine-tuning device for medical analysis, including:

[0047] An acquisition module that acquires a sparse mask for describing the positions of multiple training parameters of the large language model; the large language model data is used to analyze medical data, and the importance value of the training parameters is greater than each other parameter in the large language model;

[0048] A first training module that trains the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter; the adaptive low-rank adapter is the large language model after training the parameters according to the adaptive low-rank adaptation algorithm;

[0049] A second training module that trains the large language model according to the sparse mask to obtain a sparse adapter; the sparse adapter is the large language model after training the parameters according to the sparse mask;

[0050] A construction module that constructs a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and performs a minimization solution on the joint optimization objective to obtain a final adaptive low-rank adapter and a final sparse adapter; the joint optimization objective is used to describe the accuracy of the adaptive low-rank adapter and the sparse adapter for medical data analysis;

[0051] A merging module merges the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain the final large language model.

[0052] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned fine-tuning method of the large language model for medical analysis is implemented.

[0053] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned fine-tuning method of the large language model for medical analysis is implemented.

[0054] The above solution of the present application has the following beneficial effects:

[0055] In the embodiment of the present application, a sparse mask for describing the positions of multiple training parameters of the large language model is obtained, then the large language model is trained according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter, and then the large language model is trained according to the sparse mask to obtain a sparse adapter. Then, a joint optimization objective of the adaptive low-rank adapter and the sparse adapter is constructed, and the joint optimization objective is minimized to obtain the final adaptive low-rank adapter and the final sparse adapter. Finally, the final adaptive low-rank adapter and the final sparse adapter are merged with the large language model to obtain the final large language model. Among them, the sparse matrix describes the positions of multiple training parameters, which is convenient for training only the training parameters, reduces the scale of parameters participating in training, uses the adaptive low-rank adaptation algorithm for parameter training, realizes low-rank decomposition during parameter training, and avoids resource overhead caused by large-scale parameter updates. Merging the training results of the two trainings with the large language model can improve the accuracy of fine-tuning the large language model while effectively reducing the consumption of computing resources. Using the final large language model for medical analysis reduces the computing resources required for medical analysis.

[0056] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation part. Description of the Drawings

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0058] Figure 1Flowchart of the fine-tuning method for large language models for medical analysis provided by an embodiment of the present application;

[0059] Figure 2 Schematic diagram of the specific process of the sparsification fine-tuning method provided by an embodiment of the present application;

[0060] Figure 3 Schematic diagram of the structure of the large language model fine-tuning device for medical analysis provided by an embodiment of the present application;

[0061] Figure 4 Terminal device provided by an embodiment of the present application. Detailed implementation manners

[0062] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0063] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0064] It should also be understood that the term " / and" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0065] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when...", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0066] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0067] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0068] In view of the problem of high computational resource consumption in existing medical analysis, an embodiment of this application provides a fine-tuning method for a large language model for medical analysis. This large language model fine-tuning method obtains a sparse mask for describing the positions of multiple training parameters of the large language model, then trains the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter, then trains the large language model according to the sparse mask to obtain a sparse adapter, then constructs a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and minimizes the joint optimization objective to obtain a final adaptive low-rank adapter and a final sparse adapter. Finally, the final adaptive low-rank adapter and the final sparse adapter are merged with the large language model to obtain a final large language model. Among them, the sparse matrix describes the positions of multiple training parameters, facilitating training only on the training parameters, reducing the scale of parameters participating in training. Using the adaptive low-rank adaptation algorithm for parameter training realizes low-rank decomposition during parameter training, avoiding resource overhead caused by large-scale parameter updates. Merging the training results of the two trainings with the large language model can effectively reduce computational resource consumption while improving the accuracy of large language model fine-tuning. Using the final large language model for medical analysis reduces the computational resources required for medical analysis.

[0069] Next, an exemplary description will be given of the fine-tuning method for a large language model for medical analysis provided in this application.

[0070] As Figure 1 shown, the fine-tuning method for a large language model for medical analysis provided in this application includes the following steps:

[0071] Step 11, obtain a sparse mask for describing the positions of multiple training parameters of the large language model.

[0072] Large language model data is used to analyze medical data (such as analyzing blood pressure changes, heart rate changes). If the large language model is a Generative Pre-trained Transformer (GPT), this large language model can analyze the historical medical data of the research object (such as the disease history, visit frequency, etc. of the research object) and historical physical data (such as the blood pressure value, heart rate, etc. of the research object) to summarize and predict the future health indicators of the research object (such as blood pressure value, heart rate, etc.), so that medical staff can implement corresponding measures based on the model output results. The importance value of the training parameters is greater than each other parameter in the large language model.

[0073] In some embodiments of the present application, the step of obtaining the sparse mask for describing the positions of multiple training parameters of the large language model includes:

[0074] The first step is to calculate the importance values of all parameters in the large language model.

[0075] Specifically, through the formula:

[0076]

[0077] Calculate the importance value matrix of all parameters

[0078] where N represents the number of samples in the training dataset, y represents the output of the large language model, represents the gradient operation on the parameter vector, that is, taking the partial derivative with respect to θ, x r represents the r-th sample, p θ (y|x r ) represents the output distribution generated on y after inputting x r into the large language model, and θ represents the parameter vector composed of all parameters of the large language model.

[0079] Exemplarily, the samples in the above training dataset are the historical medical data and historical physical data of the research object.

[0080] The second step is to sort all the importance values in descending order, and take the top multiple importance values as the target importance values, and take the parameters corresponding to each target importance value as the training parameters.

[0081] The third step is to generate a sparse mask for describing the positions of all training parameters.

[0082] Exemplarily, the TopK-Mask function can be used to obtain the position information of all training parameters, and all the position information is combined into one data to obtain the sparse mask.

[0083] In some embodiments of the present application, a sparse fine-tuning method SpA (Sparse Adaptive) can also be used to obtain a sparse mask for describing the positions of multiple training parameters of a large language model. SpA adopts a strategy of adaptive sparse mask generation. It analyzes the gradient changes of the weights during the training process, dynamically identifies which parameters are most critical to the task effect, and preferentially updates these parameters. SpA pre-computes a sparse subset of the existing parameters for updating and keeps the subset fixed during multiple iterative training processes. Doing so has many benefits. It not only avoids the increase in the total model size but also ensures that this process is model-independent by avoiding manual definition of the mask. Since the sparse mask is pre-computed and fixed, the computational and memory overhead of updating the sparse mask is saved during the training process.

[0084] The above SpA method is exemplarily described below with a specific example.

[0085] The specific process of the sparse fine-tuning method SpA is as Figure 2 shown. After starting, the input fully connected weights and other parameters are obtained, a sub-dataset for generating the mask, the loss function and density are generated, the weight gradient list (and importance value list) is initialized to zero, and it is judged whether the sub-data has been traversed. If not, the passed loss function is used to calculate the gradients (i.e., importance values) according to the dataset, the fully connected weights, and the remaining parameters, and the gradients are accumulated into the weight gradient list. If so, the obtained weight gradient list is traversed, the positions of the top k elements with the largest gradients are respectively obtained, a sparse mask is generated, and the obtained sparse mask is returned to end.

[0086] Step 12, training the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter.

[0087] The adaptive low-rank adapter is a large language model after training the parameters according to the adaptive low-rank adaptation algorithm.

[0088] In some embodiments of the present application, the step of training the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter includes:

[0089] The first step is to update each weight matrix in the large language model according to the adaptive low-rank adaptation algorithm to obtain an updated weight matrix.

[0090] First, perform singular value decomposition on the incremental matrix of the weight matrix to obtain the left singular vector, right singular vector, and singular value matrix of the incremental matrix.

[0091] It should be noted that the weight matrix is the parameter matrix to be trained in the large language model. For example, for a large language model based on Transformer, it consists of L stacked blocks, and each block contains two sub-modules: the multi-head attention sub-module (MHA, Multi-Head Attention) and the feed-forward network sub-module (FFN, Feed-Forward Network). MHA executes the attention function in h heads in parallel:

[0092] MHA(X) = Concat(head1,…,head h )W o

[0093]

[0094] where W o ∈R d×d is the output projection matrix, is the query, key, and value projection matrix of head i . d h is usually set to d / h. Another important sub-module is an FFN, which consists of two linear transformations with a ReLU activation function between the two linear transformations: where Finally, residual connections are used, followed by layer normalization.

[0095] In the above expressions, matrices such as W o , etc. are all weight matrices of the large language model, that is, the parameter matrices that need to be updated and trained when fine-tuning the large language model.

[0096] Then, for each element in the left singular vector, the corresponding elements in the right singular vector and the singular value matrix are used as the target elements, and the element is combined with the two target elements to form a triple, and the target element in the singular value matrix is used as the importance score of the triple.

[0097] Then, the iteration count is incremented by one, and all triples are pruned according to the importance scores of all triples to obtain multiple new triples.

[0098] The initial iteration count is 0.

[0099] Specifically, through the formula:

[0100]

[0101] calculate the singular value matrix of the k-th weight matrix at the (t + 1)-th iteration

[0102] Among them, represents the intermediate singular value matrix of the k-th weight matrix in the t-th iteration, represents all importance scores, k = 1, 2,..., K, where K represents the total number of weight matrices, represents the singular value pruning result, represents the intermediate singular value corresponding to the i-th triple of the k-th weight matrix in the t-th iteration, represents the importance score corresponding to the i-th triple of the k-th weight matrix in the t-th iteration, b (t) represents the budget of the remaining singular values in the t-th iteration, S (t) represents the set of all importance scores in the t-th iteration, represents the singular value matrix of the k-th weight matrix in the t-th iteration, η represents the learning rate, represents the training objective in the t-th iteration:

[0103]

[0104] Among them, represents the training cost in the t-th iteration, represents the left singular vectors of all weight matrices in the t-th iteration, ε t represents the singular value matrices of all weight matrices in the t-th iteration, represents the right singular vectors of all weight matrices in the t-th iteration, γ represents the regularization coefficient, and n represents the number of weight matrices, represents the regularizer:

[0105]

[0106] Among them, represents the left singular vector of the k-th weight matrix in the t-th iteration, represents the right singular vector of the k-th weight matrix in the t-th iteration, I represents the regularization parameter,

[0107] Take the triples corresponding to the non-zero elements in the singular value matrix of the k-th weight matrix in the (t + 1)-th iteration as new triples.

[0108] Calculate the budget values of all new triples, and determine whether the budget values of all new triples are less than or equal to the preset budget value;

[0109] If so, calculate the final increment matrix based on all new triples, and update the weight matrix according to the final increment matrix to obtain the updated weight matrix.

[0110] Specifically, elements belonging to the left singular vector in all new triples are used to form a new left singular vector, elements belonging to the right singular vector in all new triples are used to form a new right singular vector, and elements belonging to the singular value matrix in all new triples are used to form a new singular value matrix. Then, through the reverse singular value decomposition algorithm, the final incremental matrix is calculated based on the new left singular vector, the new right singular vector, and the new singular value matrix.

[0111] Otherwise, use the multiple new triples as all triples, return the iteration count incremented by one, and prune all triples according to the importance scores of all triples to obtain the step of multiple new triples.

[0112] It should be noted that the above budget value is used to describe the number of parameters that can be updated, and the budget value can be calculated using the method of calculating the budget in the Adaptive Low-Rank Adaptation (AdaLoRA) algorithm.

[0113] In the second step, substitute all updated weight matrices into the large language model to obtain the adaptive low-rank adapter.

[0114] Specifically, replace the corresponding original weight matrix in the large language model with the updated weight matrix to obtain the adaptive low-rank adapter.

[0115] It can be understood that the above process is the calculation process of the adaptive low-rank algorithm.

[0116] It is worth mentioning that AdaLoRA introduces an adaptive rank adjustment mechanism. Specifically, AdaLoRA dynamically adjusts the rank size of the low-rank matrix for each layer according to the gradient norm of that layer. Layers with a large gradient norm usually indicate that the layer makes a greater contribution to the task, so a higher rank will be assigned; conversely, layers with a smaller gradient will be assigned a lower rank. Through this adaptive adjustment, AdaLoRA can allocate computing resources more flexibly, improving the performance of the model on complex tasks while maintaining computing efficiency.

[0117] Step 13, train the large language model according to the sparse mask to obtain the sparse adapter.

[0118] The sparse adapter is the large language model after training the parameters according to the sparse mask.

[0119] Specifically, according to the positions of all training parameters in the large language model described by the sparse mask, only update all training parameters to obtain the sparse adapter.

[0120] Exemplarily, model training methods such as the gradient descent method can be used to update all training parameters to obtain the sparse adapter.

[0121] It should be noted that when the method of this application is run on some devices and systems, the sparse matrix (i.e., the matrix corresponding to all parameters in the large language model, where the elements corresponding to the training parameters in the sparse matrix are non-zero, and the elements corresponding to other parameters that are not training parameters are zero) needs to be stored. The traditional sparse matrix is stored in the coordinate (COO, coordinate format) format, which stores the non-zero elements in the matrix in the form of coordinates. Two integer arrays of length n are used to represent the row index and column index respectively, and a real number array of length n is used to represent the non-zero elements of the matrix, where n is the number of non-zero elements in the sparse matrix. Although this storage is simple, intuitive, and easy to construct, this format will cause the problem of storage redundancy, and the query speed will also be relatively slow.

[0122] The method of this application stores the sparse matrix in the Compressed Sparse Row (CSR) format. The CSR format is an improvement over the COO format. This format requires the matrix elements to be stored in row order. It uses three lists to represent an m*n sparse matrix with nnz non-zero values: a value list of size nnz, which stores all the non-zero values in the sparse matrix; a row offset list of size m+1, which is just the position of the first non-zero element in each row of the value list; and a column index list of size nnz, which contains the column indices of each corresponding element in the value list. In addition, it also includes an additional row index list of size m, which sorts the rows according to the non-zero element count.

[0123] It is worth mentioning that by using a sparse mask to screen out important parameters in the parameters of the large language model for update and not updating other parameters, the computational burden can be reduced.

[0124] Step 14, construct a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and minimize the joint optimization objective to obtain the final adaptive low-rank adapter and the final sparse adapter.

[0125] The joint optimization objective is used to describe the accuracy of the adaptive low-rank adapter and the sparse adapter in performing medical data analysis.

[0126] Specifically, the joint optimization objective is:

[0127]

[0128] where, Δ L represents all parameter perturbations between the adaptive low-rank adapter and the large language model, and Δ S represents all parameter perturbations between the sparse adapter and the large language model, represents the data set, Denote the remaining parameters in the large language model that are not updated. Denote the loss function, s.t. denotes the constraint, and F denotes the number of parameters in the large language model. Denote Δ L The parameter perturbation of the f-th parameter in Denote Δ S The parameter perturbation of the f-th parameter in, d denotes the perturbation density, r denotes a constant, and m f Denote the number of rows of the sparse matrix, and n f Denote the number of columns of the sparse matrix.

[0129] Exemplarily, the above dataset is the historical medical data and historical physical data of multiple research objects. The gradient descent method, Newton's method, etc. can be used to minimize the joint optimization objective. When the joint optimization objective is not minimized, steps 12 and 13 will be returned to retrain the adaptive low-rank adapter and sparse adapter at this time. When the joint optimization objective is minimized, the adaptive low-rank adapter at this time is used as the final adaptive low-rank adapter, and the sparse adapter is used as the final sparse adapter.

[0130] It should be noted that assume N represents the large language model (LLM), and let Denote the sequence containing All weight parameters of, where Let the vector Denote The remaining parameters of (including biases, normalization parameters, etc.) are concatenated into a single vector. Given the dataset And the loss function In Full fine-tuning on can be expressed by the formula as solving the following optimization problem:

[0131]

[0132] Considering that the LLM usually contains billions of parameters, performing full fine-tuning may be slow and computationally expensive. This usually makes it impossible to execute on ordinary computing devices. Let Δ = {Δ1, Δ2,..., Δ F} contain the perturbations to the original weight parameters, and Is defined as the sum of their elements, that is, Similarly, let the vector Denote The perturbation of. Then the adapted parameters are found by solving the following optimization problem:

[0133]

[0134] Where is a set of constraints on perturbations, which is to reduce the memory requirements and computational complexity of the optimization problem. Without these constraints, this adapter is equivalent to full fine-tuning.

[0135] Since it is observed that the difference between the original parameters and the fully fine-tuned parameters is approximately low-rank, low-rank adaptation restricts the perturbations in full fine-tuning to a low-rank form, and its optimization objective is:

[0136]

[0137] where r is a fixed small number, and this method reduces the number of trainable weights in the f-th layer from m f n f to r(m f +n f ), thus achieving faster and more memory-saving fine-tuning.

[0138] ADALoRA, on the other hand, changes the way of fixed-rank allocation for each weight matrix in low-rank adaptation (LoRA) to dynamically allocate budgets according to the importance of each weight matrix, so that more budgets can be allocated to weight matrices with higher importance, thus achieving better performance than LoRA. Its optimization objective is the same as LoRA, but the perturbations restricted for full fine-tuning change from two low-rank matrices to the form of singular value decomposition.

[0139] The optimization objective of the sparse adapter is:

[0140]

[0141] where d < 1 represents the perturbation density, ||·||0 represents the l0 norm, and only the case where each perturbation has a fixed set of positions throughout the training process is considered. In this way, the number of trainable parameters is reduced by d times.

[0142] When facing complex downstream tasks with the LoRA method, these methods usually cannot achieve an accuracy level comparable to full fine-tuning. This problem occurs because this technique often filters out key information during normalization in the fine-tuning stage. When performing singular value decomposition on the parameter update Δ * of the large model, this filtering problem becomes particularly obvious. Although Δ * exhibits a low-rank structure during the update, it is not strictly low-rank. The characteristic of this difference is that there are a part of singular values, although their magnitudes are relatively small, but they are still non-zero values.

[0143] Robust principal component analysis provides an alternative method to effectively decompose a matrix into two components: a low-rank matrix and a sparse matrix. Compared with individual low-rank methods, this decomposition provides a finer approximation for fine-tuning updates.

[0144] Based on this, the above optimization objectives can be combined to obtain a combined optimization objective.

[0145] Step 15: Merge the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain the final large language model.

[0146] Specifically, the change amounts of parameters between the final adaptive low-rank adapter and the large language model, and the change amounts of parameters between the final sparse adapter and the large language model can be averaged, and then the average value is summed with the original parameters in the large language model to obtain the final large language model.

[0147] It should be noted that after obtaining the final large language model, the final large language model is used to analyze medical data. For example, the final GPT model is used to process the historical medical data and historical physical data of the target person who needs to have their health condition analyzed to obtain the final health condition analysis result of the target person. This final health condition analysis result can describe the future health status of the target person, such as blood pressure value, heart rate, etc. at a future moment.

[0148] It is worth mentioning that the sparse matrix describes the positions of multiple training parameters, facilitating training only on the training parameters, reducing the scale of parameters participating in training. Using the adaptive low-rank adaptation algorithm for parameter training realizes low-rank decomposition during parameter training, avoiding resource overhead caused by large-scale parameter updates. Combining the training results of the two types of training with the large language model can improve the accuracy of fine-tuning the large language model while effectively reducing computational resource consumption. Using the final large language model for medical analysis reduces the computational resources required for medical analysis.

[0149] In addition, for the method of fine-tuning the large language model for complex tasks in a specific field in this application, combining the flexibility and memory advantages of the low-rank adaptation of AdaLoRA with the extremely high sparsity level of sparse adaptation SPA, the updates of full fine-tuning are obtained through a low-rank matrix plus a sparse matrix to achieve a better fit than existing methods. Especially in the case of complex tasks, AdaLoRA pays more attention to enhancing the model's expressive ability by dynamically adjusting the rank of the low-rank matrix, while SpA optimizes and screens out important parameters through a sparse mask, thereby achieving the high-precision processing ability of the fine-tuned large language model for complex tasks. The combination of the two can achieve dual optimization, enabling more efficient fine-tuning of the large language model on devices with limited resources and being effectively applied to medical analysis.

[0150] An exemplary description of the large language model fine-tuning device for medical analysis provided in this application will be given below.

[0151] As Figure 3 shown, an embodiment of this application provides a large language model fine-tuning device for medical analysis. The large language model fine-tuning device 300 for medical analysis includes:

[0152] An acquisition module 301 that acquires a sparse mask for describing the positions of multiple training parameters of the large language model; the large language model data is used to analyze medical data, and the importance value of the training parameters is greater than each other parameter in the large language model;

[0153] A first training module 302 that trains the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter; the adaptive low-rank adapter is the large language model after training the parameters according to the adaptive low-rank adaptation algorithm;

[0154] A second training module 303 that trains the large language model according to the sparse mask to obtain a sparse adapter; the sparse adapter is the large language model after training the parameters according to the sparse mask;

[0155] A construction module 304 that constructs a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and performs a minimization solution on the joint optimization objective to obtain a final adaptive low-rank adapter and a final sparse adapter; the joint optimization objective is used to describe the accuracy of the adaptive low-rank adapter and the sparse adapter in performing medical data analysis;

[0156] A merging module 305 that merges the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain a final large language model.

[0157] It should be noted that the information interaction, execution process, etc. between the above-mentioned device / units, due to being based on the same concept as the method embodiment of this application, for their specific functions and the technical effects brought, please refer to the method embodiment part specifically, and will not be elaborated here.

[0158] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the division of the above-mentioned functional units and modules is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment and will not be elaborated here.

[0159] As Figure 4 shown, an embodiment of the present application provides a terminal device. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 4 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, it implements the steps in any of the above method embodiments.

[0160] Specifically, when the processor D100 executes the computer program D102, it obtains a sparse mask for describing the positions of multiple training parameters of the large language model, then trains the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter, and then trains the large language model according to the sparse mask to obtain a sparse adapter. Then, a joint optimization objective of the adaptive low-rank adapter and the sparse adapter is constructed, and the joint optimization objective is minimized to obtain a final adaptive low-rank adapter and a final sparse adapter. Finally, the final adaptive low-rank adapter and the final sparse adapter are merged with the large language model to obtain a final large language model. Among them, the sparse matrix describes the positions of multiple training parameters, which is convenient for training only the training parameters, reduces the scale of parameters participating in the training, uses the adaptive low-rank adaptation algorithm for parameter training, realizes low-rank decomposition during parameter training, and avoids resource overhead caused by large-scale parameter updates. Merging the training results of the two trainings with the large language model can improve the accuracy of the large language model fine-tuning while effectively reducing the consumption of computing resources. Using the final large language model for medical analysis reduces the computing resources required for medical analysis.

[0161] The so-called processor D100 may be a central processing unit (CPU, Central Processing Unit), and this processor D100 may also be other general-purpose processors, digital signal processors (DSP, Digital Signal Processor), application-specific integrated circuits (ASIC, Application Specific Integrated Circuit), field-programmable gate arrays (FPGA, Field-Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0162] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk equipped on the terminal device D10, a smart media card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or will be output.

[0163] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be realized.

[0164] The embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, it enables the terminal device to execute and realize the steps in the above-mentioned various method embodiments.

[0165] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the large language model fine-tuning method device / terminal device for medical analysis, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium cannot be an electrical carrier signal and a telecommunication signal.

[0166] In the above embodiments, the descriptions of the various embodiments each have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0167] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0168] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A fine-tuning method for large language models for medical analysis, characterized in that, Comprising: Obtaining a sparse mask for describing the positions of multiple training parameters of a large language model; The large language model data is used for analyzing medical data, and the importance value of the training parameters is greater than each other parameter in the large language model; Training the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter; The adaptive low-rank adapter is a large language model obtained by training the parameters according to the adaptive low-rank adaptation algorithm; Training the large language model according to the sparse mask to obtain a sparse adapter; The sparse adapter is a large language model obtained by training the parameters according to the sparse mask; Constructing a joint optimization objective for the adaptive low-rank adapter and the sparse adapter, and minimizing the joint optimization objective to obtain a final adaptive low-rank adapter and a final sparse adapter; The joint optimization objective is used to describe the accuracy of the adaptive low-rank adapter and the sparse adapter in performing medical data analysis; Merging the final adaptive low-rank adapter and the final sparse adapter with the large language model to obtain a final large language model.

2. The method for fine-tuning a large language model according to claim 1, wherein The obtaining a sparse mask for describing the positions of multiple training parameters of a large language model includes: Calculating the importance values of all parameters in the large language model; Sorting all the importance values in descending order, taking the first multiple importance values as target importance values, and taking the parameters corresponding to each target importance value as training parameters; Generating a sparse mask for describing the positions of all training parameters.

3. The large language model fine-tuning method according to claim 2, wherein, The calculating the importance values of all parameters in the large language model includes: Through the formula: Calculate the importance value matrix of all parameters Among them, N represents the number of samples in the training dataset, y represents the output of the large language model, represents the gradient operation on the parameter vector, x r represents the r-th sample, p θ (y|x r ) represents the output distribution generated on y after inputting x r into the large language model, and θ represents the parameter vector composed of all parameters of the large language model.

4. The large language model fine-tuning method according to claim 1, wherein The training the large language model according to the adaptive low-rank adaptation algorithm to obtain an adaptive low-rank adapter includes: Respectively for each weight matrix in the large language model, updating the weight matrix according to the adaptive low-rank adaptation algorithm to obtain an updated weight matrix; Substituting all the updated weight matrices into the large language model to obtain an adaptive low-rank adapter.

5. The large language model fine-tuning method according to claim 4, wherein The updating the weight matrix according to the adaptive low-rank adaptation algorithm to obtain an updated weight matrix includes: Performing singular value decomposition on the increment matrix of the weight matrix to obtain the left singular vector, right singular vector and singular value matrix of the increment matrix; Respectively for each element in the left singular vector, taking the elements corresponding to the element in the right singular vector and the singular value matrix as target elements, combining the element with the two target elements to form a triple, and taking the target element in the singular value matrix as the importance score of the triple; Incrementing the number of iterations by one, pruning all the triples according to the importance scores of all the triples to obtain multiple new triples; Calculating the budget values of all the new triples, and determining whether the budget values of all the new triples are less than or equal to a preset budget value; If so, calculating a final increment matrix based on all the new triples, and updating the weight matrix according to the final increment matrix to obtain an updated weight matrix; Otherwise, use the multiple new triples as the all triples, increment the iteration count by one, and perform pruning on all triples according to the importance scores of all triples to obtain the multiple new triples.

6. The method for fine-tuning a large language model according to claim 5, wherein The step of performing pruning on all triples according to the importance scores of all triples to obtain multiple new triples includes: By the formula: Calculate the singular value matrix of the k-th weight matrix at the (t + 1)-th iteration Among them, represents the intermediate singular value matrix of the k-th weight matrix in the t-th iteration, represents all importance scores, where k = 1, 2,..., K, and K represents the total number of weight matrices, represents the singular value pruning result, represents the intermediate singular value corresponding to the i-th triple of the k-th weight matrix in the t-th iteration, represents the importance score corresponding to the i-th triple of the k-th weight matrix in the t-th iteration, b (t) represents the budget of the remaining singular values in the t-th iteration, S (t) represents the set of all importance scores in the t-th iteration, represents the singular value matrix of the k-th weight matrix in the t-th iteration, and η represents the learning rate, represents the training objective in the t-th iteration: Among them, represents the training cost of the t-th iteration, represents the left singular vectors of all weight matrices in the t-th iteration, and ε t represents the singular value matrix of all weight matrices in the t-th iteration, represents the right singular vectors of all weight matrices in the t-th iteration, γ represents the regularization coefficient, and n represents the number of weight matrices, represents the regularizer: Among them, represents the left singular vector of the k-th weight matrix in the t-th iteration, represents the right singular vector of the k-th weight matrix in the t-th iteration, and I represents the regularization parameter, Take the triplets corresponding to the non-zero elements in the singular value matrix of the k-th weight matrix in the (t + 1)-th iteration as new triplets. ​ 7. The large language model fine-tuning method according to claim 1, characterized in that The step of training the large language model according to the sparse mask to obtain a sparse adapter includes: According to the positions of all training parameters in the large language model described by the sparse mask, only update all training parameters to obtain a sparse adapter.

8. The method for fine-tuning a large language model according to claim 1, wherein The joint optimization objective is: Among them, Δ L represents all parameter perturbations between the adaptive low-rank adapter and the large language model, and Δ S represents all parameter perturbations between the sparse adapter and the large language model, represents the dataset, represents the remaining parameters in the large language model that are not updated, represents the loss function, s.t. represents the constraint, F represents the number of parameters in the large language model, represents Δ L the parameter perturbation of the f-th parameter in represents Δ S the parameter perturbation of the f-th parameter described in f d represents the perturbation density, r represents a constant, m f represents the number of rows of the sparse matrix, and n 9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the fine-tuning method for a large language model for medical analysis according to any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the fine-tuning method for a large language model for medical analysis according to any one of claims 1 to 8.

Citation Information

Cited By

  • Multi-mode controllable large model identity tag implantation and transmission method and device based on low-rank fusion, and electronic equipment

    CN121278693A

  • Ultrasonic large model sparse tensor optimization training and pushing acceleration method based on new generation supercomputing

    CN121413782A

  • Ultrasonic large model sparse tensor optimization training acceleration method based on new generation supercomputer

    CN121413782B