Fine-tuning Method and System of a Large Model for Medical Licensing Examination Tutoring Based on Small-Sample Learning
Through a small sample learning method, the oral knowledge examination tutoring model is carefully adjusted, and the existing model has solved the problem of low accuracy in answering questions, achieving higher accuracy and lighter model.
Patent Information
- Application Number
- CN202410967300.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-18
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-18
AI Technical Summary
The existing large prophecy model with oral knowledge is low in accuracy when answering questions and cannot effectively apply knowledge for tutoring.
The precision adjustment method of the medical examination tutoring large model based on small sample learning is used to optimize the accuracy and reasoning speed of the model by obtaining oral knowledge test questions, embedding and clustering of words, calculating attention scores, training and pruning of the model.
It improves the accuracy of the model when answering questions, reduces the annotation cost of sample extraction, reduces the number of parameters of the model, makes the model lighter and easier to deploy and apply.
Smart Images

Figure CN119151737B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of question-and-answer models, and particularly to a fine-tuning method and system for a large model for medical licensing examination tutoring based on few-shot learning. Background Art
[0002] The oral licensing examination for dentists, as an important assessment in the medical field, aims to comprehensively and objectively evaluate whether applicants for dental qualifications truly possess the professional knowledge and skills necessary for practicing dentistry. The oral licensing examination for dentists covers multiple aspects of dentistry, including the diagnosis, treatment, and prevention of oral diseases, as well as knowledge in multiple professional fields such as oral and maxillofacial surgery, prosthodontics, and orthodontics. Candidates need to systematically master this professional knowledge and be able to apply it proficiently in clinical practice to provide high-quality oral medical services to patients.
[0003] In the written examination of the oral licensing examination for dentists, a total of 600 multiple-choice questions need to be answered by candidates within two and a half hours. These questions cover the following three aspects: (1) Basic theoretical knowledge of dentistry: including basic knowledge such as oral anatomy, oral physiology, and oral pathology. (2) Clinical knowledge of dentistry: including diagnostic and treatment knowledge of clinical subjects such as oral medicine, oral and maxillofacial surgery, prosthodontics, and orthodontics. (3) Comprehensive medical knowledge: including medical microbiology, medical immunology, pharmacology, medical psychology, medical ethics, and health regulations. These examinations cover a wide range of subjects, have a large span, and the knowledge points are fragmented, and the examination time is short, so it is extremely difficult to pass this examination.
[0004] In order to pass the written examination, most candidates improve their scores by attending offline tutoring classes or doing a large number of exercises. However, this method has the following disadvantages. First, attending offline tutoring classes or doing a large number of exercises requires candidates to invest a lot of time and energy. This high-intensity learning method is likely to make candidates feel tired, resulting in low learning efficiency. Second, offline tutoring classes and exercise sets are often designed for the general public, and the content design is relatively broad, making it difficult to provide personalized teaching according to the specific situation of each candidate. Therefore, candidates will inevitably spend a lot of time on the content they have already mastered and ignore their weak links that they have not mastered.
[0005] In recent years, large language models have shown great potential in various fields. Especially in the field of education, the ability to support personalized teaching models has begun to emerge. Specifically, through existing large model training techniques, it is easy to use a large number of oral professional books and materials as input data to construct a large prediction model with oral knowledge. However, this model cannot answer test questions well. For example, when asking the model to select from four options, the model will give the wrong answer. But when asking it to further analyze the test questions, it can give the correct analysis. That is, although the existing model has oral knowledge, it cannot apply this knowledge well to achieve tutoring. Summary of the Invention
[0006] Aiming at the above deficiencies in the prior art, the fine-tuning method and system for the large model of medical licensing examination tutoring based on few-shot learning provided by the present invention solve the problem of low accuracy of the existing large prediction model with oral knowledge in answering questions.
[0007] In order to achieve the above invention purpose, the technical solution adopted by the present invention is as follows:
[0008] In the first aspect, a fine-tuning method for the large model of medical licensing examination tutoring based on few-shot learning is provided, which includes the steps:
[0009] S1. Obtain oral knowledge examination questions, and convert the examination questions into word embedding vectors by using word embedding;
[0010] S2. Cluster the word embedding vectors, calculate the attention scores of each word embedding vector and the cluster center, and select the word embedding vectors whose attention scores meet the preset conditions as the data set;
[0011] S3. Train the large model of medical licensing examination tutoring by using the data set, and store the weights of the adapter during the training process. The large model of medical licensing examination tutoring is composed of Transformer layers with embedded adapters;
[0012] S4. According to the importance index and pruning threshold of the weights, perform pruning operations on the weights of the adapter. Then, perform inference tests on the large model of medical licensing examination tutoring before and after pruning to obtain the inference time and test scores before and after pruning;
[0013] S5. According to the inference time and test scores before and after pruning, calculate the accuracy decline index and acceleration performance index of the large model of medical licensing examination tutoring before and after pruning;
[0014] S6. Judge whether the accuracy decline index and acceleration performance index or the number of pruning times meet the corresponding preset conditions. If so, obtain the fine-tuned large model of medical licensing examination tutoring. Otherwise, enter step S7;
[0015] S7. Update the pruning threshold according to the accuracy decline index and the acceleration performance index, and then return to step S4.
[0016] Further, the expression for pruning the weights of the adapter is:
[0017]
[0018]
[0019] where, W s is the pruned weight matrix; M is the sparse matrix; W is the weight matrix before pruning; m and n are the total number of rows and columns of the weight matrix of the adapter respectively; W de is the weight parameter at the d-th row and e-th column in the weight matrix of the adapter; w de_matic is W de 's importance index; therhold init is the pruning threshold; W d is the weight parameter of the d-th row in the weight matrix of the adapter; W de_f is the value of the weight parameter at the f-th word embedding vector; F is the total number of word embedding vectors in the dataset.
[0020] Further, the expressions for calculating the accuracy decline index and the acceleration performance index are:
[0021]
[0022]
[0023] where, is the accuracy decline index; and are the test scores of the large model for medical licensing exam tutoring before and after pruning respectively; and are the inference times of the large model for medical licensing exam tutoring before and after pruning respectively; is the acceleration performance index.
[0024] Further, the expression for updating the pruning threshold is:
[0025]
[0026] where, thershold update and thershold init are the pruning thresholds after and before update respectively; a and b are both weight coefficients.
[0027] Further, the Transformer layer includes an embedding layer, a multi-head attention, a feed-forward layer, and an adapter connected in sequence, and layer normalization, two feed-forward layers, an adapter, and a normalization layer connected in sequence. The output of the embedding layer and the output of the adjacent adapter are superimposed and then input into the layer normalization; a dropout layer is provided between the adapter and the hidden layer;
[0028] The adapter includes an input layer and an output layer, and a plurality of hidden layers are provided between the input layer and the output layer. The hidden layers map the input features to a high-dimensional representation space through a plurality of linear transformations and non-linear activation functions, and generate a final output.
[0029] Further, the expression for the multi-head attention to process the output of the embedding layer is:
[0030]
[0031] where z is the output of the embedding layer; W Q 、W K 、W V 、W O are all weight matrices of the Transformer layer; d h is the dimension of W K ; G m is the mask matrix; Softmax(.) is the Softmax function; q is the q-th head of the multi-head attention; T is the transpose;
[0032] The expression for the feed-forward layer to process the input feature vector is:
[0033] FNN(o) = W D ReLu(W U o)
[0034] where o is the input feature vector of the feed-forward layer; W U o is the intermediate representation obtained by o through linear transformation; W U is the weight matrix from the input layer to the hidden layer; ReLU is the rectified linear unit; FNN(o) is the final output obtained by linear transformation of the intermediate representation after passing through the ReLU activation function; W D is the weight matrix from the hidden layer to the output layer;
[0035] The expression for the layer normalization to process the input feature vector is:
[0036]
[0037] Among them, LayerNorm(p) is the output of layer normalization; p is the feature vector input to layer normalization; μ and σ are the mean and standard deviation of p respectively; γ and β are the learnable scaling factor and bias term respectively;
[0038] The expression for the adapter to process the input feature vector is:
[0039] Adapter(t) = Dropout(W2 · (ReLU(W1 · t)))
[0040] Among them, Adapter(t) is the output of the adapter; t is the feature vector input to the adapter; W1 and W2 are the weight matrices of the two linear transformations of the adapter; Dropout is the dropout operation.
[0041] Furthermore, before training the large model for medical licensing examination tutoring using the dataset, it is necessary to freeze the weight information of the non - adapter part in the Transformer layer, and only update the weights of the adapter during the training process.
[0042] Furthermore, the methods for clustering word embedding vectors include:
[0043] A1. Randomly select K clustering centers from all word embedding vectors and perform K - means clustering:
[0044]
[0045] Among them, μ j is the j - th clustering center selected by K - means; c i is the index of the word embedding vector x i to the cluster; ||x i - μ j || 2 is the square of the Euclidean distance from the word embedding vector x i to μ j ;
[0046] A2. Calculate the mean of all word embedding vectors within each cluster as the new clustering center:
[0047]
[0048] Among them, \C j \ is the number of data points in the cluster C j ; is the new clustering center;
[0049] A3. Determine whether the change in the clustering center of each current cluster compared to its clustering center during the previous clustering is less than a preset threshold. If so, complete the clustering; otherwise, update and return to step A1.
[0050] The method for calculating the attention scores of each word embedding vector and the cluster center includes:
[0051] B1. Calculate the cosine similarity between each word embedding vector and the cluster center:
[0052]
[0053] where S ji is the cosine similarity between the word embedding vector x i and the cluster center μ j of the cluster C j ; ||·|| is the Euclidean norm of the vector;
[0054] B2. Calculate the attention scores of each word embedding vector according to the cosine similarity:
[0055]
[0056] where A ij is the attention score of x i relative to the cluster center μ j ; S ki is the cosine similarity between x i and the cluster center μ k ; k is a variable, 1 ≤ k ≤ K, and exp(.) is the natural exponential function.
[0057] Furthermore, the method for selecting word embedding vectors whose attention scores meet the preset conditions as the data set includes:
[0058] C1. Sort the word embedding vectors in each cluster in descending order according to their attention scores, and select the word embedding vector with the highest attention score in each cluster;
[0059] C2. Sort all the selected word embedding vectors in descending order according to their attention scores, and select the preset number of word embedding vectors with the highest attention scores to form the data set.
[0060] In a second aspect, a fine-tuning system for a large model for medical licensing examination tutoring based on few-shot learning is provided, which includes a word embedding module, a small sample extraction component for oral examination questions, and a lightweight model fine-tuning component connected in sequence. The lightweight model fine-tuning component includes a large model training module, a pruning and inference testing module, an index calculation module, a judgment module, and a threshold update module connected in sequence;
[0061] The word embedding module is used to obtain the oral knowledge examination questions and convert the examination questions into word embedding vectors by using word embedding;
[0062] The oral examination question small sample extraction component is used to cluster the word embedding vectors, calculate the attention scores of each word embedding vector and the clustering center, and select the word embedding vectors whose attention scores meet the preset conditions as the data set;
[0063] The large model training module is used to train the medical licensing examination tutoring large model with the data set and store the weights of the adapter during the training process. The medical licensing examination tutoring large model is composed of Transformer layers with embedded adapters;
[0064] The pruning and inference testing module is used to perform pruning operations on the weights of the adapter according to the importance index of the weights and the pruning threshold, and then perform inference testing on the medical licensing examination tutoring large model before and after pruning to obtain the inference time and test scores before and after pruning;
[0065] The index calculation module is used to calculate the accuracy decline index and acceleration performance index of the medical licensing examination tutoring large model before and after pruning according to the inference time and test scores before and after pruning;
[0066] The judgment module is used to judge whether the accuracy decline index, the acceleration performance index or the number of pruning times meet the corresponding preset conditions. If so, the fine-tuned medical licensing examination tutoring large model is obtained; otherwise, it enters the threshold update module;
[0067] The threshold update module is used to update the pruning threshold according to the accuracy decline index and the acceleration performance index, and then return to the pruning and inference testing module.
[0068] The beneficial effects of the present invention are as follows: When extracting small samples in this solution, based on the application of clustering and attention scores, not only the most representative sample data is found from the massive data, but also the key information in each category of samples is noticed by using the attention mechanism. Using these key information to better select representative samples greatly reduces the annotation cost of sample extraction. When extracting samples, a small number of representative oral examination questions are extracted for subsequent fine-tuning, which not only greatly reduces the cost of manual screening of questions, but also improves the generalization ability and training efficiency of the model for different data.
[0069] When fine-tuning the model in this solution, the number of parameters of the model is reduced through lightweight operations (pruning operations), the efficiency of data and the model is improved, and it has wide applicability, making the model lighter and more convenient for deployment and application. In addition, through multiple loop iterations during the pruning process, the pruning threshold is continuously updated, and an excellent model that takes into account both accuracy and inference speed can be found. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 It is a flowchart of a fine-tuning method for a medical licensing examination tutoring large model based on small sample learning.
[0071] Figure 2 It is the schematic diagram of the lightweight model fine-tuning component.
[0072] Figure 3 It is the principle block diagram of the large model fine-tuning system for medical licensing examination tutoring based on few-shot learning. Specific implementation manners
[0073] The specific implementation manners of the present invention will be described below to facilitate those skilled in the art of this technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of this technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0074] Refer to Figure 1 , Figure 1 which shows the large model fine-tuning method for medical licensing examination tutoring based on few-shot learning; as Figure 1 shown, this method S includes steps S1 to S7.
[0075] In step S1, obtain the oral knowledge examination questions and convert the examination questions into word embedding vectors by using word embedding;
[0076] In step S2, cluster the word embedding vectors, calculate the attention scores of each word embedding vector and the cluster center, and select the word embedding vectors whose attention scores meet the preset conditions as the data set;
[0077] In implementation, the method preferably used for clustering the word embedding vectors in this solution includes:
[0078] A1. Randomly select K cluster centers from all the word embedding vectors and perform K-means clustering:
[0079]
[0080] where μ j is the j-th cluster center selected by K-means; c i is the index of the word embedding vector x i to the cluster; ||x i -μ j || 2 is the square of the Euclidean distance of the word embedding vector x i to μ j ;
[0081] A2. Calculate the mean value of all the word embedding vectors within each cluster as the new cluster center:
[0082]
[0083] Among them, \C j \ is the number of data points in cluster C j ; is the new cluster center;
[0084] A3. Determine whether the change in the cluster center of each current cluster compared to its cluster center during the previous clustering is less than a preset threshold. If so, complete the clustering; otherwise, update and return to step A1.
[0085] This solution uses K-means to replace the randomly selected K cluster centers, making the cluster centers more representative and improving the convergence speed and effect. By iteratively updating each cluster center and reassigning the test question vectors until the cluster centers are stable or a preset number of iterations is reached. Finally, each test question vector is assigned a cluster label indicating the cluster it belongs to.
[0086] The method for calculating the attention score of each word embedding vector and the cluster center includes:
[0087] B1. Calculate the cosine similarity between each word embedding vector and the cluster center:
[0088]
[0089] Among them, S ji is the cosine similarity between the word embedding vector x i and the cluster center μ j of cluster C j ; ||·|| is the Euclidean norm of the vector;
[0090] B2. Calculate the attention score of each word embedding vector based on the cosine similarity:
[0091]
[0092] Among them, A ij is the attention score of x i relative to the cluster center μ j ; S ki is the cosine similarity between x i and the cluster center μ k ; k is a variable, 1 ≤ k ≤ K, and exp(.) is the natural exponential function.
[0093] In this solution, the cosine similarity is used to measure the directional relationship between two vectors rather than their absolute distance in space, which is very suitable for measuring the semantic correlation between word embeddings. Secondly, the similarity matrix obtained by the above calculation is normalized through the softmax function to obtain the attention weights (scores) of each word vector for each cluster center. This step ensures that the sum of all weights is 1, giving the relative importance of each word in different clustering dimensions. Finally, by multiplying the word embedding vector of each word by its corresponding attention score, a weighted word representation can be generated. This weighted representation strengthens the word information most relevant to the cluster center in the text, reduces the influence of noise, and thus improves the performance of the model in subsequent tasks.
[0094] The method of selecting word embedding vectors that meet the preset conditions as the data set includes:
[0095] C1. Sort the word embedding vectors in each cluster in descending order according to their attention scores, and select the word embedding vector with the highest attention score in each cluster;
[0096] C2. Sort all the selected word embedding vectors in descending order according to their attention scores, and select the preset number of word embedding vectors with the highest attention scores to form the data set. In this solution, the preset number is preferably equal to 5.
[0097] The operation of adopting step C1 ensures that representative test questions in each cluster participate in the subsequent selection process. Step C2 aggregates the representative test questions in all clusters and sorts them again. In this process, the focus is on the representativeness of the test questions on a global scale to ensure that the small number of finally selected test questions can comprehensively cover the diversity and representativeness of the entire test question bank.
[0098] When selecting the data set in this solution, through hierarchical sorting and screening, it is ensured that the most representative test questions are efficiently and accurately extracted from a large-scale test question bank, thus optimizing the construction and usage efficiency of the data set.
[0099] In step S3, the data set is used to train the medical licensing examination tutoring large model, and the weights of the adapter during the training process are stored. The medical licensing examination tutoring large model consists of Transformer layers with embedded adapters. Before using the data set to train the medical licensing examination tutoring large model, it is necessary to freeze the weight information of the non-adapter part in the Transformer layer, and only update the weights of the adapter during the training process.
[0100] Such as Figure 2As shown, in implementation, preferably, the Transformer layer of this solution includes an embedding layer, a multi-head attention, a feed-forward layer, and an adapter that are connected in sequence, and a layer normalization, two feed-forward layers, an adapter, and a normalization layer that are connected in sequence. The output of the embedding layer and the output of the adjacent adapter are superimposed and then input into the layer normalization; a dropout layer is provided between the adapter and the hidden layer.
[0101] Among them, the feed-forward layer is a feed-forward neural network (FFN) layer, which consists of two fully connected layers and a non-linear transformation. In order to maintain the lower limit of the model performance, a residual connection is also designed in the feed-forward neural network (FFN) layer to ensure that the model performance will not degrade.
[0102] In an embodiment of the present invention, the multi-head attention mechanism first performs a linear transformation on the input sequence to obtain query, key, and value vectors. Then, the query, key, and value of each head are respectively input into the attention mechanism to calculate the attention weights of each head. Next, the attention weights of each head are multiplied by the corresponding values and weighted and summed to obtain the output of each head. Finally, the outputs of all heads are concatenated together and passed through another linear transformation to obtain the final output. The expression for the multi-head attention to implement the above process is:
[0103]
[0104] Among them, z is the output of the embedding layer; W Q , W K , W V , W O are all Transformer layer weight matrices; d h is the dimension of W K so that the model can learn information from different representation subspaces; G m is a mask matrix; Softmax(.) is the Softmax function; q is the q-th head of the multi-head attention; T is the transpose;
[0105] The mask matrix is used to control the attention allocation in the attention mechanism. The role of the mask matrix is to mask some positions when calculating the attention weights, so that the model can ignore certain positions or limit the attention range according to the requirements of the task when processing the sequence. The role of the Softmax function is to convert the attention scores (attention weights) into an attention distribution, so that the sum of the attention weights is 1, and the attention weight of each position is between 0 and 1. This can ensure that the importance of each position is reasonably considered when the model performs weighted summation.
[0106] The expression for the feed-forward layer to process the input feature vector is:
[0107] FNN(o) = W D ReLu(W U o)
[0108] where o is the feature vector input to the feed-forward layer; W U o is the intermediate representation obtained by linearly transforming o; W U is the weight matrix from the input layer to the hidden layer; ReLU is the rectified linear unit, a non-linear activation function used to introduce non-linearity into the network; it sets the negative part of the input W U x to zero, enabling the network to learn more complex functions. FNN(o) is the final output obtained by linearly transforming the intermediate representation after passing through the ReLU activation function; W D is the weight matrix from the hidden layer to the output layer;
[0109] Layer Normalization is a normalization technique used in neural networks, aiming to accelerate the training process and improve the generalization ability of the model. The expression for processing the input feature vector is:
[0110]
[0111] where LayerNorm(p) is the output of layer normalization; p is the feature vector input to layer normalization; μ and σ are the mean and standard deviation of p respectively; γ and β are the learnable scaling factor and bias term respectively;
[0112] Refer again to Figure 2 , the adapter includes an input layer and an output layer, and several hidden layers are arranged between the input layer and the output layer. The hidden layers map the input features to a high-dimensional representation space through several linear transformations and non-linear activation functions and produce the final output.
[0113] The expression for the adapter to process the input feature vector is:
[0114] Adapter(t) = Dropout(W2 · (ReLU(W1 · t)))
[0115] where Adapter(t) is the output of the adapter; t is the feature vector input to the adapter; W1 and W2 are the weight matrices of the two linear transformations of the adapter respectively; Dropout is the dropout operation.
[0116] In this solution, compared with a normal feed-forward neural network, the adapter adds dropout connections between the hidden layers. The purpose of doing this is to introduce a certain degree of randomness during the fine-tuning process, prevent the model from overfitting, and improve the generalization ability of the model.
[0117] In step S4, according to the importance index of the weights and the pruning threshold, pruning operations are performed on the weights of the adapter. After that, the large model for medical licensing examination tutoring before and after pruning is subjected to inference tests to obtain the inference time and test scores before and after pruning.
[0118] During implementation, the preferred expression for performing pruning operations on the weights of the adapter in this solution is:
[0119]
[0120]
[0121] Among them, W s is the weight matrix after pruning, which has the same shape as W, and only the elements corresponding to 1 in M are retained. M is a sparse matrix with the same shape as W, where the part greater than the initial threshold is 1 and the part less than the initial threshold is 0. W is the weight matrix before pruning; m and n are the total number of rows and columns of the weight matrix of the adapter respectively; W de is the weight parameter at the d-th row and e-th column in the weight matrix of the adapter; w de_matic is the importance index of W de ; therhold init is the pruning threshold; W d is the weight parameter of the d-th row in the weight matrix of the adapter; W de_f is the value of the weight parameter at the f-th word embedding vector; F is the total number of word embedding vectors in the dataset.
[0122] In step S5, according to the inference time and test scores before and after pruning, calculate the accuracy drop index and acceleration performance index of the large model for medical licensing examination tutoring before and after pruning:
[0123]
[0124]
[0125] Among them, acc is the accuracy drop index; score dense and score spasity are the test scores of the large model for medical licensing examination tutoring before and after pruning respectively; t dense and t spasity are the inference times of the large model for medical licensing examination tutoring before and after pruning respectively; speedup is the acceleration performance index.
[0126] In step S6, determine whether the accuracy drop index and acceleration performance index or the number of pruning times meet the corresponding preset conditions. If so, obtain the fine-tuned large model for medical licensing examination tutoring; otherwise, enter step S7;
[0127] The specific operation of determining whether the accuracy degradation index, the acceleration performance index, or the pruning times meet the corresponding preset conditions is as follows: When both the accuracy degradation index and the acceleration performance index are less than their corresponding thresholds, the pruning operation is terminated, or when the pruning times are greater than or equal to the preset times, the pruning operation is terminated.
[0128] In step S7, the pruning threshold is updated according to the accuracy degradation index and the acceleration performance index, and then step S4 is returned. The expression for updating the pruning threshold is:
[0129]
[0130] where, thershold update and thershold init are the updated and the pre-updated pruning thresholds respectively; both a and b are weight coefficients.
[0131] As Figure 3 shown, this solution also provides a fine-tuning system for the medical licensing examination tutoring large model based on few-shot learning, which includes a word embedding module, an oral examination question few-shot extraction component, and a lightweight model fine-tuning component connected in sequence. The lightweight model fine-tuning component includes a large model training module, a pruning and inference testing module, an index calculation module, a judgment module, and a threshold update module connected in sequence;
[0132] The word embedding module is used to obtain oral knowledge examination questions and convert the examination questions into word embedding vectors by using word embedding;
[0133] The oral examination question few-shot extraction component is used to cluster the word embedding vectors, calculate the attention scores of each word embedding vector and the cluster center, and select the word embedding vectors whose attention scores meet the preset conditions as the data set;
[0134] The large model training module is used to train the medical licensing examination tutoring large model by using the data set and store the weights of the adapter during the training process. The medical licensing examination tutoring large model consists of Transformer layers with embedded adapters;
[0135] The pruning and inference testing module is used to prune the weights of the adapter according to the importance index of the weights and the pruning threshold, and then perform inference testing on the medical licensing examination tutoring large model before and after pruning to obtain the inference time and test scores before and after pruning;
[0136] The index calculation module is used to calculate the accuracy degradation index and the acceleration performance index of the medical licensing examination tutoring large model before and after pruning according to the inference time and test scores before and after pruning;
[0137] The judgment module is used to judge whether the accuracy decline index, the acceleration performance index, or the pruning times meet the corresponding preset conditions. If so, the fine-tuned medical licensing examination tutoring large model is obtained; otherwise, it enters the threshold update module.
[0138] The threshold update module is used to update the pruning threshold according to the accuracy decline index and the acceleration performance index, and then return to the pruning and inference test module.
[0139] In summary, the fine-tuning method of this solution can search for representative test questions as samples for fine-tuning, which not only reduces the amount of data required for fine-tuning, but also achieves the effect of multiple samples. When training, combined with the pruning strategy, the number of model parameters is reduced, making the model lighter and more convenient for deployment and application.
Claims
1. A fine-tuning method for a large model for medical examination tutoring based on small sample learning, characterized in that: Includes steps: S1. Obtain oral knowledge test questions and convert the test questions into word embedding vectors using word embedding; S2. Cluster the word embedding vectors and calculate the attention score between each word embedding vector and the cluster center. Select the word embedding vectors whose attention scores meet the preset conditions as the data set. S3, training a medical examination coaching model using the data set and storing the weights of the adapter during the training process, wherein the medical examination coaching model is composed of a Transformer layer embedded with the adapter; S4. Prune the weights of the adapter according to the importance index of the weight and the pruning threshold, and then perform reasoning tests on the medical examination tutoring model before and after pruning to obtain the reasoning time and test scores before and after pruning; S5. Calculate the accuracy drop index and acceleration performance index of the medical examination tutoring model before and after pruning based on the inference time and test score before and after pruning; S6, judging whether the accuracy drop index and the acceleration performance index or the number of pruning times meet the corresponding preset conditions, if so, obtaining the fine-tuned medical examination tutoring model, otherwise proceeding to step S7; S7. Update the pruning threshold according to the accuracy drop index and the acceleration performance index, and then return to step S4.
2. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 1 is characterized in that: The expression for pruning the adapter weights is: Among them, W s is the weight matrix after pruning; M is a sparse matrix; W is the weight matrix before pruning; m and n are the total number of rows and columns of the adapter's weight matrix, respectively; W de is the weight parameter of the dth row and eth column in the adapter’s weight matrix; w de_matic W de The importance of the index; init is the pruning threshold; W d is the weight parameter of the dth row in the adapter’s weight matrix; W de_f is the value of the weight parameter at the fth word embedding vector; F is the total number of word embedding vectors in the dataset.
3. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 1 is characterized in that: The expressions for calculating the accuracy drop index and the acceleration performance index are: Among them, acc is the accuracy drop index; score dense and score spasity are the test scores of the medical examination tutoring model before and after pruning; t dense and t spasity They are the inference time of the large model for medical examination tutoring before and after pruning; speedup is the acceleration performance indicator.
4. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 3 is characterized in that: The expression for updating the pruning threshold is: Among them, thershold update and Thershold init are the pruning thresholds before and after updating respectively; a and b are weight coefficients.
5. The method for fine-tuning a large model for tutoring medical examinations based on small sample learning according to any one of claims 1 to 4, characterized in that: The Transformer layer includes an embedding layer, a multi-head attention, a feedforward layer and an adapter connected in sequence, and a layer normalization, two feedforward layers, an adapter and a normalization layer connected in sequence. The output of the embedding layer and the output of the adapter adjacent to it are superimposed and the input layer is normalized; A dropout layer is provided between the adapter and the hidden layer; The adapter includes an input layer and an output layer, wherein a plurality of hidden layers are arranged between the input layer and the output layer, and the hidden layers map the input features to a high-dimensional representation space through a plurality of linear transformations and nonlinear activation functions, and generate a final output.
6. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 5 is characterized in that: The expression for multi-head attention processing the output of the embedding layer is: Where z is the output of the embedding layer; W Q , W K , W V , W O are all Transformer layer weight matrices; d h W K Dimension; G m is the mask matrix; Softmax(.) is the Softmax function; q is the qth head of the multi-head attention; T is the transpose; The expression for the feedforward layer to process the input feature vector is: FNN(o)=W D ReLu (W U of) Where o is the feature vector of the input feedforward layer; W U o is the intermediate representation of o obtained by linear transformation; W U is the weight matrix from the input layer to the hidden layer; ReLU is the rectified linear unit; FNN(o) is the final output obtained by linear transformation of the intermediate representation after the ReLU activation function; W D is the weight matrix from the hidden layer to the output layer; The expression for layer normalization to process the input feature vector is: Where LayerNorm(p) is the output of layer normalization; p is the normalized feature vector of the input layer; μ and σ are the mean and standard deviation of p, respectively; γ and β are the learnable scaling factor and bias term, respectively; The expression used by the adapter to process the input feature vector is: Adapter(t)=Dropout(W2·(ReLU(W1·t))) Where Adapter(t) is the output of the adapter; t is the feature vector of the input adapter; W1 and W2 are the weight matrices of the two linear transformations of the adapter respectively; Dropout is the dropout operation.
7. The method for fine-tuning a large model for tutoring medical examinations based on small sample learning according to any one of claims 1 to 4 and 6, characterized in that: Before using the dataset to train the large model for medical examination tutoring, the weight information of the non-adapter part in the Transformer layer needs to be frozen, and only the weight of the adapter is updated during the training process.
8. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 1, characterized in that: Methods for clustering word embedding vectors include: A1. Randomly select K cluster centers from all word embedding vectors and perform K-means clustering: Among them, μ j The jth cluster center selected by K-means; c i is the word embedding vector x i Index into the cluster; ||x i -μ j || 2 is the word embedding vector x i to μ j The square of the Euclidean distance; A2. Calculate the mean of all word embedding vectors in each cluster as the new cluster center: Among them, \C j \ is cluster C j The number of data points; is the new cluster center; A3: Determine whether the change between the current cluster center of each cluster and the cluster center of the last clustering is less than the preset threshold. If so, complete the clustering. Otherwise, update And return to step A1; The method of calculating the attention score of each word embedding vector and cluster center includes: B1. Calculate the cosine similarity between each word embedding vector and the cluster center: Among them, S ji is the word embedding vector x i and cluster C j The cluster center μ j The cosine similarity of ; ||·|| is the Euclidean norm of the vector; B2. Calculate the attention score of each word embedding vector based on cosine similarity: Among them, A ij For x i Relative to the cluster center μ j Attention score; S ki For x i and cluster center μ k The cosine similarity of , k is a variable, 1≤k≤K, and exp(.) is a natural exponential function.
9. The method for fine-tuning a large model for medical examination tutoring based on small sample learning according to claim 1, characterized in that: Methods for selecting word embedding vectors whose attention scores meet preset conditions as data sets include: C1. Sort the word embedding vectors in each cluster in descending order according to their attention scores, and select the word embedding vector with the highest attention score in each cluster; C2. Sort all selected word embedding vectors in descending order according to their attention scores, and select a preset number of word embedding vectors with the highest attention scores to form a data set.
10. A fine-tuning system for a large model of medical examination tutoring based on small sample learning, characterized in that: It includes a word embedding module, an oral test question small sample extraction component and a lightweight model fine-tuning component connected in sequence, wherein the lightweight model fine-tuning component includes a large model training module, a pruning and reasoning test module, an indicator calculation module, a judgment module and a threshold update module connected in sequence; The word embedding module is used to obtain oral knowledge test questions and convert the test questions into word embedding vectors using word embedding; The oral test small sample extraction component is used to cluster word embedding vectors and calculate the attention score between each word embedding vector and the cluster center, and select the word embedding vectors whose attention scores meet the preset conditions as the data set; The large model training module is used to train the medical examination tutoring large model using the data set and store the weights of the adapter during the training process, wherein the medical examination tutoring large model is composed of a Transformer layer embedded with an adapter; The pruning and reasoning test module is used to prune the weights of the adapter according to the importance index of the weights and the pruning threshold, and then perform reasoning tests on the medical examination tutoring model before and after pruning to obtain the reasoning time and test scores before and after pruning; The indicator calculation module is used to calculate the accuracy drop index and acceleration performance index of the medical examination tutoring model before and after pruning based on the inference time and test score before and after pruning; The judgment module is used to judge whether the accuracy drop index and the acceleration performance index or the number of pruning times meet the corresponding preset conditions. If so, the fine-tuned medical examination tutoring model is obtained, otherwise, the threshold update module is entered; The threshold update module is used to update the pruning threshold according to the accuracy drop index and the acceleration performance index, and then return to the pruning and reasoning test module.
Citation Information
Patent Citations
Generative large model modeling method, system and equipment based on knowledge graph
CN117688974A
Iterative random pruning method and system based on meta transfer learning
CN117787377A