Electronic medical record text normalization method based on large model fine tuning
By building a medical-specific corpus and a hierarchical strength confirmation model in a large language model and adaptively inserting Lora blocks, we solved the adaptability and computational overhead issues of electronic medical record text normalization technology in the medical field, and achieved more efficient text normalization and understanding.
Patent Information
- Application Number
- CN202510777260.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-11
Smart Images

Figure CN120690362A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical information processing and artificial intelligence technology application, and relates to an electronic medical record text normalization method based on large model fine-tuning. Background Art
[0002] Existing electronic medical record text normalization technologies primarily rely on simple logical rewriting methods such as regular expression matching. These methods often lack deep understanding and adaptability when dealing with complex and ever-changing medical texts, resulting in limited normalization effectiveness. In recent years, the emergence of large language models (LLMs) has brought breakthroughs in text processing, but their application in the medical field still faces adaptability and high cost issues. While Lora (Low-Rank Adaptation) fine-tuning technology can reduce training overhead, its traditional approach typically only inserts fine-tuning parameters in fixed layers of the model, which cannot fully address the adaptation needs of texts of varying complexity and is not combined with efficient training strategies. Summary of the Invention
[0003] To address the above-mentioned technical problems, the present invention adopts a method for normalizing electronic medical record text based on large-scale model fine-tuning, comprising: obtaining an unnormalized electronic medical record text x in the medical field, inputting the text x into a trained large-scale language model, and obtaining a normalized electronic medical record text; the training process of the large-scale language model includes:
[0004] S1. Build a medical field-specific corpus, which contains unstandardized electronic medical record text x k and its standardized electronic medical record text k ; where k represents the index of the unnormalized electronic medical record text;
[0005] S2. Obtain a pre-trained large language model, build a Lora block for each layer of the pre-trained large language model; and build a hierarchical strength confirmation model.
[0006] S3, convert the unstandardized electronic medical record text x k Input the hierarchical strength confirmation model and get the text x k Predicted layer strength G k ;
[0007] S4, according to the level strength G k Insert the corresponding Lora block into each layer of the pre-trained large language model to obtain a large language model with inserted Lora blocks;
[0008] S5. Convert the unstandardized electronic medical record text to k Input the large language model inserted into the Lora block and get the predicted text x k Standardized electronic medical record text;
[0009] S6. According to the predicted text x k Standardized electronic medical record text and its corresponding standardized electronic medical record text y k Calculate the loss function value, update the parameters of the Lora block and layer strength confirmation model according to the loss function value, and obtain the trained Lora block when the loss function value is minimized;
[0010] S7. Insert the trained Lora block into the corresponding layer of the pre-trained large language model to obtain the trained large language model.
[0011] Beneficial effects:
[0012] 1. The present invention fine-tunes the pre-trained large language model through Lora fine-tuning technology, which can improve the adaptability of the large language model in the task of normalizing electronic medical record texts in the medical field, enabling the model to have a deeper understanding of the medical record text content and accurately perform normalization operations, thereby improving the accuracy and utilization efficiency of medical record information; 2. The present invention uses a hierarchical strength confirmation model to adaptively select the hierarchical strength of the Lora block inserted in the pre-trained large language model, thereby improving the accuracy and utilization efficiency of medical record information while reducing unnecessary computational overhead; 3. The present invention uses medical dictionaries to assist in context complexity analysis to improve the ability to discriminate complex texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A flowchart of a method for normalizing electronic medical record text based on large model fine-tuning provided by an embodiment of the present invention;
[0014] Figure 2 Lora principle diagram provided for an embodiment of the present invention;
[0015] Figure 3 This is a graph showing the change trend of the training loss provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0017] like Figure 1 As shown, the present invention adopts a method for normalizing electronic medical record text based on large model fine-tuning, including: obtaining an unnormalized electronic medical record text X in the medical field, inputting the text X into a trained large language model, and obtaining a normalized electronic medical record text;
[0018] The training process of a large language model includes:
[0019] S1. Build a medical field-specific corpus, which contains multiple unstandardized electronic medical record texts x k and its standardized electronic medical record text k ; where k represents the index of the unnormalized electronic medical record text;
[0020] In order to train a large model for the task of standardizing electronic medical records in the medical field, it is first necessary to build a corpus specific to the medical field.
[0021] Unstandardized electronic medical record text data in the medical field usually contains information such as the patient's medical history, diagnosis, and treatment, but there may be differences in the expression. For example, "ever vaccinated" and "denied" in "every item has been vaccinated with hepatitis B vaccine and BCG vaccine, and denies a history of infectious diseases such as tuberculosis and hepatitis" are not clearly marked for each item, and "none" in "no subcutaneous nodules, scars, subcutaneous bleeding, rash, bruises, or ecchymosis" is not repeatedly marked for each item.
[0022] The standardized electronic medical record text follows unified standards in expression. For example, verbs such as "ever vaccinated" and "denied" are clearly marked before each piece of information, and negative words such as "none" are repeatedly marked for each item.
[0023] S2. Obtain a pre-trained large language model, build a Lora block for each layer of the pre-trained large language model; and build a hierarchical strength confirmation model.
[0024] The pre-trained large language model is Chatglm-6B, which has been trained on large amounts of text data and possesses powerful text understanding and generation capabilities. However, due to the specialized nature of the medical field, directly using this model for electronic medical record standardization may not be effective. Therefore, the present invention designs and employs adaptive multi-layer LoRa fine-tuning technology to perform targeted training on this model.
[0025] like Figure 2 As shown, the principle of Lora fine-tuning is to add a bypass next to the pre-trained large language model to perform a dimensionality reduction and then dimensionality increase operation. During training, the parameters of the pre-trained language model are fixed, and only the dimensionality reduction matrix A and the dimensionality increase matrix B are trained. The input and output dimensions of the model remain unchanged, and the parameters of BA and the pre-trained large language model are superimposed at the output. 2 Random Gaussian distribution N(0,δ 2) Initialize A with a zero matrix and initialize B with a zero matrix. This ensures that the newly added pathway BA = 0 at the start of training, thus having no impact on the model results. During inference, simply add the trained matrix product BA to the original weight matrix W as the new weight parameter to replace W in the pre-trained large language model.
[0026] For convenience of representation, the present invention uses the matrix D to represent the matrix product BA in the Lora fine-tuning technology, and the weight of the Lora block corresponding to each layer of the pre-trained large language model is expressed as: ΔW i =D i W pre,i +b i , D i =B i A i , where ΔW i is the weight of the Lora block corresponding to the i-th layer of the pre-trained large language model, D i is the Lora block ΔW i The low-rank matrix, low-rank matrix D i rank r i =2i,b i is the Lora block ΔW i The bias vector, i is the index of the layer of the pre-trained large language model, A i is the dimension reduction matrix, B i is a dimension-raising matrix.
[0027] S3, convert the unstandardized electronic medical record text x k Input the hierarchical strength confirmation model and get the text x k Predicted layer strength G k ;
[0028] Specifically, the hierarchical strength confirmation model includes: a pre-trained Transformer encoder, an aggregation module, a medical dictionary auxiliary analysis module, and a gating module; the hierarchical strength confirmation model processes unstandardized electronic medical record text by:
[0029] S31. Convert the unstandardized electronic medical record text to k Input the pre-trained Transformer encoder to get the text embedding vector H k ;
[0030] S32. Embed the text into vector H k Input aggregation module to obtain the global context vector F k ;
[0031] The aggregation module is the average pooling module (MeanPooling) or the extraction class token ([CLS]token), the global context vector Fk Carrying semantic information of the entire text context, including complexity and ambiguity features.
[0032] S33, convert the unstandardized electronic medical record text into k Input the medical dictionary auxiliary analysis module to obtain the text x k The encoding vector N of the number of medical terms k ;
[0033] In order to more accurately represent the complexity of the context, the medical dictionary auxiliary analysis is introduced to count the number of medical terms appearing in the text. Specifically, the medical dictionary auxiliary analysis module processes the unstandardized electronic medical record text by: obtaining a medical dictionary, which includes multiple medical terms; converting the unstandardized electronic medical record text into k Match the medical terms in the medical dictionary to get the text x k The number of medical terms for text x k The number of medical terms is encoded to obtain the encoding vector N k .
[0034] The more medical terms appear in the text, the more professional and complex the text content is, which dynamically affects the Lora insertion strategy.
[0035] S34, the global context vector F k and the encoding vector N of the number of medical terms k Splice and get the complexity feature vector C k ;
[0036] S35, the complexity characteristic vector C k Input the gating module and get the text x k Predicted hierarchical strength.
[0037] For complex texts (high context complexity), we choose to insert Lora in the deep layer and increase the size of the A matrix and b vector; for simple texts (low context complexity), we choose to insert Lora in the shallow layer and reduce the size of A and b.
[0038] The specific approach is: the gating module is based on the context complexity feature vector C k The distribution characteristics of each layer generate a set of gating coefficients: G k =[g k,1 ,g k,2 ,...g k,L ]=softmax(W g ·C k +b g ); where g k,i ∈[0,1] is the text x kThe predicted strength of the Lora block inserted into the i-th layer of the pre-trained large language model, L is the number of layers of the pre-trained large language model, W g 、b g is the gate weight, G k For text x k The predicted pre-trained large language model is inserted into the hierarchical strength matrix of the Lora block, i.e., the gating coefficient.
[0039] This formula dynamically maps the context complexity to the insertion position of Lora to achieve adaptive matching between the model structure and the task difficulty. k When the semantic complexity or density of professional terms is higher (e.g., high frequency of medical terms, C k The vector norm is large), the gating module allocates higher strength in the deep layer (close to the output) of the pre-trained large language model, that is, increasing the insertion strength of the deep Lora and the scale of the low-rank matrix and bias vector; when the complexity feature vector C k When the complexity is low (e.g. C k If the norm of the representation is small or the professional terms are sparse), the gating module allocates higher strength to the shallow layer (close to the input) of the pre-trained large language model, that is, increases the insertion strength of the shallow Lora, reduces the size of the low-rank matrix and bias vector, thereby improving inference efficiency and saving computing resources.
[0040] S4, according to the level strength G k Insert the corresponding Lora block into each layer of the pre-trained large language model to obtain the large language model M with the Lora block inserted lora ;
[0041] According to the level strength G k Inserting each Lora block into the corresponding layer of the pre-trained large language model is represented as: W lora,i =W pre,i +g k,i ΔW i ; where ΔW i For the pre-trained large language model M pre The Lora block corresponding to the i-th layer, W pre,i For the pre-trained large language model M pre The weight of the i-th layer, W lora,i Insert the Lora block into the large language model M lora The weight of the i-th layer.
[0042] By using the above-mentioned adaptive multi-layer Lora fine-tuning technology to conduct targeted training on the model, the accuracy of text normalization can be improved and the amount of computation can be reduced.
[0043] S5. Convert the unstandardized electronic medical record text in the medical field to k Input the large language model inserted into the Lora block and get the predicted text x k Standardized electronic medical record text;
[0044] S6. According to the predicted text x k Standardized electronic medical record text and its corresponding standardized electronic medical record text y k Calculate the loss function value, update the parameters of the Lora block and layer strength confirmation model according to the loss function value, and obtain the trained Lora block when the loss function value is minimized;
[0045] The parameters of the Lora block and the layer strength confirmation model are updated according to the loss function value, including: freezing the parameters of the pre-trained large language model and the pre-trained Transform encoder, and updating the parameters D of the Lora block according to the loss function value. i 、b i and the parameter W of the gating module g 、b g .
[0046] Preferably, the loss function is a cross entropy loss function.
[0047] like Figure 3 As shown in the figure, during the training process, as the number of iterations increases, the loss converges quickly, which improves the accuracy of normalization while shortening the training time.
[0048] As shown in Table 1, the performance of the fine-tuned large language model is better than that of the large language model before fine-tuning (i.e., pre-training).
[0049] Table 1 Performance of large language model before and after fine-tuning
[0050]
[0051] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation plans of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for normalizing electronic medical record text based on large model fine-tuning, characterized by: include: Obtain unstandardized electronic medical record text x in the medical field, input text x into the trained large language model, and obtain standardized electronic medical record text; The training process of a large language model includes: S1. Build a medical field-specific corpus, which contains unstandardized electronic medical record text x k and its standardized electronic medical record text k ; where k represents the index of the unnormalized electronic medical record text; S2. Obtain a pre-trained large language model, build a Lora block for each layer of the pre-trained large language model; and build a hierarchical strength confirmation model. S3, convert the unstandardized electronic medical record text x k Input the hierarchical strength confirmation model and get the text x k Predicted layer strength G k ; S4, according to the level strength G k Insert the corresponding Lora block into each layer of the pre-trained large language model to obtain a large language model with inserted Lora blocks; S5. Convert the unstandardized electronic medical record text to k Input the large language model inserted into the Lora block and get the predicted text x k Standardized electronic medical record text y′ k ; S6. According to the predicted text x k Standardized electronic medical record text y′ k and its corresponding standardized electronic medical record text y k Calculate the loss function value, update the parameters of the Lora block and layer strength confirmation model according to the loss function value, and obtain the trained Lora block when the loss function value is minimized; S7. Insert the trained Lora block into the corresponding layer of the pre-trained large language model to obtain the trained large language model.
2. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 1 is characterized in that: The weight of the Lora block corresponding to each layer of the pre-trained large language model is expressed as: ΔW i =D i W pre,i +b i , where ΔW i is the weight of the Lora block corresponding to the i-th layer of the pre-trained large language model, D i is the low-rank matrix of the Lora block, b i is the bias vector of the Lora block, and i is the index of the layer of the pre-trained large language model.
3. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 2 is characterized in that: Low-rank matrix D i rank r i =2i.
4. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 1 is characterized in that: The hierarchical strength confirmation model includes: a pre-trained encoder, an aggregation module, a medical dictionary auxiliary analysis module, and a gating module; the hierarchical strength confirmation model is used to verify the accuracy of the unstandardized electronic medical record text x k Processing includes: S31. Convert the unstandardized electronic medical record text to k Input the pre-trained encoder to obtain the text embedding vector; S32, input the text embedding vector into the aggregation module to obtain a global context vector; S33, convert the unstandardized electronic medical record text into k Input the medical dictionary auxiliary analysis module to obtain the text x k The encoding vector of the number of medical terms; S34, concatenating the global context vector and the encoding vector of the number of medical terms to obtain a complexity feature vector; S35. Input the complexity feature vector into the gating module to obtain the text x k Predicted layer strength G k .
5. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 4 is characterized in that: The aggregation module is an average pooling module.
6. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 4 is characterized in that: Medical dictionary auxiliary analysis module for non-standardized electronic medical record text x k The processing includes: obtaining a medical dictionary, which includes multiple medical terms; converting the unstandardized electronic medical record text into k Match the medical terms in the medical dictionary to get the text x k The number of medical terms for text x k The number of medical terms is feature encoded to obtain the encoding vector.
7. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 4 is characterized in that: The gating module processes the complexity feature vector including: G k =[g k,1 ,g k,2 ,...g k,L ]=softmax(W g ·C k +b g ) Among them, W g 、b g is the gate weight, C k For text x k The complexity feature vector of L is the number of layers of the pre-trained large language model, g k,i ∈[0,1] is the text x k The predicted strength of the Lora block inserted into the i-th layer of the pre-trained large language model.
8. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 4 is characterized in that: Updating the parameters of the Lora block and the hierarchical strength confirmation model according to the loss function value includes: freezing the parameters of the pre-trained large language model and the pre-trained encoder, and updating the parameters of the Lora block and the gating module according to the loss function value.
9. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 1 is characterized in that: Inserting the corresponding Lora block into each layer of the pre-trained large language model includes: IN lora,i =In pre,i +g k,i ΔW i ; Where ΔW i is the weight of the Lora block corresponding to the i-th layer of the pre-trained large language model, W pre,i is the weight of the i-th layer of the pre-trained large language model, W lora,i is the weight of the i-th layer of the large language model with the Lora block inserted, G k =[g k,1 ,g k,2 ,...g k,L ], g k,i The strength of the Lora block inserted into the i-th layer of the pre-trained large language model, L is the number of layers of the pre-trained large language model.
10. The method for normalizing electronic medical record text based on large model fine-tuning according to claim 1, characterized in that: The loss function is the cross entropy loss function.
Citation Information
Patent Citations
Text normalization method and device, storage medium and electronic equipment
CN115758990A
Fine adjustment method and device for point cloud classification
CN116778133A
Disposal plan generation method based on case
CN117151095A
Fine-grained emotion recognition method, electronic equipment and storage medium
CN117370736A
Automatic voice processing method based on language large model and electronic equipment
CN119091864A