ICD automatic coding method and system based on large language model transfer learning
By adopting the transfer learning method of large language model in low-resource language environments, we construct multilingual ICD automatic coding instructions and use the language transfer ICD automatic coding model for ICD automatic coding, which solves the problem of lack of task-related training data in low-resource language environments, and improves the accuracy and universality of ICD encoding.
Patent Information
- Application Number
- CN202411887436.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-20
AI Technical Summary
In low-resource language environments, the lack of task-related training data makes traditional automatic ICD encoding technology less adaptable and has high training costs. Especially in low-resource language environments such as Chinese, ICD encoding data annotation relies on manual labor, is inefficient and prone to errors.
Using a transfer learning method based on large language model, by obtaining medical data samples and ICD encoding sets containing diagnostic texts, using large language models to identify the type of text to be encoded, and constructing multilingual ICD automatic encoding instructions, using language transfer ICD automatic encoding model for ICD automatic encoding, and finally verify and align the output results according to the encoding rules of low-resource languages.
It greatly reduces the cost of training models from scratch to obtain ICD encoding capabilities in low-resource language environments. Through transfer learning, the model quickly obtains task capabilities, improves the accuracy and universality of ICD encoding, and reduces the dependence of manual annotation.
Smart Images

Figure CN120012715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of medical event recognition and entity linking, and in particular relates to an ICD automatic encoding method based on large language model transfer learning. Background Art
[0002] In the medical field, standardized coding of diseases and surgeries is a core component of information management. As a widely used medical coding standard worldwide, the International Classification of Diseases (ICD) coding system has not only played a positive role in promoting disease diagnosis, standardization of treatment processes, data analysis and international exchanges, but also provided basic support for medical insurance reimbursement, health management and medical research. ICD coding includes two main parts: disease diagnosis and surgical operation. Each ICD code is usually composed of English letters and numbers and matches the corresponding medical name. Due to its importance in medical informatization, the ICD coding system is widely used in major hospitals, medical insurance institutions and public health fields.
[0003] In order to solve these problems, automatic coding technology has become an important direction to improve the efficiency and accuracy of medical information processing. With the rapid development of natural language processing (NLP) technology and artificial intelligence, automatic coding technology has gradually been widely used. Traditional deep learning-based automatic coding methods usually require a large amount of manually annotated high-quality data for training, which makes them less adaptable and has high training costs. Especially in low-resource language environments, it is particularly difficult to build an effective automatic coding system due to the lack of sufficient task-related data.
[0004] Most of the existing public datasets are based on high-resource languages (such as English), and research on low-resource languages (such as Chinese) is relatively scarce. In some regions, in particular, ICD coding data annotation still relies on manual work, which is inefficient and prone to errors. Therefore, how to overcome the problem of data scarcity in low-resource language environments and improve the accuracy and universality of automatic coding has become a technical problem that needs to be solved urgently.
[0005] Based on the above problems, a method and system for ICD automatic encoding based on large language model transfer learning is proposed. Summary of the invention
[0006] The purpose of the present invention is to provide a method and system for ICD automatic encoding based on large language model transfer learning, aiming to solve the problem of lack of task-related training data in low-resource language environments.
[0007] In order to solve the above technical problems, the present invention is achieved through the following technical solutions:
[0008] A method for automatic ICD coding based on large language model transfer learning is used for data annotation of ICD coding of electronic medical record data in low-resource languages, the method comprising:
[0009] Acquire a medical data sample containing a diagnosis text and an ICD code set, wherein the medical data sample has a correct ICD code corresponding thereto;
[0010] For the input electronic medical record data in a low-resource language, a large language model is used to identify the type of text to be encoded, wherein the step guides the large language model to perform text classification by assembling prompt words through a template, wherein the template includes a role definition label, a coding type label, an example label, and an input electronic medical record data label in a low-resource language;
[0011] Based on the content to be encoded, construct a multilingual ICD automatic encoding instruction, which includes the content to be encoded input in a low-resource language and its corresponding high-resource language translation content, with the translated high-resource language content in front and the input low-resource language content in the back, and the language tags are marked respectively;
[0012] Based on the content to be encoded, multilingual ICD automatic encoding instructions are constructed; the language transfer ICD automatic encoding large model is used to perform ICD automatic encoding on the content to be encoded, and the ICD encoding results are output; based on the encoding rules of low-resource languages, the output ICD encoding results are verified and aligned, and the final encoding results are output.
[0013] On the one hand, an ICD automatic encoding system based on large language model transfer learning is provided, including:
[0014] The electronic medical record classification module is used to determine what type of coding is required based on the input low-resource language electronic medical records;
[0015] A translation module, used to translate the input electronic medical records in a low-resource language into a high-resource language;
[0016] A multi-language instruction building module is used to build multi-language ICD automatic coding instructions according to the input resource language electronic medical record;
[0017] Training module, used to conduct cross-language transfer training for large language models;
[0018] The deployment module is used to deploy large language models and make them accessible through APIs;
[0019] The code mapping module is used to map the results generated by the large language model to obtain ICD codes that meet the application environment of low-resource languages;
[0020] The user interaction module is used to provide coding services to doctors and coders after the large language model is deployed.
[0021] On the one hand, a readable storage medium is provided, which stores a computer program of a method or system for ICD automatic encoding based on large language model transfer learning, and the computer program can implement the method for ICD automatic encoding based on large language model transfer learning when executed in a computing device.
[0022] Beneficial effects:
[0023] The present invention is based on natural language processing technology, performs medical event recognition and entity linking on the original electronic medical record, uses a public data set of a high-resource language, translates it into a low-resource language through a translation model, constructs a parallel corpus, and conducts transfer learning training on the large language model; in the practical stage, it is necessary to perform certain transformations and alignments according to the actual rules of ICD coding in different language environments. This method takes advantage of the powerful language capabilities of the large language model, greatly reducing the cost of training the model from scratch to obtain ICD coding capabilities in a low-resource language environment. Through transfer learning, the model can quickly acquire task capabilities without the need for annotation. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for describing the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0025] Figure 1 It is a flow chart of the ICD automatic encoding method based on large language model transfer learning of the present invention;
[0026] Figure 2 It is an example of the process of collecting and collating data, translating it, and building a parallel corpus;
[0027] Figure 3 This is an example of the transformation and mapping process when the ICD versions of high-resource languages and low-resource languages are inconsistent;
[0028] Figure 4 It is an example of the transformation process of different encoding rules of the present invention;
[0029] Figure 5 It is a relationship diagram of various modules of the automatic encoding system of the present invention. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] In the specific implementation of this application, a method of ICD automatic encoding based on large language model transfer learning is proposed to solve the application problem of traditional ICD encoding in low-resource language environments. This method aims to improve the accuracy and universality of ICD encoding through the technical advantages of transfer learning.
[0032] In order to make it easier for readers to understand the entire content of this disclosure, the following terms appearing in this disclosure are now explained. It should be noted that the term explanations herein are only to assist readers in understanding and do not constitute a limitation on the technical solutions of this disclosure;
[0033] Large language model: refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. It usually refers to a language model with more than billions of parameters. ;
[0034] Transfer learning: refers to the deep learning training method that applies the knowledge or patterns learned in a certain field or task to different but related fields or problems;
[0035] like Figure 1 As shown, in some disclosures, the method steps include:
[0036] Step S1: Input electronic medical record data in a low-resource language and use a large language model to identify the type of text to be coded;
[0037] This step uses a specific template assembly prompt word to guide the large language model to perform text classification. The structure of the template assembly prompt word is as follows:
[0038] "{Role definition label}: {Role definition statement that sets the large language model as an encoding type classifier}\n{Encoding type label}: {Encoding type shape set or sequence}\n{Example label}: \n{Example text and classification result}\n{Input electronic medical record data label in low-resource language}: {Sentence that guides the large model to classify}\n{Input electronic medical record data in low-resource language}\n{Classification result output label}:"
[0039] Among them, {sample text and classification results} can contain a single example, multiple examples, or no examples. When there are examples, each example contains a statement that guides the large model to perform classification.
[0040] In some specific implementations, taking the migration from the English language environment (high resources) to the Chinese language environment (low resources) as an example, the specific content of the template is determined as follows based on the structural composition of the above template assembly prompt words:
[0041] "You are an excellent medical text classification expert. You are familiar with various types of electronic medical records. Now you need to select the type of electronic medical record data in [Data to be classified] from the classification provided in [Coding type]. You can refer to the examples provided in [Example] to complete your task.
[0042] [Encoding type]
[0043] Disease diagnosis codes, surgical operation codes
[0044] [Example]
[0045] Example 1: {Sample text 1}
[0046] Output 1: {classification result}
[0047] Example 2: {Sample text 2}
[0048] Output 2: {Classification result}
[0049]
Data to be classified
[0050] {Enter the data to be classified}
[0051] [Output]"
[0052] In step S2, the system constructs a multilingual ICD automatic encoding instruction based on the content to be encoded. The instruction includes the content to be encoded input in the low-resource language and its corresponding high-resource language translation content to form an instruction. The specific construction method is as follows:
[0053] In step S2, based on the content to be encoded, a multilingual ICD automatic encoding instruction is constructed, which includes the content to be encoded input in the low-resource language and the multilingual instruction for its corresponding high-resource language translation; in the multilingual instruction, the translated high-resource language content is in front, and the input low-resource language content is in the back, and the language tags are marked respectively, and the specific form is as follows:
[0054] "{High resource language instruction label}:{High resource language coding instruction}\n{Low resource language instruction label}:{Low resource language coding instruction}\n{High resource language text label}:\n{High resource language electronic medical record text}:{Low resource language text label}:\n{Low resource language electronic medical record text}\n"
[0055] In some specific implementations, taking the migration from the English language environment (high resources) to the Chinese language environment (low resources) as an example, the specific content of the template is determined as follows based on the structural composition of the above template assembly prompt words:
[0056] Instruction template: "Please extract the name of the surgical operation from the following medical record text:\nPlease extract the name of the surgical operation from the following medical record text:\nEnglish Text:\n{English medical record translated from the input Chinese medical record}\nChinese text:\n{input Chinese medical record text}\n".
[0057] This construction method ensures that during the coding process, the system can understand and process medical data in low-resource languages (such as Chinese) while referring to the translated data in high-resource languages (such as English). This dual-language instruction helps the large language model accurately understand and generate ICD coding results.
[0058] Step S3: Use the language transfer ICD automatic coding model to perform ICD automatic coding on the content to be coded, and output the ICD coding result;
[0059] The language transfer ICD automatic encoding large model training process adopts the cross-language transfer learning method, and its training process includes collecting and arranging data and translating, building parallel corpus and model fine-tuning. The details are as follows;
[0060] Collect and translate data: Select appropriate high-resource language datasets and translate the data into low-resource languages.
[0061] Constructing parallel corpus: By means of parallel corpus, the ICD coded data of high-resource languages are aligned with the data of low-resource languages.
[0062] Model fine-tuning: Use the constructed parallel corpus to fine-tune the large model to adapt to low-resource language environments.
[0063] A further solution of the present invention is: in collecting and collating data and translating, collecting and collating public data sets in the field of ICD coding in a high-resource language environment, and using a closed-source or open-source large language model with strong translation capabilities to translate the source language field task data set into a low-resource language, including input and output;
[0064] In constructing the parallel corpus, each piece of data in the high-resource language dataset is organized together with the corresponding translated data to form a training corpus for large language model transfer learning, where the input part is organized in the same way as the multilingual instruction construction method in step S2, and the output part is composed only of the translated low-resource language content, and the specific form is as follows:
[0065] Input: "{high resource language instruction label}:{high resource language encoding instruction}\n{low resource language instruction label}:{low resource language encoding instruction}\n{high resource language text label}:\n{high resource language electronic medical record text}:{low resource language text label}:\n{low resource language electronic medical record text}\n"
[0066] Output: "{Encoding result that complies with the encoding rules in a low-resource language environment}"
[0067] A further technical solution of the present invention is: in the process of collecting, collating and translating data, it is necessary to judge the consistency of the ICD coding version, industry requirements and actual situation used in the ICD coding task of medical records in the high-resource language data set and the low-resource language environment. If they are consistent, the input data is directly translated, and the output result is directly mapped using a standard dictionary. If they are inconsistent, it is necessary to transform according to the actual situation and then use the standard dictionary for mapping; the standard dictionary refers to the mapping dictionary formed by directly associating the ICD coding set issued by the national functional department or industry association of the low-resource language with the ICD coding set of the high-resource language through the ICD code or indirectly associating after transformation;
[0068] For example, Figure 2 As shown in the figure, the process of collecting and translating data and building parallel corpora includes:
[0069] The English dataset uses the MIMIC dataset, which contains anonymous patient data collected in the Intensive Care Unit (ICU) of Beth Israel Deaconess Medical Center since 2001. The database contains the patient's discharge records and their corresponding ICD codes, which are translated into Chinese and the corresponding ICD codes are converted into Chinese codes;
[0070] Specifically, when processing the ICD code in the MIMIC library, since the current ICD code version in the Chinese environment is not completely consistent with the English code, the Chinese ICD code version needs to be matched with the English ICD code according to certain rules to form a mapping dictionary.
[0071] For example, Figure 3As shown, the Chinese ICD code and the English ICD code are mapped. Taking the surgical operation code as an example, the version of the Chinese ICD code is the National Clinical Version 3.0 of the Surgical Operation Code issued by the National Health Commission of China, and the version of the English ICD code is ICD-9-CM-3 issued by the Centers for Disease Control and Prevention of the United States.
[0072] Step S4: Based on the encoding rules of the low-resource language, verify and align the output ICD encoding results and output the final encoding results. The specific steps are as follows;
[0073] Based on the encoding rules of low-resource languages, the output ICD encoding results are verified and aligned, and the final encoding results are output. If there are differences in the actual application of the data set in the high-resource language environment and the low-resource language environment, the ICD encoding version, industry requirements and actual situation used in the medical record ICD encoding task of the high-resource language data set and the low-resource language environment are judged for consistency according to the transformation method in step 3. If they are consistent, the model can be used directly to generate results. If not, according to the correspondence between the encoding rules in different language environments, the results that meet the encoding rules of the high-resource language environment are mapped or transformed to the results that meet the encoding rules of the low-resource language environment. The output results are post-processed to make them consistent with the actual application of the low-resource language environment.
[0074] For example, Figure 4 As shown in the method, since the ICD coding rules in the Chinese environment are different from those in the English environment, it is necessary to post-process the coding results according to the correspondence between the coding rules in the Chinese and English environments, and map the coding results to the Chinese application environment for practical application or further processing;
[0075] Specifically, the output ICD coding results that are inconsistent with the high-resource language coding rules are processed. The processing method is to correspond the Chinese ICD coding version with the ICD coding of the high-resource language according to certain rules for the inconsistent ICD coding of the high-resource language and the low-resource language, and map the high-resource language coding results to the low-resource language application environment.
[0076] like Figure 5 As shown, the present invention provides an ICD automatic coding system based on large language model transfer learning, which aims to solve the application problem of traditional ICD coding in low-resource language environments. The system improves the accuracy and universality of ICD coding through cross-language transfer learning technology, and supports automatic coding of electronic medical records in low-resource language environments. The core functions of the system include: electronic medical record classification, translation, instruction construction, transfer learning training, result mapping and other modules.
[0077] Detailed description of system modules;
[0078] The electronic medical record classification module is used to determine what type of coding is required based on the input low-resource language electronic medical records;
[0079] A translation module, used to translate the input electronic medical records in a low-resource language into a high-resource language;
[0080] A multi-language instruction building module is used to build multi-language ICD automatic coding instructions according to the input resource language electronic medical record;
[0081] The training module is used to perform cross-language transfer training on large language models.
[0082] The deployment module is used to deploy large language models and make them accessible through APIs;
[0083] The code mapping module is used to map the results generated by the large language model to obtain ICD codes that meet the application environment of low-resource languages;
[0084] The user interaction module is used to provide coding services to doctors and coders after the large language model is deployed.
[0085] In the description of this specification, the description with reference to the terms "one embodiment", "example", "specific example", etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0086] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and use the present invention well. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for automatic ICD coding based on large language model transfer learning, used for data annotation of ICD coding of electronic medical record data in low-resource languages, characterized in that: The method comprises: Acquire a medical data sample containing a diagnosis text and an ICD code set, wherein the medical data sample has a correct ICD code corresponding thereto; For the input electronic medical record data in a low-resource language, a large language model is used to identify the type of text to be encoded, wherein the step guides the large language model to perform text classification by assembling prompt words through a template, wherein the template includes a role definition label, a coding type label, an example label, and an input electronic medical record data label in a low-resource language; Based on the content to be encoded, construct a multilingual ICD automatic encoding instruction, which includes the content to be encoded input in a low-resource language and its corresponding high-resource language translation content, with the translated high-resource language content in front and the input low-resource language content in the back, and the language tags are marked respectively; Based on the content to be encoded, multilingual ICD automatic encoding instructions are constructed; the language transfer ICD automatic encoding large model is used to perform ICD automatic encoding on the content to be encoded, and the ICD encoding results are output; based on the encoding rules of low-resource languages, the output ICD encoding results are verified and aligned, and the final encoding results are output.
2. The ICD automatic encoding method based on large language model transfer learning according to claim 1 is characterized in that: The template assembly prompt word structure consists of the following parts: {Set the large language model as the role definition statement of the encoding type classifier}\n{Encoding type label}\n{Encoding type shape set or sequence}\n{Example label}\n{Example text and classification result}\n{Input electronic medical record data label in low-resource language}\n{Input electronic medical record data in low-resource language}\n{Classification result output label}\n” Among them, {sample text and classification results} can contain a single example, multiple examples, or no examples. When there are examples, each example contains a statement that guides the large model to perform classification.
3. The ICD automatic encoding method based on large language model transfer learning according to claim 1 is characterized in that: The specific format of the multilingual ICD automatic coding instruction is: {High resource language encoding instructions}\n{Low resource language encoding instructions}\n{High resource language text label}:\n{High resource language electronic medical record text}:{Low resource language text label}:\n{Low resource language electronic medical record text}\n.
4. The ICD automatic encoding method based on large language model transfer learning according to claim 1 is characterized in that: The training of the language transfer ICD automatic encoding large model adopts the method of cross-language transfer learning, and its training process includes collecting and arranging data and translating, building parallel corpus and model fine-tuning; The model fine-tuning is to use the constructed parallel corpus to fine-tune the large model to adapt to the low-resource language environment.
5. The ICD automatic encoding method based on large language model transfer learning according to claim 4 is characterized in that: The specific steps of collecting, collating and translating data are as follows: Select a dataset in a high-resource language and use a large language model to translate that dataset into a low-resource language; The translated low-resource language data is aligned with the high-resource language data to form a cross-language training corpus, where the input part is constructed using multilingual instruction templates, while the output part consists only of the encoding results of the low-resource language for transfer learning training.
6. The ICD automatic encoding method based on large language model transfer learning according to claim 4 is characterized in that: When constructing the parallel corpus, each piece of data in the high-resource language data set is organized together with the corresponding translated data to form a training corpus for large language model transfer learning, and the organization method of the input part is the same as the multilingual instruction construction method, and the output part is composed of the translated low-resource language content.
7. The ICD automatic encoding method based on large language model transfer learning according to claim 4 is characterized in that: The method further includes judging the consistency of the ICD coding versions, industry requirements and actual conditions used in the medical record ICD coding task in the high-resource language dataset and the low-resource language environment, and if they are consistent, directly translating the input data and directly mapping the output results using a standard dictionary; If they are inconsistent, they will be transformed according to the actual situation and then mapped using the standard dictionary.
8. The ICD automatic encoding method based on large language model transfer learning according to claim 1 is characterized in that: The encoding rules based on the low-resource language are used to verify and align the output ICD encoding results. When the final encoding results are output, if there are differences between the data set in the high-resource language environment and the actual application in the low-resource language environment, the output results are post-processed according to the following transformation method to make them conform to the actual application of the low-resource language environment. The specific transformation method is as follows: The consistency of the ICD coding versions, industry requirements and actual conditions used in the ICD coding task of medical records in high-resource language datasets and low-resource language environments is judged. If they are consistent, the model can be used directly to generate results. If they are inconsistent, the results that conform to the coding rules of the high-resource language environment are mapped or transformed to results that conform to the coding rules of the low-resource language environment based on the correspondence between the coding rules in different language environments.
9. An ICD automatic encoding system based on large language model transfer learning, characterized in that: include: The electronic medical record classification module is used to determine what type of coding is required based on the input low-resource language electronic medical records; A translation module, used to translate the input electronic medical records in a low-resource language into a high-resource language; A multi-language instruction building module is used to build multi-language ICD automatic coding instructions according to the input resource language electronic medical record; Training module, used to conduct cross-language transfer training for large language models; The deployment module is used to deploy large language models and make them accessible through APIs; The code mapping module is used to map the results generated by the large language model to obtain ICD codes that meet the application environment of low-resource languages; The user interaction module is used to provide coding services to doctors and coders after the large language model is deployed.
10. A readable storage medium, characterized in that: The storage medium stores a computer program of the method or system for ICD automatic encoding based on large language model transfer learning according to any one of claims 1 to 8, and the computer program can implement the method for ICD automatic encoding based on large language model transfer learning according to any one of claims 1 to 8 when executed in a computing device.
Citation Information
Patent Citations
Sequence labeling method utilizing cross-language information
CN111274829A
Disease classification ICD automatic coding method and device based on comparative learning
CN116822579A
Automatic coding method and system for surgical operation records based on large language model
CN119092046A