Automatic ICD coding method and system

Through a multi-option verification method based on a large language model, clinical text is processed and standardized ICD codes are generated, which solves the problems of poor complexity and generalization of ICD coding tasks, and achieves more efficient and accurate ICD coding.

CN119943238APending Publication Date: 2025-05-06RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411700538.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, ICD encoding tasks are complex and prone to omissions, errors and hallucinations, manual encoding is time-consuming, labor-intensive and inaccurate, and the existing automation methods have poor generalization in different data sets and languages.

Method used

ICD encoding method based on large language model and multi-option verification is adopted, and standardized ICD code codes are finally output by processing clinical text, pre-screening ICD code options, generating prompt words and inputting LLM for inference.

Benefits of technology

It simplifies the difficulty of ICD encoding tasks, improves the accuracy and generalization of ICD code verification, and reduces the error rate and time-consuming of manual encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943238A_ABST
    Figure CN119943238A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic ICD coding method which comprises the following steps: S1, processing a clinical text, extracting text information of each part in an EHR, and processing the text information into an input text; s2, generating ICD code options through pre-screening of a classification model; s3, predicting cue words based on the ICD code of multi-option verification, wherein the cue words comprise an instruction, an input text, options and an output template; s4, inputting the designed cue word into the LLM for reasoning, and outputting each token in sequence by the LLM through an autoregression mode and decoding the token into a text character; and S5, processing the free text output by the LLM, and outputting a standard ICD code. By adopting the method, the difficulty of an ICD coding task is simplified by adjusting a task form, and the advantages of a large language model in a medical long text reasoning task are fully played.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to an automatic ICD coding method and system. Background Art

[0002] The International Classification of Diseases (ICD) code is a unified coding method used in hospitals and other medical institutions. It classifies diseases according to their etiology, pathology, clinical manifestations, and anatomical location. It also contains unified codes for surgical, diagnostic, and therapeutic procedures. ICD codes use a combination of letters and numbers to represent specific diseases or diagnoses. ICD codes have a variety of uses, such as reporting diseases and health conditions, assisting in medical reimbursement decisions, and collecting morbidity and mortality statistics.

[0003] Electronic Health Records (EHRs) are an important part of medical data management. EHRs record the full cycle of patient data in the hospital, including key medical information such as outpatient visits, admissions, diagnoses, examinations, tests, treatments, and discharges. These clinical medical record texts can be associated with diagnostic codes in the ICD, which represent the diagnosis and procedural information of the visit. In medical institutions, ICD coders manually assign appropriate ICD codes by viewing EHRs. Such manual coding is time-consuming, labor-intensive, and prone to errors. Manual coding often encounters the following difficulties: the hierarchical structure of ICD codes makes it difficult to distinguish diseases at the same level; when doctors write diagnostic descriptions, they often use abbreviations and synonyms, which can easily cause ambiguity with the descriptions of ICD codes; in many cases, multiple closely related diagnostic descriptions should be mapped to a specific ICD code, and inexperienced coders may code each disease separately.

[0004] In the field of medical artificial intelligence, the application of natural language processing (NLP) algorithms is gradually maturing, especially in the field of ICD coding automation. By using deep learning technology and combining large medical data sets, the performance of automatic ICD coding tasks has been improved. In recent years, language models represented by the Transformer architecture have become the dominant force in NLP research and have also been applied to ICD coding tasks. For example, the pretrained language model-ICD (PLM-ICD) algorithm treats the automatic ICD coding task as a multi-label text classification problem and optimizes the ICD coding task based on the pretrained model in the medical text field. The multi-label classification method represented by PLM-ICD mainly includes a text feature extraction network and a classification layer. The text feature extraction network uses pre-trained models such as BERT. The number of classifications in the classification layer is usually set to the number of ICD codes in the training set. In practical applications, in order to adapt to data sets with different label types and data sets in different languages, this method needs to be fine-tuned on the corresponding data sets. This will result in the generalization of the classification method in different data sets and different languages. The model fine-tuned on data set A has poor effect on data set B.

[0005] Unlike classification models such as PLM-ICD, with the development of generative artificial intelligence, more and more methods have begun to try to apply large language models (LLM) represented by ChatGPT and GPT-4 to ICD coding tasks. LLM has significant advantages in processing long input texts and generalization of various downstream tasks, and has good performance in some tasks in the medical field such as medical question answering (MedQA) and the United States Medical Licensing Examination (USMLE). However, due to the large number of ICD coding types and complex tasks, LLM directly generates text to process ICD coding tasks and encounters some new problems. Based on LLM and the method of directly generating ICD codes, all ICD codes are directly generated from the input text through the next token prediction in LLM. In essence, the direct generation paradigm transforms the code classification problem into a code generation problem. Unlike the classification paradigm where each ICD code is regarded as an independent category, in the LLM decoding process, one ICD code may correspond to multiple Token mappings, which makes LLM-generated ICD codes prone to omissions, errors, and hallucinations. Summary of the invention

[0006] The object of the present invention is to provide an automatic ICD encoding method and system to solve the problems raised in the above background technology.

[0007] To achieve the above object, one aspect of the present invention provides an automatic ICD coding method, comprising the following steps:

[0008] Step S1, processing clinical text, extracting text information of each part in EHR, and processing it into input text;

[0009] Step S2, pre-screening and generating ICD code options through the classification model;

[0010] Step S3, predicting prompt words of ICD codes based on multiple-option verification, the prompt words including instructions, input text, options and output templates;

[0011] Step S4, input the designed prompt words into LLM for reasoning, and LLM outputs each token in turn through autoregression and decodes it into text characters;

[0012] Step S5, processing the free text output by LLM and outputting a standardized ICD code.

[0013] Furthermore, the clinical text described in step S1 includes admission diagnosis, chief complaint, allergy history, family history, social history, past medical history, physical examination, relevant test results, hospitalization process, discharge diagnosis, discharge instructions, discharge medication, discharge arrangements, and follow-up instructions.

[0014] Furthermore, in step S1, the private information in the clinical text is anonymized, and all clinical texts are sequentially concatenated into a large text as input.

[0015] Furthermore, in step S2, the classification model PLM-ICD is used for preliminary screening, and the scores are sorted from high to low, and the top k ICD codes are selected as options, where top k is set to 50.

[0016] Furthermore, in step S2, a training data set is constructed for fine-tuning of non-English medical record text data.

[0017] Furthermore, the LLM model in step S4 includes commercial models and open source models. The commercial models include ChatGPT and GPT-4o; the open source models include Qwen and Llama.

[0018] Further, outputting the ICD code includes the following steps:

[0019] Step S501, extracting the code name and verification result in the LLM original output by regular expression matching;

[0020] Step S502, filtering according to the verification result of each ICD code option;

[0021] Step S503, checking whether the filtered ICD code is a legal ICD code according to the ICD code library;

[0022] Step S504, output all legal ICD codes.

[0023] Another aspect of the present invention provides an automatic ICD coding system, comprising a text processing module, an option generation module, a prompt word module, a model reasoning module, and an output module, wherein:

[0024] A text processing module is used to extract text information from various parts of EHR and process it into input text;

[0025] Option generation module, which pre-screens ICD code options through classification models;

[0026] A prompt word module is used for predicting prompt words of ICD codes based on multiple-option verification. The prompt words include instructions, input text, options and output templates.

[0027] Model inference module: input the designed prompt words into LLM for inference. LLM outputs tokens in sequence through autoregression and decodes them into text characters.

[0028] The output module processes the free text output by LLM and outputs standardized ICD codes.

[0029] Compared with the prior art, the present system and method have the following advantages:

[0030] 1. This paper proposes an ICD encoding method based on a large language model and multiple-choice verification. By adjusting the task form, the difficulty of the ICD encoding task is simplified, and the advantages of the large language model in the medical long text reasoning task are fully utilized;

[0031] 2. The present invention directly adopts the existing classification method to help the initial screening of ICD code options. Compared with random screening, it not only speeds up the processing speed, but also improves the accuracy of the final ICD code verification;

[0032] 3. The present invention does not require large-scale training, fully utilizes the zero-sample reasoning capability of LLM, and has better generalization in cross-dataset reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 A flow chart of an automatic ICD coding method.

[0034] Figure 2 Schematic diagram of the principle of an automatic ICD coding method.

[0035] Figure 3 This is a schematic diagram of the prompt word template.

[0036] Figure 4 Schematic diagram of the ICD code output process. DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] like Figure 1 , 2 are respectively a flow chart and a schematic diagram of the method of the present invention. The embodiment of the present invention provides a method, and the specific steps are as follows:

[0039] Step S1, clinical text processing.

[0040] First, the clinical text in the EHR is extracted, including admission diagnosis, chief complaint, allergy history, family history, social history, past medical history, physical examination, test results, hospitalization process, discharge diagnosis, discharge instructions, discharge medication, discharge arrangements, follow-up instructions, etc. Then, the private information in the text is anonymized, and all clinical texts are spliced ​​into a large text in sequence as input.

[0041] Step S2, option generation.

[0042] Option generation provides preliminary selection for subsequent ICD code verification. The ICD code types are large in scale. For example, the ICD-10-CM code types developed by the U.S. Department of Health and Human Services have 68,000 types. In order to speed up the option generation, the classification model PLM-ICD is used for preliminary screening.

[0043] PLM-ICD outputs a classification score {x1, x2, x3...x n}, sort by score from high to low, select the topk ICD codes as options, here topk is set to 50. It should be noted that the original PLM-ICD model was trained on the English datasets MIMIC-2 and MIMIC-3. If it is to be migrated to non-English medical record text data, it is necessary to build a training dataset for fine-tuning.

[0044] Step S3, prompt word filling.

[0045] In order to make full use of the powerful context processing and generalization capabilities of LLM in clinical coding tasks, this patent proposes a prompt word based on multiple-choice verification. The prompt word template is as follows: Figure 3 As shown below.

[0046] {note}, {candidate1}, {candidate2}...{candidate50}, etc. are all placeholders. After receiving the input clinical text and the ICD code options after the initial screening, the placeholder {note} is filled with the clinical text.

[0047] {candidate1}~{candidate50} are filled with ICD code options in sequence. In addition to the code, the name description of the corresponding disease needs to be supplemented. For example, to fill in ICD code L03.116 in {candidate1}, the corresponding disease description is "Cellulitis of left lower limb", then when filling in, {candidate1} is replaced with "L03.116-Cellulitis of left lower limb", and all subsequent candidates are filled in sequence.

[0048] Step S4, model reasoning.

[0049] When the prompt words are filled, they are input into the LLM model for reasoning. LLM outputs each token in turn through autoregression and decodes it into output text. In terms of LLM model selection, you can use commercial models such as ChatGPT, GPT-4o, etc., or open source models such as Qwen, Llama, etc. For commercial models, reasoning is performed by calling their application service interfaces. For open source models, reasoning can be performed after deployment on the local GPU server.

[0050] Step S5, output processing.

[0051] The output of LLM reasoning is free text, which is then processed into ICD code. The specific process is as follows: Figure 4 As shown. The following steps are included:

[0052] Step S501, regular expression extraction, extracting the code name and verification result in the original output of LLM through regular expression matching;

[0053] Step S502, exclusion filtering, filtering according to the verification result of each ICD code option, for the positive result of the verification result of "Confirmed", retain the option, and exclude the option of the verification result of "Ruled out";

[0054] Step S503, code check, based on the ICD code library, check whether each filtered ICD code is a legal ICD code;

[0055] Step S504, output all legal ICD codes.

[0056] Table 1 is a comparison of the experimental results of this method and the direct generation method on data set A. The models called by both are OpenAI's GPT-4o. Due to the complexity of the ICD encoding task itself, the Micro F1 and Macro F1 of the direct generation method are both low. Micro F1 and Macro F1 are used to evaluate the indicators of multi-classification tasks. The multi-option verification method proposed in the present invention can effectively improve the indicators Micro-F1 from 0.2401 to 0.4428, and Macro-F1 from 0.0923 to 0.2614. It can be seen that the present invention can effectively improve the accuracy of ICD encoding by comparing the method based on direct generation of large models.

[0057] Table 1

[0058]

[0059] Table 2 is a comparison of the zero-sample reasoning effect of this method and PLM-ICD on data set B. This method has a great advantage in both Micro F1 and Macro F1. It can be seen that the present invention can effectively improve the zero-sample reasoning effect compared with classification methods such as PLM-ICD.

[0060] Table 2

[0061]

[0062] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An automatic ICD coding method, characterized in that: The following steps are involved: Step S1, processing clinical text, extracting text information of each part in EHR, and processing it into input text; Step S2, pre-screening and generating ICD code options through the classification model; Step S3, predicting prompt words of ICD codes based on multiple-option verification, the prompt words including instructions, input text, options and output templates; Step S4, input the designed prompt words into LLM for reasoning, and LLM outputs each token in turn through autoregression and decodes it into text characters; Step S5, processing the free text output by LLM and outputting a standardized ICD code.

2. An automatic ICD coding method according to claim 1, characterized in that: The clinical text described in step S1 includes admission diagnosis, chief complaint, allergy history, family history, social history, past medical history, physical examination, relevant test results, hospitalization process, discharge diagnosis, discharge instructions, discharge medication, discharge arrangements, and follow-up instructions.

3. An automatic ICD coding method according to claim 1, characterized in that: In step S1, the private information in the clinical text is anonymized, and all clinical texts are sequentially concatenated into a large text as input.

4. The automatic ICD coding method according to claim 1, characterized in that: In step S2, the classification model PLM-ICD is used for preliminary screening and scoring, and the scores are sorted from high to low, and the top k ICD codes are selected as options, and the top k is set to 50.

5. The automatic ICD coding method according to claim 1, characterized in that: In step S2, for non-English medical record text data, a training data set is constructed for fine-tuning.

6. The automatic ICD coding method according to claim 1, characterized in that: The LLM model in step S4 includes commercial models and open source models. The commercial models include ChatGPT and GPT-4o; the open source models include Qwen and Llama.

7. An automatic ICD coding method according to claim 1, characterized in that: Outputting ICD codes involves the following steps: Step S501, extracting the code name and verification result in the LLM original output by regular expression matching; Step S502, filtering according to the verification result of each ICD code option; Step S503, checking whether the filtered ICD code is a legal ICD code according to the ICD code library; Step S504, output all legal ICD codes.

8. An automatic ICD coding system, characterized in that: It includes text processing module, option generation module, prompt word module, model reasoning module and output module, among which: A text processing module is used to extract text information from various parts of EHR and process it into input text; Option generation module, which pre-screens ICD code options through classification models; A prompt word module is used for predicting prompt words of ICD codes based on multiple-option verification. The prompt words include instructions, input text, options and output templates. Model inference module: input the designed prompt words into LLM for inference. LLM outputs tokens in sequence through autoregression and decodes them into text characters. The output module processes the free text output by LLM and outputs standardized ICD codes.

Citation Information

Cited By

  • Personalized medical document oriented large language model agent construction method and system

    CN121096508A