ICD coding prediction method and device based on large language model
By using an ICD encoding prediction method based on a large language model, combined with hierarchical prediction and prompt word restrictions, the problem of uninterpretable ICD encoding prediction results is solved, achieving higher usability and accuracy, and facilitating manual inspection and processing.
Patent Information
- Application Number
- CN202510992038.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-10-31
AI Technical Summary
Existing ICD encoding prediction methods employ an end-to-end prediction approach, resulting in uninterpretable results that are difficult to meet the needs of manual analysis.
An ICD coding prediction method based on a large language model is adopted. Through ICD knowledge base module, disease pattern prediction module, specific diagnosis prediction module, accurate diagnosis prediction module and coding inspection module, combined with hierarchical prediction and prompt word restriction, the prediction reason and the credibility of the code are output, which can be easily checked by humans.
It significantly improves the usability and accuracy of the ICD coding prediction system, provides reliable prediction results, and facilitates manual inspection and subsequent processing.
Smart Images

Figure CN120878263A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing, and specifically provides an ICD encoding prediction method and apparatus based on a large language model. Background Technology
[0002] As natural language processing and deep learning technologies continue to develop, ICD encoding prediction methods are also constantly improving. Traditional deep learning algorithms treat the ICD encoding grammar problem as a text classification problem, generally using Word2Vec, FastText, or BERT for word embedding and RNN, LSTM, and other algorithms for text classification.
[0003] This method uses an end-to-end prediction approach, inputting patient case information and directly outputting predictive codes; the intermediate process is uninterpretable. This makes manual analysis of the prediction results extremely inconvenient. Summary of the Invention
[0004] This invention addresses the shortcomings of the prior art by providing a highly practical ICD encoding prediction method based on a large language model.
[0005] A further technical objective of this invention is to provide a reasonably designed, safe, and applicable ICD encoding prediction device based on a large language model.
[0006] The technical solution adopted by this invention to solve its technical problem is:
[0007] An ICD encoding prediction method based on a large language model has the following steps:
[0008] S1, the ICD knowledge base module organizes the contents of the International Statistical Classification in a tree structure;
[0009] S2. Predict disease patterns;
[0010] S3. Construct specific diagnostic prediction modules;
[0011] S4, Accurate Diagnosis and Prediction Module;
[0012] S5. Perform encoding checks;
[0013] S6. Perform primary and secondary identification.
[0014] Furthermore, in step S1, after organizing according to a tree structure, the data is vectorized, inputting case information and outputting content related to the case information filtered from the ICD knowledge base.
[0015] Furthermore, in step S2, a hierarchical prediction method is adopted, which calls the large language model to predict the patient's diagnosis within a specific disease pattern based on the case information, and outputs the first letter of the prediction code and the basis for the prediction.
[0016] Furthermore, in step S3, after obtaining the disease pattern, content related to the disease model is filtered from the knowledge base and case information and passed to the large model. The predicted content is restricted to specific diagnostic codes and related explanations under the disease pattern by prompt words.
[0017] Furthermore, in step S4, after obtaining a specific diagnosis of the disease, content related to the specific diagnosis is filtered from the knowledge base and case information and passed to the large model. The predicted content is limited to the accurate diagnosis code and related explanation of the disease pattern by prompt words.
[0018] Furthermore, in step S5, the reflective capabilities of the large model are used to check the credibility of its own encoding, and prompt words are used to restrict its output to numbers between 0 and 1.
[0019] Furthermore, in step S5, the primary diagnosis is the main reason for the patient's current medical visit, which must exist and be unique; other diagnoses are other diseases or conditions that the patient also has, which are not the primary treatment goals.
[0020] The primary and secondary identification module distinguishes between primary and other diagnoses by calling a large model and returns the diagnostic codes in JSON format for easy subsequent processing.
[0021] An ICD encoding prediction device based on a large language model includes: at least one memory and at least one processor;
[0022] The at least one memory is used to store a machine-readable program;
[0023] The at least one processor is configured to invoke the machine-readable program to execute an ICD encoding prediction method based on a large language model.
[0024] Compared with existing technologies, the ICD encoding prediction method and apparatus based on a large language model of the present invention have the following outstanding advantages:
[0025] This invention leverages the natural language understanding capabilities of a large language model, utilizes the national clinical version of the disease classification code knowledge base, and combines a hierarchical prediction arrangement to output the rationale behind the ICD coding prediction while performing the prediction. This facilitates manual review of the prediction results and significantly improves the practical usability of the ICD coding prediction system. This method not only optimizes the usability and reliability of the ICD coding prediction system during diagnosis and treatment but also incorporates a knowledge base to enhance prediction accuracy, having a profound impact on accelerating information technology development and benefiting both medical professionals and patients, as well as the development of the healthcare industry. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart illustrating an ICD encoding prediction method based on a large language model. Detailed Implementation
[0028] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] The following is a preferred embodiment:
[0030] like Figure 1 As shown, the ICD encoding prediction method based on a large language model in this embodiment has the following steps:
[0031] S1, the ICD knowledge base module organizes the contents of the International Statistical Classification in a tree structure;
[0032] The content of the "National Clinical Version 2.0 of Disease Classification and Codes" is organized into a tree structure and vectorized. Inputting case information outputs content related to the case information, filtered by the ICD knowledge base module. This module allows adjustment of the TopK filtering and relevance threshold.
[0033] S2. Predict disease patterns;
[0034] Because the National Clinical Version 2.0 of the "Disease Classification and Codes" contains 37,294 disease codes, directly predicting the complete code would result in an excessively large search space, affecting accuracy. Therefore, a hierarchical prediction method is adopted. The disease pattern prediction module calls a large language model to predict which disease pattern the patient's diagnosis is likely to fall under based on case information, outputting the first letter of the predicted code and the basis for the prediction, such as "E" indicating endocrine, nutritional, and metabolic diseases.
[0035] S3. Construct specific diagnostic prediction modules;
[0036] After obtaining the disease pattern, content related to the disease model is filtered from the knowledge base and case information and passed to the large model. The predicted content is restricted to specific diagnostic codes and related explanations under the disease pattern by prompt words.
[0037] S4, Accurate Diagnosis and Prediction Module;
[0038] After obtaining a specific diagnosis of the disease, content related to that specific diagnosis is filtered from the knowledge base and case information and passed to the large model. The predicted content is limited by prompt words to be the accurate diagnostic code and related explanation for that disease pattern.
[0039] S5. Perform encoding checks;
[0040] Because there are exclusion clauses in ICD coding, for example, "E00" represents congenital iodine deficiency syndrome, but subclinical iodine deficiency hypothyroidism must be excluded. The coding check module uses the reflective capabilities of the large model to check the reliability of its own coding and uses prompts to limit its output to numbers between 0 and 1.
[0041] S6. Perform primary and secondary identification;
[0042] A patient often has multiple diagnoses. The primary diagnosis is the main reason for the patient's current medical visit; it must exist and be unique. Other diagnoses are other diseases or conditions that the patient also has, and are not the primary treatment goals. The primary and secondary diagnosis identification module distinguishes between the primary and other diagnoses by calling a larger model and returns the diagnosis code in JSON format for subsequent processing.
[0043] The data was primarily based on the content of the "National Clinical Version 2.0 of Disease Classification and Codes," with other relevant materials collected and organized into a tree structure. Each node contains a partial code (such as the second and third digits of the code) and its explanation. After the tree structure was constructed, the text was vectorized using an embedding model, ultimately forming a vector tree diagram.
[0044] The coding prediction process requires sequentially calling the disease pattern prediction module, the specific diagnosis prediction module, the accurate diagnosis prediction module, and the coding check module. Within each module, prompts are used to guide the larger model to complete the corresponding tasks. The prompts are in Markdown format.
[0045] As the final output of the algorithm, the language understanding ability of the large model is first invoked, using prompt words to guide it in self-reflection and check the accuracy of the output. Additionally, prompt words guide the large model to output the final recognition result in JSON format, separating multiple other diagnostic codes with semicolons, such as:
[0046] {
[0047] "Main Diagnostic Code": "J81.x00x002"
[0048] "Other Diagnostic Codes": "I50.907; I50.903; I25.103; I20.000;"
[0049] Based on the above method, an ICD encoding prediction device based on a large language model in this embodiment includes: at least one memory and at least one processor;
[0050] The at least one memory is used to store a machine-readable program;
[0051] The at least one processor is configured to invoke the machine-readable program to execute an ICD encoding prediction method based on a large language model.
[0052] The above-described specific embodiments are merely specific examples of the present invention. The patent protection scope of the present invention includes, but is not limited to, the above-described specific embodiments. Any technical solution that conforms to the above-described specific embodiments of the present invention and any appropriate changes or substitutions made by those skilled in the art should fall within the patent protection scope of the present invention.
[0053] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for predicting ICD encodings based on a large language model, characterized in that, It has the following steps: S1, the ICD knowledge base module organizes the contents of the International Statistical Classification in a tree structure; S2. Predict disease patterns; S3. Construct specific diagnostic prediction modules; S4, Accurate Diagnosis and Prediction Module; S5. Perform encoding checks; S6. Perform primary and secondary identification.
2. The ICD encoding prediction method based on a large language model according to claim 1, characterized in that, In step S1, after organizing the data according to a tree structure, it is vectorized, inputting case information and outputting content related to the case information filtered by the ICD knowledge base.
3. The ICD encoding prediction method based on a large language model according to claim 1, characterized in that, In step S2, a hierarchical prediction method is adopted, which calls a large language model to predict the patient's diagnosis within a specific disease pattern based on the case information, and outputs the first letter of the prediction code and the basis for the prediction.
4. The ICD encoding prediction method based on a large language model according to claim 3, characterized in that, In step S3, after obtaining the disease pattern, content related to the disease model is filtered from the knowledge base and case information and passed to the large model. The predicted content is restricted to specific diagnostic codes and related explanations under the disease pattern by prompt words.
5. The ICD encoding prediction method based on a large language model according to claim 4, characterized in that, In step S4, after obtaining a specific diagnosis of the disease, content related to the specific diagnosis is filtered from the knowledge base and case information and passed to the large model. The predicted content is limited to the accurate diagnosis code and related explanation of the disease pattern by prompt words.
6. The ICD encoding prediction method based on a large language model according to claim 5, characterized in that, In step S5, the reflective ability of the large model is used to check the credibility of its own encoding, and prompt words are used to limit its output to numbers between 0 and 1.
7. The ICD encoding prediction method based on a large language model according to claim 6, characterized in that, In step S5, the primary diagnosis is the main reason for the patient's current medical visit, which must exist and be unique; other diagnoses are other diseases or conditions that the patient also has, which are not the primary treatment goals. The primary and secondary identification module distinguishes between primary and other diagnoses by calling a large model and returns the diagnostic codes in JSON format for easy subsequent processing.
8. An ICD encoding prediction device based on a large language model, characterized in that, include: At least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to perform the method according to any one of claims 1 to 7.