Automatic ICD code conversion method of outpatient and emergency diagnosis and treatment information page and medium

By using neural network recognition model in the electronic medical record system, ICD code is automatically converted into national standard code, and through similarity verification, the problem of low efficiency and poor accuracy of converting ICD code to national standard code in the electronic medical record system is solved, and the efficiency and accuracy of filling out disease codes are improved.

CN120199398APending Publication Date: 2025-06-24四川互慧软件有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510266231.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing electronic medical record system can only accept ICD codes and cannot be directly converted into China's national standard codes, resulting in low efficiency in filling out disease codes in the outpatient and emergency diagnosis and treatment information page and prone to errors.

Method used

By using a neural network-based recognition model in the electronic medical record system, the ICD encoding is automatically converted into national standard encoding, and the accuracy of the conversion results is verified by calculating the similarity of disease names.

Benefits of technology

It improves the efficiency of filling in Chinese code codes and disease names on the information page, enhances the accuracy of encoding conversion, and reduces the need for manual review.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199398A_ABST
    Figure CN120199398A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic ICD code conversion method of an outpatient and emergency treatment information page and a medium, and relates to the technical field of medical information identification. First medical information and second medical information of a patient are acquired based on an electronic medical record system; the first medical information comprises medicine information, allergy history and chief complaint information, and the second medical information comprises a disease name and an ICD code; in response to the acquired first medical information and second medical information, performing identification according to a preset first identification rule, and taking the identified disease name as a second disease name; and calculating the similarity between the first disease name and the second disease name as a first similarity, and if the first similarity reaches a first similarity threshold, editing the disease code in the information page based on the national standard code corresponding to the second disease name. Compared with the prior art, the filling efficiency and accuracy of filling the national standard codes and the corresponding disease names in the information pages are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the applicant's previous application. The application number of the previous application is CN202411806231.1, and the application title is a method and medium for automatic conversion of disease codes on the outpatient and emergency diagnosis and treatment information page. Technical Field

[0002] The present invention relates to the technical field of medical information recognition, and specifically relates to a method and medium for automatic conversion of ICD codes on the outpatient and emergency diagnosis and treatment information page. Background Art

[0003] The outpatient (emergency) diagnosis and treatment information page (hereinafter referred to as the "information page", with a total of 72 fields) is an information summary reflecting the patient's current visit process formed by the hospital based on the outpatient (emergency) medical record and various information generated during the patient's outpatient (emergency) visit in the hospital, including patient basic information, visit process information, diagnosis and treatment information, and expense information, etc.

[0004] Disease codes need to be filled in the information page. In China, the Diagnosis-related Groups (DRGs) are currently used for the allocation and control of medical insurance funds. In the DRGs scheme, the disease codes use the classification codes of the ICD-10 type (hereinafter referred to as ICD codes), while in the currently issued new regulations, it is clearly stated that the disease codes need to use the classification codes of the "National Clinical Version of Disease Classification and Codes" (hereinafter referred to as national standard codes).

[0005] In the current electronic medical record system adopted by hospitals, the classification codes that can be accepted and edited for the classification of diseases are only ICD codes. Therefore, how to convert ICD codes to national standard codes when reading from the electronic medical record system to the information page has become an urgent problem to be solved. Summary of the Invention

[0006] The technical problem to be solved by this application is to provide a method and medium for automatic conversion of disease codes on the outpatient and emergency diagnosis and treatment information page, which has the characteristic of automatically converting ICD codes to national standard codes when reading disease codes from the medical information system to the information page.

[0007] In a first aspect, in one embodiment, a method for automatic conversion of disease codes on the outpatient and emergency diagnosis and treatment information page is provided, including: Obtaining the first medical information and the second medical information of the patient based on the electronic medical record system; the first medical information includes drug information, allergy history, and chief complaint information, and the second medical information includes disease name and ICD code; taking the disease name in the second medical information as the first disease name, and taking the ICD code in the second medical information as the first ICD code; In response to the acquired first medical information and second medical information, perform identification according to a preset first identification rule, and use the identified disease name as the second disease name. The first identification rule is used to identify the national standard code and the corresponding disease name; Calculate the similarity between the first disease name and the second disease name as the first similarity. If the first similarity reaches the first similarity threshold, edit the disease code in the information page based on the national standard code corresponding to the second disease name.

[0008] In one embodiment, the step of in response to the acquired first medical information and second medical information, performing identification according to a preset first identification rule includes: In response to the acquired first medical information and second medical information, identify the national standard code and the corresponding disease name based on the national standard code identification model; the national standard code identification model is a neural network model trained in combination with a loss function.

[0009] In one embodiment, the training method of the national standard code identification model includes: Collect multiple medical text data containing the first medical information and the second medical information, and perform annotation of the national standard code and the corresponding disease name; Convert the multiple medical text data with the national standard code and the corresponding disease name annotated into a format suitable for input to the Huatuo large model to obtain annotated data; Divide the annotated data into a training set, a validation set, and a test set; Based on the preset training parameters, use the training set as the input of the Huatuo large model for training, evaluate the performance of the Huatuo large model during training based on the validation set and adjust the hyperparameters, and evaluate the generalization ability of the Huatuo large model based on the test set, so as to obtain a trained and improved Huatuo large model; Use the trained and improved Huatuo large model as the national standard code identification model.

[0010] In one embodiment, the step of if the first similarity reaches the first similarity threshold, then edit the disease code in the information page based on the national standard code corresponding to the second disease name includes: If the first similarity reaches the first similarity threshold, edit the national standard code corresponding to the second disease name as the disease code in the information page.

[0011] In one embodiment, before obtaining the second disease name, it further includes: in response to the acquired first medical information, perform identification according to a preset second identification rule, use the identified disease name as the third disease name, and use the identified ICD code as the second ICD code. The second identification rule is used to identify the ICD code and the corresponding disease name; Calculate the similarity between the first disease name and the third disease name as the second similarity. If the second similarity reaches the second similarity threshold, then proceed to the step of obtaining the second disease name.

[0012] In one embodiment, it further includes: in response to the acquired first medical information, performing identification according to a preset second identification rule, where the second identification rule is used to identify the ICD code and the corresponding disease name, taking the identified disease name as the third disease name, and taking the identified ICD code as the second ICD code; If the first similarity reaches the first similarity threshold, then editing the disease code in the information page based on the national standard code corresponding to the second disease name includes: calculating the similarity between the first disease name and the third disease name as the second similarity. If the second similarity reaches the second similarity threshold and the first similarity reaches the first similarity threshold, then editing the disease code in the information page based on the national standard code corresponding to the second disease name.

[0013] In one embodiment, the performing identification according to a preset second identification rule in response to the acquired first medical information includes: In response to the acquired first medical information, identifying the ICD code and the corresponding disease name based on the ICD code recognition model; the ICD code recognition model is a neural network model trained in combination with a loss function.

[0014] In one embodiment, the training method of the ICD code recognition model includes: Collecting multiple medical text data containing the first medical information and performing annotation of the ICD code and the corresponding disease name; Converting the multiple medical text data with the annotated ICD code and the corresponding disease name into a format suitable for input to the Huatuo large model to obtain annotated data; Dividing the annotated data into a training set, a validation set, and a test set; Based on the preset training parameters, using the training set as the input to the Huatuo large model for training, evaluating the performance of the Huatuo large model during training based on the validation set and adjusting the hyperparameters, and evaluating the generalization ability of the Huatuo large model based on the test set, so as to obtain a trained improved Huatuo large model; Taking the trained improved Huatuo large model as the ICD code recognition model.

[0015] In one embodiment, calculating the similarity between the first disease name and the third disease name, and / or calculating the similarity between the first disease name and the second disease name includes: Respectively vectorizing the two disease names for which the similarity needs to be calculated to obtain a first vector and a second vector; Respectively calculating the norm of the first vector and the norm of the second vector; Calculate the product of the magnitudes of the first vector and the second vector as the first product; Calculate the product of the first vector and the second vector as the second product; Calculate the ratio of the second product to the first product. The closer the ratio is to 1, the higher the similarity.

[0016] In a second aspect, in one embodiment, a computer-readable storage medium is provided. A program is stored in the medium, and the program can be loaded and executed by a processor to perform the disease coding automatic conversion method described in any one of the above embodiments.

[0017] The beneficial effects of the present invention are as follows: Based on the solution of the present application, the ICD code is automatically converted into the national standard code to achieve automatic conversion of disease coding, and an information page is edited based on the automatically converted national standard code, thereby improving the filling efficiency of the national standard code and disease name in the information page. Moreover, by calculating the similarity between the first disease name and the second disease name, the accuracy of filling in the national standard code and its corresponding disease name in the information page is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a schematic flowchart of a method for automatically converting disease coding on an outpatient and emergency diagnosis and treatment information page according to an embodiment of the present application; Figure 2 is a schematic flowchart of a method for training a national standard code recognition model according to an embodiment of the present application; Figure 3 is a schematic flowchart of a method for calculating the similarity between two disease names according to an embodiment of the present application; Figure 4 is a schematic flowchart of a method for training an ICD code recognition model according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific embodiments. Similar elements in different embodiments are labeled with related similar element numbers. In the following embodiments, many details are described to make the present application better understood. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, and methods. In some cases, some operations related to the present application are not shown or described in the specification to avoid the core part of the present application being overwhelmed by excessive description. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the description in the specification and the general technical knowledge in the art.

[0020] In addition, the features, operations, or characteristics described in the specification can be combined in any suitable manner to form various embodiments. At the same time, the steps or actions in the method description can also be reordered or adjusted in a manner that is obvious to those skilled in the art. Therefore, the various sequences in the specification and drawings are only for clearly describing a certain embodiment and do not mean a necessary sequence, unless it is stated that a certain sequence must be followed.

[0021] The serial numbers assigned to the components herein, such as "first", "second", etc., are only used to distinguish the objects described and do not have any sequential or technical meaning.

[0022] For the convenience of explaining the inventive concept of this application, the following briefly describes the disease coding conversion technology.

[0023] Currently, in the electronic medical record system edited by doctors, only ICD codes and their corresponding disease names can be accepted. If you want to fill in the national standard codes and their corresponding disease classifications on the information page, the most direct method is for doctors to manually fill them in according to their professional knowledge of disease classification. However, this filling method is inefficient, limited by doctors' professional knowledge of disease classification, prone to filling errors, and requires a large amount of additional manual review, causing a lot of inconvenience and waste of resources.

[0024] In view of this, an embodiment of this application provides a method and medium for automatically converting disease codes on the outpatient and emergency diagnosis and treatment information page. In the solution, the first medical information and the second medical information of the patient are obtained based on the electronic medical record system; among them, the first medical information includes drug information, allergy history, and chief complaint information, and the second medical information includes the first disease name and the first ICD code; the disease name in the second medical information is used as the first disease name, and the ICD code in the second medical information is used as the first ICD code; in response to the obtained first medical information and second medical information, the second disease name is obtained by identifying according to a preset first identification rule; the similarity between the first disease name and the second disease name is calculated as the first similarity. If the first similarity reaches the first similarity threshold, the disease code on the information page is edited based on the national standard code corresponding to the second disease name. In this way, through automatic identification, the efficiency and accuracy of filling in the national standard code and its corresponding disease name on the information page are improved.

[0025] Please refer to Figure 1 , an embodiment of this application provides a method for automatically converting disease codes on the outpatient and emergency diagnosis and treatment information page, including: Step S10, obtain the first medical information and the second medical information of the patient based on the electronic medical record system. Among them, the first medical information includes drug information, allergy history, and chief complaint information, and the second medical information includes disease name and ICD code. Use the disease name in the second medical information as the first disease name, and use the ICD code in the second medical information as the first ICD code.

[0026] The information in the electronic medical record system is edited by doctors. For the first medical information, it is not limited to the drug information used by the patient, the patient's allergy history, and the patient's chief complaint information, and may also include other information recorded by doctors in the electronic medical record that helps to judge the disease name. Since only ICD codes and their corresponding disease names are accepted in the current electronic medical record system, it is impossible to directly extract national standard codes and their corresponding disease names from the electronic medical record.

[0027] Step S20, in response to the obtained first medical information and second medical information, perform identification according to a preset first identification rule, and use the identified disease name as the second disease name. Among them, the first identification rule is used to identify the national standard code and the corresponding disease name.

[0028] The applicant found in the research that although there is a certain correspondence between ICD codes as international standard codes and national standard codes as China's standard codes, this correspondence is not an absolute one-to-one relationship. It is also necessary to further classify according to the specific condition, and there are also some diseases with detailed classifications that do not have ICD codes, and the corresponding disease names are also different. Therefore, it is not realistic to directly convert ICD codes into national standard codes.

[0029] In view of this, comprehensively consider the first medical information and the second medical information to identify the national standard code and the disease name corresponding to the national standard code.

[0030] In one embodiment, in response to the obtained first medical information and second medical information, the national standard code and the corresponding disease name are identified based on the national standard code recognition model. Among them, the national standard code recognition model is a neural network model trained in combination with a loss function.

[0031] Huatuo large model, also known as Huatuo GPT, is an existing open-source medical large model of ChatGPT, which has good natural language understanding ability of medical information. Patients can interact with Huatuo GPT in a dialogue mode to achieve functions such as department triage and auxiliary diagnosis, providing assistance to doctors.

[0032] In one embodiment, based on the Huatuo large model suitable for the medical field, the Huatuo large model is trained and fine-tuned to obtain a national standard code recognition model suitable for identifying the national standard code and the disease name corresponding to the national standard code. Please refer to Figure 2, the training method of the national standard coding recognition model includes: Step S201, collect multiple medical text data containing first medical information and second medical information, and perform national standard coding and annotation of corresponding disease names.

[0033] The medical text data can be extracted from historical electronic medical records. Based on the extracted medical text data, national standard coding and annotation of the disease names corresponding to the national standard coding are performed based on the "National Clinical Version of Disease Classification and Codes" standard.

[0034] The annotation work requires professional medical knowledge and a rigorous attitude to ensure the accuracy and consistency of the annotation. Both the national standard coding that conforms to the "National Clinical Version of Disease Classification and Codes" standard and the consistent corresponding disease names need to be annotated.

[0035] Those skilled in the art can understand that the given medical text data needs to cover all types of national standard coding and disease names, and the medical text data corresponding to each type of national standard coding and disease name should be as much as possible to ensure the accuracy of the trained model.

[0036] Step S202, convert multiple medical text data with annotated national standard coding and corresponding disease names into a format suitable for input to the Huatuo large model to obtain annotated data.

[0037] Those skilled in the art can understand that the text data needs to be converted into a vector representation that the model can understand through word segmentation, word vector conversion, etc.

[0038] Step S203, divide the annotated data into a training set, a validation set, and a test set.

[0039] The training set is used for model training, the validation set is used to evaluate the model's performance and adjust hyperparameters during training, and the test set is used to finally evaluate the model's generalization ability.

[0040] Step S204, based on the preset training parameters, use the training set as the input to the Huatuo large model for training, evaluate the performance of the Huatuo large model and adjust hyperparameters during training based on the validation set, and evaluate the generalization ability of the Huatuo large model based on the test set, so as to obtain a trained and improved Huatuo large model.

[0041] Appropriate training parameters can be set according to the requirements of the model and computing resources, such as learning rate, batch size, number of training epochs, etc. The learning rate determines the step size of model parameter update, the batch size affects the efficiency and stability of model training, and the number of training epochs determines the number of iterations of model training.

[0042] During the training process, according to the output of the model and the labeled true results, calculate the loss function of the model, and use an optimization algorithm to update the parameters of the model to reduce the value of the loss function. In one embodiment, during the training process, the weight file of the model can be saved regularly so that the training can be resumed in case of training interruption or problems.

[0043] In one embodiment, use the validation set to evaluate the trained model, calculate metrics such as the accuracy and recall rate of the model on the validation set, and evaluate the performance of the model for entity extraction. According to the evaluation results, analyze the problems and deficiencies of the model.

[0044] In one embodiment, the model can be optimized according to the results of model evaluation. The training parameters can be adjusted, the amount of training data can be increased, the quality of data annotation can be improved, etc. to improve the performance of the model. If the model has problems of overfitting or underfitting, regularization techniques, increasing data diversity and other methods can be used to solve them.

[0045] Step S205, use the trained improved Huatuo large model as the national standard code recognition model.

[0046] In the embodiment of the present application, based on the Huatuo large model suitable for the medical field, the Huatuo large model is trained and fine-tuned to obtain a national standard code recognition model suitable for extracting national standard codes and corresponding disease names.

[0047] Based on the national standard code recognition model obtained from the above embodiments, the corresponding national standard code and corresponding disease name can be recognized based on the electronic medical record data.

[0048] Step S30, calculate the similarity between the first disease name and the second disease name as the first similarity. If the first similarity reaches the first similarity threshold, then edit the disease code in the information page based on the national standard code corresponding to the second disease name.

[0049] The applicant found in the research that when doctors edit ICD codes in the electronic medical record system, the probability of making mistakes is relatively high. If the ICD codes given by doctors are incorrect, it will lead to incorrect national standard codes and corresponding disease names obtained based on the first medical information and the second medical information.

[0050] In view of this, calculate the similarity between the first disease name and the second disease name as the first similarity, and verify whether the disease code in the information page can be edited based on the national standard code corresponding to the second disease name according to whether the first similarity reaches the similarity threshold.

[0051] In one embodiment, if the first similarity reaches the first similarity threshold, it is considered that the ICD code is correct, and the national standard code and the corresponding disease name given are also correct. Then, the national standard code corresponding to the second disease name is edited as the disease code in the information page. If the first similarity does not reach the first similarity threshold, it is considered that the ICD code is probably incorrect, and the result that the first similarity test fails is returned for manual inspection of the ICD code.

[0052] Through the above embodiment scheme, the ICD code is automatically converted into the national standard code to realize the automatic conversion of the disease code, and the information page is edited based on the automatically converted national standard code, thereby improving the filling efficiency of the national standard code and the disease name in the information page. Moreover, by calculating the similarity between the first disease name and the second disease name, the accuracy of filling in the national standard code and its corresponding disease name in the information page is improved.

[0053] Regarding the incorrect ICD code, currently in each hospital, to avoid this problem, a large amount of manpower needs to be allocated to manually check whether the code is incorrect.

[0054] In view of this, to further avoid incorrect ICD codes, in response to the acquired first medical information, it is identified according to a preset second identification rule, the identified disease name is used as the third disease name, and the identified ICD code is used as the second ICD code. Among them, the second identification rule is used to identify the ICD code and the corresponding disease name. Then, the similarity between the first disease name and the third disease name is calculated as the second similarity, and it is determined whether the second similarity reaches the second similarity threshold.

[0055] In one embodiment, before obtaining the second disease name, the second similarity is calculated first, and it is determined whether the second similarity reaches the second similarity threshold. If the second similarity reaches the second similarity threshold, then the step of obtaining the second disease name is entered. In this way, the correctness of the ICD code in the electronic medical record is verified first. If the first round of verification is correct, then the step of determining whether the first similarity reaches the first similarity threshold is entered, so as to achieve the second round of verification. If both rounds of verification are correct, it means that the ICD code in the electronic medical record is correct. Finally, the disease code in the information page is edited based on the national standard code corresponding to the second disease name.

[0056] If it is found in the first round of verification that the second similarity does not reach the second similarity threshold, it is considered that the ICD code is probably incorrect, and the result that the second similarity test fails is returned for manual inspection of the ICD code.

[0057] In another embodiment, the above first-round verification and second-round verification are not in a specific order. That is, if the second similarity reaches the second similarity threshold and the first similarity reaches the first similarity threshold, then edit the disease code in the information page based on the national standard code corresponding to the second disease name. If either the second similarity threshold or the first similarity threshold does not reach the corresponding preset threshold, it is considered that the ICD code is probably incorrect, and the result of the corresponding similarity test failure is returned for manual inspection of the ICD code.

[0058] In one embodiment, please refer to Figure 3 , calculating the similarity between the first disease name and the third disease name, and / or calculating the similarity between the first disease name and the second disease name may include: Step S100, vectorize the two disease names for which similarity needs to be calculated to obtain a first vector and a second vector.

[0059] Taking the calculation of the similarity between the first disease name and the second disease name as an example, vectorize the first disease name to obtain a first vector A, and vectorize the second disease name to obtain a second vector B.

[0060] In one embodiment, a pre-trained Word2Vec model can be used. Input the disease name into the model to obtain the corresponding word vector. For example, the word vector corresponding to "cold" may be [0.1, 0.2, -0.3,..., 0.4], and the word vector corresponding to "flu" may be [0.2, 0.3, -0.2,..., 0.5], etc. These word vectors can reflect the semantic similarity between disease names. For example, the word vectors of "cold" and "flu" are close in the vector space because they are semantically similar and both belong to respiratory infection diseases.

[0061] Step S200, calculate the norm of the first vector and the norm of the second vector respectively.

[0062] In one embodiment, step S200 can be expressed as: , , where represents the norm of the first vector, represents the norm of the second vector, represents the value of the i-th element in the first vector, represents the value of the i-th element in the second vector, and n represents the total number of elements in the first vector or the second vector, 1 ≤ i ≤ n.

[0063] Those skilled in the art can understand that, in order to compare the similarity between the first vector and the second vector, the forms and the number of elements included in the first vector and the second vector are the same.

[0064] Step S300, calculate the product of the modulus lengths of the first vector and the second vector as the first product.

[0065] Step S400, calculate the product of the first vector and the second vector as the second product.

[0066] Step S500, calculate the ratio of the second product to the first product. The closer the ratio is to 1, the higher the similarity.

[0067] In one embodiment, step S500 can be expressed as: , where, represents the similarity value.

[0068] For the obtained similarity value, when the value is 1, it is considered that the two vectors are exactly the same; when the value is -1, it is considered that the two vectors are exactly opposite; when the value is 0, it indicates that the two vectors are orthogonal. The closer the value is to 1, the more similar the two disease names are.

[0069] Those skilled in the art can understand that the second similarity threshold and the first similarity threshold can be the same or different. Those skilled in the art can understand that the second similarity threshold and the first similarity threshold can be set based on actual requirements. In one embodiment, both the first similarity threshold and the second similarity threshold can be set to 0.95.

[0070] Based on the above embodiment, by calculating the similarity between the first disease name and the third disease name, the accuracy of filling in the national standard code and its corresponding disease name in the information page is further improved.

[0071] In one embodiment, in response to the acquired first medical information, the ICD code and the corresponding disease name are recognized based on the ICD code recognition model. Among them, the ICD code recognition model is a neural network model trained in combination with a loss function.

[0072] In one embodiment, based on the Huatuo large model suitable for the medical field, the Huatuo large model is trained and fine-tuned to obtain a national standard code recognition model suitable for recognizing ICD codes and the disease names corresponding to the ICD codes. Please refer to Figure 4 , and the training method of the ICD code recognition model includes: Step S211, collect multiple medical text data containing the first medical information, and perform annotation of ICD codes and corresponding disease names.

[0073] The medical text data can be extracted from historical electronic medical records. Based on the extracted medical text data and the ICD-10 international standard, ICD coding and annotation of the disease names corresponding to the ICD codes are performed.

[0074] The annotation work requires professional medical knowledge and a rigorous attitude to ensure the accuracy and consistency of the annotation. Both the ICD codes that conform to the ICD-10 international standard and the consistent corresponding disease names need to be annotated.

[0075] As can be understood by those skilled in the art, the provided medical text data needs to cover all types of ICD codes and disease names, and the medical text data corresponding to each type of ICD code and disease name should be as much as possible to ensure the accuracy of the trained model.

[0076] Step S212: Convert multiple medical text data with annotated ICD codes and corresponding disease names into a format suitable for input to the Huatuo large model to obtain annotated data.

[0077] As can be understood by those skilled in the art, the text data needs to be converted into a vector representation that the model can understand through word segmentation, word vector conversion, etc.

[0078] Step S213: Divide the annotated data into a training set, a validation set, and a test set.

[0079] The training set is used for training the model, the validation set is used to evaluate the performance of the model and adjust the hyperparameters during the training process, and the test set is used to finally evaluate the generalization ability of the model.

[0080] Step S214: Based on the preset training parameters, use the training set as the input to the Huatuo large model for training, evaluate the performance of the Huatuo large model and adjust the hyperparameters based on the validation set during the training process, and evaluate the generalization ability of the Huatuo large model based on the test set, so as to obtain a trained and improved Huatuo large model.

[0081] Step S215: Use the trained and improved Huatuo large model as an ICD code recognition model.

[0082] In the embodiment of the present application, based on the Huatuo large model suitable for the medical field, the Huatuo large model is trained and fine-tuned to obtain an ICD code recognition model suitable for extracting ICD codes and corresponding disease names.

[0083] Based on the ICD code recognition model obtained from the above embodiment, the corresponding ICD codes and corresponding disease names can be recognized based on the electronic medical record data.

[0084] In an embodiment of the present application, a computer-readable storage medium is provided. A program is stored on the storage medium, and the stored program includes a method that can be loaded and processed by a processor to implement any of the above embodiments.

[0085] Those skilled in the art can understand that all or part of the functions of the above methods can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium may include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions are implemented by a computer executing the program. For example, the program is stored in the memory of the device, and when the processor executes the program in the memory, all or part of the above functions can be implemented. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in a storage medium such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, and saved to the memory of the local device by downloading or copying, or the system of the local device is updated. When the processor executes the program in the memory, all or part of the functions in the above embodiments can be implemented.

[0086] The above uses specific examples to elaborate on the present invention, which is only used to help understand the present invention and is not intended to limit the present invention. For those skilled in the art of the present invention, based on the idea of the present invention, several simple deductions, deformations or substitutions can be made.

Claims

1. A method for automatically converting ICD codes of outpatient and emergency treatment information pages, characterized in that: include: Acquiring first medical information and second medical information of the patient based on the electronic medical record system; The first medical information includes drug information, allergy history and chief complaint information, and the second medical information includes disease name and ICD code; the disease name in the second medical information is used as the first disease name, and the ICD code in the second medical information is used as the first ICD code; In response to the acquired first medical information and second medical information, identification is performed according to a preset first identification rule, and the identified disease name is used as the second disease name, wherein the first identification rule is used to identify the national standard code and the corresponding disease name; Calculate the similarity between the first disease name and the second disease name as the first similarity, and if the first similarity reaches the first similarity threshold, edit the disease code in the information page based on the national standard code corresponding to the second disease name; The method also includes: before obtaining the second disease name, it also includes: in response to the obtained first medical information, identifying according to a preset second identification rule, using the identified disease name as the third disease name, and using the identified ICD code as the second ICD code, the second identification rule is used to identify the ICD code and the corresponding disease name; calculating the similarity between the first disease name and the third disease name as the second similarity, if the second similarity reaches the second similarity threshold, then entering the step of obtaining the second disease name.

2. The ICD code automatic conversion method according to claim 1, characterized in that: The step of identifying the acquired first medical information and second medical information according to a preset first identification rule includes: In response to the acquired first medical information and second medical information, the national standard code and the corresponding disease name are identified based on the national standard code recognition model; the national standard code recognition model is a neural network model trained in combination with a loss function.

3. The ICD code automatic conversion method according to claim 1, characterized in that: The training method of the national standard coding recognition model includes: Collect multiple medical text data including first medical information and second medical information, and mark them with national standard codes and corresponding disease names; Convert multiple medical text data with national standard codes and corresponding disease names into a format suitable for the input of the Hua Tuo model to obtain labeled data; Dividing the labeled data into a training set, a validation set, and a test set; Based on the preset training parameters, the training set is used as the input of the Hua Tuo model for training. The performance of the Hua Tuo model is evaluated and the hyperparameters are adjusted based on the validation set during the training process. The generalization ability of the Hua Tuo model is evaluated based on the test set, thereby obtaining a trained and improved Hua Tuo model. The trained improved Hua Tuo model is used as the national standard coding recognition model.

4. The ICD code automatic conversion method according to claim 1, characterized in that: If the first similarity reaches the first similarity threshold, the disease code in the information page is edited based on the national standard code corresponding to the second disease name, including: If the first similarity reaches a first similarity threshold, the national standard code corresponding to the second disease name is edited as the disease code in the information page.

5. The ICD code automatic conversion method according to claim 1, characterized in that: The step of identifying the acquired first medical information according to a preset second identification rule includes: In response to the acquired first medical information, the ICD code and the corresponding disease name are identified based on an ICD code recognition model; the ICD code recognition model is a neural network model trained in combination with a loss function.

6. The ICD code automatic conversion method according to claim 5, characterized in that: The training method of the ICD code recognition model includes: Collect multiple medical text data containing first medical information, and annotate ICD codes and corresponding disease names; Convert multiple medical text data with ICD codes and corresponding disease names annotated into a format suitable for the Hua Tuo model input to obtain annotated data; Dividing the labeled data into a training set, a validation set, and a test set; Based on the preset training parameters, the training set is used as the input of the Hua Tuo model for training. The performance of the Hua Tuo model is evaluated and the hyperparameters are adjusted based on the validation set during the training process. The generalization ability of the Hua Tuo model is evaluated based on the test set, thereby obtaining a trained and improved Hua Tuo model. The trained improved Hua Tuo model is used as the ICD coding recognition model.

7. The ICD code automatic conversion method according to claim 1, characterized in that: Calculating the similarity between the first disease name and the third disease name, and / or calculating the similarity between the first disease name and the second disease name, including: Vectorize the two disease names whose similarity needs to be calculated respectively to obtain a first vector and a second vector; Calculate the modulus of the first vector and the modulus of the second vector respectively; Calculate the product of the modulus of the first vector and the modulus of the second vector as a first product; calculating a product of the first vector and the second vector as a second product; The ratio of the second product to the first product is calculated. The closer the ratio is to 1, the higher the similarity.

8. A computer-readable storage medium, characterized in that: The medium stores a program, which can be loaded by a processor and execute the ICD code automatic conversion method as described in any one of claims 1 to 7.