A method for reconstructing a disease course from insurance claims

By cleaning insurance claim documents and establishing a correlation judgment model, the problem that disease grouping methods fail to consider the disease treatment process is solved, enabling accurate cost estimation of disease course, reducing distortion of treatment costs, and improving the accuracy of cost prediction.

CN115456800BActive Publication Date: 2026-05-12SHANGHAI SHANGYONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI SHANGYONG TECH CO LTD
Filing Date
2022-09-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing disease grouping methods fail to consider the development and evolution of diseases during treatment, resulting in an inability to accurately estimate the cost of the entire course of a disease when assessing the disease risk and cost risk of insured individuals. This leads to deviations in treatment costs and economic losses for insurance companies.

Method used

By cleaning and standardizing insurance claim documents, a model for judging the correlation between claim documents and diagnoses is established. Machine learning is used to reconstruct the disease course, including initializing the course of the disease, allocating unallocated records, merging related courses of the disease, and iteratively executing until the patient's course of the disease and cost information are output.

Benefits of technology

It achieves accurate cost estimation throughout the entire course of a disease, reduces distortion of treatment costs, improves the accuracy of cost prediction, captures misallocated costs, and reduces economic losses for insurance companies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115456800B_ABST
    Figure CN115456800B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of medical management, and provides a method for restoring a disease course through insurance claim documents, which comprises the following steps: according to a disease insurance claim document of an insured person, cleaning and standardizing claim record data, taking the claim record as a training sample, establishing a diagnosis correlation judgment model of the claim document through machine learning, taking a disease group corresponding to a diagnosis name in the claim document as a seed event, initializing a disease course through the seed event, assigning a claim record without a disease course in the initialization process to a disease course with the highest correlation, correlating disease courses in different disease groups according to the correlation judgment model of the claim document and the diagnosis, merging the disease courses in the same disease group, iteratively executing the above steps until the disease course grouping is unchanged, and outputting patient disease course cost information. The application can estimate the cost of the whole disease course of a patient, and avoid losses caused by the deviation of actual treatment cost estimation in insurance claim.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical management technology, and in particular to a method for reconstructing the course of a disease using insurance claim documents. Background Technology

[0002] Existing disease grouping methods are mainly based on disease diagnosis or treatment procedures during a single visit. Representative methods include DRG (Diagnosis Related Groups) and APG (Outpatient Grouping by Capacity and Outpatient Cases).

[0003] The aforementioned disease grouping methods focus on a single treatment session and do not consider the development and evolution of the disease during treatment. When assessing the patient's disease risk and cost risk, they cannot estimate the cost of the entire disease course, easily leading to a shift of treatment costs before / after hospitalization, resulting in deviations in actual treatment costs. For example, suppose someone is hospitalized for appendicitis and then hospitalized again at another hospital for recovery due to surgical wound infection. Grouping methods such as DRG and APG can only capture the individual cost and risk information of the two hospitalizations, failing to indicate that the post-discharge wound infection hospitalization is a sequela of appendicitis surgery. Therefore, this results in an estimation bias of the appendicitis treatment cost and the insured person's health status.

[0004] In insurance claims, the current disease grouping methods have errors in estimating and calculating the actual treatment costs for insured individuals, resulting in significant economic losses for insurance companies. Summary of the Invention

[0005] This invention primarily addresses the technical problem of existing technologies failing to consider the development and evolution of disease treatment processes, thus hindering the estimation of costs throughout the entire course of a disease when assessing the disease risk and cost risk of insured individuals. It proposes a method to reconstruct the disease course using insurance claim documents, thereby enabling the estimation of costs for the entire course of the patient's disease based on the information provided in the insurance claim documents.

[0006] This invention provides a method for reconstructing the course of a disease using insurance claim documents, comprising:

[0007] S1. Clean and standardize the claim record data based on the insured's disease insurance claim documents;

[0008] S2. Using the cleaned and standardized claims records as training samples, a model for judging the correlation between claims documents and diagnoses is established through machine learning.

[0009] S3. Group the diseases corresponding to the diagnosis names in the claim documents as seed events, and initialize the disease course through the seed events;

[0010] S4. Assign the claim records that were not assigned to a disease course during the initialization of the disease course to the disease course with the highest relevance.

[0011] S5. Based on the correlation judgment model between the claim documents and the diagnosis, link the course of disease in different disease groups and merge the course of disease in the same disease group;

[0012] S6. Iterate through steps S4 and S5 until there are no changes in the disease course groups, then output the patient's disease course cost information.

[0013] Further, step S1 includes: sorting the claims records by time; mapping the diagnostic data in the claims records to ICD diagnostic codes, and assigning the ICD diagnostic codes to corresponding disease groups according to preset ICD diagnostic coding rules; and mapping the drugs, medical consumables, and medical services in the claims records to standard drug catalogs, consumable catalogs, and medical service catalogs, respectively.

[0014] Further, step S2 includes: using claim documents with the same invoice number and the same diagnosis as training samples, using the disease classification corresponding to the diagnosis information as the text topic, and using the items involved in the claim documents as the content in the text; the items include: medicines, consumables, and medical services; using the training samples to train a latent Dirichlet distribution model to form a model for judging the correlation between claim documents and diagnoses.

[0015] Further, step S3 includes: defining the items and diagnostic information associated with the same invoice number and its diagnosis as a disease course, wherein the disease course includes: disease group information, item information and start and end time information.

[0016] Furthermore, step S3 also includes: symptom grouping in the disease group is not used as a seed event to participate in the initialization of the disease course; the symptom grouping includes: symptoms, signs, clinical and laboratory abnormalities, and factors that affect health status and contact with healthcare institutions.

[0017] Further, step S4 includes: sequentially searching for claim records not included in the medical course in ascending chronological order; determining the relationship between the claim record or record group and the medical course, and processing the claim records not included in the medical course according to the relationship; and applying the processed claim record information to update the start time and duration of the medical course.

[0018] Furthermore, determining the relationship between the claim record or record group and the course of illness, and processing the claim records not included in the course of illness based on the relationship, includes: querying the course number of claim records that share an invoice number with the claim record or record group; if a result is found, assigning the claim record to this course number; if no result is found, searching with the claim record as the center by setting a time window; using the claim document and diagnosis correlation judgment model to determine whether the topic corresponding to the searched record is consistent with the course of illness; if consistent, assigning the record to the course number of the course of illness; if inconsistent, determining whether the record is consistent with other courses of illness.

[0019] Furthermore, the time window setting is achieved by querying a pre-established knowledge base and setting different search time windows for different items or disease course groups according to the theoretical incubation period and disease course information.

[0020] This invention provides a method for reconstructing the course of a disease using insurance claim documents. By utilizing the information provided in the claim documents, the method reconstructs the course of the insured's disease, more accurately reconstructing the costs required for the entire treatment process, and improving the accuracy and objectivity of cost prediction. For example, after the implementation of DRG and other bundled payment policies for hospitalization, medical institutions are incentivized to shift some costs before and after hospitalization to circumvent the constraints of bundled payment, causing distortions in hospitalization costs and resulting in significant deviations in cost prediction for insurance risk control. By applying the method of reconstructing the treatment process according to the course of the disease, this shifted cost can be captured, enabling a more accurate estimate of the cost and duration of disease treatment. Attached Figure Description

[0021] Figure 1 This is a flowchart of the method for reconstructing the course of a disease using insurance claim documents, as per the present invention.

[0022] Figure 2 This is a technical rendering of the invention, which uses insurance claim documents to reconstruct the course of a disease.

[0023] Figure 3 This is a logic diagram of the method for reconstructing the course of a disease using insurance claim documents, as described in this invention.

[0024] Figure 4 This is a schematic diagram illustrating the disease progression relationship between different disease groups in the embodiment. Detailed Implementation

[0025] To make the technical problems solved by this invention, the technical solutions adopted, and the technical effects achieved clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings, not all of them.

[0026] like Figure 1 , Figure 3 As shown, the method for reconstructing the course of a disease from insurance claim documents provided in this embodiment of the invention includes:

[0027] S1. Clean and standardize the claim record data based on the insured's disease insurance claim documents;

[0028] Specifically, data cleaning and preparation, including claims record data cleaning and standardization, includes the following steps:

[0029] S11. As shown in Table 1, a row of records is called a claim record. A single insured person is a record of the same individual. Claim records are sorted in ascending order by date and time.

[0030] Table 1

[0031] Invoice Number Date and Time Project Name Project Amount Main diagnosis name Sub-diagnosis name Diagnostic doctor Source of documents D0001 2019-08-20 Quick-acting heart-saving pills 300 / / / D Pharmacy A0002 2020-01-01 ambulance fee 100 Chest pain, cause to be determined / Zhao Hospital A A0002 2020-01-01 Electrocardiogram A 20 Chest pain, cause to be determined Zhao Hospital A A0002 2020-01-01 Myocardial enzyme 3 200 Chest pain, cause to be determined Zhao Hospital A A0002 2020-01-02 nitroglycerin 200 angina pectoris hypertension money Hospital A B0001 2020-02-02 nitroglycerin 200 / / / B Pharmacy C0001 2020-03-01 nitroglycerin 200 Myocardial infarction hypertension Sun C Hospital C0001 2020-03-01 Stent implantation 8000 Myocardial infarction hypertension Sun C Hospital C0001 2020-03-07 Inpatient nursing 2000 Myocardial infarction hypertension Sun C Hospital C0002 2020-04-01 ambulance fee 100 Difficulty breathing / plum C Hospital C0002 2020-04-01 <![CDATA[ Furosemium Semi]]> 20 Acute left heart failure Stent implantation status plum C Hospital C0002 2020-04-01 Uradil 800 Acute left heart failure Stent implantation status plum C Hospital

[0032] S12. Map the diagnoses in the claims record to ICD diagnostic codes, and further map the diagnoses in the claims record to disease groups using preset rules for mapping ICD diagnostic codes to disease groups. Examples of predicted disease classification rules are shown in Table 2, which are mainly divided according to the dimensions of disease etiology, location, and treatment cost, and are generally set to about 500-600 groups.

[0033] Table 2

[0034] ICD encoding ICD Name Disease grouping Disease group name I21 Acute myocardial infarction HRT003 Myocardial infarction I21.0 Acute transmural myocardial infarction of the anterior wall HRT003 Myocardial infarction I21.000 Acute transmural myocardial infarction of the anterior wall HRT003 Myocardial infarction I21.000x005 Acute anterior apical myocardial infarction HRT003 Myocardial infarction I21.001 Acute anterior myocardial infarction HRT003 Myocardial infarction I21.002 Acute anterior wall myocardial infarction HRT003 Myocardial infarction

[0035] S13. If the same claim record involves primary and secondary diagnoses, both primary and secondary diagnoses must be cleaned and mapped to the corresponding disease categories.

[0036] S14. Map the drugs, medical consumables, and medical services in the claims records to the standard drug catalog, consumable catalog, and medical service catalog, respectively (the catalogs here only need to be standard, including but not limited to the catalog standards issued by the National Healthcare Security Administration). The cleaned claims document is shown in Table 3:

[0037] Table 3

[0038]

[0039] S2. Using the cleaned and standardized claims records as training samples, a model for judging the correlation between claims documents and diagnoses is established through machine learning.

[0040] Specifically, S21 uses claim documents with the same invoice number and diagnosis as training samples, takes the disease category corresponding to the diagnosis information as the text topic, and takes items such as medicines, consumables, and medical services mentioned in the documents as "words" in the text. See Table 4 for a sample of the corpus processing:

[0041] Table 4

[0042]

[0043] In Table 4, the claim documents can be organized as follows: {Subject: Myocardial infarction, Content: [Nitroglycerin, Stent implantation, Hospitalization care]}.

[0044] S22. Train the Latent Dirichlet Allocation (LDA) model to form a correlation model between claim documents and diagnoses, which will be used for subsequent correlation calculations between claim documents and diagnoses.

[0045] S23. In practical applications, the model selection here is not limited to LDA. Other deep learning models can also be selected. The main purpose is to build a topic recognition model that can take document items as input and get disease classification as output.

[0046] S24. This model for processing claims documents and the correlation model with diagnosis and the subsequent course of illness are two separate models and do not need to be run in sequence. This model can be trained independently.

[0047] S3. Use the disease group corresponding to the diagnosis name in the claim document as a seed event, and initialize the disease process through the seed event;

[0048] Specifically, S31, claims items and diagnoses associated with the same invoice number are grouped together to define a course of illness. The course of illness mainly includes disease group information, item information, and start and end time information. For example, a course of illness can be initially established through the myocardial infarction disease group appearing in invoice number c0001 in Table 3. The course of illness content is {myocardial infarction: [nitroglycerin, stent implantation, hospitalization care], course duration: {start: 2020-03-01, end: 2020-03-07}, duration 8 days}.

[0049] S32. Symptom-related groups within disease groups (corresponding to the ICD chapter: Symptoms, Signs and Clinical and Laboratory Abnormalities) and (corresponding to the ICD chapter: Factors Affecting Health Status and Contact with Healthcare Institutions) are not used as seed events to initiate the establishment of disease courses. For example, the first three items of A0002 in Table 3 above are not used to establish disease courses because the disease group is a symptom group. The disease course corresponding to A0002 is only {Coronary Artery Disease: [Nitroglycerin], course duration: {Start: 2020-01-02, End: 2020-01-02}, duration 1 day}.

[0050] S33. Taking Table 3 as an example, the final generated medical records and the usage of claim documents are shown in Table 5:

[0051] Table 5

[0052] Invoice Number Date and Time Project Name Project Amount …… Disease grouping Group Name Is it a seed? Disease progress number D0001 2019-08-20 Quick-acting heart-saving pills 300 …… / no A0002 2020-01-01 ambulance fee 100 …… SYM002 Chest pain symptoms no A0002 2020-01-01 electrocardiogram 20 …… SYM002 Chest pain symptoms no A0002 2020-01-01 Myocardial enzymes 200 …… SYM002 Chest pain symptoms no A0002 2020-01-02 nitroglycerin 200 …… HRT010 Coronary heart disease yes HRT010-1 B0001 2020-02-02 Nitroglycerin 200 …… / no C0001 2020-03-01 Nitroglycerin 200 …… HRT003 Myocardial infarction yes HRT003-1 C0001 2020-03-01 Stent implantation 8000 …… HRT003 Myocardial infarction yes HRT003-1 C0001 2020-03-07 Inpatient nursing 2000 …… HRT003 Myocardial infarction yes HRT003-1 C0002 2020-04-01 ambulance fee 100 …… SYM003 respiratory symptoms no C0002 2020-04-01 <![CDATA[ Furosemium Semi]]> 20 …… HRT009 Heart failure yes HRT009-1 C0002 2020-04-01 Uradil 800 …… HRT009 Heart failure yes HRT009-1

[0053] The preliminary disease course information is shown in Table 6. In actual implementation, the disease course data structure is based on linked lists, with the disease course information for different disease groups forming a single linked list.

[0054] Table 6

[0055]

[0056] S4. Assign claims records that were not assigned to a specific disease course during the initialization of the disease course to the disease course with the highest relevance;

[0057] Specifically, S41, search for claims records not included in the medical record in ascending chronological order, such as the first record in Table 6 with invoice number D0001. If there are a series of records with the same invoice number, the items with the same invoice number will be grouped together for subsequent allocation. For example, the three records with invoice number A0002 in Table 6 will be grouped together.

[0058] S42. To determine the relationship between a claim record or record group and the medical history, first query the medical history number of the claim record that shares the same invoice number. If a result is found, assign the claim record to this medical history number. For example, record A0002 can be assigned to HRT010-1 using its invoice number. If no medical history information for the same invoice number is found, search forward and backward around that record. The default search time window is one month. For example, for record D0001, the search time is from July 20, 2019 to September 20, 2019. In this case, no match was found. However, if the time window is adjusted to six months, the disease course HRT010-1 can also be found as a candidate disease course. After finding a candidate disease course, we apply the topic judgment model trained in step 2 above to determine whether the topic corresponding to record D0001 (or record group) is consistent with HRT010-1. If they are consistent, the record is assigned to the consistent disease course. If they are inconsistent, the next candidate disease course is judged. The candidate diseases are arranged in ascending order of the time interval with the record to be assigned, that is, those that are close in time are given priority for assignment. The above time window settings can also be set by querying a pre-established knowledge base and setting different search time windows for different items or disease course groups according to the theoretical incubation period, disease course, and other information of the disease. In addition, in order to speed up the calculation process and reduce repeated matching judgments, during the search for candidate diseases, this solution will first judge the majority of item types in the claims record or record group. For example, if the items are mainly drugs and medical services, the search will be prioritized forward; if the items are mainly examinations and tests, the search will be prioritized backward.

[0059] S43. Apply the allocated claims record information to update the start time and duration of the illness. The final result is shown in Table 7. The bold text indicates the allocation and update results.

[0060] Table 7

[0061]

[0062] The updated course of illness information is shown in Table 8:

[0063] Table 8

[0064]

[0065] Additional explanation: A disease group can contain multiple disease courses (the example here is unrelated to the data in the table above). For instance, assuming the patient subsequently experiences a myocardial infarction in May, the course of the myocardial infarction could be:

[0066]

[0067]

[0068] S5. Based on the correlation judgment model between claim documents and diagnoses, link the course of disease in different disease groups and merge the course of disease in the same disease group;

[0069] Specifically, step S51 determines whether there is an overlap (inclusion) relationship between the disease courses of different disease groups. Taking the disease course information obtained in the previous step as an example, firstly, an observation time window is added before and after the start and end times of each disease course. The default time window is 30 days for non-chronic diseases and 90 days for chronic diseases. Different observation time windows can also be set for different disease groups by querying the preset knowledge base. Here, the example uses the default time window. For example, the "start" time of the "myocardial infarction" group is updated from "2020-03-01" to "2020-01-31", and the "end" time is updated from "2020-03-07" to "2020-04-06". Similarly, the start-end time of the "heart failure" group is updated to "2020-03-02" and "2020-05-01". At this time, it is easy to find that there is an overlap in time between the two groups.

[0070] S52. Determine whether overlapping disease courses need to be merged: Group diseases according to the chronological order of their occurrence, then search the preset disease-disease relationship table. If complications or disease progression relationships exist, merge the two overlapping disease courses into the disease course group with the earlier start time. The preset disease-disease relationships are shown in Table 9:

[0071] Table 9

[0072] First disease group A The second disease group B B is a complication of A. B represents the progress of A. INF003 RSP011 yes no EXT015 BLD003 no yes HRT003 HRT009 yes no

[0073] S53. Determine if there is any overlap (inclusion) between disease courses within disease groups. Add observation windows before and after the course of disease within each disease group, with a default value of 15 days (2 weeks). Alternatively, a lookup table can be created to set different observation windows for different disease groups, and then the course of disease within the same disease group can be determined from beginning to end to determine if there is a merging relationship. Figure 4 As shown, after merging the disease courses between disease groups in step 5, the disease course A-1 of disease group A was extended from the original 5 weeks to 7 weeks. At this point, there is an overlap after the observation windows of both ends of disease course A-1 and disease course A-2 are extended by 2 weeks. Therefore, we extend disease course A-1 to the end time of disease course A-2.

[0074] S54. Update the disease progress information, such as deleting disease progress B-1 and disease progress A-2 as described above, and extending disease progress A-1 to the end of disease progress A-2.

[0075] S6. Iterate through steps S4 and S5 until there are no changes in the disease course groups, then output the patient's disease course cost information.

[0076] Specifically, after the above steps, due to the change in the length of the disease course, the disease course that could not be searched in the observation window in the original step S4 is moved to the observation window range. Therefore, this step makes the most of the claims records by repeating steps S4 and S5.

[0077] In addition to using default settings and pre-set query tables during initialization, the time window parameters used in the allocation, association, and merging of the disease groups mentioned in the preceding steps can also be automatically updated through model iteration.

[0078] One-third of the mean length of disease course for different patients within each disease group is used as the search time window parameter for each disease group. The new parameters are then applied to reconstruct the disease course and update the search time window parameters according to steps S4-S6.

[0079] Knowing the patient's demographic parameters such as age and gender, the length of each disease course in different disease groups for different patients can be used as the dependent variable, while the patient's demographic parameters and the existence of other disease groups excluding their current disease group can be used as independent variables. By establishing a linear regression or generalized linear model, different search time windows can be provided for different disease groups in different individual claims records in a more personalized way. For example, for patients in the technical effects section below, taking the first disease group "varicose veins of the lower extremities" as an example, there are two disease courses of 11 months and 4 months respectively. We can establish the following linear regression to make recommendations for future search time windows.

[0080] Simulation experiment:

[0081] Distributed Parallelism: Since this method reconstructs an individual's medical history based on their claims records, the operation of different individuals is independent. In practice, the data can be split according to the individual as the data dimension, and the data of different individuals can be distributed to different computers for parallel processing.

[0082] Technical effects such as Figure 2 As shown, Figure 2 Each row represents a disease group, and the corresponding color bar within that group indicates the time elapsed for that group's disease course. Disease recurrence and treatment are reflected in multiple disease courses within the same group. Based on this reconstructed data, we can more accurately obtain information on disease treatment costs and recurrence, as well as evaluate the patient's health status at each time cross-section, such as in... Figure 2 In September 2017, the patient was suffering from eight diseases simultaneously and was in a poor condition, while in August 2018, the patient was in a relatively better condition. Furthermore, for chronic diseases such as hypertension, this data processing makes it easy to observe the patient's treatment status. For example, the patient experienced a treatment interruption in August 2018, indicating the need for active intervention.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications to the technical solutions described in the foregoing embodiments, or equivalent substitutions for some or all of the technical features, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for reconstructing the course of a disease from insurance claim documents, characterized in that, The method includes: S1. Clean and standardize the claim record data based on the insured's disease insurance claim documents; S2. Using the cleaned and standardized claims records as training samples, a model for judging the correlation between claims documents and diagnoses is established through machine learning. This includes: using claims documents with the same invoice number and the same diagnosis as training samples; using the disease classification corresponding to the diagnosis information as the text topic; and using the items involved in the claims documents as the content of the text. The items include: medicines, consumables, and medical services. The latent Dirichlet distribution model is trained using the training samples to form a model for judging the correlation between claims documents and diagnoses. S3. The disease group corresponding to the diagnosis name in the claim document is used as a seed event, and the disease course is initialized through the seed event, including: defining the items and diagnostic information associated with the same invoice number and its diagnosis as a disease course, the disease course including: disease group information, item information and start and end time information; the symptom group in the disease group is not used as a seed event to participate in the initialization of the disease course; the symptom group includes: symptoms, signs, clinical and laboratory abnormalities, and factors that affect health status and contact with healthcare institutions; S4. Assigning the claim records that were not assigned to a disease course during the initialization of the disease course to the disease course with the highest relevance, including: sequentially searching for claim records that were not included in the disease course in ascending order of time; determining the relationship between the claim record or record group and the disease course, and processing the claim records that were not included in the disease course according to the relationship; and applying the processed claim record information to update the start time and duration of the disease course. The process of determining the relationship between a claim record or record group and a medical course, and processing claim records not included in a medical course based on the relationship, includes: querying the medical course number of a claim record that shares an invoice number with the claim record or record group; if a result is found, assigning the claim record to this medical course number; if no result is found, searching using a time window centered on the claim record; and using the claim document and diagnosis correlation judgment model to determine whether the topic corresponding to the searched record is consistent with the medical course. If consistent, assigning the record to the medical course number of the medical course; if inconsistent, determining whether the record is consistent with other medical courses. S5. Based on the correlation judgment model between the claim documents and the diagnosis, link the course of disease in different disease groups and merge the course of disease in the same disease group; S6. Iterate through steps S4 and S5 until there are no changes in the disease course groups, then output the patient's disease course cost information.

2. The method for reconstructing the course of a disease from insurance claim documents according to claim 1, characterized in that, Step S1 includes: Sort the claims records by time; The diagnostic data in the claims records is mapped to ICD diagnostic codes, and the ICD diagnostic codes are assigned to the corresponding disease groups according to the preset ICD diagnostic coding rules. The claims records for medicines, medical consumables, and medical services are mapped to the standard drug catalog, consumable catalog, and medical service catalog, respectively.

3. The method for reconstructing the course of a disease from insurance claim documents according to claim 2, characterized in that, The time window is set by querying a pre-established knowledge base and setting different search time windows for different items or disease course groups according to the theoretical incubation period and disease course information.