Electronic medical record generation method and system based on BART framework
Through the pre-trained model based on the BART framework, the age and time difference at the time of medical treatment were introduced, which solved the problem of insufficient time information processing, generation ability and interpretability of the Med-BERT model, and achieved stronger time pattern learning and task generation capabilities, and improved the accuracy and credibility of the model in the generation of electronic medical records.
Patent Information
- Application Number
- CN202511061387.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-08-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the Med-BERT model has shortcomings in processing time information, generation ability, interpretability and data set limitations, which affects its application effect in clinical scenarios.
A pre-trained model based on the BART framework is adopted to introduce age and time difference at the time of treatment, and a logically coherent electronic medical record is generated through generative capabilities, the data set type is extended, and time information is explicitly captured.
The model's learning ability of time mode is improved, the ability and interpretability of generating tasks is enhanced, the flexibility and adaptability of the model is improved, and the generalization ability in different clinical scenarios is enhanced.
Smart Images

Figure CN120564941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic medical records, and in particular to a method and system for generating electronic medical records based on a BART framework. Background Art
[0002] In recent years, with the widespread use of electronic health record (EHR) data, deep learning (DL)-based predictive models have made significant progress in clinical applications. By analyzing large amounts of patient data, these models can achieve early disease prediction, assist in diagnosis, and recommend treatment options. Med-BERT, pre-trained on large-scale EHR datasets, can be used for disease prediction tasks. Med-BERT has demonstrated excellent accuracy in predicting heart failure and pancreatic cancer incidence in diabetic patients.
[0003] Although Med-BERT has achieved significant performance improvements in disease prediction tasks, it still has some limitations and shortcomings: 1. Lack of temporal information: While the Med-BERT model can process structured EHR data, it does not explicitly account for the time difference between two visits. In EHR data, time difference is a crucial clinical metric. For example, the time between two visits can influence disease progression and prognosis. Failure to account for time difference may lead to insufficient model learning of time-related patterns, such as rapid disease progression or long-term stability, thus affecting the model's predictive accuracy and generalization ability.
[0004] 2. Limited generative capabilities: The Med-BERT model is primarily an encoder model based on the Transformer architecture, primarily used for classification tasks such as disease prediction. However, in clinical scenarios, in addition to classification tasks, many other types of tasks exist, such as generating patient medical record summaries, generating possible diagnostic recommendations, and generating subsequent treatment plans. Med-BERT's insufficient performance in generative tasks limits its potential for broader clinical applications. Furthermore, generative tasks typically require models to possess stronger reasoning capabilities and interpretability, and Med-BERT is relatively weak in these areas.
[0005] 3. Insufficient model interpretability: Although Med-BERT learns contextual information through pre-training tasks, its output is primarily a classification result, lacking a detailed representation of the reasoning process. This makes it difficult for clinicians to understand the model's decision-making basis, reducing the model's interpretability and credibility. This lack of interpretability may cause clinicians to be cautious about using the model, thus limiting its application in real-world clinical settings.
[0006] 4. Dataset Limitations: The Med-BERT model's pre-training dataset primarily comes from Cerner HealthFacts®. While large in scale, the data source is relatively limited. Furthermore, the pre-training tasks primarily focus on diagnosis codes and length of stay prediction, lacking the use of other types of data, such as laboratory test results. These dataset limitations may restrict the model's generalization capabilities across diverse clinical scenarios, particularly when processing multimodal data. Summary of the Invention
[0007] The present invention aims to address the aforementioned issues. To this end, the present invention provides a method and system for generating electronic medical records based on the BART framework. This method utilizes a pre-trained model based on the BART framework and incorporates age at appointment and time difference between appointments, enabling the model to better capture temporal information, thereby improving the accuracy and relevance of generated medical records. Furthermore, through generative capabilities, the model can generate future electronic medical records based on multiple existing electronic medical records. The records are logically coherent, helping clinicians understand and trust the model's predictions.
[0008] The present invention provides a method for generating electronic medical records based on the BART framework, and the technical solution adopted is as follows: comprising the following steps: S1: Obtain multiple electronic medical records and extract age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information from the electronic medical records; S2: Encode age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information as input sequences; S3: Input the input sequence into a pre-trained model based on the BART framework to generate a future electronic medical record; the pre-trained model based on the BART framework includes an encoder and a decoder.
[0009] Furthermore, the age at the time of consultation is calculated in months.
[0010] Furthermore, step S2 includes: Age at presentation was coded as a positional embedding; Generate a visit ID based on the diagnosis date and then encode it into a visit embedding; The visit time difference was generated based on the diagnosis date and then embedded as a marker along with the visit type, diagnosis code, abnormal laboratory test result code, and surgery information code; Concatenate the token embedding, visit embedding, and location embedding into the input sequence.
[0011] Furthermore, the encoding process of the tag embedding is: The diagnosis codes in each electronic medical record were sorted according to their priority, and then concatenated after the visit type, followed by the abnormal laboratory test result code, surgical information, and visit time difference; Using the dictionary, the concatenation results are encoded into token embeddings.
[0012] Furthermore, starting from the second electronic medical record, the time difference between it and the previous electronic medical record is calculated based on the diagnosis date, and mapped according to the following rules to obtain the consultation time difference: Less than 30 days: Gap_1, Greater than or equal to 30 days and less than 60 days: Gap_2, Greater than or equal to 60 days and less than 180 days: Gap_3, Greater than or equal to 180 days and less than 360 days: Gap_4, Greater than or equal to 360 days: Gap_5.
[0013] Furthermore, the plurality of electronic medical records are at least three electronic medical records within one year; The future electronic medical records are electronic medical records within the next year.
[0014] Furthermore, the content of the future electronic medical record includes diagnosis date range, diagnosis code, abnormal laboratory test result code and surgery information.
[0015] Furthermore, the future electronic medical records are electronic medical records of one or more of Parkinson's disease, diabetes, chronic kidney disease, hypertension, cerebrovascular disease, cognitive impairment diseases, and asthma.
[0016] Furthermore, the encoder adopts a bidirectional Transformer structure, which includes 6 Transformer layers; the decoder adopts an autoregressive Transformer structure, which includes 6 Transformer layers.
[0017] The present invention also provides an electronic medical record generation system based on the BART framework, which adopts the following technical solutions: including: an electronic medical record acquisition module, a sequence encoding module and an electronic medical record generation module, The electronic medical record acquisition module is used to acquire multiple electronic medical records and extract the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code and surgery information from the electronic medical records; The sequence encoding module is used to encode the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code and surgical information into an input sequence; The electronic medical record generation module is used to input the input sequence into a pre-trained model based on the BART framework to generate future electronic medical records; the pre-trained model based on the BART framework includes an encoder and a decoder.
[0018] The above one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects: 1. This paper extracts the age and time difference at the time of consultation from electronic medical records, and introduces the time interval token into the pre-trained model based on the BART framework. This can explicitly incorporate time information into the input of the model, thereby enhancing the model's ability to learn time patterns.
[0019] 2. This paper adopts a pre-trained model based on the BART framework. By adding a decoder part, the model is transformed into a generative language model, which can expand its capabilities in generation tasks while improving the flexibility, adaptability and interpretability of the model.
[0020] 3. By introducing generative capabilities, the present invention enables the model to generate logically coherent text or sequences to demonstrate its reasoning process, thereby improving the interpretability and credibility of the model.
[0021] 4. The present invention extracts age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code and surgical information from electronic medical records. By expanding the sources and types of data sets and designing more diverse pre-training tasks, the generalization ability and adaptability of the model can be improved.
[0022] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0024] Figure 1 It is a flow chart of the method provided by the present invention.
[0025] Figure 2 It is a structural block diagram of the system provided by the present invention.
[0026] Reference numerals: 1. Electronic medical record acquisition module; 2. Sequence coding module; 3. Electronic medical record generation module. DETAILED DESCRIPTION
[0027] To make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the present invention. Obviously, the embodiments described are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0028] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiment of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0029] The following combination Figure 1 and Figure 2 The present invention is further described in detail, and a method and system for generating electronic medical records based on the BART framework are described as follows: In this embodiment, Figure 1 As shown, a method for generating an electronic medical record based on a BART framework is provided, comprising the following steps: S1: Obtain multiple electronic medical records and extract age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information from the electronic medical records.
[0030] The age at the time of consultation is calculated in months, the diagnosis code is in ICD-10 format, the abnormal laboratory test result code is in LOINC format, and the surgical information is in ICD-10PCS format.
[0031] A diagnosis record contains age at diagnosis, diagnosis date, visit type, and one or more of the following: diagnosis code, abnormal laboratory test result code, and surgical procedure information. Therefore, age at diagnosis, diagnosis date, and visit type can be extracted from the diagnosis record, while at least one of the following: diagnosis code, abnormal laboratory test result code, and surgical procedure information can be extracted.
[0032] The multiple electronic medical records are at least three electronic medical records of the same patient within one year. The prediction result is the next electronic medical record within one year after the last electronic medical record.
[0033] S2: Encode the age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information into an input sequence. The specific process is: S21: Coding Tag Embedding: Generate visit time difference based on diagnosis date, and then encode it into tag embedding together with visit type, diagnosis code, abnormal laboratory test result code and surgery information.
[0034] The time difference between the first electronic medical record and the previous one is recorded as Gap_0. Starting from the second electronic medical record, the time difference between it and the previous electronic medical record is calculated based on the diagnosis date, and mapped according to the following rules to obtain the time difference between the two electronic medical records: Less than 30 days: Gap_1, Greater than or equal to 30 days and less than 60 days: Gap_2, Greater than or equal to 60 days and less than 180 days: Gap_3, Greater than or equal to 180 days and less than 360 days: Gap_4, Greater than or equal to 360 days: Gap_5.
[0035] The diagnosis codes in each electronic medical record were sorted according to their priority. Then, they were spliced in the order of visit type, diagnosis code, abnormal laboratory test result code, surgery information, and visit time difference to obtain the splicing results.
[0036] If one or two of the diagnosis code, abnormal laboratory test result code, and procedure information are missing data, the corresponding content will be left blank. For example, if an electronic medical record only contains procedure information but no diagnosis code or abnormal laboratory test result code, the records will be joined in the order of visit type, procedure information, and visit time difference.
[0037] Using the dictionary, the concatenation results are encoded into token embeddings.
[0038] This embodiment also adjusts the length of the embedded markers to a fixed length, for example, 512. If the length is insufficient, it is padded with zeros, and if the length exceeds 512, it is cut.
[0039] S22: Encoding visit embeddings: Sort multiple electronic medical records according to the chronological order of diagnosis date, generate visit IDs from 1 to N, and then encode the visit IDs into visit embeddings.
[0040] S23: Encode positional embedding: Encode the number of months of the patient's age at the time of the visit as a positional embedding. For example, if the patient's age is 69 years and 5 months, the number of months is 833.
[0041] S24: Concatenate the token embedding, visit embedding, and location embedding into the input sequence.
[0042] This example extracts the age at the time of visit from electronic medical records and encodes it as a positional embedding. This reflects the actual interval between visits and helps the model better capture long-term temporal dependencies between multiple visits. Extracting diagnosis codes, abnormal laboratory test result codes, and surgical information from electronic medical records helps improve the accuracy of model-generated electronic medical records for important diseases. Common important diseases include Parkinson's disease, diabetes, chronic kidney disease, hypertension, cerebrovascular disease, cognitive impairment, and asthma.
[0043] S3: Input the input sequence into the pre-trained model based on the BART framework to generate future electronic medical records.
[0044] The future electronic medical record is the next electronic medical record within the next year. The content of the future electronic medical record includes a diagnosis date range, diagnosis code, abnormal laboratory test result code, and surgical information. Preferably, the future electronic medical record is an electronic medical record for one or more of Parkinson's disease, diabetes, chronic kidney disease, hypertension, cerebrovascular disease, cognitive impairment, and asthma.
[0045] The pre-training model based on the BART framework includes an encoder and a decoder. The encoder adopts a bidirectional Transformer structure and includes 6 Transformer layers. The decoder adopts an autoregressive Transformer structure and includes 6 Transformer layers.
[0046] The pre-trained model (BART-EMR) based on the BART framework constructed in this example is specifically designed to process structured electronic health record (EMR) data to generate electronic medical records for patients within the next year. The model, through the contextual embeddings learned during the pre-training phase, is able to capture the complex features and relationships in EMR data, significantly improving performance in the medical record generation task during the fine-tuning phase. The following is the complete implementation process and details of the BART-EMR model.
[0047] 1. Data Preprocessing (1) Patient selection: A total of 90,394 patients in Miami, USA, who visited the hospital three times or more during the three years from 2019 to 2021, were selected.
[0048] (2) Data cleaning: Ensure that the patient's diagnostic information is complete and in the correct chronological order. For each visit, the following information was extracted: age at visit (calculated in months), diagnosis date, diagnosis code (ICD-10 format), abnormal laboratory test result code (LOINC format), surgical information (ICD-10 PCS format), and the type of each visit.
[0049] (3) Data sorting: Sort the patient's electronic medical records in chronological order and generate visit IDs from 1 to N.
[0050] Sort the diagnosis codes in each electronic medical record by priority. If the visit includes a code for an abnormal laboratory test result, it will be added after the diagnosis code. If the visit includes surgical information, it will be added after the code for the abnormal laboratory test result.
[0051] Starting from the second electronic medical record, calculate the time difference between two consecutive medical records and map them according to the following rules: Less than 30 days: Gap_1, Greater than or equal to 30 days and less than 60 days: Gap_2, Greater than or equal to 60 days and less than 180 days: Gap_3, Greater than or equal to 180 days and less than 360 days: Gap_4, Greater than or equal to 360 days: Gap_5.
[0052] (4) Tokenizer: Build a dictionary based on the sorted data of all patients according to independent medical codes, so that each code corresponds one-to-one with a token in the dictionary.
[0053] (5) Data padding: According to the length of each sample, 512 marks are added to fill the samples with insufficient length with 0, and the samples with length exceeding 512 are cut.
[0054] 2. BART-EMR Model Architecture (1) Input layer: Token Embedding: A low-dimensional representation of each diagnosis code (including visit type, diagnosis code, abnormal laboratory test result code, surgery information, and visit time difference).
[0055] Position Embedding: It uses the age at the time of medical consultation (calculated in months) to accurately reflect the time interval between medical records.
[0056] Visit Embedding: Represented by visit ID, it is used to distinguish the location of each visit.
[0057] (2) Encoder: It adopts BART's bidirectional Transformer structure, which consists of 6 layers, each of which contains two main modules: a multi-head self-attention mechanism and a feedforward neural network.
[0058] Multi-Head Self-Attention Mechanism: Self-attention mechanism: allows the model to dynamically pay attention to other tokens in the sequence as it processes each token, thereby capturing global dependencies.
[0059] Input: Input sequence, including token embeddings, visit embeddings, and location embeddings.
[0060] Query, Key, Value: The input sequence is transformed into query, key, and value through three different linear transformations.
[0061] Attention score calculation: The attention score is obtained by calculating the dot product between the query and the key, and is normalized by scaling and softmax function.
[0062] Multi-head mechanism: The input is divided into multiple heads, each head performs self-attention calculation independently, and finally the results of these heads are spliced together, which increases the model's expressive power and ability to learn different subspaces.
[0063] Number of Heads: The BART-EMR model has 12 heads.
[0064] Feed-Forward Neural Network: Each encoder layer contains a feedforward neural network that performs a nonlinear transformation on the label at each position to increase the fitting ability of the model.
[0065] It contains two linear transformations with a ReLU activation function in the middle.
[0066] Input: Output of the multi-head self-attention mechanism.
[0067] Output: Embedding representation of the encoder.
[0068] (3) Decoder: It uses the BART autoregressive Transformer structure, which consists of 6 layers.
[0069] Multi-Head Self-Attention Mechanism: Self-attention mechanism: allows the model to dynamically pay attention to the generated tokens when generating each token, thereby capturing the dependencies in the generation process.
[0070] Input: The embedding representation of the decoder, including token embedding and time interval embedding.
[0071] Query, Key, Value: The input sequence is transformed into query, key, and value through three different linear transformations.
[0072] Attention score calculation: The attention score is obtained by calculating the dot product between the query and the key, and is normalized by scaling and softmax function.
[0073] Multi-head mechanism: The input is divided into multiple heads, each head performs self-attention calculation independently, and finally the results of these heads are spliced together, which increases the model's expressive power and ability to learn different subspaces.
[0074] Number of Heads: The BART-EMR model has 12 heads.
[0075] Encoder-Decoder Attention Mechanism: This allows the decoder to dynamically attend to the encoder’s output as it generates each token, thereby capturing the contextual information learned by the encoder.
[0076] Input: The decoder's query (Query) and the encoder's key (Key) and value (Value).
[0077] Attention score calculation: The attention score is obtained by calculating the dot product between the query and the key, and is normalized by scaling and softmax function.
[0078] Multi-head mechanism: The input is divided into multiple heads, each head performs attention calculation independently, and finally the results of these heads are spliced together, which increases the model's expressive power and ability to learn different subspaces.
[0079] Number of Heads: The BART-EMR model has 12 heads.
[0080] Feed-Forward Neural Network: Each decoder layer contains a feed-forward neural network that performs a nonlinear transformation on the label at each position to increase the fitting ability of the model.
[0081] It contains two linear transformations with a ReLU activation function in the middle.
[0082] Input: Output of the encoder-decoder attention mechanism.
[0083] (4) Output layer: In the generation phase, the model outputs a generated sequence of tokens that is used to generate electronic medical records for patients within the next year.
[0084] Existing models do not explicitly consider the time interval between visits, resulting in missing temporal information, limited generalization, and insufficient clinical interpretability. This embodiment uses time interval embedding to mark the time interval between two visits during each different visit.
[0085] Enhanced representation of temporal information: By introducing time interval tokens, the model can explicitly learn and utilize temporal information, thereby better capturing the temporal patterns of disease progression. For example, multiple visits within a short period of time may indicate worsening of the disease, while long intervals may indicate stable disease.
[0086] Improving the clinical interpretability of the model: Time interval tokens can provide clinicians with more intuitive explanations, helping them understand the model's predictions. For example, the model can clearly indicate the contribution of time intervals to the prediction results, thereby increasing clinicians' trust in the model.
[0087] Improve the generalization ability of the model: The introduction of time intervals can enable the model to better adapt to the temporal characteristics in different datasets, thereby improving its generalization ability in different clinical scenarios.
[0088] Support for more complex clinical tasks: For example, when predicting a patient's time to readmission or disease recurrence, time interval information is crucial. Introducing time interval tokens can enable the model to better handle such tasks.
[0089] The model of this embodiment adopts an encoder-decoder structure and is a generative language model.
[0090] Support for more types of clinical tasks: By adding the decoder part, the model can handle generation tasks: generating the patient's electronic medical records within the next year, which greatly expands the application scope of the model.
[0091] Improve model flexibility and adaptability: Generative models can generate multiple possible outputs based on the context of the input, which makes the model more flexible in dealing with complex clinical problems.
[0092] Supporting personalized medicine: Generative models can generate personalized medical recommendations based on individual patient characteristics. For example, based on a patient's medical history and current health status, they can generate treatment plans or health management recommendations tailored to that patient. This personalized capability is crucial for improving the quality and efficiency of healthcare services.
[0093] Promoting cross-domain applications: Generative language models have achieved tremendous success in natural language processing, such as in text generation and machine translation. Introducing this capability into healthcare can promote cross-domain applications, such as the development of intelligent medical assistants and medical chatbots.
[0094] 3. Pre-training phase (1) Pre-training task: Masked Language Model (MLM): predict masked diagnosis codes.
[0095] (2) Specific operation: 80% probability of replacing the code with the mask: [MASK]. 10% probability of replacing the code with a random code. 10% probability of keeping the code unchanged.
[0096] (3) Optimization algorithm: Use Adam optimizer, learning rate of 1e-5, and dropout rate of 0.1.
[0097] 4. Fine-tuning stage (1) Fine-tuning tasks Input: Patient's electronic medical record for the last year; Output: Generate electronic medical records of patients within the next year; Task Objective: Generate electronic medical records for patients within the next year, including diagnosis date range, diagnosis codes, abnormal laboratory test result codes, and surgical information.
[0098] (2) Fine-tuning process Data Split: The data of patients with 3 or more visits in 2019 were divided into training and test sets with a ratio of 8:2. The data of patients with 3 or more visits in 2020 were used as the validation set.
[0099] Fine-tuning phase: Updates the model's encoder and decoder parameters. The encoder feeds the source text into the encoder, generating a contextual representation. The decoder uses autoregression to gradually generate the target text. At each time step, the decoder predicts the next token based on the generated portion and the encoder output. The encoder-decoder attention mechanism is used to incorporate contextual information from the encoder into the decoder.
[0100] 5. Performance Evaluation (1) Evaluation indicators: Loss (loss value): used to evaluate the degree of training convergence.
[0101] BLEUScore (Bilingual Evaluation Substitute): This tool is used to assess the accuracy of generated medical records. It calculates a score by comparing the n-gram matches between the generated text and the reference text. A higher BLEUScore score indicates a greater similarity between the generated text and the reference text.
[0102] ROUGEScore (recall, precision, and F1 score): used to evaluate the quality of generated medical records. The score is calculated by comparing the overlap between the generated summary and the reference summary.
[0103] (2) Experimental results: Pre-training phase: Training set: Loss: 0.4428266507945955; Validation set: Loss: 0.44441801811015846.
[0104] Fine-tuning phase: Training set: Loss: 0.13100771605968475 BLEUScore: 0.45; ROUGEScore: 0.55.
[0105] Validation set: Loss: 0.1380409449338913 BLEUScore: 0.43; ROUGEScore: 0.53.
[0106] Table 1 shows the BART-EMR model's performance in classifying important diseases. Evaluation metrics include: Positive Predictive Value (PPV): The probability of a positive test result indicating the patient actually has the disease. Higher values are preferred; Negative Predictive Value (NPV): The probability of a negative test result indicating the patient does not have the disease. Higher values are preferred; Sensitivity: The ability to identify all patients who actually have the disease, also known as recall; Specificity: The ability to exclude patients without the disease; and Accuracy: The proportion of correctly classified patients (a combination of sensitivity and specificity).
[0107] Table 1
[0108] This example demonstrates very high negative predictive values and specificities for most diseases (both between 0.93 and 0.99), demonstrating the model's high accuracy in excluding those without the disease. Accuracy is generally high, exceeding 0.93 for most diseases and reaching a high of 0.9943 (for Parkinson's disease), demonstrating the model's overall reliability.
[0109] The model was most effective in identifying Parkinson's disease: all indicators were high (positive predictive value 0.8113, negative predictive value 0.9961, sensitivity 0.6754, specificity 0.9981, and accuracy 0.9943), indicating that the model can both identify patients and rarely misjudge healthy individuals, making it very suitable for providing diagnostic recommendations for Parkinson's disease.
[0110] The identification of chronic diseases (such as diabetes, chronic kidney disease, and hypertension) is relatively stable: the sensitivity of diabetes (without complications) is 0.8392, and the positive predictive value is also high.
[0111] Although the sensitivity of chronic kidney disease was slightly lower (0.6663), its specificity and accuracy were high.
[0112] The sensitivity of essential hypertension is as high as 0.7781 and is stable in common diseases.
[0113] This example also compares the prediction accuracy of Med-BERT for various diseases, as shown in Table 2. This example extracts age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical procedure information from electronic medical records and encodes them as input sequences, which serve as model input and help improve the model's prediction accuracy.
[0114] Table 2
[0115] This example generates a future electronic medical record based on three electronic medical records of a patient. The content is as follows: (1) The first electronic medical record Basic information: Age: 69 years and 0 months.
[0116] Type of visit: Outpatient.
[0117] Consultation time: 2019-04-20.
[0118] Abnormal laboratory test data: 1920-8: aspartate aminotransferase; 2345-7: glucose; 48642-3: glomerular filtration rate; 718-7: hemoglobin; 3094-0: urea nitrogen.
[0119] (2) Second electronic medical record Basic information: Age: 69 years and 4 months.
[0120] Type of visit: Outpatient.
[0121] Consultation time: 2019-08-21.
[0122] Diagnosis: I50.32: Chronic diastolic heart failure; J44.1: Acute exacerbation of chronic obstructive pulmonary disease; severe stage.
[0123] (3) The third electronic medical record Basic information: Age: 69 years and 5 months.
[0124] Type of visit: Hospitalization.
[0125] Consultation time: 2019-09-07.
[0126] Diagnosis: D64.89: Anemia; E66.01: Obesity; G47.33: Obstructive sleep apnea; E11.22: Chronic kidney disease; I25.10: Atherosclerotic heart disease.
[0127] Abnormal laboratory test data: 1988-5: C-reactive protein (CRP); 2345-7: glucose; 48642-3: glomerular filtration rate; 2502-3: iron saturation; 2951-2: sodium.
[0128] (4) The next electronic medical record generated (future electronic medical record) Diagnosis date range: within 30 days (calculated from September 7, 2019).
[0129] Diagnosis: E78.5: Hyperlipidemia; I13.0: Hypertensive heart disease and chronic kidney disease with heart failure; I25.10: Atherosclerotic heart disease; I27.20: Pulmonary hypertension; N18.9: Chronic kidney disease.
[0130] Abnormal laboratory test data: 48642-3: Glomerular filtration rate; 6598-7: Troponin; 6690-2: White blood cells.
[0131] Surgery: 4A023N6: Cardiac sampling and pressure measurement; B2111ZZ: Coronary angiography.
[0132] The BART-EMR model, through contextual embeddings learned during the pre-training phase, effectively captures the complex features and relationships in EMR data. During fine-tuning, the BART-EMR model significantly improved performance on medical record generation tasks, particularly with limited data. By introducing time interval markers, the model better captures temporal information, thereby improving the accuracy and relevance of generated medical records. Furthermore, through its generative capabilities, the model can generate logically coherent medical record content, helping clinicians understand and trust the model's predictions.
[0133] This embodiment also provides an electronic medical record generation system based on the BART framework, such as Figure 2 As shown, the adopted technical solution is as follows: including: an electronic medical record acquisition module 1, a sequence encoding module 2 and an electronic medical record generation module 3.
[0134] The electronic medical record acquisition module is configured to acquire multiple electronic medical records and extract the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code, and surgical information from each of the electronic medical records. The multiple electronic medical records are at least three electronic medical records within a year.
[0135] The sequence encoding module is used to encode the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code and operation information into an input sequence.
[0136] Age at presentation was coded as a positional embedding; Generate a visit ID based on the diagnosis date and then encode it into a visit embedding; Generate the visit time difference based on the diagnosis date, and then embed it into a token along with the visit type, diagnosis code, abnormal laboratory test result code, and surgical information code. The diagnosis codes in each electronic medical record are sorted by priority and then concatenated after the visit type, followed by the abnormal laboratory test result code, surgical information, and visit time difference. This is encoded into a token embed using a dictionary. Concatenate the token embedding, visit embedding, and location embedding into the input sequence.
[0137] The electronic medical record generation module is configured to input an input sequence into a pre-trained model based on the BART framework to generate a future electronic medical record. The pre-trained model based on the BART framework includes an encoder and a decoder. The encoder utilizes a bidirectional Transformer architecture with six Transformer layers, while the decoder utilizes an autoregressive Transformer architecture with six Transformer layers. The future electronic medical record is for the next year. The content of the future electronic medical record includes diagnosis date range, diagnosis code, abnormal laboratory test result code, and surgical information.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for generating electronic medical records based on the BART framework, characterized in that: The following steps are involved: S1: Obtain multiple electronic medical records and extract age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information from the electronic medical records; S2: Encode age at visit, diagnosis date, visit type, diagnosis code, abnormal laboratory test result code, and surgical information as input sequences; S3: Input the input sequence into a pre-trained model based on the BART framework to generate a future electronic medical record; the pre-trained model based on the BART framework includes an encoder and a decoder.
2. The method for generating an electronic medical record based on the BART framework according to claim 1, wherein: The age at the time of consultation is calculated in months.
3. The method for generating electronic medical records based on the BART framework according to claim 1 or 2, characterized in that: Step S2 includes: Age at presentation was coded as a positional embedding; Generate a visit ID based on the diagnosis date and then encode it into a visit embedding; The visit time difference was generated based on the diagnosis date and then embedded as a marker along with the visit type, diagnosis code, abnormal laboratory test result code, and surgery information code; Concatenate the token embedding, visit embedding, and location embedding into the input sequence.
4. The method for generating an electronic medical record based on the BART framework according to claim 3, wherein: The encoding process of token embedding is: The diagnosis codes in each electronic medical record were sorted according to their priority, and then concatenated after the visit type, followed by the abnormal laboratory test result code, surgical information, and visit time difference; Using the dictionary, the concatenation results are encoded into token embeddings.
5. The method for generating electronic medical records based on the BART framework according to claim 4, characterized in that: Starting from the second electronic medical record, calculate the time difference between it and the previous electronic medical record based on the diagnosis date, and map them according to the following rules to obtain the consultation time difference: Less than 30 days: Gap_1, Greater than or equal to 30 days and less than 60 days: Gap_2, Greater than or equal to 60 days and less than 180 days: Gap_3, Greater than or equal to 180 days and less than 360 days: Gap_4, Greater than or equal to 360 days: Gap_5.
6. The method for generating electronic medical records based on the BART framework according to claim 1, wherein: The multiple electronic medical records are at least 3 electronic medical records within one year; The future electronic medical records are electronic medical records within the next year.
7. The method for generating electronic medical records based on the BART framework according to claim 1 or 6, characterized in that: The content of the future electronic medical record includes diagnosis date range, diagnosis code, abnormal laboratory test result code and surgical information.
8. The method for generating electronic medical records based on the BART framework according to claim 7, wherein: The future electronic medical records are electronic medical records of one or more of Parkinson's disease, diabetes, chronic kidney disease, hypertension, cerebrovascular disease, cognitive impairment diseases, and asthma.
9. The method for generating electronic medical records based on the BART framework according to claim 1, wherein: The encoder adopts a bidirectional Transformer structure, which contains 6 Transformer layers; the decoder adopts an autoregressive Transformer structure, which contains 6 Transformer layers.
10. An electronic medical record generation system based on the BART framework, characterized in that: The method for generating an electronic medical record based on the BART framework according to any one of claims 1 to 9 comprises: an electronic medical record acquisition module, a sequence encoding module and an electronic medical record generation module. The electronic medical record acquisition module is used to acquire multiple electronic medical records and extract the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code and surgery information from the electronic medical records; The sequence encoding module is used to encode the age at the time of consultation, diagnosis date, consultation type, diagnosis code, abnormal laboratory test result code and surgical information into an input sequence; The electronic medical record generation module is used to input the input sequence into a pre-trained model based on the BART framework to generate future electronic medical records; the pre-trained model based on the BART framework includes an encoder and a decoder.
Citation Information
Patent Citations
Medical record text generation method and device, computer equipment and storage medium
CN115862794A
Interpretable sequence prediction method based on space-time perception
CN116451102A
Medical image diagnosis report generation method, device, equipment, medium and product
CN119905191A
Artificial intelligence-based safe privacy zone management server and method and privacy detection device
KR1020250061975A
Method for operating medical artificial intelligence model, and electronic device for performing same
WO2024096307A1