Cardiovascular disease medical record generation method based on large language model
Through the improved large language model and dynamic information interaction attention mechanism, the existing electronic medical record system has solved the problem of incomplete information recording and logical fragmentation in the diagnosis and treatment of cardiovascular diseases, and achieved high-quality medical record generation and causal relationship recognition, which is suitable for high-risk cardiovascular disease management.
Patent Information
- Application Number
- CN202510485526.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-22
AI Technical Summary
In the diagnosis and treatment of cardiovascular diseases, the existing electronic medical record system has complicated information records, uneven data quality, lack of structured processing and causal logical relationships, making it difficult to accurately identify the internal logic between cause, examination, diagnosis and treatment, and the long text processing capacity is insufficient, resulting in incomplete medical record records and logical separation.
The improved large language model combined with the dynamic information interaction attention mechanism is adopted to construct a causal network and timeline through local aggregation, proximity diffusion and global association optimization, and automatically identify the causal path between cause-diagnosis-treatment to generate structured medical records.
The semantic modeling ability, causal logic reasoning ability and time series understanding ability of medical records are improved. The generated medical records are more in line with doctors' cognition in terms of time logic and dimensions, and are suitable for high-risk cardiovascular disease management, achieving high-quality disease course modeling and accurate identification of medical causal relationships.
Smart Images

Figure CN120356599A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cardiovascular diseases, and more specifically, to a method for generating medical records of cardiovascular diseases based on large language models. Background Art
[0002] Cardiovascular diseases are one of the main causes of death globally. Their complex pathological mechanisms and variable clinical manifestations make the diagnosis and treatment process full of challenges. As the core tool for medical information management, electronic medical records play a crucial role in the diagnosis, treatment, and follow-up of cardiovascular diseases. However, there are still many problems in the information recording, structured processing, and intelligent reasoning of existing electronic medical records, seriously affecting the efficiency and accuracy of clinical decision-making. Specifically, the current electronic medical record system mainly relies on manual filling and traditional text processing methods. The data entry is cumbersome, and doctors need to manually input a large amount of text information. This not only takes a lot of time and effort for doctors to input information but also easily causes problems such as incomplete medical record records and uneven data quality. At the same time, a large amount of information in electronic medical records exists in an unstructured or semi-structured form, lacking a unified standard and being difficult to support subsequent data mining, semantic analysis, and medical reasoning. In addition, although the application of natural language processing (NLP) technology in the medical field is constantly expanding, existing NLP methods still face many challenges such as shallow structure modeling, insufficient semantic understanding, and weak causal logic association, especially in the parsing of complex cardiovascular medical records. Specifically, current NLP methods mainly rely on rule-based and shallow feature-based models in medical text processing, lacking the ability to model medical causal relationships and being difficult to accurately identify the internal logic between causes, examinations, diagnoses, and treatments. At the same time, cardiovascular medical records usually exhibit characteristics such as spanning multiple time periods, multiple stages, and being information-intensive. Insufficient long text processing capabilities often lead to information omission or logical fragmentation. In addition, time nodes in the course of disease development are of great significance to the evolution of the disease, but most electronic medical record systems fail to effectively construct a time axis for the course of the disease, resulting in a lack of continuity and traceability in the development process of the disease.
[0003] The prior art discloses a method for generating a nursing medical record, which includes the following method steps: collecting a nursing medical record template library and patient text data, wherein the patient text data includes multiple semantic units; using a bidirectional text vectorization model, by mapping multiple semantic units in the patient text data into a vector space, generating multiple text word vectors for capturing the semantic information of words in the patient text data; and converting all templates in the nursing medical record template library into multiple template library word vectors for capturing the semantic information of the templates; according to the multiple text word vectors, by calculating the arithmetic mean of the multiple text word vectors, obtaining a text average word vector for capturing the overall semantic information; and calculating the cosine similarity between the text average word vector and the multiple template library word vectors as the template matching similarity score, and comparing the cosine similarity between the text average word vector and the multiple template library word vectors as the template matching similarity score. The cosine similarity between the overall semantic information of the average word vector of the text and the semantic information of all templates is compared, and the template with the largest template matching similarity score is selected as the nursing medical record template to be used; the nursing medical record template to be used is parsed, and each field in the nursing medical record template to be used is parsed to generate a field word vector; the cosine similarity between the field word vector and multiple text word vectors is calculated as the field filling similarity score, which is used to locate the semantic matching information related to the field in the text; and according to the size of the field filling similarity score, a list of semantic relevance between each field in the nursing medical record template to be used and multiple text word vectors is generated; according to the semantic relevance list, the corresponding fields of the nursing medical record template are filled to improve the filling accuracy of the semantic requirements of specific fields. It adopts a bidirectional text vectorization model and cosine similarity calculation, which is mainly based on field-level matching filling. It has insufficient dependency modeling capabilities for long texts and is difficult to handle multi-segment logical associations in cardiovascular medical records, such as disease progression, recurrence risk, and comorbidities. In addition, semantic understanding is mainly based on the averaging of word vectors, without in-depth analysis of global semantics, making it difficult to optimize the complex structure of cardiovascular medical records, resulting in a lack of coherence in the description of the course of disease. In addition, this method lacks real-time reasoning and adjustment capabilities when dealing with dynamic changes in high-risk cardiovascular events (such as acute myocardial infarction). Therefore, there is an urgent need for an intelligent medical record generation method that is oriented to major cardiovascular disease scenarios and has strong semantic modeling capabilities, causal logic reasoning capabilities, and time series understanding capabilities. Summary of the invention
[0004] In order to overcome at least one of the defects described in the above-mentioned prior art, the present invention provides a method for generating medical records of cardiovascular diseases based on a large language model. The present invention can improve the semantic modeling ability, causal logic reasoning ability and time series understanding ability of medical record generation.
[0005] In order to solve the above technical problems, the technical solution of the present invention is as follows:
[0006] Obtaining medical records of patients with cardiovascular diseases and preprocessing the medical records of the patients;
[0007] Input the preprocessed patient medical record into an improved large language model to output a semi-structured medical record, where the improved large language model is obtained by replacing the self-attention mechanism with a dynamic information interaction attention mechanism;
[0008] Input the semi-structured medical record into a causal network to extract the causal relationship between paragraphs and output a causal chain;
[0009] Build a timeline based on the semi-structured medical record and the extracted causal relationship;
[0010] Fuse the causal chain and the timeline to output a structured medical record.
[0011] Furthermore, the preprocessing includes: segmenting the patient medical record by category using syntactic dependency relations and a medical term dictionary to obtain the preprocessed patient medical record, where the categories include symptoms, examinations, diagnoses, treatments, and medications.
[0012] Furthermore, inputting the preprocessed patient medical record into an improved large language model to output a semi-structured medical record specifically includes:
[0013] Input the preprocessed patient medical record into an encoder for semantic encoding, extract the latent representation of each medical entity in each segment, and construct a corresponding feature vector according to the latent representation;
[0014] Input the feature vector into a dynamic information interaction attention mechanism to perform local aggregation, adjacent diffusion, and global association attention calculations respectively to obtain corresponding attention results;
[0015] Obtain an attention weight distribution based on the three types of attention results and a value vector;
[0016] Input the attention weight distribution into a decoder for semantic parsing to output a semi-structured medical record.
[0017] Furthermore, inputting the feature vector into a dynamic information interaction attention mechanism to perform local aggregation attention calculation to obtain the attention result corresponding to local aggregation specifically includes:
[0018] Define the initial matching weight P of medical terms k , expressed as:
[0019]
[0020] Use a dynamic weight strategy to adjust the initial matching weight and output the first matching weight P1, expressed as:
[0021] P1 = αP k+(1-α)·softmax(βP k )
[0022] where α and β are adaptive adjustment parameters;
[0023] Adjust the initial matching weight using weighted cosine similarity and output the second matching weight P2, expressed as:
[0024]
[0025] where γ is a balance parameter, and tanh(q i -k j ) is used to capture the offset information between the query vector and the key vector;
[0026] Perform weighted summation on the first matching weight P1 and the second matching weight P2, and output the attention result P corresponding to local aggregation local , expressed as:
[0027] P local =λP1+(1-λ)·P2
[0028] where λ∈[0,1] is the attention fusion control parameter.
[0029] Furthermore, the step of inputting the feature vector into the dynamic information interaction attention mechanism to perform proximity diffusion attention calculation to obtain the attention result corresponding to proximity diffusion specifically includes:
[0030] Construct a segment propagation guidance term using the structural relationship Expressed as:
[0031]
[0032] where PosEnc(i) represents the encoding of the token based on position i, w k =1 / (d ik +ε) is the distance-based weight adjustment factor, d ik is the distance between token i and the k-th paragraph, and E k [i] is the position of the medical entity represented by PosEnc(i), and E k [i]=PosEnc(i);
[0033] Based on the segment propagation guidance term Calculate the attention result P inter corresponding to proximity diffusion, expressed as:
[0034]
[0035] where W q ,Ws ,W k ,W p is a learnable linear mapping matrix.
[0036] Furthermore, the step of inputting the feature vector into the dynamic information interaction attention mechanism to perform global correlation attention calculation and obtain the attention result corresponding to the global correlation specifically includes:
[0037] Construct a global guidance vector G based on the distribution shift of medical entities between paragraphs q [i], expressed as:
[0038]
[0039] where U k [i] represents the global position encoding of the paragraph where the i-th key is located, L is the total number of medical record paragraphs, and μ m is the global offset importance weight of paragraph m to paragraph i;
[0040] Based on the global guidance vector G q [i] and the query vector to establish a global attention query representation expressed as:
[0041]
[0042] where R q ,R g is a trainable linear transformation matrix used to control the fusion degree between the original semantics and the global offset; represents element-wise multiplication;
[0043] According to the global attention query representation and the key vector to obtain the attention result P corresponding to the global correlation global , expressed as:
[0044]
[0045] where R k is the key vector mapping matrix to ensure dimension alignment with the query vector.
[0046] Furthermore, the step of inputting the semi-structured medical record into the causal network to extract the causal relationship between paragraphs and output the causal chain specifically includes:
[0047] Use the YAKE algorithm to calculate the importance scores of each medical entity in the semi-structured medical record in the semi-structured medical record, and sort the importance scores from largest to smallest;
[0048] According to the sorting results, p medical entities with scores higher than 0.5 are selected, and the p medical entities are clustered into q medical factors using the clustering method. Each medical factor corresponds to a node, and each medical factor is binarized;
[0049] Use conditional independence tests to calculate the causal relationships of the binarized medical factors. If two medical factors are independent under the given variable conditions, the direct causal relationship between these two medical factors is removed, and the causal direction of these two medical factors is determined according to the V-structure rule. Otherwise, the causal direction is determined according to sampling, thus forming a partially directed graph;
[0050] Sample the corresponding causal directions for the edges with undetermined causal directions in the partially directed graph with equal probabilities, and remove the bidirectional edges that appear during the sampling process. Repeat the sampling process until every causal direction of all the edges with undetermined causal directions is traversed, and output m candidate causal graphs;
[0051] Based on the Bayesian information criterion, calculate the matching scores of each candidate causal graph with the data X, and take the candidate causal graph corresponding to the highest score as the final causal graph;
[0052] Use a search algorithm to search the final causal graph to obtain a causal chain that meets the start-end conditions.
[0053] Further, the determining of the causal direction of these two medical factors according to the V-structure rule specifically includes: when it is calculated using conditional independence tests that the first medical factor and the second medical factor are independent under the given variable conditions, and there is a third medical factor that makes the first medical factor and the second medical factor related, then the third medical factor serves as the intermediate point between the first medical factor and the second medical factor, and the causal direction is: the first medical factor points to the third medical factor, and the second medical factor points to the third medical factor.
[0054] Further, the establishing of the timeline based on the semi-structured medical records and the extracted causal relationships specifically includes:
[0055] Use the BiLSTM and conditional random field model to extract medical events from the semi-structured medical records to obtain a medical event set, and standardize the medical event set in combination with the SNOMED-CT terminology library to obtain a standardized medical event set;
[0056] Use predefined time-dependent inference rules to sort the events in the standardized medical event set, and use the time causal inference scoring model Score(A→B) to score the sorted events. Delete the event sorting with a score lower than 0.5, and output the first event time sorting, where the time causal inference scoring model Score(A→B) is expressed as:
[0057] Score(A→B) = λ1·C(A,B) + λ2·T(A,B) + λ3·R(A,B)
[0058] Among them, C(A,B) is the causal relationship, T(A,B) is the temporal rationality, R(A,B) is the context consistency, and λ1, λ2, λ3 are the weights corresponding to the causal relationship, temporal rationality, and context consistency respectively;
[0059] Predict the events lacking time in the time sorting of the first event, output the corresponding predicted time information, and update the time sorting of the first event based on the predicted time information, and output the second event time sorting;
[0060] Use a graph neural network to perform propagation optimization on the second event time sorting, and output the third event time sorting;
[0061] Use the time window prediction method to optimize the third event time sorting, output the fourth event time sorting, and output according to time for the fourth time sorting to obtain a timeline.
[0062] Furthermore, the time window prediction method is determined by the following formula:
[0063] [T min , T max = [T mean - kσ, T mean + kσ]
[0064] Among them, T min is the minimum value of the time interval, that is, the start time; T max is the maximum value of the time interval, that is, the end time; T mean is the time mean statistically based on historical data; σ is the standard deviation, indicating the fluctuation range of time; k is an adjustment factor.
[0065] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0066] (1). Through the dynamic information interaction attention mechanism, the present invention effectively solves the problem of insufficient long-term dependence modeling in the prior art when dealing with long text medical records by means of local aggregation, adjacent diffusion, and global association collaborative optimization. Among them, local aggregation focuses on local semantic information, adjacent diffusion captures cross-block relationships, and global association fuses local and global information, so as to achieve accurate modeling of semantic dependencies in complex cardiovascular medical records;
[0067] (2). By constructing a causal network that combines with the SNOMED-CT medical ontology, it breaks through the traditional rule-based information extraction method based on field matching, and can automatically identify causal paths such as cause-diagnosis-treatment from semi-structured texts, improving the medical logic accuracy and interpretability of medical record generation. Compared with existing systems that rely on templates or shallow semantic matching, this method has stronger reasoning and adaptability, can automatically analyze the evolution of complex medical conditions and the basis of clinical decisions, is applicable to application scenarios with high requirements for precise reasoning such as high-risk cardiovascular disease management, and helps to achieve precise modeling and automatic identification of medical causal relationships;
[0068] (3). Also, based on the method of constructing the medical record timeline using temporal causal reasoning and graph attention networks, it can infer the causal relationships in the development of medical records over time, making the generated medical record text more in line with doctors' cognition in terms of time logic and dimensions. At the same time, different from the methods in the prior art that rely on simple timestamps or manual annotations for sorting, the present invention can, in the absence of explicit time annotations, perform automatic sorting and logical correction through event cross-mapping, time constraints, and prediction, thereby achieving high-quality disease course modeling. This mechanism not only ensures the correctness and completeness of the medical record history but also excludes the adverse effects of disease factors in different periods. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a flowchart of a method for generating a medical record of cardiovascular disease based on a large language model according to Embodiment 1 of the present invention;
[0070] Figure 2 is a flowchart of a method for generating a medical record of cardiovascular disease based on a large language model according to Embodiment 1 of the present invention;
[0071] Figure 3 is a diagram of the dynamic information interaction attention mechanism according to Embodiment 1 of the present invention;
[0072] Figure 4 is a flowchart of the causal network processing according to Embodiment 1 of the present invention;
[0073] Figure 5 is a flowchart of the timeline processing according to Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] The drawings are only for illustrative purposes and should not be construed as limiting the patent;
[0075] For better illustration of this embodiment, some components in the drawings are omitted, enlarged, or reduced, and do not represent the dimensions of the actual product;
[0076] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0077] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0078] Embodiment 1
[0079] As Figures 1-3 shown, a method for generating medical records of cardiovascular diseases based on a large language model according to a preferred embodiment of an embodiment of the present invention includes the following steps:
[0080] S1: Obtain the medical records of patients with cardiovascular system diseases and preprocess the medical records of patients;
[0081] S2: Input the preprocessed medical records of patients into an improved large language model to output semi-structured medical records, wherein the improved large language model is obtained by replacing the self-attention mechanism with a dynamic information interaction attention mechanism;
[0082] S3: Input the semi-structured medical records into a causal network to extract the causal relationship between paragraphs and output a causal chain;
[0083] S4: Establish a timeline based on the semi-structured medical records and the extracted causal relationship;
[0084] S5: Integrate the causal chain and the timeline to output structured medical records.
[0085] In this embodiment, through the dynamic information interaction attention mechanism, with the collaborative optimization of local aggregation, neighboring diffusion, and global association, the problem of insufficient modeling of long-term dependencies in the prior art when dealing with long-text medical records is effectively solved. Among them, local aggregation focuses on local semantic information, neighboring diffusion captures cross-block relationships, and global association fuses local and global information, thereby achieving accurate modeling of semantic dependencies in complex cardiovascular medical records. Also, by constructing a causal network combined with the SNOMED-CT medical ontology, it breaks through the traditional rule-based information extraction method based on field matching, and can automatically identify causal paths such as cause-diagnosis-treatment from semi-structured text, improving the medical logic accuracy and interpretability of medical record generation. Compared with existing systems that rely on templates or shallow semantic matching, this method has stronger reasoning and adaptability capabilities, can automatically analyze complex disease evolutions and clinical decision-making bases, is applicable to application scenarios with high requirements for accurate reasoning such as high-risk cardiovascular disease management, and helps to achieve accurate modeling and automatic identification of medical causal relationships. Also, based on the medical record timeline construction method of temporal causal reasoning and graph attention network, it can reason about the causal relationships of medical record development in time, making the generated medical record text more in line with doctors' cognition in terms of time logic and dimensions. At the same time, different from the methods in the prior art that rely on simple timestamps or manual annotations for sorting, the present invention can perform automatic sorting and logical correction through event cross-mapping, time constraints, and predictions in the absence of explicit time annotations, thereby achieving high-quality course modeling. This mechanism not only ensures the correctness and completeness of the medical record history but also excludes the adverse effects of disease factors in different periods.
[0086] In the specific implementation process, in S1, the preprocessing includes: segmenting the patient's medical record by category using syntactic dependency relationships and a medical term dictionary to obtain the preprocessed patient's medical record, where the categories include symptoms, examinations, diagnoses, treatments, medications, and optionally, also include: causes and reexaminations, etc.
[0087] In one embodiment, in S2 includes:
[0088] S2.1: Input the preprocessed patient's medical record into the encoder for semantic encoding, extract the latent representations of each medical entity in each segment, and construct corresponding feature vectors according to the latent representations;
[0089] S2.2: Input the feature vectors into the dynamic information interaction attention mechanism, perform local aggregation, neighboring diffusion, and global association attention calculations respectively, and obtain corresponding attention results;
[0090] S2.3: Obtain the attention weight distribution based on the three types of attention results and the value vector;
[0091] S2.4: Input the attention weight distribution into the decoder for semantic parsing to output a semi-structured medical record. The dynamic information interaction attention mechanism collaboratively optimizes through local aggregation, adjacent diffusion, and global association, effectively solving the problem of insufficient long-term dependence modeling in the prior art when processing long-text medical records. Among them, local aggregation focuses on local semantic information, adjacent diffusion captures cross-block relationships, and global association fuses local and global information, thereby achieving accurate modeling of semantic dependencies in complex cardiovascular medical records.
[0092] Further, S2.2 specifically includes:
[0093] 1. Obtain the attention result corresponding to local aggregation, specifically including:
[0094] 1). Define the initial matching weight P of medical terms k , to optimize the medical concept parsing of local medical record texts, in this embodiment, an asymmetric local attention mapping matrix is introduced to define the initial matching weight P k , expressed as:
[0095]
[0096] Among them, W k is a trainable weight matrix, used to calculate the matching degree between medical entities and examination data, which helps to solve problems such as over-concentration of information and lack of dynamic adjustment in medical text parsing, and is particularly suitable for text parsing of cardiovascular examination results;
[0097] 2). Use a dynamic weight strategy to adjust the initial matching weight to output the first matching weight P1, query vector q i and key vector k j are usually matched using dot product similarity, but generally it is defaulted that all query vectors have the same importance for all key vectors, that is, the query-key weight is fixed, without considering the semantic changes between different medical terms and examination data. And the softmax mechanism will make the attention weights concentrate on a small number of specific values, resulting in the model being unable to reasonably utilize all relevant information. Therefore, in this embodiment, a dynamic weight strategy is used to allow the query vector to partially inherit the key vector information, avoid information loss, and make the query-key matching more flexible. The first matching weight P1 is expressed as:
[0098] P1 = αP k +(1 - α)·softmax(βP k )
[0099] Among them, α and β are adaptive adjustment parameters. Dynamically adjusting the matching method of the query-key ensures the asymmetry of the local matching matrix, which is more in line with the structural characteristics of cardiovascular medical entities, beneficial to local aggregation and focusing on local semantic information, so as to achieve accurate modeling of semantic dependencies in complex cardiovascular medical records;
[0100] 3). Adjust the initial matching weight using weighted cosine similarity and output the second matching weight P2. The traditional method uses simple cosine similarity or dot product to calculate the match. Although this method is effective, it cannot capture the non-linear relationship between the query vector and the key vector, which may lead to information loss at the semantic level and over-rely on vector normalization. Therefore, in this embodiment, a non-linear transformation is introduced, and weighted cosine similarity is used to enhance medical entity matching. The second matching weight P2 is expressed as:
[0101]
[0102] Among them, γ is a balance parameter, and tanh(q i -k j ) is used to capture the offset information between the query vector and the key vector; especially when there are slight changes in the matching of some medical terms under different examinations, this deviation can be better modeled. This method can effectively enhance the robustness of local features and avoid over-relying on traditional attention calculations.
[0103] 4). Perform weighted summation on the first matching weight P1 and the second matching weight P2, and output the attention result P local corresponding to local aggregation, which is expressed as:
[0104] P local = λP1+(1 - λ)·P2
[0105] Among them, λ∈[0,1] is the attention fusion control parameter, which is used to balance between normalized scoring and semantic matching enhancement.
[0106] 2. Obtain the attention result corresponding to adjacent diffusion, specifically including:
[0107] 1). Construct a segment propagation guiding term using the structural relationship In this embodiment, in order to aggregate the information of other paragraphs and more fully explore the internal connections in cardiovascular disease medical records, a cross-sentence and cross-paragraph information completion method is used to ensure that medical concepts in adjacent paragraphs can interact with each other, such as calculating the relationship between examinations and diagnoses, and between diagnoses and treatments. By identifying cross-paragraph medical entities, information loss caused by syntactic changes is avoided, so as to ensure the logical coherence of different paragraphs of medical texts. In this embodiment, an adjacent segment information propagation mechanism is introduced, and first, a segment propagation guiding term is constructed using the structural relationship Segment propagation guiding term Describes the distance relationship between a certain medical entity and the medical entities in adjacent paragraphs, as a guiding feature for adjacent paragraph attention, and the segment propagation guiding item It is expressed as:
[0108]
[0109] Among them, PosEnc(i) represents the encoding of the token based on position i, and w k = 1 / (d ik + ε) is the distance-based weight adjustment factor, d ik is the distance between token i and the k-th paragraph, and E k [i] is the position of the medical entity represented by PosEnc(i), and E k [i] = PosEnc(i);
[0110] 2). Based on the above-mentioned segment propagation guiding item Calculate the attention result P inter corresponding to adjacent diffusion, which is expressed as:
[0111]
[0112] Among them, W q , W s , W k , W p are learnable linear mapping matrices. In the above formula, the left term is a linear combination of the query vector and its paragraph distance feature; the right term is a combination of the key vector and its position embedding, which comprehensively considers semantic similarity and inter-segment structure information, effectively improving the model's ability to model the entity association between adjacent paragraphs.
[0113] 3. Obtain the attention result corresponding to global association, specifically including:
[0114] 1). Construct a global guiding vector G q [i] based on the distribution shift of medical entities between paragraphs. The global association attention module aims to make up for the problem that the adjacent diffusion strategy can only handle adjacent paragraphs, and further enhance the model's ability to model the dependence relationship of medical entities between non-consecutive but semantically related paragraphs. Especially in medical record texts, certain medical entity relationships often span multiple paragraphs and have long-distance dependencies. Therefore, a global cross-segment attention mechanism is introduced, and the global guiding vector G q [i] is expressed as:
[0115]
[0116] Among them, U k [i] represents the global position encoding of the paragraph where the i-th key is located, L is the total number of medical record paragraphs, and μ mis the global offset importance weight of paragraph m with respect to paragraph i, where μ m is expressed as follows:
[0117]
[0118] 2). Based on the global guiding vector G q [i] and the query vector, establish the global attention query representation which is expressed as:
[0119]
[0120] where R q , R g are trainable linear transformation matrices used to control the degree of fusion between the original semantics and the global offset; represents element-wise multiplication;
[0121] 3). According to the global attention query representation and the key vector, obtain the attention result P corresponding to the global association global , which is expressed as:
[0122]
[0123] where R k is the key vector mapping matrix to ensure dimension alignment with the query vector. The global association injects the cross-segment distance relationship guidance into the attention weights, realizing the dynamic association matching of non-adjacent but semantically related medical entities in multi-paragraph texts, thereby achieving precise modeling of semantic dependencies in complex cardiovascular medical records.
[0124] Furthermore, after obtaining the attention results corresponding to local aggregation, adjacent diffusion, and global association, these attention results are uniformly processed as follows:
[0125]
[0126] Then, perform unified softmax normalization to obtain the weight vector:
[0127]
[0128] where d represents the dimension of the hidden state, and finally it is used to calculate the corresponding information feature with the value vector v:
[0129] output = ∑p ij ·v j
[0130] This output will continue to be processed by subsequent processing modules such as the generator of large language models to generate high-quality semi-structured data. In S2, the dynamic information interaction modeling strategy optimizes the attention mechanism in the large language model, enabling the model to simultaneously focus on short-range, near-range, and far-range medical entity dependencies, and improving the accuracy of cross-span reasoning and entity matching in medical scenarios.
[0131] In one embodiment, S3 includes:
[0132] S3.1: Calculate the importance scores of each medical entity in the semi-structured medical record using the YAKE algorithm, and sort the importance scores from largest to smallest. The method for calculating the importance scores is as follows:
[0133] S(w j , d i ) = YAKE(w j , d i )
[0134] where w j is a medical term, d i is the d i th semi-structured medical record, and S(w j , d i ) is the importance score of the medical term w j in the semi-structured medical record d i ;
[0135] S3.2: According to the sorting result, select p medical entities with scores higher than 0.5, and use the clustering method to cluster the p medical entities into q medical factors. The clustering method can be k-means. After clustering, the q medical factors are as follows:
[0136] {C1, C2,..., C q}, q ≤ p
[0137] Each medical factor corresponds to a node. Then, perform binary processing on each medical factor as follows:
[0138]
[0139] C i is a binary variable.
[0140] S3.3: Use conditional independence testing to calculate the causal relationships of the medical factors after binary processing. If two medical factors are independent under the given variable conditions, remove the direct causal relationship between the two medical factors, and determine the causal direction of the two medical factors according to the V-structure rule. Otherwise, determine the causal direction according to sampling, thereby forming a partial directed graph. The conditional independence testing is as follows:
[0141] If two medical factors A and B are independent given a certain condition, then the edge between A and B should be removed in the causal graph.
[0142]
[0143] For example, a high-salt diet (A) directly causes hypertension (C), and hypertension (C) directly increases the risk of heart disease (B). Then, after controlling for hypertension (C), the high-salt diet (A) and heart disease (B) are independent, and the direct connection between A and B should be removed.
[0144] Meanwhile, in this embodiment, the V-structure rule is used to determine the causal direction of these two medical factors, specifically including: when it is calculated by conditional independence test that the first medical factor and the second medical factor are independent given a certain variable, and there is a third medical factor that makes the first medical factor and the second medical factor correlated, then the third medical factor serves as the intermediate point between the first medical factor and the second medical factor, and the causal direction is: the first medical factor points to the third medical factor, and the second medical factor points to the third medical factor. Specifically, if A and C are independent, and B makes them correlated:
[0145] A→B←C
[0146] Then B is the causal intermediate point, that is, A→B and B←C cannot be removed. All edges (such as V-structure) whose directions are determined by conditional independence and structural rules are marked as "determined edges" with fixed directions.
[0147] For example, medical factors: A. Hypertension (independent risk factor); C. Smoking (independent risk factor); B. Myocardial infarction (common effect);
[0148] Structural relationship: Hypertension (A)→Myocardial infarction (B), Smoking (C)→Myocardial infarction (B);
[0149] Independence: In the case where myocardial infarction (B) is not observed, hypertension (A) and smoking (C) are usually independent risk factors. For example, in the healthy population, whether a person has hypertension has no direct association with whether they smoke.
[0150] Conditional correlation: When myocardial infarction (B) is observed, hypertension (A) and smoking (C) become correlated. This is because both are risk factors for myocardial infarction. In patients with myocardial infarction, the distribution of hypertension and smoking may show a correlation (for example, smokers among patients with myocardial infarction may have a higher prevalence of hypertension). Therefore, the causal graph structure should be: A→B←C.
[0151] S3.4: Sample the causal directions corresponding to the edges with undetermined causal directions in the partial directed graph with equal probability, and remove the bidirectional edges that appear during the sampling process. Repeat the sampling process until every causal direction of all the edges with undetermined causal directions is traversed, and output m candidate causal graphs. The probability distribution of the equal-probability sampling is as follows:
[0152]
[0153] where, during sampling, for bidirectional edges represents an unobserved common cause. Therefore, delete such edges in each sampled graph. Repeat the sampling process until every causal direction of all the edges with undetermined causal directions is traversed, and output m candidate causal graphs {G1, G2,..., G m}.
[0154] S3.5: Calculate the matching score between each candidate causal graph and the data X based on the Bayesian Information Criterion, and take the candidate causal graph corresponding to the highest score as the final causal graph. Among them, the data X is the data set used to fit the causal graph, that is, the causal factors extracted from the medical records. The formula for calculating the matching score is as follows:
[0155]
[0156] where, L(G q , X) is the maximum log-likelihood estimate (MLE), representing the fitting degree of the causal graph G q to the data X, k is the number of free parameters (i.e., the number of causal relationships), and n is the number of data samples;
[0157] Selection of the final causal graph:
[0158]
[0159] That is, select the causal graph with the highest BIC score as the final output.
[0160] S3.6: Use a search algorithm to search the final causal graph to obtain a causal chain that meets the start-end conditions. The causal chain is represented as follows:
[0161] Given a causal graph G, in the causal graph: V is the causal node (medical factor, such as "hypertension", "ST-segment elevation", "myocardial infarction"), and E is the causal relationship (directed edge)
[0162] A causal chain C is a directed path in the graph G:
[0163] C = {A1 → A2 →... → A n}, A i ∈ V
[0164] Among them, the search algorithm is a priority search algorithm, with the starting point being the cause of disease and / or examination items, and the ending point being, for example, diseases, diagnoses, and treatment results.
[0165] In one embodiment, S4 specifically includes:
[0166] S4.1: Use the BiLSTM and conditional random field model to extract medical events from the semi-structured medical records, obtain a medical event set, and standardize the medical event set in combination with the SNOMED-CT terminology library to obtain a standardized medical event set; specifically, the clinical course of cardiovascular diseases involves multiple key time nodes, including initial symptoms, examinations, diagnoses, treatments, postoperative recoveries, etc. Since electronic medical record data is usually distributed in different time periods and some medical records lack clear timestamps, constructing a complete course timeline can ensure the reasonable chronological order of medical record events. In this embodiment, key medical events are first extracted from the semi-structured medical records, including examinations (ECG, coronary angiography), diagnoses (myocardial infarction, hypertension), treatments (PCI procedure, antihypertensive drugs), postoperative recoveries (reexamination, follow-up), etc., to form a medical event set. Event recognition is performed through entity recognition using the BiLSTM + conditional random field (CRF) model. The formula is as follows:
[0167]
[0168] Among them, x is the input medical text feature vector, W is the trainable parameter matrix of the CRF, b is the bias term, and y is the predicted medical event category. Then, event standardization is carried out in combination with the SNOMED-CT terminology library to match similar medical terms and ensure a unified data format
[0169] S4.2: Use predefined time-dependent inference rules to sort the standardized medical event set, and use the time causal inference scoring model Score(A→B) to score the sorted events, delete the event sorting with a score lower than 0.5, and output the first event time sorting. Among them, the time causal inference scoring model Score(A→B) is expressed as:
[0170] Score(A→B) = λ1·C(A,B) + λ2·T(A,B) + λ3·R(A,B)
[0171] Among them, C(A,B) is the causal relationship, T(A,B) is the time rationality, R(A,B) is the context consistency, and λ1, λ2, λ3 are the weights corresponding to the causal relationship, time rationality, and context consistency respectively; the context consistency score R(A,B) is as follows:
[0172]
[0173] The input of the time score T(A,B) is two timestamps t A , t B , and a predefined reasonable time range [L, U]. The formula is as follows:
[0174]
[0175] where: μ represents the ideal time delay, and σ represents the allowable fluctuation range. A piecewise scoring rule can also be used. If t B - t A ∈[L, U], the score is 1. If it exceeds the range, it decreases linearly.
[0176] The time-dependent inference rule is: the examination should be earlier than the diagnosis (T 检查 < T 诊断 ), the diagnosis should be earlier than the treatment (T 诊断 < T 治疗 ), and there should be a postoperative follow-up after the treatment (T 治疗 < T 随访 ).
[0177] S4.3: Predict the events lacking time in the first event time sorting, output the corresponding predicted time information, and update the first event time sorting based on the predicted time information, and output the second event time sorting; this embodiment proposes event cross-mapping for predicting the missing timestamps:
[0178] T unknown = T known + μ event + σ event · ∈
[0179] where: T unknown represents the time of the event to be speculated; T known is the time of the known related event; μ event is the average time interval obtained from historical statistics; σ event is the time standard deviation; ε follows the standard normal distribution N(0, 1) and is used to simulate individual differences.
[0180] S4.4: Use a graph neural network to propagate and optimize the second event time sorting, and output the third event time sorting; this embodiment uses the graph neural network GAT for time inference optimization to ensure that the disease course time logic conforms to medical laws.
[0181]
[0182] where:
[0183] T iThe optimized timestamp for event i; N(i) is the set of historical events related to event i; W is the trainable time propagation matrix; α ij is the time weight of the influence of event j on event i.
[0184] S4.5: Optimize the third event time sorting using the time window prediction method, output the fourth event time sorting, and output the fourth time sorting by time to obtain the timeline. The time window prediction method is determined by the following formula:
[0185] [T min , T max = [T mean - kσ, T mean + kσ]
[0186] where T min is the minimum value of the time interval, i.e., the start time; T max is the maximum value of the time interval, i.e., the end time; T mean is the time mean statistically calculated from historical data; σ is the standard deviation, representing the fluctuation range of time; k is the adjustment factor. The loss function for training the entire model is as follows:
[0187]
[0188] The entire model includes an improved large language model, a causal network, and a timeline establishment network.
[0189] In one embodiment, S5 includes: fusing the output of the causal network and the output of the timeline through calling the DeepSeek large model, and using functions such as dialogue and language text generation of DeepSeek to guide the model to generate medical record text content that conforms to clinical expression norms, has medical professionalism and readability, and ensures that the generated results meet the actual application scenarios of doctor reading and auxiliary decision-making.
[0190] Embodiment 2
[0191] The embodiment of the present invention also provides a computer-readable storage medium, on which a program of a method for generating a medical record of cardiovascular diseases based on a large language model proposed in Embodiment 1 is stored. When the program is executed, the steps of the method for generating a medical record of cardiovascular diseases based on a large language model are executed.
[0192] The same or similar reference numerals correspond to the same or similar components; the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this patent;
[0193] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the claims of the present invention.
Claims
1. A medical record generation method for cardiovascular diseases based on large language models, characterized in that, The method includes the following steps: Obtain the medical records of patients with cardiovascular diseases and preprocess the medical records of the patients; Input the preprocessed medical records of the patients into an improved large language model to output semi-structured medical records, wherein the improved large language model is obtained by replacing the self-attention mechanism with a dynamic information interaction attention mechanism; Input the semi-structured medical records into a causal network to extract the causal relationship between paragraphs and output a causal chain; Establish a timeline based on the semi-structured medical records and the extracted causal relationship; Fuse the causal chain and the timeline to output a structured medical record.
2. The medical record generation method for cardiovascular diseases based on large language models according to claim 1, wherein, The preprocessing includes: segmenting the medical records of the patients by category using syntactic dependency relations and a medical term dictionary to obtain the preprocessed medical records of the patients, wherein the categories include symptoms, examinations, diagnoses, treatments, and medications.
3. A medical record generation method for cardiovascular diseases based on a large language model according to claim 1, characterized in that, Inputting the preprocessed medical records of the patients into an improved large language model to output semi-structured medical records specifically includes: Inputting the preprocessed medical records of the patients into an encoder for semantic encoding, extracting the latent representations of each medical entity in each segment, and constructing corresponding feature vectors according to the latent representations; Inputting the feature vectors into a dynamic information interaction attention mechanism to perform local aggregation, proximity diffusion, and global association attention calculations respectively to obtain corresponding attention results; Obtaining an attention weight distribution based on the three types of attention results and a value vector; Inputting the attention weight distribution into a decoder for semantic parsing to output semi-structured medical records.
4. A method for generating medical records of cardiovascular diseases based on a large language model according to claim 3, characterized in that, The step of inputting the feature vectors into a dynamic information interaction attention mechanism to perform local aggregation attention calculation to obtain the attention result corresponding to local aggregation specifically includes: Define the initial matching weight P of medical terms k , expressed as: Adjusting the initial matching weight using a dynamic weight strategy to output a first matching weight P1, expressed as: P1 = αP k +(1 - α)·softmax(βP k ) where α and β are adaptive adjustment parameters; Adjusting the initial matching weight using a weighted cosine similarity to output a second matching weight P2, expressed as: where γ is the balance parameter, and tanh(q i -k j ) is used to capture the offset information between the query vector and the key vector; Perform a weighted sum of the first matching weight P1 and the second matching weight P2, and output the attention result P corresponding to local aggregation local , expressed as: P local = λP1 + (1 - λ)·P2 where λ ∈ [0,1] is an attention fusion control parameter.
5. The method for generating a medical record of cardiovascular disease based on a large language model according to claim 3, wherein, The step of inputting the feature vectors into a dynamic information interaction attention mechanism to perform proximity diffusion attention calculation to obtain the attention result corresponding to proximity diffusion specifically includes: Constructing segment propagation guiding items using structural relationships Expressed as: Among them, PosEnc(i) represents the encoding of the token based on position i, w k = 1 / (d ik + ε) is the weight adjustment factor based on distance, d ik is the distance between token i and the k-th paragraph, E k [i] is the position of the medical entity represented by PosEnc(i), and E k [i] = PosEnc(i); Based on the segment propagation guiding item Calculate the attention result P corresponding to the adjacent diffusion inter , expressed as: Among them, W q , W s , W k , W p is a learnable linear mapping matrix.
6. The method for generating medical records of cardiovascular diseases based on large language models according to claim 3, wherein, The step of inputting the feature vectors into a dynamic information interaction attention mechanism to perform global association attention calculation to obtain the attention result corresponding to global association specifically includes: Constructing a global guiding vector G based on the distribution shift of medical entities between paragraphs q [i], expressed as: Among them, U k [i] represents the global position encoding of the paragraph where the i-th key is located, L is the total number of medical record paragraphs, and μ m is the global offset importance weight of paragraph m to paragraph i; Based on the global guiding vector G q Establish a global attention query representation with [i] and the query vector Expressed as: where, R q , R g is a trainable linear transformation matrix for controlling the degree of fusion between the original semantics and the global offset; represents element-wise multiplication; According to the global attention query representation and the key vector, obtain the attention result P corresponding to the global association global , which is expressed as: Among them, R k is the key vector mapping matrix, ensuring alignment with the dimension of the query vector.
7. A method for generating medical records of cardiovascular diseases based on large language models according to claim 1, characterized in that The step of inputting the semi-structured medical records into a causal network to extract the causal relationship between paragraphs and output a causal chain specifically includes: Calculating the importance scores of each medical entity in the semi-structured medical records using the YAKE algorithm and sorting the importance scores from largest to smallest; According to the sorting result, selecting p medical entities with scores higher than 0.5, clustering the p medical entities into q medical factors using a clustering method, corresponding each medical factor to a node, and performing binary processing on each medical factor; Calculate the causal relationships of each medical factor after binarization using conditional independence tests. If two medical factors are independent under the given variable conditions, remove the direct causal relationship between the two medical factors and determine the causal direction of the two medical factors according to the V-structure rule. Otherwise, determine the causal direction based on sampling, thereby forming a partial directed graph; Sample the corresponding causal directions for the edges with undetermined causal directions in the partial directed graph with equal probability and remove the bidirectional edges that appear during the sampling process. Repeat the sampling process until every causal direction of all the edges with undetermined causal directions is traversed, and output m candidate causal graphs; Based on the Bayesian information criterion, calculate the matching scores of each candidate causal graph with the data X, and take the candidate causal graph corresponding to the highest score as the final causal graph; Use a search algorithm to search the final causal graph to obtain a causal chain that meets the start-end conditions.
8. A method for generating medical records of cardiovascular diseases based on large language models according to claim 7, characterized in that, The determination of the causal direction of the two medical factors according to the V-structure rule specifically includes: when it is calculated by conditional independence test that the first medical factor and the second medical factor are independent under the given variable conditions, and there is a third medical factor that makes the first medical factor and the second medical factor related, then the third medical factor serves as the intermediate point between the first medical factor and the second medical factor, and the causal direction is: the first medical factor points to the third medical factor, and the second medical factor points to the third medical factor.
9. A medical record generation method for cardiovascular diseases based on a large language model according to claim 1, characterized in that, The establishment of a timeline based on the semi-structured medical records and the extracted causal relationships specifically includes: Use a BiLSTM and conditional random field model to extract medical events from the semi-structured medical records to obtain a medical event set, and standardize the medical event set in combination with the SNOMED-CT terminology library to obtain a standardized medical event set; Use predefined time-dependent inference rules to sort the events in the standardized medical event set, and use the time causal inference scoring model Score(A→B) to score the sorted events, and delete the event sorting with a score lower than 0.5, and output the first event time sorting, where the time causal inference scoring model Score(A→B) is expressed as: Score(A→B) = λ1·C(A,B) + λ2·T(A,B) + λ3·R(A,B) where C(A,B) is the causal relationship, T(A,B) is the time rationality, R(A,B) is the context consistency, and λ1, λ2, λ3 are the weights corresponding to the causal relationship, time rationality, and context consistency respectively; Predict the events lacking time in the first event time sorting, output the corresponding predicted time information, and update the first event time sorting based on the predicted time information to output the second event time sorting; Use a graph neural network to perform propagation optimization on the second event time sorting to output the third event time sorting; Use the time window prediction method to optimize the third event time sorting to output the fourth event time sorting, and output the timeline according to the time of the fourth time sorting.
10. The method for generating a medical record of cardiovascular diseases based on a large language model according to claim 9, characterized in that, The time window prediction method is determined by the following formula: [T min , T max = [T mean -kσ, T mean +kσ] Among them, T min is the minimum value of the time interval, i.e., the start time; T max is the maximum value of the time interval, i.e., the end time; T mean is the time mean value statistically obtained from historical data; σ is the standard deviation, representing the fluctuation range of time; k is the adjustment factor.
Citation Information
Cited By
Medical record knowledge graph construction method and device based on event triggering and medium
CN121460217A