Patient characterization similarity recognition method and system based on knowledge graph
By constructing a dynamic medical knowledge graph and diffusion model, combined with spatiotemporal attention networks and sparse coding, the problem of insufficient spatiotemporal modeling in existing patient representation similarity recognition methods is solved, and efficient and accurate patient similarity recognition and model adaptability are achieved.
Patent Information
- Application Number
- CN202510690735.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing patient representation similarity recognition methods based on knowledge graphs cannot dynamically model the spatiotemporal evolution of the patient's disease course, lack explicit modeling of time dependence, resulting in insufficient sensitivity to acute events, and lack of sparse regularization during feature extraction, which affects diagnostic accuracy and model generalization.
By constructing a dynamic medical knowledge graph, combining diffusion models and meta-learning, we generate multimodal synthetic data that conforms to the real distribution, dynamically adjust the sparse coding strategy, utilize spatiotemporal attention networks and medical prior rules to capture the spatiotemporal dependencies of nodes and edges, and construct a spatiotemporal-aware dynamic sparse encoder to optimize feature compression and explanatory analysis.
It enhances the model's sensitivity to emergencies, improves diagnostic accuracy and model generalization capabilities, reduces noise interference, improves data security and model adaptability, and adapts to the data distribution of different institutions.
Smart Images

Figure CN120598019A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of knowledge graph technology, and specifically refers to a patient representation similarity recognition method and system based on knowledge graph. Background Art
[0002] With the widespread use of electronic health records, how to effectively extract valuable information from large amounts of patient data has become an important research direction. Traditional patient similarity analysis methods often rely on simple statistical indicators or specific disease characteristics, which are difficult to fully and accurately reflect the patient's comprehensive condition.
[0003] However, existing methods for similarly identifying patient representations based on knowledge graphs still have certain flaws. Existing methods for similarly identifying patient representations based on knowledge graphs are based solely on static entities and relationships and cannot dynamically model the spatiotemporal evolution of a patient's disease course. Existing medical knowledge graphs mostly focus on static entity relationships such as diseases and drugs and lack explicit modeling of time dependencies. Similarity matching is based on static medical records or single-point data, without considering dynamic edge weights and historical states between nodes, resulting in insufficient sensitivity to acute events. Existing knowledge graph reasoning often relies on pure data-driven reasoning and does not use the probability distribution or entity relationships in clinical guidelines as constraints, which can easily lead to conclusions that contradict medical logic. Fixed basis matrices and sparsity are used and cannot be dynamically adjusted according to medical scenarios, resulting in over-compression of key lesion areas and affecting diagnostic accuracy. The lack of sparsity regularization and KL divergence constraints in feature extraction results in a large amount of noise in the compressed features, reducing the reliability of similarity identification. Due to insufficient data, grassroots hospitals rely on artificial generation or simple interpolation to generate synthetic data, resulting in significant differences in distribution from real data and affecting model generalization. Therefore, a method and system for similarly identifying patient representations based on knowledge graphs are proposed. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for identifying similarity of patient representations based on knowledge graphs to solve the problems raised in the above background technology.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying similar patient representations based on a knowledge graph, comprising the following steps:
[0006] S1. Acquire and preprocess multimodal data information and extract medical prior rules;
[0007] S2. Generate multimodal synthetic data that conforms to the real distribution through the diffusion model, and constrain it through medical prior rules;
[0008] S3. Build a dynamic medical knowledge graph and quickly adapt it to small amounts of data from different institutions and patient groups through meta-learning;
[0009] S4. By constructing a patient's disease trajectory diagram, encoding multimodal medical data into spatiotemporal dynamic features, and calculating the spatiotemporal attention weights between nodes, and then combining medical priors to update the node state and generate the hidden state at the next moment;
[0010] S5. Based on the spatiotemporal dynamic characteristics of patient representation, a spatiotemporal-aware dynamic sparse encoder is constructed, and the compression strategy is dynamically adjusted according to the medical scenario;
[0011] S6. Explain the similarity basis based on feature importance analysis and form a closed-loop optimization with clinical feedback.
[0012] Among them, the S2 obtains the multimodal data preprocessed by S1 and inputs them into the encoder respectively, dynamically fuses the multimodal features through the double cross attention mechanism, and introduces the medical prior constraints as conditional input in the diffusion process of the diffusion model. Through the dual-granularity conditional guidance strategy, the global prior and local prior are modeled at the same time, and the medical prior rules extracted by S1 are converted into conditional vectors that can be processed by the model. The entities, relationships and attributes of the knowledge graph are encoded into vectors, and the probability distribution in the clinical guidelines is used as the generation condition. The gradient signal of the classifier is introduced in the diffusion process, and the regularization loss is used to force the generated data to match the distribution of the real data under the medical prior constraints.
[0013] Said S2 gradually adds noise to the multimodal data x0 in the forward diffusion process to generate a noise sequence x1, x2, ..., x T , and finally get the Gaussian noise distribution. In the reverse denoising process, the noise x is gradually reduced by iterative denoising. T Restore to x0 and associate different modal features through the cross-modal attention mechanism; minimize the mean square error between the predicted noise and the real noise, combine MMD regularization and classifier consistency loss, start from Gaussian noise, and reversely generate synthetic data through the trained diffusion model. During the generation process, medical prior rules are applied in real time for constraints, and the similarity between the generated data and the real data is evaluated through FID. According to the verification results of the generated data, the hyperparameters of the diffusion model are adjusted.
[0014] Among them, S3 builds a dynamic medical knowledge graph and quickly adapts to a small amount of data from different institutions and patient groups through meta-learning; integrates the original data of S1, the synthetic data of S2 and the historical data of the institution, extracts entities from the text, extracts relationships based on rules and deep learning, trains the meta-model on the historical data of multiple institutions, learns common knowledge across tasks, encodes institutional features as conditional inputs for meta-learning, and fine-tunes the meta-model for new institutions or patient groups with only a small amount of data.
[0015] The steps of constructing the patient's disease trajectory diagram in S4 include: obtaining multimodal synthetic data and dynamic medical knowledge graph, embedding the encoded medical entities through the knowledge graph, mapping entities such as drugs, diseases, symptoms, etc. into low-dimensional vectors through knowledge graph embedding, so that the semantic relationship between entities can be directly inferred through the distance between vectors, and the node states in the dynamic medical knowledge graph are fused according to the spatiotemporal characteristics. Contains the hidden state of the previous moment Then, the spatial features, time coding, historical state linear combination and nonlinear activation function in the dynamic medical knowledge graph are used to capture the complex spatiotemporal relationship through the weight matrix, and the spatiotemporal features are integrated with the historical hidden state. The implementation formula is:
[0016]
[0017] In the formula, represents the spatiotemporal dynamic characteristics of node i at time t, and also represents the hidden state of node i at time t. Represents the spatial features of node i at time t, through W s Mapped to the latent space, T(t) represents the time code, W s Indicates that the spatial features Mapping to the latent space, W t Indicates mapping the time code T(t) to the hidden space, and U indicates mapping the historical hidden state of the previous moment Recursively fuse to the current state, σ represents the nonlinear activation function;
[0018] in, represents the hidden state of node i at time t-1, and b represents the bias term, which is used to adjust the output range of the nonlinear activation function.
[0019] The step of calculating the spatiotemporal attention weight in S4 includes: assuming the historical hidden state of the neighbor node is According to the spatiotemporal dynamic characteristics and the historical hidden states of neighboring nodes Calculate the spatiotemporal attention weight of node i to neighbor j Spatiotemporal attention weights The implementation formula is:
[0020]
[0021] In the formula, represents the spatiotemporal attention weight of node i’s neighbor node j at time t, N i represents the neighbor set of node i, Indicates that the hidden state h of node i is transformed intoi t Mapped to the attention calculation space, Δt represents the time interval between nodes j and i, λ represents the time attenuation coefficient, and σ represents the nonlinear activation function to solve the gradient disappearance problem;
[0022] Where a represents the similarity of the calculated node pair, W represents the weight matrix that maps the hidden state to the attention calculation space, and the value of λ is adjusted according to the data information. For acute diseases, a larger value can be set to emphasize the importance of recent associations.
[0023] The steps of dynamically updating the edge weight of the patient's disease trajectory graph include: setting the historical edge weight to Combining spatiotemporal attention weights and the historical edge weight is The edge weights are dynamically updated through the forgetting coefficient, and the implementation formula is:
[0024]
[0025] In the formula, represents the dynamic edge weight between nodes i and j at time t, θ represents the forgetting coefficient, which ranges from [0, 1] and controls the degree of retention of historical weights. When θ = 0, it completely depends on the current attention weight, and when θ = 1, it completely inherits the historical weight.
[0026] Among them, θ determines the proportion of historical information retained. For example, for chronic diseases, due to their long-term impact, a larger θ value may need to be set; while for acute diseases, the opposite is true.
[0027] The step S4 of updating the node and generating the hidden state at the next moment includes: assuming that the medical prior is According to the dynamic edge weight Weighted aggregate hidden states of neighbor node j Combined with medical priors Generate the next hidden state The implementation formula is:
[0028]
[0029] In the formula, Indicates that dynamic edge weights Aggregate the hidden states of neighboring node j to capture spatiotemporal dependencies, β represents the strength of the control medical prior; the output Again as the dynamic encoding input of spatiotemporal node features.
[0030] Wherein, the step of constructing the spatiotemporal perception dynamic coefficient encoding in S5 includes: obtaining spatiotemporal dynamic characteristics Construct spatiotemporal perception dynamic coefficient coding, the dynamic coefficient coding implementation formula is:
[0031]
[0032] In the formula, X t represents the original data at time t, D t represents the dynamic basis matrix, Represents the spatiotemporal dynamic characteristics, Z t represents the sparse coefficient, ∈t represents the error term; the original data X t Through the dynamic basis matrix D t and the sparse coefficient Z t , forming a sparse code.
[0033] in, Represents the dynamic basis matrix, which is dynamically generated by the current spatiotemporal dynamic features. The column vector of the dynamic basis matrix represents the basis vector, which is used for linear combination to represent the original data. Expressed as NN D Represents a fully connected network, including two hidden layers, ct represents splicing in the specified dimension; Z t Represents the sparse coefficient, which means the original data is in the dynamic basis matrix D t The linear combination coefficient under NN Z Represents a fully connected network, with the input being The output dimension is K×N, and Sct represents the sparse activation function.
[0034] The step S5, dynamic compression strategy optimization, includes: learning the optimal D by minimizing the reconstruction error, sparse regularization and KL divergence constraint. t and Z t , the dynamic compression strategy optimization implementation formula is:
[0035]
[0036] In the formula, represents the data reconstruction error, λ(t)||Z t ||1 represents the spatiotemporal dynamic sparse regularization term, β(t) represents the KL divergence weight, which controls the sparsity constraint strength. represents KL divergence, ρ(t) represents target sparsity, Represents the actual average activation of the hidden layer, the optimized D t and Z t Further updates
[0037] Among them, the patient representation similarity recognition system based on knowledge graph includes a multimodal data acquisition and preprocessing module, a multimodal data enhancement and synthesis module, a dynamic medical knowledge graph construction and adaptation module, a spatiotemporal dynamic disease course modeling and feature extraction module, a spatiotemporal perception dynamic sparse coding and compression module, and an explainability analysis and clinical feedback module.
[0038] Compared with the prior art, the present invention has the following beneficial effects:
[0039] 1. This invention captures the spatiotemporal dependencies of nodes and edges through a spatiotemporal graph attention network. It dynamically reflects the evolution of the disease through the time decay coefficient and historical edge weights. The spatiotemporal attention weights combine the historical status of neighboring nodes and medical priors, integrating clinical logic to avoid isolated analysis of a single time point. The forgetting coefficient balances current attention with historical weights, enhancing the model's sensitivity to emergencies.
[0040] 2. This invention uses a dynamic sparse encoder to compress high-dimensional data into low-dimensional spatiotemporal features using a basis matrix and sparse coefficients. The KL divergence constraint dynamically adjusts the sparsity according to the medical scenario, optimizing resource utilization while ensuring accuracy. The reconstruction error and sparse regularization term are minimized to ensure that the compressed features retain key information and reduce noise interference. The spatiotemporal features are used as input to provide efficient feature representation for the disease trajectory diagram, accelerating the inference speed.
[0041] 3. The present invention generates data through inverse denoising using a diffusion model, and combines preprocessed data to generate a large amount of synthetic data that conforms to the real distribution, alleviating the problem of medical data scarcity. The probability distribution of clinical guidelines and the entity relationship of the knowledge graph are encoded into conditional vectors through a dual-granularity conditional guidance strategy, forcing the generated data to conform to medical logic. The dual cross-attention mechanism integrates multimodal features, and ensures the consistency of generated data between modalities through MMD regularization and classifier consistency loss, thereby improving the reliability of downstream tasks. Furthermore, through FID evaluation and hyperparameter adjustment, it can quickly respond to the difference between generated data and real data, adapt to the data distribution of different institutions, and improve the generalization ability of the model.
[0042] 4. The present invention integrates multi-source data through dynamic knowledge graphs, continuously updates medical knowledge through entity alignment and ontology expansion, and pre-trains the meta-learning framework on historical data from multiple institutions to learn common knowledge. Only a small amount of new institutional data is needed to fine-tune model parameters, reducing data annotation costs. Through federated learning, cross-institutional collaboration updates the model without sharing original data, thereby improving data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 The operation process of the patient characterization similarity recognition method based on the knowledge graph of the present invention Figure 1 ;
[0044] Figure 2 The operation process of the patient characterization similarity recognition method based on the knowledge graph of the present invention Figure 2 ;
[0045] Figure 3 The operation process of the patient characterization similarity recognition method based on the knowledge graph of the present invention Figure 3 ;
[0046] Figure 4 Schematic diagram of the structure of the patient characterization similarity recognition system based on knowledge graph of the present invention. DETAILED DESCRIPTION
[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0048] Example
[0049] See also Figures 1-4 As shown, the present invention provides a technical solution: comprising the following steps:
[0050] S1. Acquire and preprocess multimodal data information and extract medical prior rules;
[0051] S2. Generate multimodal synthetic data that conforms to the real distribution through the diffusion model, and constrain it through medical prior rules;
[0052] S3. Build a dynamic medical knowledge graph and quickly adapt it to small amounts of data from different institutions and patient groups through meta-learning;
[0053] S4. By constructing a patient's disease trajectory diagram, encoding multimodal medical data into spatiotemporal dynamic features, and calculating the spatiotemporal attention weights between nodes, and then combining medical priors to update the node state and generate the hidden state at the next moment;
[0054] S5. Based on the spatiotemporal dynamic characteristics of patient representation, a spatiotemporal-aware dynamic sparse encoder is constructed, and the compression strategy is dynamically adjusted according to the medical scenario;
[0055] S6. Explain the similarity basis based on feature importance analysis and form a closed-loop optimization with clinical feedback.
[0056] Among them, the S2 obtains the multimodal data preprocessed by S1 and inputs them into the encoder respectively, dynamically fuses the multimodal features through the double cross attention mechanism, and introduces the medical prior constraints as conditional input in the diffusion process of the diffusion model. Through the dual-granularity conditional guidance strategy, the global prior and local prior are modeled at the same time, and the medical prior rules extracted by S1 are converted into conditional vectors that can be processed by the model. The entities, relationships and attributes of the knowledge graph are encoded into vectors, and the probability distribution in the clinical guidelines is used as the generation condition. The gradient signal of the classifier is introduced in the diffusion process, and the regularization loss is used to force the generated data to match the distribution of the real data under the medical prior constraints.
[0057] Said S2 gradually adds noise to the multimodal data x0 in the forward diffusion process to generate a noise sequence x1, x2, ..., x T , and finally get the Gaussian noise distribution. In the reverse denoising process, the noise x is gradually reduced by iterative denoising. T Restore to x0 and associate different modal features through the cross-modal attention mechanism; minimize the mean square error between the predicted noise and the real noise, combine MMD regularization and classifier consistency loss, start from Gaussian noise, and reversely generate synthetic data through the trained diffusion model. During the generation process, medical prior rules are applied in real time for constraints, and the similarity between the generated data and the real data is evaluated through FID. According to the verification results of the generated data, the hyperparameters of the diffusion model are adjusted.
[0058] Among them, S3 builds a dynamic medical knowledge graph and quickly adapts to a small amount of data from different institutions and patient groups through meta-learning; integrates the original data of S1, the synthetic data of S2 and the historical data of the institution, extracts entities from the text, extracts relationships based on rules and deep learning, trains the meta-model on the historical data of multiple institutions, learns common knowledge across tasks, encodes institutional features as conditional inputs for meta-learning, and fine-tunes the meta-model for new institutions or patient groups with only a small amount of data.
[0059] The steps of constructing the patient's disease trajectory diagram in S4 include: obtaining multimodal synthetic data and dynamic medical knowledge graph, embedding the encoded medical entities through the knowledge graph, mapping entities such as drugs, diseases, symptoms, etc. into low-dimensional vectors through knowledge graph embedding, so that the semantic relationship between entities can be directly inferred through the distance between vectors, and the node states in the dynamic medical knowledge graph are fused according to the spatiotemporal characteristics. Contains the hidden state of the previous moment Then, the spatial features, time coding, historical state linear combination and nonlinear activation function in the dynamic medical knowledge graph are used to capture the complex spatiotemporal relationship through the weight matrix, and the spatiotemporal features are integrated with the historical hidden state. The implementation formula is:
[0060]
[0061] In the formula, represents the spatiotemporal dynamic characteristics of node i at time t, and also represents the hidden state of node i at time t. Represents the spatial features of node i at time t, through W s Mapped to the latent space, T(t) represents the time code, W s Indicates that the spatial features Mapping to the latent space, W t Indicates mapping the time code T(t) to the hidden space, and U indicates mapping the historical hidden state of the previous moment Recursively fuse to the current state, σ represents the nonlinear activation function;
[0062] in, represents the hidden state of node i at time t-1, and b represents the bias term, which is used to adjust the output range of the nonlinear activation function.
[0063] The step of calculating the spatiotemporal attention weight in S4 includes: assuming the historical hidden state of the neighbor node is According to the spatiotemporal dynamic characteristics h i t and the historical hidden states of neighboring nodes Calculate the spatiotemporal attention weight of node i to neighbor j Spatiotemporal attention weights The implementation formula is:
[0064]
[0065] In the formula, represents the spatiotemporal attention weight of node i’s neighbor node j at time t, N i represents the neighbor set of node i, Represents the hidden state of node i through the learnable matrix W Mapped to the attention calculation space, Δt represents the time interval between nodes j and i, λ represents the time attenuation coefficient, and σ represents the nonlinear activation function to solve the gradient disappearance problem;
[0066] Where a represents the similarity of the calculated node pair, W represents the weight matrix that maps the hidden state to the attention calculation space, and the value of λ is adjusted according to the data information. For acute diseases, a larger value can be set to emphasize the importance of recent associations.
[0067] The steps of dynamically updating the edge weight of the patient's disease trajectory graph include: setting the historical edge weight to Combining spatiotemporal attention weights and the historical edge weight is The edge weights are dynamically updated through the forgetting coefficient, and the implementation formula is:
[0068]
[0069] In the formula, represents the dynamic edge weight between nodes i and j at time t, θ represents the forgetting coefficient, which ranges from [0, 1] and controls the degree of retention of historical weights. When θ = 0, it completely depends on the current attention weight, and when θ = 1, it completely inherits the historical weight.
[0070] Among them, θ determines the proportion of historical information retained. For example, for chronic diseases, due to their long-term impact, a larger θ value may need to be set; while for acute diseases, the opposite is true.
[0071] The step S4 of updating the node and generating the hidden state at the next moment includes: assuming that the medical prior is According to the dynamic edge weight Weighted aggregate hidden states of neighbor node j Combined with medical priors Generate the next hidden state The implementation formula is:
[0072]
[0073] In the formula, Indicates that dynamic edge weights Aggregate the hidden states of neighboring node j to capture spatiotemporal dependencies, β represents the strength of the control medical prior; the output Again as the dynamic encoding input of spatiotemporal node features.
[0074] Wherein, the step of constructing the spatiotemporal perception dynamic coefficient encoding in S5 includes: obtaining spatiotemporal dynamic characteristics Construct spatiotemporal perception dynamic coefficient coding, the dynamic coefficient coding implementation formula is:
[0075]
[0076] In the formula, X t represents the original data at time t, D t represents the dynamic basis matrix, Represents the spatiotemporal dynamic characteristics, Z t represents the sparse coefficient, ∈t represents the error term; the original data X t Through the dynamic basis matrix D t and the sparse coefficient Z t , forming a sparse code.
[0077] in, Represents the dynamic basis matrix, which is dynamically generated by the current spatiotemporal dynamic features. The column vector of the dynamic basis matrix represents the basis vector, which is used for linear combination to represent the original data. Expressed as NN D Represents a fully connected network, including two hidden layers, ct represents splicing in the specified dimension; Z t Represents the sparse coefficient, which means the original data is in the dynamic basis matrix D t The linear combination coefficient under NN Z Represents a fully connected network, with the input being The output dimension is K×N, and Sct represents the sparse activation function.
[0078] The step S5, dynamic compression strategy optimization, includes: learning the optimal D by minimizing the reconstruction error, sparse regularization and KL divergence constraint. t and Z t , the dynamic compression strategy optimization implementation formula is:
[0079]
[0080] In the formula, represents the data reconstruction error, λ(t)||Z t ||1 represents the spatiotemporal dynamic sparse regularization term, β(t) represents the KL divergence weight, which controls the sparsity constraint strength. represents KL divergence, ρ(t) represents target sparsity, Represents the actual average activation of the hidden layer, the optimized D t and Z t Further updates
[0081] Among them, the patient representation similarity recognition system based on knowledge graph includes a multimodal data acquisition and preprocessing module, a multimodal data enhancement and synthesis module, a dynamic medical knowledge graph construction and adaptation module, a spatiotemporal dynamic disease course modeling and feature extraction module, a spatiotemporal perception dynamic sparse coding and compression module, and an explainability analysis and clinical feedback module.
[0082] How it works: It collects medical data from multiple modalities and performs preprocessing operations such as cleaning and standardization to remove noise and outliers. It then extracts medical prior rules from knowledge sources such as medical literature, clinical guidelines, and expert experience. These rules incorporate knowledge and experience in the medical field.
[0083] By inputting the preprocessed multimodal data into the encoder separately, the features of different modalities are dynamically integrated through the double cross attention mechanism. The double cross attention mechanism can capture the correlation information between different modal data. In the forward diffusion process, noise is gradually added to the multimodal data to generate a noise sequence, and finally a Gaussian noise distribution is obtained; in the reverse denoising process, the original data is gradually restored from the noise through iterative denoising; in the reverse denoising process, different modal features are associated through the cross-modal attention mechanism to minimize the mean square error between the predicted noise and the real noise. At the same time, combined with constraints such as MMD regularization and classifier consistency loss, starting from Gaussian noise, synthetic data is reversely generated through the trained diffusion model, and medical prior constraints are introduced in the diffusion process. The paper uses a bundle as a conditional input and simultaneously models global and local priors through a dual-granularity conditional guidance strategy. The extracted medical prior rules are converted into conditional vectors that the model can process. The entities, relationships, and attributes of the knowledge graph are encoded as vectors, and the probability distribution of clinical guidelines is used as the generation condition. The gradient signal of the classifier is introduced during the diffusion process, and a regularization loss is used to force the distribution of generated data to match that of real data under medical prior constraints. Simultaneously, the similarity between generated and real data is evaluated through FID, and the hyperparameters of the diffusion model are adjusted based on the verification results of the generated data. The paper also integrates raw data, synthetic data, and historical institutional data to extract entities from text, and performs relationship extraction based on rules and deep learning to construct a dynamic medical knowledge graph. Knowledge graphs can represent the relationships between medical entities, and meta-models can be trained on historical data from multiple institutions to learn general knowledge across tasks. Institutional characteristics are encoded as conditional inputs for meta-learning. For new institutions or patient groups, the meta-model is fine-tuned with only a small amount of data. Multimodal synthetic data and dynamic medical knowledge graphs are obtained to construct a patient's disease trajectory diagram. Each node represents the patient's clinical status at a certain point in time. The encoded medical entities are embedded in the knowledge graph, and node features are fused based on spatial and temporal features. The node state contains the hidden state of the previous moment. The spatial features, time encoding, and historical state are linearly combined through the weight matrix. The bias b and activation function capture the complex spatiotemporal relationship. The spatiotemporal features are fused with the historical state. Based on the spatiotemporal dynamic characteristics and the historical hidden states of neighboring nodes, the spatiotemporal attention weight of node i to neighbor j is calculated. The hidden state of node i is mapped to the attention calculation space through a learnable matrix. The time interval and time decay coefficient between nodes j and i are considered to calculate the similarity of the node pair, and a nonlinear activation function is applied to obtain the spatiotemporal attention weight. The hidden state of neighbor node j is weighted and aggregated according to the dynamic edge weight, and the hidden state of the next moment is generated in combination with medical priors.The output is used again as the dynamic encoding input of spatiotemporal node features to obtain spatiotemporal dynamic features and construct spatiotemporal-aware dynamic coefficient encoding. The original data is sparsely coded through the dynamic basis matrix and sparse coefficients. The optimal dynamic basis matrix and sparse coefficients are learned by minimizing the reconstruction error, sparse regularization and KL divergence constraints. The optimized dynamic basis matrix and sparse coefficients are further updated to realize dynamic adjustment of compression strategies according to medical scenarios. Based on feature importance analysis, the basis of patient representation similarity is explained to help doctors understand the decision-making process of the model. The output results of the model are combined with clinical feedback to form a closed-loop optimization, and the parameters and strategies of the model are adjusted according to clinical feedback.
[0084] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
[0085] The present invention and its embodiments are described above. This description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. In short, if a person skilled in the art is inspired by this and, without departing from the purpose of the present invention, designs structures and embodiments similar to this technical solution without inventiveness, they shall fall within the scope of protection of the present invention.
Claims
1. A patient representation similarity recognition method based on knowledge graph, characterized by: The following steps are involved: S1. Acquire and preprocess multimodal data information and extract medical prior rules; S2. Generate multimodal synthetic data that conforms to the real distribution through the diffusion model, and constrain it through medical prior rules; S3. Build a dynamic medical knowledge graph and quickly adapt it to small amounts of data from different institutions and patient groups through meta-learning; S4. Construct a patient's disease trajectory diagram by encoding multimodal medical data into spatiotemporal dynamic features, calculating the spatiotemporal attention weights between nodes, and then combining medical priors to update the node state and generate the hidden state at the next moment; S5. Based on the spatiotemporal dynamic characteristics of patient representation, a spatiotemporal-aware dynamic sparse encoder is constructed, and the compression strategy is dynamically adjusted according to the medical scenario; S6. Explain the similarity basis based on feature importance analysis and form a closed-loop optimization with clinical feedback.
2. The method for identifying similar patient representations based on knowledge graph according to claim 1, characterized in that: The step of constructing the patient's disease trajectory diagram in S4 includes: obtaining multimodal synthetic data and a dynamic medical knowledge graph, fusing the node states in the dynamic medical knowledge graph according to spatiotemporal features, and the node states including the hidden state at the previous moment. Then, the spatial features, time coding, historical state linear combination and nonlinear activation function in the dynamic medical knowledge graph are used to capture the complex spatiotemporal relationship through the weight matrix, and the spatiotemporal features are integrated with the historical hidden state. The implementation formula is: In the formula, represents the spatiotemporal dynamic characteristics of node i at time t, and also represents the hidden state of node i at time t. Represents the spatial features of node i at time t, through W s Mapped to the latent space, T(t) represents the time code, W s Indicates that the spatial features Mapping to the latent space, W t Indicates mapping the time code T(t) to the hidden space, and U indicates mapping the historical hidden state of the previous moment Recursively fuse to the current state, σ represents a nonlinear activation function.
3. The method for identifying similar patient representations based on knowledge graph according to claim 2, characterized in that: The step of calculating the spatiotemporal attention weight in S4 includes: assuming the historical hidden state of the neighbor node is Through spatiotemporal dynamic characteristics and the historical hidden states of neighboring nodes Calculate the spatiotemporal attention weight of node i to neighbor j Spatiotemporal attention weights The implementation formula is: In the formula, represents the spatiotemporal attention weight of node i’s neighbor node j at time t, N i represents the neighbor set of node i, Represents the hidden state of node i through the learnable matrix W Mapped to the attention calculation space, Δt represents the time interval between nodes j and i, λ represents the time decay coefficient, a represents the similarity of the calculated node pair, and σ represents the nonlinear activation function; The steps of dynamically updating the edge weight of the patient's disease trajectory graph include: setting the historical edge weight to Combining spatiotemporal attention weights and the historical edge weight is The edge weights of the patient's disease trajectory graph are dynamically updated through the forgetting coefficient θ, and the implementation formula is: In the formula, represents the dynamic edge weight of the patient's disease trajectory graph between nodes i and j at time t, and θ represents the forgetting coefficient.
4. The method for identifying similar patient representations based on knowledge graph according to claim 3, characterized in that: The step S4, updating the node and generating the hidden state at the next moment, includes: assuming that the medical prior is Dynamic edge weights based on the patient's disease trajectory graph Weighted aggregate hidden states of neighbor node j Combined with medical priors Generate the next hidden state The implementation formula is: In the formula, Indicates that dynamic edge weights Aggregate the hidden states of neighboring node j to capture spatiotemporal dependencies, β represents the strength of the control medical prior; the output Again as the dynamic encoding input of spatiotemporal node features.
5. The patient representation similarity recognition system based on knowledge graph implemented according to the method of claim 1 includes a multimodal data acquisition and preprocessing module, a multimodal data enhancement and synthesis module, a dynamic medical knowledge graph construction and adaptation module, a spatiotemporal dynamic disease course modeling and feature extraction module, a spatiotemporal perception dynamic sparse coding and compression module, and an interpretability analysis and clinical feedback module.