Clinical auxiliary decision-making system based on big data
Through the big data clinically assisted decision-making system, combined with multimodal medical data, cross-modal feature fusion and knowledge graph matching are solved, the problem of traditional methods ignoring deep medical characteristics is achieved, more accurate recommendations of similar cases are achieved, and the effectiveness of precision medicine is improved.
Patent Information
- Application Number
- CN202510168799.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The traditional recommended methods for similar cases are based only on surface feature matching, ignoring deep-level medical characteristics such as imaging data, pathological results and genetic information, making it difficult to accurately characterize the patient's condition and affecting the effect of precision medicine.
Using a clinically assisted decision-making system based on big data, data is extracted and standardized from multimodal medical data sources through the medical data acquisition module, and cross-modal feature fusion and disease representation module are used to align and fusion across modal features to generate high-dimensional disease representation vectors, and match them through the knowledge graph to screen cases with high similarity.
By combining deep-seated features such as imaging and genes, more accurate representations of the disease can be generated, which improves the accuracy and adaptability of recommendations for similar cases and enhances the application capabilities of precision medicine.
Smart Images

Figure CN120108696A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and in particular to a clinical decision support system based on big data. Background Art
[0002] In the modern medical system, the clinical decision support system (CDSS) relies on clinical guidelines, expert experience and patient data to provide doctors with diagnostic advice, medical advice reminders and rational medication tips; in the diagnosis and treatment of complex cases, CDSS can help doctors quickly screen possible diseases, reduce the risk of misdiagnosis and improve the standardization of medical services.
[0003] In terms of recommending similar cases, traditional methods mainly rely on structured data, such as age, gender, main symptoms, and laboratory test results for matching. This method ignores deep medical characteristics such as imaging data, pathological results, and genetic information, resulting in the inability to accurately characterize the patient's condition. In addition, most traditional methods use feature similarity calculations based on rule libraries, which makes it difficult to capture the complexity of the patient's condition, resulting in the recommended cases may not fully match the patient's actual condition.
[0004] To solve these problems, some traditional methods introduce expert knowledge or optimize rule bases. Doctors can manually adjust matching rules or optimize them in combination with the diagnosis and treatment pathways of specific diseases. However, this method is still limited by the lag in knowledge base updates and the difficulty in generalizing rules, making it difficult to adapt to the needs of personalized medicine. In the face of some rare diseases, due to the lack of sufficient case accumulation, the matching method based on the rule base cannot provide effective recommendations. Therefore, in the absence of in-depth pathological feature correlation analysis, the accuracy and adaptability of traditional similar case recommendation methods are still greatly restricted. There is an urgent need for a clinical decision-making support method based on big data and multimodal fusion to enhance the application capabilities of CDSS in precision medicine. Summary of the invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a clinical decision support system based on big data to solve the problem that traditional similar case recommendations are based only on surface feature matching, ignoring deep features such as images and genes, making it difficult to accurately characterize the condition and affecting the effect of precision medicine.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] The embodiment of the present invention provides a clinical decision support system based on big data, which includes:
[0009] Medical data collection module, which is responsible for extracting and standardizing data from multimodal medical data sources, including electronic medical records (EMR), electronic health records (EHR), laboratory results (LIS), medical imaging (PACS), etc.;
[0010] The cross-modal feature fusion and condition representation module fuses the standardized data collected in the medical data collection module to generate a high-dimensional condition representation vector for condition matching;
[0011] In the cross-modal feature fusion and disease condition representation module:
[0012] A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment.
[0013] Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images.
[0014] Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space.
[0015] Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation;
[0016] The knowledge graph is constructed using RDF triples based on clinical guidelines and historical cases to match the disease representation vector with the knowledge graph;
[0017] The similar case screening module screens highly similar cases from the historical case database based on the condition representation and knowledge graph matching results.
[0018] As a preferred solution of the clinical decision support system based on big data described in the present invention, the auxiliary method of the clinical decision support system is:
[0019] Step S1, extracting patient information from electronic medical records and medical images, and performing standardization, denoising and normalization processing, the patient information includes structured data and unstructured data;
[0020] Step S2, using a bidirectional LSTM deep learning framework based on a self-attention mechanism, the medical record text structured data extracted in step S1 is fused with the image-based unstructured data to generate a representation of the patient's condition;
[0021] The method of fusing the medical record text structured data extracted in step S1 with the image unstructured data includes:
[0022] A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment.
[0023] Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images.
[0024] Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space.
[0025] Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation;
[0026] Step S3, using a method based on RDF triples to construct a knowledge graph including diseases, symptoms, examinations and treatments, and matching the patient's condition representation;
[0027] Step S4, based on the semantic matching result generated in step S3, a high-dimensional disease vector representation is formed to match similar cases;
[0028] Step S5, dynamically adjust the weight of the recommended cases based on the patient's current diagnosis and treatment stage, including initial diagnosis, treatment, and recovery.
[0029] As a preferred solution of the clinical decision support system based on big data described in the present invention, in step S1, the steps of standardization, denoising and normalization processing include:
[0030] Structured text data including age, sex, main symptoms, laboratory test results, and diagnosis were extracted from electronic medical records;
[0031] Extract unstructured data from medical imaging data and use convolutional neural networks for deep feature learning and feature extraction;
[0032] By adopting standardization, data cleaning and dimensionality reduction methods, different modal data are uniformly encoded to form a standardized feature vector.
[0033] As a preferred solution of the clinical decision support system based on big data described in the present invention, the step of using a bidirectional LSTM deep learning framework based on a self-attention mechanism to fuse the medical record text structured data extracted in step S1 with the image-based unstructured data to generate a patient's condition representation is as follows:
[0034] The medical record text structured data X extracted in step S1 t and image-based unstructured data X i Extract features and convert them into vector representations that can be input into the model.
[0035] Perform bidirectional LSTM encoding of medical record text data, and the decoding process is expressed as:
[0036]
[0037] Among them, h t represents the hidden state matrix of the medical record text data, BiLSTM(·) is a bidirectional long short-term memory network, X t Represents the input medical record text data, and Represent the forward and backward LSTM respectively, represents the concatenation operation of the forward and backward LSTM hidden states,
[0038] ResNet feature extraction is performed on image data, and the extraction formula is:
[0039] f i =ResNet(X i ),
[0040] Among them, f i represents the deep feature vector of image data, ResNet(·) represents the pre-trained ResNet50 convolutional neural network for feature extraction, and X i Represents the input image data,
[0041] The self-attention mechanism is used to calculate the feature alignment weight and the attention representation of the medical record text features. The calculation formula is:
[0042]
[0043] Among them, H′ t is the attention feature matrix of text data, Q t =W q h t is the query matrix of text data, K t =W k h t is the key matrix of text data, V t =W v h t is the value matrix of text data, W q ,W k ,W v is the trainable weight matrix, d k is the dimension of the key matrix, softmax(·) represents the normalized attention weight calculation,
[0044] Calculate the attention representation of the image data features, the calculation formula is:
[0045]
[0046] Among them, F′ i is the attention feature matrix of the image data, Q i =W′ q f iis the query matrix of the image data, K i =W′ k f i is the key matrix of the image data, V i =W′ v f i is the value matrix of image data, W′ q ,W′ k ,W′ v is the trainable weight matrix of image features;
[0047] The cross-modal interactive attention mechanism CMA is used to construct a fusion representation of medical record text and image data, and the interactive attention matrix of text to image is calculated. The calculation formula is:
[0048]
[0049] Among them, A t→i is the attention weight of the text on the image, W c is the trainable cross-modal transformation matrix, d c is the cross-modal feature dimension,
[0050] Calculate the interactive attention matrix of the image to the text. The calculation formula is:
[0051]
[0052] Among them, A i→t is the attention weight of the image to the text,
[0053] Fusion of cross-modal features:
[0054]
[0055] Among them, Z is the final fusion representation of the patient's condition, W z is the cross-modal feature transformation matrix, σ(·) is the nonlinear activation function, Represents a vector concatenation operation;
[0056] Variational autoencoder VAE is used for latent variable modeling, and the encoded condition is expressed as:
[0057] μ,σ 2 =f enc (Z),
[0058] Among them, μ is the mean of the encoding distribution, σ 2 is the variance of the encoding distribution, f enc (·) represents the encoder of VAE,
[0059] The sampled latent variables are:
[0060] Among them, z′ is the latent variable after reparameterization, represents the standard normal distribution,
[0061] The reconstruction condition is expressed as: Z′=f dec (z′),
[0062] Among them, Z′ is the optimized disease condition representation, f dec (·) denotes the decoder of VAE.
[0063] As a preferred solution of the clinical decision support system based on big data described in the present invention, in step S3, the knowledge graph is dynamically expanded by combining the reasoning method based on graph neural network GNN.
[0064] At the same time, a semantic similarity calculation model based on the attention mechanism is used to match the patient condition representation generated in step S2 with the knowledge graph.
[0065] As a preferred solution of the clinical decision support system based on big data described in the present invention, the steps of constructing a knowledge graph including diseases, symptoms, examinations and treatments by using a method based on RDF triples are as follows:
[0066] RDF triples are used to represent medical knowledge, where each triple T is defined as:
[0067] T=(h,r,t),
[0068] Among them, h represents the head entity, such as disease, symptom or examination item, r represents the relationship, such as cause, manifest as or need to be examined, and t represents the tail entity, such as specific symptoms, examination methods or treatment methods.
[0069] The knowledge graph constructed includes:
[0070] Disease related relationships:
[0071] T d ={(d i ,r di ,d j )},
[0072] Among them, T d represents the set of disease-related triples, d i and d j is a disease entity, r di Represents the relationship between diseases, such as concurrent or secondary;
[0073] Symptoms related to:
[0074] T s ={(d i ,r si,s j )},
[0075] Among them, T s A set of triples representing diseases and symptoms, s j is a symptom entity, r si represents a "behaves as" class relationship,
[0076] Relationship between examination and treatment:
[0077] T c ={(d i ,r ci ,c j ),(d i ,r ti ,t j )},
[0078] Among them, T c A set of triples representing examination and treatment, c j For inspection items, t j For treatment, ci Represents need to be checked, ti Representatives recommend treatment;
[0079] The TransE knowledge graph embedding method is used to map knowledge triples into a low-dimensional space. The mapping process is as follows:
[0080] t≈h+r,
[0081] Among them, h, r, t represent the vector representation of the head entity, relation, and tail entity respectively.
[0082] The optimization goal is:
[0083]
[0084] in, is the TransE loss function, γ is the loss interval hyperparameter, and (h′, r, t′) is the negative sample.
[0085] As a preferred solution of the clinical decision support system based on big data described in the present invention, the step of dynamically expanding the knowledge graph by combining the reasoning method based on the graph neural network GNN is as follows:
[0086] The graph neural network GNN is used for dynamic knowledge reasoning expansion. The message passing update formula of the graph neural network is:
[0087]
[0088] in, represents the embedding vector of node v at layer l+1, is the embedding of neighbor node u in layer l, is the set of neighbor nodes of node v, W g ,b g is the trainable parameter of GNN, σ(·) is the nonlinear activation function,
[0089] The disease diagnosis path reasoning based on GNN is formulated as follows:
[0090]
[0091] Among them, P(d) represents the inference probability of the disease, α r is the relationship weight, f GNN (h, t) is the knowledge graph similarity calculated by GNN;
[0092] The step of matching the patient condition representation generated in step S2 with the knowledge graph using a semantic similarity calculation model based on an attention mechanism includes:
[0093] Calculate the matching degree between the patient's condition representation and the knowledge graph entity, the knowledge graph entity E = {e 1 ,e 2 ,…,e n}After being embedded in the knowledge graph, it is represented as a low-dimensional vector e,
[0094] Calculate the matching degree between the patient's condition representation and the knowledge graph entity. The calculation formula is:
[0095] S(Z′,e)=cos(W s Z′,W e e),
[0096] Among them, S(Z′,e) represents the semantic matching score between the patient’s condition representation Z′ and the knowledge graph entity e, cos(·) is the cosine similarity calculation function, and W s is the feature transformation matrix of the patient's condition, Z′ is the condition representation after VAE optimization, and W e is the feature transformation matrix of the knowledge graph entity, e is the vector representation of a medical entity in the knowledge graph, which comes from the knowledge graph embedding,
[0097] The attention mechanism is used to calculate the weighted matching score and the attention distribution:
[0098]
[0099] Among them, β e is the attention weight of entity e on the patient's condition representation Z′. The numerator exp(S(Z′,e)) calculates the exponential mapping of the matching score between the patient's condition representation Z′ and entity e. The denominator ∑ e′∈Eexp(S(Z′,e′)) normalizes all knowledge graph entities. This formula uses softmax normalization.
[0100] Calculate the weighted final matching score:
[0101] S final =∑ e∈E β e S(Z′,e),
[0102] Among them, S final is the final semantic matching score of the patient’s condition representation Z′ in the knowledge graph, β e is the attention weight of entity e, S(Z′,e) is the matching score between the patient's condition representation Z′ and entity e;
[0103] The loss function is optimized using a contrastive learning-based method. The loss function based on contrastive learning is:
[0104]
[0105] Among them, L sim is the semantic matching loss function, (Z′,e + ) is a positive sample pair, indicating the correct patient condition representation and knowledge graph entity matching pair, and the numerator part exp(S(Z′,e + )) represents the matching score of the positive sample, and the denominator ∑ e∈E exp(S(Z′,e)) normalizes all possible matching entities, and the loss function uses cross entropy loss.
[0106] As a preferred solution of the clinical decision support system based on big data described in the present invention, in step S4, the method of matching similar cases is as follows:
[0107] A multi-layer similarity calculation framework is used to screen highly similar cases in the historical case database, including:
[0108] Surface feature matching: rough screening based on age, gender, symptoms and laboratory test results,
[0109] Condition representation matching: Calculate similarity based on the condition semantic vector optimized by GNN to select cases that are more consistent with the current condition.
[0110] Knowledge-enhanced matching: Combined with the disease-treatment pathway reasoning of the knowledge graph, cases on similar disease progression pathways are matched.
[0111] As a preferred solution of the clinical decision support system based on big data described in the present invention, the steps of forming a high-dimensional disease vector representation based on the semantic matching result generated in step S3 and matching similar cases are as follows:
[0112] Combine the patient's condition representation Z′ and the knowledge graph matching result S final , a multi-layer feature fusion strategy is used to construct a disease vector representation, which is expressed as:
[0113] V p =W v Z+W s S final ,
[0114] Among them, V p is the high-dimensional vector representation of the patient’s condition, W v is the transformation matrix of the disease representation, Z′ is the disease representation after VAE optimization, and W s is the transformation matrix of the matching score, S final is the final semantic matching score,
[0115] Perform nonlinear activation mapping:
[0116] V′ p =σ(W p V p +b p ),
[0117] Among them, V′ p is the final disease vector after nonlinear mapping, W p is the nonlinear transformation matrix, b p is the bias term, σ(·) is the nonlinear activation function;
[0118] For historical case database To find the most similar case, a multi-layer similarity calculation framework is used here to perform surface feature matching. The matching process is as follows:
[0119] S basic (V′ p , V d )=cos(V′ p , V d ),
[0120] Among them, S basic (V′ p , V d ) is the surface feature matching score between the patient's condition vector and the historical case vector, V d is the disease vector of historical case d, cos(·) is the cosine similarity calculation,
[0121] Symptoms match:
[0122] S deep (V′ p , V d ) = f GNN (V′ p , V d ),
[0123] Among them, S deep (V′ p , V d ) is the deep semantic matching score of the patient’s condition after GNN optimization, f GNN (·) is the similarity calculation function based on graph neural network GNN
[0124] Knowledge Enhanced Matching:
[0125] S kg (V′ p , V d )=∑ (h,r,t)∈T β r S deep (V′ p , V d ),
[0126] Among them, S kg (V′ p , V d ) is the similarity of disease treatment path reasoning combined with knowledge graph, β r is the path reasoning weight, which is learned from the knowledge graph;
[0127] Calculate the final matching score using the following formula:
[0128] S final (V′ p , V d )=λ 1 S basic +λ 2 S deep +λ 3 S kg ,
[0129] Among them, S final (V′ p , V d ) is the final matching score between the patient's condition vector and case d, λ 1 ,λ 2 ,λ 3 is the weight hyperparameter for calculating similarity at different levels,
[0130] Search for the most similar cases:
[0131]
[0132] Among them, d * For the historical case with the highest similarity, from the historical case database The case with the highest matching degree is retrieved.
[0133] As a preferred solution of the clinical decision support system based on big data described in the present invention, the step of dynamically adjusting the weight of the recommended cases in combination with the patient's current diagnosis and treatment stage, including initial diagnosis, treatment and rehabilitation, is as follows:
[0134] Define different stages of patient care, including
[0135] Initial diagnosis stage: The patient has just come to the hospital and his condition is not yet clear. He needs to match cases with similar symptoms.
[0136] Treatment stage: patients have been diagnosed and received treatment, focusing on cases similar to the current treatment path.
[0137] Recovery stage: patients are in the recovery stage, focusing on cases of recovery after similar treatment plans;
[0138] Define the weight factors for the diagnosis and treatment stage:
[0139] Λ=(λ d ,λ t ,λ r ),
[0140] Among them, λ d represents the weight of the initial diagnosis stage, λ t represents the weight of the treatment stage, λ r Represents the weight of the rehabilitation stage, and the weight satisfies the normalization constraint: d +λ t +λ r =1,
[0141] Adjust the matching weights of similar cases based on the patient's diagnosis and treatment stage:
[0142] The matching score at the initial diagnosis stage is calculated using the following formula:
[0143] S early (V′ p ,V d )=cos(V′ p ,V d ),
[0144] Among them, S early (V′ p ,V d ) is the matching score between the patient’s initial diagnosis and the historical case d, V′ p is the patient's condition vector representation, Vd is the disease vector of historical case d, cos(·) is the cosine similarity calculation function. This stage focuses on symptom similarity, so the calculation is based on the similarity of disease representation.
[0145] The matching score for the treatment phase was calculated using the formula:
[0146]
[0147] Among them, Streatment(V′ p , Vd) is the matching score of the patient's treatment stage, Td is the set of disease treatment relationship triples in the knowledge graph, (h, r, t) represents a disease treatment triple in the knowledge graph: h is the disease entity, r is the treatment relationship, t is the treatment plan entity, β r is the attention weight of the treatment path, which is learned through the cross-modal attention mechanism:
[0148]
[0149] Among them, R is the set of all possible treatment relationships, S deep (V′ p ,V d ) is calculated by GNN reasoning,
[0150] The matching score of the rehabilitation stage was calculated as follows:
[0151]
[0152] Among them, S recovery (V′ p ,V d ) is the matching score of the patient’s recovery stage, T d is the set of triples of disease treatment paths in the knowledge graph, t j The treatment plan for historical case d is: is the weight of different treatment options, obtained through knowledge graph learning:
[0153]
[0154] S kg (V′ p ,V d ) is calculated by knowledge graph reasoning;
[0155] Based on the three types of matching scores, the final case recommendation score is calculated:
[0156] S final (V′ p ,V d )=λ d Searly +λ t S treatment +λ r S recovery ,
[0157] Among them, S final (V′ p ,V d ) is the final recommended case matching score, λ d ,λ t ,λ r Dynamically adjust according to the patient's current diagnosis and treatment stage,
[0158] Select the final recommended case.
[0159]
[0160] Among them, d * The most similar case is finally recommended. A historical case database.
[0161] The beneficial effects of the present invention are as follows: the present invention extracts structured and unstructured data from electronic medical records and medical images, adopts a bidirectional LSTM deep learning framework based on a self-attention mechanism, aligns the cross-modal features of medical record text and image data, and models high-order interactive relationships through a cross-modal interactive attention mechanism CMA, so that the condition representation is more accurate and interpretable; in order to enhance the clinical rationality of condition matching, the present invention constructs a medical knowledge graph based on RDF triples, combines the graph neural network GNN to dynamically expand knowledge, and matches the patient's condition representation through a semantic similarity calculation model, thereby effectively making up for the deficiency of the traditional rule base method in capturing deep features.
[0162] The present invention adopts a multi-layer similarity calculation framework to screen similar cases, performs rough screening through surface feature matching, and then combines the GNN-optimized condition semantic vector for deep semantic matching. It matches cases with similar disease progression paths through knowledge graph reasoning to achieve comprehensive and multi-level screening of similar cases.
[0163] The present invention dynamically adjusts the matching weights in combination with the patient's current diagnosis and treatment stage, and introduces reinforcement learning to optimize the recommendation strategy, so that the system can adapt to changes in the patient's condition and improve the personalization of the recommendation.
[0164] In summary, the present invention effectively overcomes the problems of traditional methods that only rely on surface feature matching, lack of deep pathological correlation analysis, and difficulty in dynamically adjusting recommendation strategies, thereby improving the accuracy, clinical applicability and interpretability of similar case recommendations. BRIEF DESCRIPTION OF THE DRAWINGS
[0165] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0166] Figure 1 It is a schematic diagram of the framework of the clinical decision support system based on big data of the present invention. DETAILED DESCRIPTION
[0167] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0168] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0169] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0170] Example 1, reference Figure 1 , this embodiment provides a clinical decision support system based on big data, including:
[0171] The medical data collection module is responsible for extracting and standardizing data from multimodal medical data sources, including electronic medical records (EHRs) and medical images;
[0172] The cross-modal feature fusion and condition representation module fuses the standardized data collected in the medical data collection module to generate a high-dimensional condition representation vector for condition matching;
[0173] In the cross-modal feature fusion and disease condition representation module:
[0174] A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment.
[0175] Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images.
[0176] Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space.
[0177] Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation;
[0178] The knowledge graph is constructed using RDF triples based on clinical guidelines and historical cases to match the disease representation vector with the knowledge graph;
[0179] The similar case screening module screens highly similar cases from the historical case database based on the condition representation and knowledge graph matching results.
[0180] This embodiment also provides an auxiliary method for the above-mentioned clinical decision support system based on big data, including:
[0181] Step S1, extracting patient information from electronic medical records and medical images, and performing standardization, denoising and normalization processing, the patient information includes structured data and unstructured data;
[0182] In step S1, the steps of standardization, denoising and normalization processing include:
[0183] Structured text data including age, sex, main symptoms, laboratory test results, and diagnosis were extracted from electronic medical records;
[0184] Extract unstructured data from medical imaging data and use convolutional neural networks for deep feature learning and feature extraction;
[0185] Standardization, data cleaning and dimensionality reduction methods are used to uniformly encode different modal data to form standardized feature vectors;
[0186] Step S2, using a bidirectional LSTM deep learning framework based on a self-attention mechanism, the medical record text structured data extracted in step S1 is fused with the image-based unstructured data to generate a representation of the patient's condition;
[0187] The method of fusing the medical record text structured data extracted in step S1 with the image unstructured data includes:
[0188] A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment.
[0189] Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images.
[0190] Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space.
[0191] Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation;
[0192] The bidirectional LSTM deep learning framework based on the self-attention mechanism is used to fuse the medical record text structured data extracted in step S1 with the image unstructured data to generate the patient's condition representation.
[0193] The medical record text structured data X extracted in step S1 t and image-based unstructured data X i Extract features and convert them into vector representations that can be input into the model.
[0194] Perform bidirectional LSTM encoding of medical record text data, and the decoding process is expressed as:
[0195]
[0196] Among them, h t represents the hidden state matrix of the medical record text data, BiLSTM(·) is a bidirectional long short-term memory network, X t Represents the input medical record text data, and Represent the forward and backward LSTM respectively, represents the concatenation operation of the forward and backward LSTM hidden states,
[0197] ResNet feature extraction is performed on image data, and the extraction formula is:
[0198] f i =ResNet(X i ),
[0199] Among them, f i represents the deep feature vector of image data, ResNet(·) represents the pre-trained ResNet50 convolutional neural network for feature extraction, and X i Represents the input image data,
[0200] The self-attention mechanism is used to calculate the feature alignment weight and the attention representation of the medical record text features. The calculation formula is:
[0201]
[0202] Among them, H′ t is the attention feature matrix of text data, Q t =W q h t is the query matrix of text data, K t =W k h tis the key matrix of text data, V t =W v h t is the value matrix of text data, W q ,W k ,W v is the trainable weight matrix, d k is the dimension of the key matrix, softmax(·) represents the normalized attention weight calculation,
[0203] Calculate the attention representation of the image data features, the calculation formula is:
[0204]
[0205] Among them, F′ i is the attention feature matrix of the image data, Q i =W′ q f i is the query matrix of the image data, K i =W′ k f i is the key matrix of the image data, V i =W′ v f i is the value matrix of image data, W′ q ,W′ k ,W′ v is the trainable weight matrix of image features;
[0206] The cross-modal interactive attention mechanism CMA is used to construct a fusion representation of medical record text and image data, and the interactive attention matrix of text to image is calculated. The calculation formula is:
[0207]
[0208] Among them, A t→i is the attention weight of the text on the image, W c is the trainable cross-modal transformation matrix, d c is the cross-modal feature dimension,
[0209] Calculate the interactive attention matrix of the image to the text. The calculation formula is:
[0210]
[0211] Among them, A i→t is the attention weight of the image to the text,
[0212] Fusion of cross-modal features:
[0213]
[0214] Among them, Z is the final fusion representation of the patient's condition, W z is the cross-modal feature transformation matrix, σ(·) is the nonlinear activation function, Represents a vector concatenation operation;
[0215] Variational autoencoder VAE is used for latent variable modeling, and the encoded condition is expressed as:
[0216] μ,σ 2 =f enc (Z),
[0217] Among them, μ is the mean of the encoding distribution, σ 2 is the variance of the encoding distribution, f enc (·) represents the encoder of VAE,
[0218] The sampled latent variables are:
[0219] Among them, z′ is the latent variable after reparameterization, represents the standard normal distribution,
[0220] The reconstruction condition is expressed as: Z′=f dec (z′),
[0221] Among them, Z′ is the optimized condition representation, f dec (·) denotes the decoder of VAE;
[0222] Specifically, a bidirectional LSTM fusion method based on the self-attention mechanism is used here to integrate the medical record text structured data and the image unstructured data to form a unified patient condition representation:
[0223] Bidirectional LSTM and ResNet extract deep features of medical record text and image data, capturing temporal information and spatial information respectively;
[0224] The self-attention mechanism calculates the feature correlation between medical record text and image data to improve information alignment capabilities;
[0225] The cross-modal interactive attention mechanism (CMA) learns the high-order interactive relationship between text and image data to achieve more accurate cross-modal feature fusion;
[0226] Variational autoencoder (VAE) is used for latent variable modeling to improve the interpretability and robustness of disease representation;
[0227] Step S3, using a method based on RDF triples to construct a knowledge graph including diseases, symptoms, examinations and treatments, and matching the patient's condition representation;
[0228] In step S3, the knowledge graph is dynamically expanded by combining the reasoning method based on graph neural network GNN.
[0229] At the same time, a semantic similarity calculation model based on the attention mechanism is used to match the patient condition representation generated in step S2 with the knowledge graph;
[0230] Using the RDF triple-based method, the steps to construct a knowledge graph including diseases, symptoms, examinations and treatments are as follows:
[0231] RDF triples are used to represent medical knowledge, where each triple T is defined as:
[0232] T=(h,r,t),
[0233] Among them, h represents the head entity, such as disease, symptom or examination item, r represents the relationship, such as cause, manifest as or need to be examined, and t represents the tail entity, such as specific symptoms, examination methods or treatment methods.
[0234] The knowledge graph constructed includes:
[0235] Disease related relationships:
[0236] T d ={(d i ,r di ,d j )},
[0237] Among them, T d represents the set of disease-related triples, d i and d j is a disease entity, r di Represents the relationship between diseases, such as concurrent or secondary;
[0238] Symptoms related to:
[0239] T s ={(d i ,r si ,s j )},
[0240] Among them, T s A set of triples representing diseases and symptoms, s j is a symptom entity, r si represents a "behaves as" class relationship,
[0241] Relationship between examination and treatment:
[0242] T c ={(d i ,r ci ,c j ),(d i,r ti ,t j )},
[0243] Among them, T c A set of triples representing examination and treatment, c j For inspection items, t j For treatment, ci Represents need to be checked, ti Representatives recommend treatment;
[0244] The TransE knowledge graph embedding method is used to map knowledge triples into a low-dimensional space. The mapping process is as follows:
[0245] t≈h+r,
[0246] Among them, h, r, t represent the vector representation of the head entity, relation, and tail entity respectively.
[0247] The optimization goal is:
[0248]
[0249] in, is the TransE loss function, γ is the loss interval hyperparameter, (h′, r, t′) is the negative sample;
[0250] Combined with the reasoning method based on graph neural network GNN, the steps for dynamic knowledge expansion of knowledge graph are as follows:
[0251] The graph neural network GNN is used for dynamic knowledge reasoning expansion. The message passing update formula of the graph neural network is:
[0252]
[0253] in, represents the embedding vector of node v at layer l+1, is the embedding of neighbor node u in layer l, is the set of neighbor nodes of node v, W g , b g is the trainable parameter of GNN, σ(·) is the nonlinear activation function,
[0254] The disease diagnosis path reasoning based on GNN is formulated as follows:
[0255]
[0256] Among them, P(d) represents the inference probability of the disease, α r is the relationship weight, f GNN (h, t) is the knowledge graph similarity calculated by GNN;
[0257] The steps of matching the patient condition representation generated in step S2 with the knowledge graph using a semantic similarity calculation model based on an attention mechanism include:
[0258] Calculate the matching degree between the patient's condition representation and the knowledge graph entity, the knowledge graph entity E = {e 1 , e 2 , ..., e n}After being embedded in the knowledge graph, it is represented as a low-dimensional vector e,
[0259] Calculate the matching degree between the patient's condition representation and the knowledge graph entity. The calculation formula is:
[0260] S(Z′,e)=cos(W s z′,W e e),
[0261] Among them, S(Z′, e) represents the semantic matching score between the patient’s condition representation Z′ and the knowledge graph entity e, cos(·) is the cosine similarity calculation function, and W s is the feature transformation matrix of the patient's condition, Z′ is the condition representation after VAE optimization, and W e is the feature transformation matrix of the knowledge graph entity, e is the vector representation of a medical entity in the knowledge graph, which comes from the knowledge graph embedding,
[0262] The attention mechanism is used to calculate the weighted matching score and the attention distribution:
[0263]
[0264] Among them, β e is the attention weight of entity e on the patient's condition representation z′. The numerator part exp(S(Z′, e)) calculates the exponential mapping of the matching score between the patient's condition representation Z′ and entity e. The denominator part ∑ e′∈E exp(S(Z′, e′)) normalizes all knowledge graph entities. This formula uses softmax normalization.
[0265] Calculate the weighted final matching score:
[0266] S final =∑ e∈E β e S(Z′,e),
[0267] Among them, S final is the final semantic matching score of the patient’s condition representation Z′ in the knowledge graph, β e is the attention weight of entity e, S(Z′, e) is the matching score between the patient's condition representation Z′ and entity e;
[0268] The loss function is optimized using a contrastive learning-based method. The loss function based on contrastive learning is:
[0269]
[0270] Among them, L sim is the semantic matching loss function, (Z′, e + ) is a positive sample pair, indicating the correct patient condition representation and knowledge graph entity matching pair, and the numerator part exp(S(Z′, e + )) represents the matching score of the positive sample, and the denominator ∑ e∈E exp(S(Z′, e)) normalizes all possible matching entities, and the loss function uses cross entropy loss;
[0271] Specifically, in this step, a semantic similarity calculation model based on the attention mechanism is used to match the patient's condition representation with the medical entities in the knowledge graph. A matching model based on cosine similarity is used to measure the semantic similarity between the patient's condition representation and the knowledge graph entity. An attention weighting mechanism is introduced to dynamically adjust the matching weights of different medical entities to improve the matching accuracy. Contrastive learning is used to optimize the matching loss so that the model can effectively distinguish between medical entities with high matching and low matching.
[0272] Step S4, based on the semantic matching result generated in step S3, a high-dimensional disease vector representation is formed to match similar cases;
[0273] In step S4, the method of matching similar cases is as follows:
[0274] A multi-layer similarity calculation framework is used to screen highly similar cases in the historical case database, including:
[0275] Surface feature matching: rough screening based on age, gender, symptoms and laboratory test results,
[0276] Condition representation matching: Calculate similarity based on the condition semantic vector optimized by GNN to select cases that are more consistent with the current condition.
[0277] Knowledge-enhanced matching: Combined with the disease-treatment path reasoning of the knowledge graph, matching cases on similar disease progression paths,
[0278] Based on the semantic matching result generated in step S3, a high-dimensional disease vector representation is formed, and the steps for matching similar cases are as follows:
[0279] Combine the patient's condition representation Z′ and the knowledge graph matching result S final , a multi-layer feature fusion strategy is used to construct a disease vector representation, which is expressed as:
[0280] V p =W v Z′+W s S final ,
[0281] Among them, V p is the high-dimensional vector representation of the patient’s condition, W v is the transformation matrix of the disease representation, Z′ is the disease representation after VAE optimization, and W s is the transformation matrix of the matching score, S final is the final semantic matching score,
[0282] Perform nonlinear activation mapping:
[0283] V′ p =σ(W p V p +b p ),
[0284] Among them, V′ p is the final disease vector after nonlinear mapping, W p is the nonlinear transformation matrix, b p is the bias term, σ(·) is the nonlinear activation function;
[0285] For historical case database To find the most similar case, a multi-layer similarity calculation framework is used here to perform surface feature matching. The matching process is as follows:
[0286] S basic (V′ p , V d )=cos(V′ p , V d ),
[0287] Among them, S basic (V′ p , V d ) is the surface feature matching score between the patient's condition vector and the historical case vector, V d is the disease vector of historical case d, cos(·) is the cosine similarity calculation,
[0288] Symptoms match:
[0289] S deep (V′ p , V d )=f GNN (V′ p , V d ),
[0290] Among them, S deep (V′ p , Vd ) is the deep semantic matching score of the patient’s condition after GNN optimization, f GNN (·) is the similarity calculation function based on graph neural network GNN
[0291] Knowledge Enhanced Matching:
[0292] S kg (V′ p , V d )=∑ (h,r,t)∈T β r S deep (V′ p , V d ),
[0293] Among them, S kg (V′ p , V d ) is the similarity of disease treatment path reasoning combined with knowledge graph, β r is the path reasoning weight, which is learned from the knowledge graph;
[0294] Calculate the final matching score using the following formula:
[0295] S final (V′ p , V d )=λ 1 S basic +λ 2 S deep +λ 3 S kg ,
[0296] Among them, S final (V′ p , V d ) is the final matching score between the patient's condition vector and case d, λ 1 ,λ 2 ,λ 3 is the weight hyperparameter for calculating similarity at different levels,
[0297] Search for the most similar cases:
[0298]
[0299] Among them, d* is the historical case with the highest similarity, from the historical case database Retrieve the case with the highest matching degree;
[0300] Specifically, in this step, a high-dimensional disease vector representation is constructed, combining the disease representation and knowledge graph matching scores; a multi-layer similarity calculation framework is adopted, including surface feature matching, disease representation matching and knowledge enhancement matching, to improve the comprehensiveness of the matching results; deep matching based on GNN, combined with disease treatment path reasoning of the knowledge graph, further optimizes the matching results, and uses the final fusion matching score to weight the similarity calculations at different levels to improve the accuracy of the final retrieval results;
[0301] Step S5, dynamically adjusting the weight of the recommended cases in combination with the patient's current diagnosis and treatment stage, including initial diagnosis, treatment, and recovery;
[0302] Combined with the patient's current diagnosis and treatment stage, including initial diagnosis, treatment, and recovery, the steps for dynamically adjusting the weight of recommended cases are:
[0303] Define different stages of patient care, including
[0304] Initial diagnosis stage: The patient has just come to the hospital and his condition is not yet clear. He needs to match cases with similar symptoms.
[0305] Treatment stage: patients have been diagnosed and received treatment, focusing on cases similar to the current treatment path.
[0306] Recovery stage: patients are in the recovery stage, focusing on cases of recovery after similar treatment plans;
[0307] Define the weight factors for the diagnosis and treatment stage:
[0308] Λ=(λ d ,λ t ,λ r ),
[0309] Among them, λ d represents the weight of the initial diagnosis stage, λ t represents the weight of the treatment stage, λ r Represents the weight of the rehabilitation stage, and the weight satisfies the normalization constraint: d +λ t +λ r =1,
[0310] Adjust the matching weights of similar cases based on the patient's diagnosis and treatment stage:
[0311] The matching score at the initial diagnosis stage is calculated using the following formula:
[0312] S early (V′ p ,V d )=cos(V′ p ,V d ),
[0313] Among them, S early (V′ p ,V d ) is the matching score between the patient’s initial diagnosis and the historical case d, V′ p is the patient's condition vector representation, V d is the disease vector of historical case d, cos(·) is the cosine similarity calculation function. This stage focuses on symptom similarity, so the calculation is based on the similarity of disease representation.
[0314] The matching score for the treatment phase was calculated using the formula:
[0315]
[0316] Among them, S treatment (V′ p ,V d ) is the matching score of the patient’s treatment stage, T d is the set of triples of disease-treatment relations in the knowledge graph. (h, r, t) represents a disease-treatment triple in the knowledge graph: h is the disease entity, r is the treatment relation, t is the treatment plan entity, and β r is the attention weight of the treatment path, which is learned through the cross-modal attention mechanism:
[0317]
[0318] Among them, R is the set of all possible treatment relationships, S deep (V′ p ,V d Calculated by GNN reasoning,
[0319] The matching score of the rehabilitation stage was calculated as follows:
[0320]
[0321] Among them, S recovery (V′ p ,V d ) is the matching score of the patient’s recovery stage, T d is the set of triples of disease treatment paths in the knowledge graph, t j The treatment plan for historical case d is: is the weight of different treatment options, obtained through knowledge graph learning:
[0322]
[0323] S kg (V′ p ,V d ) is calculated by knowledge graph reasoning;
[0324] Based on the three types of matching scores, the final case recommendation score is calculated:
[0325] S final (V′ p ,V d )=λ d S early +λ t S treatment +λ r S recovery ,
[0326] Among them, S final (V′ p ,V d ) is the final recommended case matching score, λ d ,λ t ,λ r Dynamically adjust according to the patient's current diagnosis and treatment stage,
[0327] Select the final recommended case.
[0328]
[0329] Among them, d * The most similar case is finally recommended. It is a historical case database;
[0330] Specifically, a multi-stage matching weight adjustment is introduced here to optimize the matching strategies for the initial diagnosis, treatment and rehabilitation stages respectively. The attention weight calculation based on the knowledge graph enables the weights of different treatment plans to change dynamically. A recommendation strategy combining GNN reasoning and semantic matching is adopted to improve the matching accuracy.
[0331] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A clinical decision support system based on big data, characterized by: include, Medical data collection module, which is responsible for extracting and standardizing data from multimodal medical data sources, including electronic medical records (EMR), electronic health records (EHR), laboratory results (LIS), medical imaging (PACS), etc.; The cross-modal feature fusion and condition representation module fuses the standardized data collected in the medical data collection module to generate a high-dimensional condition representation vector for condition matching; In the cross-modal feature fusion and disease condition representation module: A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment. Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images. Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space. Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation; The knowledge graph is constructed using RDF triples based on clinical guidelines and historical cases to match the disease representation vector with the knowledge graph; The similar case screening module screens highly similar cases from the historical case database based on the condition representation and knowledge graph matching results.
2. A clinical decision support system based on big data as claimed in claim 1, characterized in that: The auxiliary modes of clinical decision support system are: Step S1, extracting patient information from electronic medical records and medical imaging systems, and performing standardization, denoising and normalization processing, wherein the patient information includes structured data and unstructured data; Step S2, using a bidirectional LSTM deep learning framework based on a self-attention mechanism, the medical record text structured data extracted in step S1 is fused with the image-based unstructured data to generate a representation of the patient's condition; The method of fusing the medical record text structured data extracted in step S1 with the image unstructured data includes: A bidirectional LSTM deep learning framework based on self-attention mechanism is used for cross-modal feature alignment. Through the cross-modal attention mechanism, we learn the high-order interactions between different texts and images. Construct a multi-layer feature fusion network to map the medical record text structured data and image unstructured data into a high-dimensional latent space. Variational autoencoder VAE and graph convolutional network GCN are used to perform hierarchical optimization of disease representation; Step S3, using a method based on RDF triples to construct a knowledge graph including diseases, symptoms, examinations, tests and treatments, and matching the patient's condition representation; Step S4, based on the semantic matching result generated in step S3, a high-dimensional disease vector representation is formed to match similar cases; Step S5, dynamically adjust the weight of the recommended cases based on the patient's current diagnosis and treatment stage, including initial diagnosis, treatment, and recovery.
3. A clinical decision support system based on big data as claimed in claim 2, characterized in that: In step S1, the steps of standardization, denoising and normalization processing include: Structured text data including age, sex, main symptoms, laboratory test results, and diagnosis were extracted from electronic medical records; Extract unstructured data from medical imaging data and use convolutional neural networks for deep feature learning and feature extraction; By adopting standardization, data cleaning and dimensionality reduction methods, different modal data are uniformly encoded to form a standardized feature vector.
4. A clinical decision support system based on big data as claimed in claim 3, characterized in that: The step of using a bidirectional LSTM deep learning framework based on a self-attention mechanism to fuse the medical record text structured data extracted in step S1 with the image-based unstructured data to generate a patient's condition representation is as follows: The medical record text structured data X extracted in step S1 t and image-based unstructured data X t Extract features and convert them into vector representations that can be input into the model. Perform bidirectional LSTM encoding of medical record text data, and the decoding process is expressed as: Among them, h t represents the hidden state matrix of the medical record text data, BiLSTM(·) is a bidirectional long short-term memory network, X t Represents the input medical record text data, and Represent the forward and backward LSTM respectively, represents the concatenation operation of the forward and backward LSTM hidden states, ResNet feature extraction is performed on image data, and the extraction formula is: f i =ResNet(X i ), Among them, f i represents the deep feature vector of image data, ResNet(·) represents the pre-trained ResNet50 convolutional neural network for feature extraction, and X i Represents the input image data, The self-attention mechanism is used to calculate the feature alignment weight and the attention representation of the medical record text features. The calculation formula is: Among them, H' t is the attention feature matrix of text data, Q t =W q h t is the query matrix of text data, K t =W k h t is the key matrix of text data, V t =W v h t is the value matrix of text data, W q ,W k ,W v is the trainable weight matrix, d k is the dimension of the key matrix, softmax(·) represents the normalized attention weight calculation, Calculate the attention representation of the image data features, the calculation formula is: Among them, F' i is the attention feature matrix of the image data, Q i =W' q f i is the query matrix of the image data, K i =W' k f i is the key matrix of the image data, V i =W' v f i is the value matrix of the image data, W' q ,W' k ,W′ v is the trainable weight matrix of image features; The cross-modal interactive attention mechanism CMA is used to construct a fusion representation of medical record text and image data, and the interactive attention matrix of text to image is calculated. The calculation formula is: Among them, A t→i is the attention weight of the text on the image, W c is the trainable cross-modal transformation matrix, d c is the cross-modal feature dimension, Calculate the interactive attention matrix of the image to the text. The calculation formula is: Among them, A i→t is the attention weight of the image to the text, Fusion of cross-modal features: Among them, Z is the final fusion representation of the patient's condition, W z is the cross-modal feature transformation matrix, σ(·) is the nonlinear activation function, Represents a vector concatenation operation; Variational autoencoder VAE is used for latent variable modeling, and the encoded condition is expressed as: m,s 2 =f enc (Z), Among them, μ is the mean of the encoding distribution, σ 2 is the variance of the encoding distribution, f enc (·) represents the encoder of VAE, The sampled latent variables are: Among them, z' is the latent variable after reparameterization, represents the standard normal distribution, The reconstruction condition is expressed as: Z′=f dec (z′), Among them, Z′ is the optimized disease condition representation, f dec (·) denotes the decoder of VAE.
5. A clinical decision support system based on big data as claimed in claim 4, characterized in that: In step S3, the knowledge graph is dynamically expanded by combining the reasoning method based on the graph neural network GNN. At the same time, a semantic similarity calculation model based on the attention mechanism is used to match the patient condition representation generated in step S2 with the knowledge graph.
6. A clinical decision support system based on big data as claimed in claim 5, characterized in that: The steps of constructing a knowledge graph including diseases, symptoms, examinations and treatments by using the RDF triple-based method are as follows: RDF triples are used to represent medical knowledge, where each triple T is defined as: T=(h,r,t), Among them, h represents the head entity, r represents the relationship, and t represents the tail entity. The knowledge graph constructed includes: Disease related relationships: T d ={(d i ,r di ,d j )), Among them, T d represents the set of disease-related triples, d i and d j is a disease entity, r di Represents the relationship between diseases; Symptoms related to: T s ={(d i ,r si ,s j )}, Among them, T s A set of triples representing diseases and symptoms, s j is a symptom entity, r si represents a "performs as" class relationship, Relationship between examination and treatment: T c ={(d i ,r ci ,c j ),(d i ,r ti ,t j )}, Among them, T c A set of triples representing examination and treatment, c j For inspection items, t j For treatment, ci Indicates that inspection is required. ti Representatives recommend treatment; The TransE knowledge graph embedding method is used to map knowledge triples into a low-dimensional space. The mapping process is as follows: t≈h+r, Among them, h, r, t represent the vector representation of the head entity, relation, and tail entity respectively. The optimization goal is: in, is the TransE loss function, γ is the loss interval hyperparameter, and (h',r,t') is the negative sample.
7. A clinical decision support system based on big data as claimed in claim 6, characterized in that: The steps of dynamically expanding the knowledge graph by combining the reasoning method based on the graph neural network GNN are as follows: The graph neural network GNN is used for dynamic knowledge reasoning expansion. The message passing update formula of the graph neural network is: in, represents the embedding vector of node v at layer l+1, is the embedding of neighbor node u in layer l, is the set of neighbor nodes of node v, W g ,b g is the trainable parameter of GNN, σ(·) is the nonlinear activation function, The disease diagnosis path reasoning based on GNN is formulated as follows: Among them, P(d) represents the inference probability of the disease, α r is the relationship weight, f GNN (h, t) is the knowledge graph similarity calculated by GNN; The step of matching the patient condition representation generated in step S2 with the knowledge graph using a semantic similarity calculation model based on an attention mechanism includes: Calculate the matching degree between the patient's condition representation and the knowledge graph entity, the knowledge graph entity E = {e1, e2, ..., e n }After being embedded in the knowledge graph, it is represented as a low-dimensional vector e, Calculate the matching degree between the patient's condition representation and the knowledge graph entity. The calculation formula is: S(Z,e)=cos(W s W',W e e), Among them, S(Z',e) represents the semantic matching score between the patient's condition representation Z' and the knowledge graph entity E, cos(·) is the cosine similarity calculation function, and W s is the feature transformation matrix of the patient's condition, Z' is the condition representation after VAE optimization, and W e is the feature transformation matrix of the knowledge graph entity, e is the vector representation of a medical entity in the knowledge graph, which comes from the knowledge graph embedding, The attention mechanism is used to calculate the weighted matching score and the attention distribution: Among them, β e is the attention weight of entity e on the patient's condition representation Z'. The numerator part exp(S(Z',e)) calculates the exponential mapping of the matching score between the patient's condition representation Z' and entity e. The denominator part ∑ e'∈E exp(S(Z',e')) normalizes all knowledge graph entities. This formula uses softmax normalization. Calculate the weighted final matching score: S final =∑ e∈E β e S(Z',e), Among them, S final is the final semantic matching score of the patient’s condition representation Z’ in the knowledge graph, β e is the attention weight of entity e, S(Z',e) is the matching score between the patient's condition representation Z' and entity e; The loss function is optimized using a contrastive learning-based method. The loss function based on contrastive learning is: Among them, L sim is the semantic matching loss function, (Z',e + ) is a positive sample pair, indicating the correct patient condition representation and knowledge graph entity matching pair, and the molecular part exp(S(Z',e + )) represents the matching score of the positive sample, and the denominator ∑ e∈E exp(S(Z',e)) normalizes all possible matching entities, and the loss function uses cross entropy loss.
8. A clinical decision support system based on big data as claimed in claim 7, characterized in that: In step S4, the method of matching similar cases is as follows: A multi-layer similarity calculation framework is used to screen highly similar cases in the historical case database, including: Surface feature matching: rough screening based on age, gender, symptoms and laboratory test results, Condition representation matching: Calculate similarity based on the condition semantic vector optimized by GNN to select cases that are more consistent with the current condition. Knowledge-enhanced matching: Combined with the disease-treatment pathway reasoning of the knowledge graph, cases on similar disease progression pathways are matched.
9. A clinical decision support system based on big data as claimed in claim 8, characterized in that: The steps of forming a high-dimensional disease vector representation based on the semantic matching result generated in step S3 and matching similar cases are as follows: Combine the patient's condition representation Z' and the knowledge graph matching result S final , a multi-layer feature fusion strategy is used to construct a disease vector representation, which is expressed as: V p =W v Z’+W s S final , Among them, V p is the high-dimensional vector representation of the patient’s condition, W v is the transformation matrix of the disease representation, Z′ is the disease representation after VAE optimization, and W s is the transformation matrix of the matching score, S final is the final semantic matching score, Perform nonlinear activation mapping: V′ p =σ(W p V p +b p ), Among them, V′ p is the final disease vector after nonlinear mapping, W p is the nonlinear transformation matrix, b p is the bias term, σ(·) is the nonlinear activation function; In order to find the most similar case in the historical case database D, a multi-layer similarity calculation framework is used here to perform surface feature matching. The matching process is: With basic (V′ p ,In d )=cos(V′ p ,In d ), Among them, S basic (V′ p , V d ) is the surface feature matching score between the patient's condition vector and the historical case vector, V d is the disease vector of historical case d, cos(·) is the cosine similarity calculation, Symptoms match: S deep (V′ p ,V d )=f GNN (V′ p ,V d ), Among them, S deep (V′ p , V d ) is the deep semantic matching score of the patient’s condition after GNN optimization, f GNN (·) is the similarity calculation function based on graph neural network GNN Knowledge Enhanced Matching: S kg (V′ p ,V d )=∑ (h,r,t)∈T β r S deep (V′ p ,V d ), Among them, S kg (V′ p , V d ) is the similarity of disease treatment path reasoning combined with knowledge graph, β r is the path reasoning weight, which is learned from the knowledge graph; Calculate the final matching score using the following formula: S final (V′ p ,V d )=λ1S basic +λ2S deep +λ3S kg , Among them, S final (V′ p , V d ) is the final matching score between the patient's condition vector and case d, λ1, λ2, λ3 are weight hyperparameters for similarity calculation at different levels, Search for the most similar cases: Among them, d * For the historical case with the highest similarity, from the historical case database The case with the highest matching degree is retrieved.
10. A clinical decision support system based on big data as claimed in claim 9, characterized in that: The steps of dynamically adjusting the weight of the recommended cases in combination with the patient's current diagnosis and treatment stage, including initial diagnosis, treatment and recovery period, are: Define different stages of patient care, including Initial diagnosis stage: The patient has just come to the hospital and his condition is not yet clear. He needs to match cases with similar symptoms. Treatment stage: patients have been diagnosed and received treatment, focusing on cases similar to the current treatment path. Recovery stage: patients are in the recovery stage, focusing on cases of recovery after similar treatment plans; Define the weight factors for the diagnosis and treatment stage: L=(l d ,l t ,l r ), Among them, λ d represents the weight of the initial diagnosis stage, λ t represents the weight of the treatment stage, λ r Represents the weight of the rehabilitation stage, and the weight satisfies the normalization constraint: d +λ t +λ r =1, Adjust the matching weights of similar cases based on the patient's diagnosis and treatment stage: The matching score at the initial diagnosis stage is calculated using the following formula: With early (V′ p ,V d )=cos(V′ p ,V d ), Among them, S early (V′ p ,V d ) is the matching score between the patient’s initial diagnosis and the historical case d, V′ p is the patient's condition vector representation, V d is the disease vector of historical case d, cos(·) is the cosine similarity calculation function, The matching score for the treatment phase was calculated using the formula: Among them, S treatment (V′ p ,V d ) is the matching score of the patient’s treatment stage, T d is the set of triples of disease-treatment relations in the knowledge graph. (h, r, t) represents a disease-treatment triple in the knowledge graph: h is the disease entity, r is the treatment relation, t is the treatment plan entity, and β r is the attention weight of the treatment path, which is learned through the cross-modal attention mechanism: Among them, R is the set of all possible treatment relationships, S deep (V′ p ,V d ) is calculated by GNN reasoning, The matching score of the rehabilitation stage was calculated as follows: Among them, S recovery (V' p ,V d ) is the matching score of the patient’s recovery stage, T d is the set of triples of disease treatment paths in the knowledge graph, t j The treatment plan for historical case d is: is the weight of different treatment options, obtained through knowledge graph learning: S kg (V′ p ,V d ) is calculated by knowledge graph reasoning; Based on the three types of matching scores, the final case recommendation score is calculated: S final (V′ p ,V d )=λ d S early +λ t S treatment +λ r S recovery , Among them, S final (V′ p ,V d ) is the final recommended case matching score, λ d ,λ t ,λ r Dynamically adjust according to the patient's current diagnosis and treatment stage, Select the final recommended case. Among them, d * The most similar case is finally recommended. A historical case database.
Citation Information
Cited By
Tumor patient clinical test matching system and method based on large language model and OCR technology
CN120913728A
A Clinical Trial Matching System and Method for Cancer Patients Based on Large Language Model and OCR Technology
CN120913728B
Retrieval intelligent sorting method and system based on medical image intelligent database
CN121256082A
A Retrieval and Intelligent Ranking Method and System Based on a Medical Image Intelligent Database
CN121256082B
Otology medicine knowledge base construction method and device and storage medium
CN121327146A