Multi-modal fusion-based explainable rare disease auxiliary diagnosis system and method
Through a multimodal fusion interpretable rare disease auxiliary diagnosis system, visual and text encoders are used to extract features and generate heat maps, solving the accuracy and interpretability of rare disease diagnosis and improving the early diagnosis rate of rare diseases.
Patent Information
- Application Number
- CN202510670891.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-15
AI Technical Summary
Most existing rare disease-assisted diagnosis systems rely on single modal information and lack multimodal data fusion mechanism, resulting in poor accuracy of diagnostic results and insufficient interpretability.
A system for interpretable rare disease assisted diagnosis based on multimodal fusion is designed, and the visual features and text features of multimodal data are extracted respectively through the visual encoder and the text encoder, and a modal fusion classifier is used for splicing diagnosis. A thermal map is generated in combination with the interpretability analysis module to provide interpretability analysis.
Accurate diagnosis results for rare diseases are achieved, and the transparency and credibility of the diagnostic process are improved, helping doctors make more accurate decisions in complex diagnostic tasks.
Smart Images

Figure CN120496811A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and in particular to an interpretable rare disease auxiliary diagnosis system and method based on multimodal fusion. Background Art
[0002] Rare diseases refer to diseases with extremely low incidence rates. Due to the large number of types and complex and diverse associated phenotypes, clinicians lack understanding of rare diseases and are often unable to accurately identify and diagnose patients with rare diseases from a large number of patients.
[0003] With the popularization of medical imaging and electronic health records (EHR), the use of artificial intelligence for auxiliary diagnosis has become an important direction. However, most existing auxiliary diagnosis systems rely on single-modal information (images or text), lack sufficient multimodal data fusion mechanisms, and are not interpretable enough. For example, Chinese patent CN119067915A proposes a method and system for auxiliary diagnosis of brain diseases based on medical imaging. This solution uses an integrated learning method to analyze and process magnetic resonance imaging to obtain output disease auxiliary diagnosis labels and corresponding category probabilities; Chinese patent CN119811649A proposes a training method for an auxiliary diagnosis model, in which each sample data in the sample data set of each preset disease includes: characteristics of multiple symptoms of a patient, related medical history characteristics, and preliminary diagnostic information for each preset disease.
[0004] However, unlike common diseases, the diagnosis of rare diseases usually requires the combination of multiple data sources. If auxiliary diagnosis relies solely on medical images or medical record text information, the diagnostic results will be less accurate, which is not conducive to assisting doctors in decision-making in complex diagnostic tasks. Summary of the Invention
[0005] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and to provide an interpretable rare disease auxiliary diagnosis system and method based on multimodal fusion. By fusing the multimodal data of patient users, using deep learning models for auxiliary diagnosis of rare diseases, and providing interpretable analysis, it can provide accurate rare disease auxiliary diagnosis results and improve the transparency and credibility of the diagnosis process.
[0006] The objectives of the present invention can be achieved through the following technical solutions: an interpretable rare disease auxiliary diagnosis system based on multimodal fusion, comprising an input module for receiving multimodal data of a patient user, the input module being connected to a visual encoder and a text encoder, respectively, the visual encoder and the text encoder being connected to a modal fusion classifier, the input module, the visual encoder, the text encoder and the modal fusion classifier being respectively connected to an interpretability analysis module, the visual encoder being used to extract visual features from the multimodal data;
[0007] The text encoder is used to extract text features from multimodal data;
[0008] The modality fusion classifier is used to combine visual features and text features, perform classification diagnosis, and output classification recognition results;
[0009] The explainability analysis module is used to generate a heat map containing the distribution of the importance of each multimodal data to the classification and recognition results.
[0010] Furthermore, the multimodal data includes electronic health record data and endoscopic image data, and the electronic health record data includes clinical symptoms and laboratory test indicators.
[0011] Furthermore, the visual encoder is based on a Transformer architecture, including an image input layer, a self-attention layer, a convolutional layer, a PCC module, and a visual feature output layer, wherein the image input layer is used to receive endoscopic image data;
[0012] The self-attention layer adopts a multi-head self-attention mechanism to capture the long-range dependencies between image features;
[0013] The convolutional layer is used to extract local features from endoscopic image data;
[0014] The PCC (pyramid pooling connector) module is used to extract multi-scale features from endoscopic image data;
[0015] The visual feature output layer is used to output key features and semantic information of the endoscopic image.
[0016] Furthermore, the text encoder is based on the BERT model, including a text input layer, a Transformer layer, and a text feature output layer. The text input layer is used to receive electronic health record data and perform word embedding encoding;
[0017] The Transformer layer uses a multi-layer Transformer to extract text features;
[0018] The text feature output layer is used to output key concepts and semantic features in electronic health records.
[0019] Furthermore, the modality fusion classifier adopts a multilayer perceptron (MLP), including a splicing input layer, a hidden layer and a classification output layer, wherein the splicing input layer is used to splice the output features of the visual encoder and the text encoder;
[0020] The hidden layer includes multiple fully connected layers, which are used to perform dimensionality reduction, nonlinear activation and feature abstraction processing on the splicing features to extract high-level semantic information;
[0021] The classification output layer is used to output the probability distribution of the diagnosis category.
[0022] Furthermore, the interpretability analysis module uses GradCAM technology to generate a heat map.
[0023] An interpretable rare disease auxiliary diagnosis method based on multimodal fusion, comprising the following steps:
[0024] S1. Collect multimodal data samples including electronic health records and endoscopic images, annotate them accordingly, and construct a dataset;
[0025] S2. Using the dataset and a combined loss function, the visual encoder, text encoder, modality fusion classifier, and interpretability analysis module are jointly optimized and trained to obtain an auxiliary diagnosis model.
[0026] S3. Input the current user's electronic health record data and endoscopic image data into the auxiliary diagnosis model, output the diagnosis result, and generate an importance distribution heat map.
[0027] Furthermore, the combined loss function includes a cross entropy loss function, a contrast loss function, a weighted cross entropy loss function and a thermal Figure 1 A consistency loss function, a cross entropy loss function and a contrast loss function are used for optimizing the training of the visual encoder and the text encoder;
[0028] The weighted cross entropy loss function is used for optimizing the training of the modality fusion classifier;
[0029] The thermal Figure 1 The consistency loss function is used for optimization training of the solvability analysis module.
[0030] Furthermore, the cross entropy loss function is specifically:
[0031]
[0032] The contrast loss function is specifically:
[0033]
[0034] Among them, y i is the true label of the sample (one-hot encoding), which is 1 if the sample belongs to category i, otherwise it is 0, p i is the probability that the sample is predicted to be category i, and N is the total number of categories;
[0035] s is the similarity label of the sample pair, s = 1 means the sample pair is similar, s = 0 means the sample pair is dissimilar, d is the Euclidean distance between two images or two texts in the feature space, m is the interval hyperparameter, and the penalty term is only generated when the distance between heterogeneous samples is less than m; max(0,md) 2 ) is used to ensure that the distance difference is not less than the interval value to avoid excessive penalty.
[0036] Furthermore, the weighted cross entropy loss function is specifically:
[0037]
[0038] Among them, w i It is the weight corresponding to the category i to which the sample belongs, which is used to increase the penalty for samples of scarce categories to ensure the training effect when the categories are unbalanced;
[0039] The thermal Figure 1 The specific consistency loss function is:
[0040]
[0041] Among them, M is the total number of samples, H j is the interpretable heat map generated by the j-th sample, is the target heatmap; IoU is the Intersection-over-Union metric between two heatmaps. The larger the value, the better the overlap between the two.
[0042] Compared with the prior art, the present invention has the following advantages:
[0043] The present invention designs an input module for receiving multimodal data from patient users. The input module is connected to a visual encoder and a text encoder to extract visual and text features from the multimodal data, respectively. A modal fusion classifier is then used to combine the visual and text features for classification and diagnosis, outputting classification and recognition results. Furthermore, an interpretability analysis module is designed to generate a heat map that contains the distribution of the importance of each multimodal data point to the classification and recognition results. By fusing multimodal data from patient users, deep learning can be used to assist in the diagnosis of rare diseases, providing interpretable analysis and improving the transparency and credibility of the diagnostic process.
[0044] In the present invention, the visual encoder is based on the Transformer architecture to extract key concepts and semantic features from endoscopic images; the text encoder is based on the BERT model to extract semantic features of clinical symptoms and laboratory indicators; the modal fusion classifier splices the visual and text features, classifies them through a multi-layer perceptron (MLP) and generates a diagnosis result. In addition, the interpretability analysis module generates a heat map through GradCAM technology to show the importance of key concepts to the final diagnosis result. The present invention can provide accurate auxiliary diagnosis results for rare diseases, and provide the credibility and decision-making basis of the diagnosis results, which helps to assist doctors in making decisions in complex diagnostic tasks and improve the early diagnosis rate of rare diseases.
[0045] The present invention adopts a combined loss function to jointly optimize the training of the visual encoder, text encoder, modality fusion classifier and interpretability analysis module. The combined loss function includes cross entropy loss, contrast loss, weighted cross entropy loss, and thermal Figure 1 Consistency loss, among which cross entropy loss is applied to the optimization training of visual encoder and text encoder respectively, to calculate the gap between the output category probability and the true label; contrast loss is applied to the optimization training of visual encoder and text encoder respectively, to ensure that similar image pairs and similar texts are closer in the feature space; weighted cross entropy loss is applied to the optimization training of modal fusion classifier, by adjusting the weight of each category to deal with unbalanced data sets; thermal Figure 1 Consistency loss is applied to the optimized training of the interpretability analysis module to improve the quality of the generated heatmaps. This effectively completes visual feature extraction, text feature extraction, image classification, and text classification tasks, ensuring the reliability of the auxiliary diagnosis model. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 Schematic diagram of the system structure of the present invention;
[0047] Figure 2 Schematic diagram of the method flow of the present invention;
[0048] Figure 3 Schematic diagram of the application architecture of the embodiment;
[0049] Description of the marks in the figure: 1. Input module, 2. Visual encoder, 3. Text encoder, 4. Modal fusion classifier, 5. Interpretability analysis module. DETAILED DESCRIPTION
[0050] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] Example
[0052] like Figure 1 As shown, an interpretable rare disease auxiliary diagnosis system based on multimodal fusion includes an input module 1 for receiving multimodal data of a patient user, the input module 1 is respectively connected to a visual encoder 2 and a text encoder 3, the visual encoder 2 and the text encoder 3 are connected to a modal fusion classifier 4, and in addition, the input module 1, the visual encoder 2, the text encoder 3 and the modal fusion classifier 4 are respectively connected to an interpretability analysis module 5, wherein the visual encoder 2 is used to extract visual features from the multimodal data;
[0053] The text encoder 3 is used to extract text features from multimodal data;
[0054] The modality fusion classifier 4 is used to combine visual features and text features, perform classification diagnosis, and output classification recognition results;
[0055] The explainability analysis module 5 is used to generate a heat map containing the distribution of the importance of each multimodal data to the classification and recognition results.
[0056] Based on the above system, an interpretable rare disease auxiliary diagnosis method based on multimodal fusion is implemented, such as Figure 2 As shown, the following steps are included:
[0057] S1. Collect multimodal data samples including electronic health records and endoscopic images, annotate them accordingly, and construct a dataset;
[0058] S2. Using the dataset and a combined loss function, the visual encoder 2, the text encoder 3, the modality fusion classifier 4, and the interpretability analysis module 5 are jointly optimized and trained to obtain an auxiliary diagnosis model.
[0059] S3. Input the current user's electronic health record data and endoscopic image data into the auxiliary diagnosis model, output the diagnosis result, and generate an importance distribution heat map.
[0060] This embodiment applies the above solution, such as Figure 3 As shown, first, a system architecture including an input module 1, a visual encoder 2, a text encoder 3, a modality fusion classifier 4, and an interpretability analysis module 5 is constructed, wherein the input module 1 is used to receive the patient's clinical symptoms, laboratory indicators, and endoscopic image data;
[0061] Visual Encoder 2, which extracts key concepts and semantic features of endoscopic images through the Transformer architecture;
[0062] Text encoder 3, extracts semantic features of clinical symptoms and laboratory index texts based on the BERT model;
[0063] Modal fusion classifier 4, used to combine visual and text features for classification and diagnosis;
[0064] Explainability analysis module 5 generates a heat map using GradCAM technology to show the importance of key concepts to the final diagnosis results.
[0065] In this embodiment, the input module 1 receives multimodal data, including:
[0066] Clinical symptoms, which are input through structured text to describe the patient's main symptoms;
[0067] Laboratory indicators, including blood tests, urine tests and other indicators;
[0068] Endoscopic images are images or video data obtained through endoscopic equipment, reflecting the patient's diseased area.
[0069] The visual encoder 2 consists of multiple Transformer layers, each of which contains a self-attention mechanism and a feed-forward neural network.
[0070] The text encoder is trained using the BERT model and extracts semantic features of the text through multiple Transformer layers.
[0071] The modality fusion classifier uses a multi-layer perceptron (MLP) model for multi-task learning to jointly train the visual encoder and the text encoder.
[0072] During the training phase, a combination of loss functions is used to simultaneously optimize the parameters of the visual encoder, text encoder, modality fusion classifier, and interpretability analysis module. The combined loss functions include cross entropy loss, contrast loss, weighted cross entropy loss, and thermal loss. Figure 1 Consistency loss. The design of the loss function mainly depends on the nature of the task, including visual feature extraction, text feature extraction, image classification and text classification tasks. The following are the definitions of loss functions for different modules:
[0073] 1. Loss Function of Visual Encoder
[0074] 1.1 Cross-Entropy Loss
[0075] Used to calculate the gap between the category probability output by the model and the actual label.
[0076]
[0077] 1.2 Contrastive Loss
[0078] It is used to learn the semantic feature vector of the image, so that similar image pairs are closer in the feature space and different image pairs are farther away.
[0079]
[0080] Among them, y i is the true label of the sample (one-hot encoding), which is 1 if the sample belongs to category i, otherwise it is 0, p i is the probability that the sample is predicted to be category i, and N is the total number of categories;
[0081] s is the similarity label of the sample pair, s = 1 means the sample pair is similar, s = 0 means the sample pair is dissimilar, d is the Euclidean distance between two images or two texts in the feature space, m is the interval hyperparameter, and the penalty term is only generated when the distance between heterogeneous samples is less than m; max(0,md) 2 ) is used to ensure that the distance difference is not less than the interval value to avoid excessive penalty.
[0082] 2. Loss Function of Text Encoder
[0083] 2.1 Cross Entropy Loss
[0084] Used to handle text classification problems, ensuring that the model can accurately classify different types of diseases or symptoms.
[0085] 2.2 Contrastive Loss
[0086] Ensure that similar texts (such as descriptions of the same symptoms) are closer in feature space, and dissimilar texts are farther apart.
[0087] The loss function calculation formula for the text encoder is the same as that for the visual encoder.
[0088] 3. Loss Function of Modal Fusion Classifier
[0089] Weighted cross entropy loss
[0090] Adjust the weight of each class to handle imbalanced datasets.
[0091]
[0092] Among them, w i It is the weight corresponding to the category i to which the sample belongs, which is used to increase the penalty for the scarce category samples to ensure the training effect when the categories are unbalanced
[0093] 4. Interpretability Analysis Loss Function
[0094] heat Figure 1 Heatmap Consistency Loss
[0095] Used to optimize the quality of the generated heatmap to make it consistent with the doctor's diagnosis results.
[0096]
[0097] Among them, M is the total number of samples, H j is the interpretable heat map generated by the j-th sample, is the target heatmap; IoU is the Intersection-over-Union metric between two heatmaps. The larger the value, the better the overlap between the two.
[0098] During the specific training process, the visual encoder is based on the Transformer architecture and is designed to extract key features and semantic information from endoscopic images. Input layer: Image input, size W\times H\times C (for example, 416x416x3).
[0099] Self-attention layer: The multi-head self-attention mechanism captures the long-range dependencies between image features.
[0100] Convolutional layer: used to extract local features and process image detail information.
[0101] PCC module: multi-scale feature extraction to enhance the balance between local and global information.
[0102] Output layer: Output feature vector with size D1, representing the semantic features of the image.
[0103] The text encoder is based on the BERT model and is designed to extract key concepts and semantic features from electronic health records. Input layer: Text input (such as a patient's symptom description) is encoded using word embeddings.
[0104] Transformer layer: Multi-layer Transformer performs feature extraction.
[0105] Output layer: Outputs the semantic features of the text, with a size of D2, representing the semantic information of the text.
[0106] In addition, the modality fusion classifier concatenates the output features of the visual encoder and the text encoder and performs classification through a multi-layer perceptron (MLP). Input layer: The concatenation of image features and text features, with a size of D1+D2.
[0107] Hidden layer: multiple fully connected layers, each layer uses ReLU activation function.
[0108] Output layer: Outputs the probability distribution of diagnostic categories, obtained through the Softmax function.
[0109] After training the auxiliary diagnosis model, the patient's medical imaging data and HER data are fed into the model in real-world applications. The trained visual encoder and text encoder output key concept predictions, and the modal fusion classifier outputs the final classification diagnosis. GradCAM technology is also used to generate heat maps, showing the areas the model focuses on during diagnosis, helping doctors understand the model's decision-making process.
[0110] In summary, this solution integrates multimodal data, including patient clinical symptoms, laboratory parameters, and endoscopic images, using deep learning models to aid in the diagnosis of rare diseases and provide interpretable analysis. Through multi-task learning and loss function optimization, it can provide accurate rare disease diagnoses and enhance the transparency and credibility of the diagnostic process. This solution helps doctors make decisions in complex diagnostic tasks, improves the early diagnosis rate of rare diseases, and has significant clinical application value.
Claims
1. An interpretable rare disease auxiliary diagnosis system based on multimodal fusion, characterized by: The invention comprises an input module (1) for receiving multimodal data of a patient user, wherein the input module (1) is respectively connected to a visual encoder (2) and a text encoder (3), wherein the visual encoder (2) and the text encoder (3) are connected to a modality fusion classifier (4), wherein the input module (1), the visual encoder (2), the text encoder (3) and the modality fusion classifier (4) are respectively connected to an interpretability analysis module (5), and wherein the visual encoder (2) is used to extract visual features from the multimodal data; The text encoder (3) is used to extract text features from multimodal data; The modality fusion classifier (4) is used to combine visual features and text features, perform classification diagnosis, and output classification recognition results; The explainability analysis module (5) is used to generate a heat map containing the importance distribution of each multimodal data to the classification recognition result.
2. The interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 1, characterized in that: The multimodal data includes electronic health record data and endoscopic image data, and the electronic health record data includes clinical symptoms and laboratory test indicators.
3. The interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 2 is characterized in that: The visual encoder (2) is based on a Transformer architecture and includes an image input layer, a self-attention layer, a convolutional layer, a PCC module, and a visual feature output layer, wherein the image input layer is used to receive endoscopic image data; The self-attention layer adopts a multi-head self-attention mechanism to capture the long-range dependencies between image features; The convolutional layer is used to extract local features from endoscopic image data; The PCC module is used to extract multi-scale features from endoscopic image data; The visual feature output layer is used to output key features and semantic information of the endoscopic image.
4. The interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 2 is characterized in that: The text encoder (3) is based on the BERT model and includes a text input layer, a Transformer layer, and a text feature output layer. The text input layer is used to receive electronic health record data and perform word embedding encoding; The Transformer layer uses a multi-layer Transformer to extract text features; The text feature output layer is used to output key concepts and semantic features in electronic health records.
5. The interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 1 is characterized in that: The modality fusion classifier (4) adopts a multi-layer perceptron MLP, including a splicing input layer, a hidden layer and a classification output layer, wherein the splicing input layer is used to splice the output features of the visual encoder (2) and the text encoder (3); The hidden layer includes multiple fully connected layers, which are used to perform dimensionality reduction, nonlinear activation and feature abstraction processing on the splicing features to extract high-level semantic information; The classification output layer is used to output the probability distribution of the diagnosis category.
6. The interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 1 is characterized in that: The explainability analysis module (5) uses GradCAM technology to generate a heat map.
7. An interpretable rare disease auxiliary diagnosis method based on multimodal fusion, applied to the interpretable rare disease auxiliary diagnosis system based on multimodal fusion according to claim 1, characterized in that: The following steps are involved: S1. Collect multimodal data samples including electronic health records and endoscopic images, annotate them accordingly, and construct a dataset; S2. Using the dataset and a combined loss function, the visual encoder (2), text encoder (3), modality fusion classifier (4), and interpretability analysis module (5) are jointly optimized and trained to obtain an auxiliary diagnosis model. S3. Input the current user's electronic health record data and endoscopic image data into the auxiliary diagnosis model, output the diagnosis result, and generate an importance distribution heat map.
8. The interpretable rare disease auxiliary diagnosis method based on multimodal fusion according to claim 7, characterized in that: The combined loss function includes a cross entropy loss function, a contrast loss function, a weighted cross entropy loss function and a heat map consistency loss function, and the cross entropy loss function and the contrast loss function are used for optimizing the training of the visual encoder (2) and the text encoder (3); The weighted cross entropy loss function is used for optimizing the training of the modality fusion classifier (4); The heat map consistency loss function is used for optimization training of the solvability analysis module.
9. The interpretable rare disease auxiliary diagnosis method based on multimodal fusion according to claim 8, characterized in that: The cross entropy loss function is specifically: The contrast loss function is specifically: Among them, y i is the true label of the sample (one-hot encoding), which is 1 if the sample belongs to category i, otherwise it is 0, p i is the probability that the sample is predicted to be category i, and N is the total number of categories; s is the similarity label of the sample pair, s = 1 means the sample pair is similar, s = 0 means the sample pair is dissimilar, d is the Euclidean distance between two images or two texts in the feature space, m is the interval hyperparameter, and the penalty term is only generated when the distance between heterogeneous samples is less than m; max(0,md) 2 ) is used to ensure that the distance difference is not less than the interval value to avoid excessive penalty.
10. The interpretable rare disease auxiliary diagnosis method based on multimodal fusion according to claim 9, characterized in that: The weighted cross entropy loss function is specifically: Among them, w i It is the weight corresponding to the category i to which the sample belongs, which is used to increase the penalty for samples of scarce categories to ensure the training effect when the categories are unbalanced; The heat map consistency loss function is specifically: Among them, M is the total number of samples, H j is the interpretable heat map generated by the j-th sample, is the target heatmap; IoU is the Intersection-over-Union metric between two heatmaps. The larger the value, the better the overlap between the two.
Citation Information
Patent Citations
Cerebral disease auxiliary diagnosis method and system based on medical image
CN119067915A
Auxiliary diagnosis model training method, disease symptom collection method and equipment
CN119811649A
Cited By
Diagnosis model capable of explaining rare craniofacial disease, diagnosis method and electronic equipment
CN122135942A