Multimodal disease data processing system based on interpretable model
By adopting a multimodal disease data processing system based on interpretability models in the medical image data processing system, the problems of multimodal data processing, labeled data dependence and model decision transparency are solved, and more efficient, reliable and transparent medical image data processing is achieved.
Patent Information
- Application Number
- CN202510083306.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
AI Technical Summary
Existing deep learning-based medical image data processing systems have challenges in multimodal data processing, labeled data dependence and model decision transparency, limiting their wide application in medical practice.
A multimodal disease data processing system based on interpretability model is adopted, and the multimodal data fusion and information extraction are realized through report extraction concept library, pre-trained image encoder, multimodal concept score calculation module, multimodal concept score correction module, linear classifier and report generation module, and the key concepts in clinical diagnostic reports are automatically extracted through large language models to reduce the dependence on professional medical knowledge annotation.
It significantly enhances the visualization and understanding ability of the model decision-making process, reduces the time and cost of the data preparation stage, improves the accuracy and reliability of multimodal medical image data processing, and enhances the transparency and user trust of the model.
Smart Images

Figure CN120126656A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and in particular, to a multi-modal disease data processing system based on an interpretability model. Background Art
[0002] In recent years, with the rapid progress of machine learning and deep neural networks (DNNs), significant development has also been achieved in the field of computer-aided diagnosis (CAD). Against this background, a large number of studies on medical image disease data processing based on deep learning have emerged, greatly improving the diagnostic performance of medical images. These deep learning-based models have shown comparable levels to radiologists in various medical data processing tasks. For example, analyzing chest X-rays (CXR) for screening, fundus photography for automatic retinopathy screening, and brain MRI for quantitative analysis of tumors and stroke injuries. Although these research works have achieved performance similar to that of human experts, there are still some significant drawbacks and challenges that limit their wide application in medical practice.
[0003] Firstly, during the clinical disease diagnosis process, patients often need to take various modalities of medical images (MRI, CT, ultrasound, etc.), and each imaging technique has its unique advantages and limitations. For example, MRI can provide detailed soft tissue images, helping to determine the exact location and size of tumors; CT imaging is very sensitive to detecting bone diseases and tumor calcifications; while (Doppler) ultrasound shows unique advantages in evaluating blood flow conditions and guiding minimally invasive surgeries. By combining medical images of different modalities, doctors can have a more comprehensive understanding of diseases and thus make more reasonable treatment decisions. Therefore, a medical data processing system should not only focus on single-modal data but also have the ability to fuse and extract information from multi-modal data.
[0004] Secondly, the training of such deep learning models highly depends on rich and high-quality labeled datasets. In the field of medical images, such datasets are often difficult to obtain. Different from ordinary natural image data, the annotation process of medical image data requires professional medical knowledge and rich practical experience, which is costly and time-consuming. If there is a lack of sufficient labeled data, the performance of deep learning models may be limited and unable to achieve the expected high accuracy. On the other hand, each medical image generated in clinical practice is accompanied by a corresponding diagnostic report, which details the disease characteristics in the image. However, these text information is often ignored and not fully utilized.
[0005] Finally, deep learning models, especially complex neural network structures, are often regarded as "black box" models. This means that although the model can give a classification result based on the input image data, the internal decision-making process lacks transparency and is difficult to intuitively explain. In the processing of medical image data, doctors not only need to know the classification result, but also need to understand how the model arrives at this conclusion. This helps doctors to conduct manual review and intervention when necessary to ensure the accuracy and reliability of the classification. However, due to the complexity of deep learning models, their decision-making processes are often difficult to explain. This lack of interpretability may limit the application of deep learning in medical practice. On the one hand, doctors may avoid using the model due to distrust. On the other hand, the lack of interpretability may also lead to more stringent reviews by regulatory agencies, further restricting the application scope of deep learning models in the medical field. To overcome this challenge, current research on model interpretability mainly focuses on two directions. One is post hoc interpretation. Such methods regard the model as a trained black box, do not attempt to open the black box, but speculate on the working principle of the model through hypothesis testing to provide a reasonable explanation. Among them, techniques based on neuron activation are commonly used means, and the most relevant regions to the output can be identified through CAM and Grad CAM techniques. However, this method often only shows strong correlations and cannot provide high-confidence causal relationships. The other direction is intrinsic explanations, and a representative method is the Concept Bottleneck Model (CBM). The data processing method of the CBM model is similar to the human thinking mode. First, it learns and understands human-interpretable concepts, and then makes the final prediction based on these identified concepts. These research works not only help to enhance doctors' trust in the model, but also promote the further development of deep learning in the field of medical imaging.
[0006] In summary, although the computer-aided medical (CAD) system based on deep learning has great potential, it also faces challenges in aspects such as dataset quality, model interpretability, and practical implementation. Summary of the Invention
[0007] The present invention aims to provide a multimodal disease data processing system based on an interpretable model to solve the problems existing in the prior art and improve the practicality and reliability of deep learning in medical image classification.
[0008] The technical solution adopted by the present invention is as follows:
[0009] The present invention provides a multimodal disease-assisted data processing system based on an interpretable model, including:
[0010] The report extraction concept library consists of several concept embeddings, which are generated based on large language models, pre-trained image encoders, and image embedding learning concept activation vectors;
[0011] The pre-trained image encoder is used to extract the feature vectors of multi-modal radiological images to obtain image embeddings;
[0012] The multi-modal concept score calculation module is used to calculate the multi-modal concept scores based on the image embeddings and the concept embeddings in the report extraction concept library;
[0013] The multi-modal concept score correction module is used to provide a window for modifying the multi-modal concept scores;
[0014] The linear classifier is used to calculate the classification result data based on the multi-modal concept scores;
[0015] The report generation module is used to generate a data processing classification report based on the classification result data and the concept embeddings in the report extraction concept library.
[0016] In a preferred embodiment of the present invention, the method for generating the report extraction concept library includes the following steps:
[0017] 1) Obtain several clinical diagnosis reports;
[0018] 2) For each clinical diagnosis report, use the large language model to extract concepts from it, use the pre-trained image encoder to encode the corresponding radiological image of the clinical diagnosis report to obtain image features, form image feature-concept pairs with the obtained concepts and image features, and process the image feature-concept pairs based on the image embedding learning concept activation vector to generate concept embeddings;
[0019] 3) Use all the concept embeddings to form the report extraction concept library.
[0020] In a preferred embodiment of the present invention, the clinical diagnosis reports in step 1) come from the fundus disease dataset that has undergone data cleaning; data cleaning includes identifying and correcting incorrect or abnormal data, and performing consistency checks on the data.
[0021] In a preferred embodiment of the present invention, the training dataset of the pre-trained image encoder includes a training set and a test set. Both the training set and the test set include several multi-modal radiological images from the fundus disease dataset. The training set uses five-fold cross-validation to train the pre-trained image encoder, and the test set is used to evaluate the final performance of the pre-trained image encoder.
[0022] In a preferred embodiment of the present invention, the pre-trained image encoder includes three encoders that correspond one-to-one to three modalities and are used to obtain image embeddings; during the training process of the pre-trained image encoder, a pre-trained auxiliary classifier composed of an attention pooling module, a self-attention mechanism module, and a classifier module is adopted. The attention pooling module is used to fuse the image embeddings from different modalities and map them into a shared feature space. The self-attention mechanism module and the classifier module are used to refine the fused features, extract the most crucial information for disease prediction, and give a classification result.
[0023] In a preferred embodiment of the present invention, cross-entropy is used as the loss function during the training process of the pre-trained image encoder.
[0024] In a preferred embodiment of the present invention, each row of the report extraction concept library is a concept embedded by the concept activation vector method; during the construction of the report extraction concept library, to learn the concept activation vector, positive and negative samples need to be divided first. For concept i, the radiological images containing concept i are used as positive samples, and the radiological images not containing concept i are used as negative samples. Then, the pre-trained image encoder φ is used to encode the radiological images to obtain the positive sample embedding vector P i and the negative sample embedding vector N i , and then a support vector machine is trained to perform binary classification on the positive sample embedding vector and the negative sample embedding vector. The normal vector of the classification boundary of the support vector machine is used as the embedding of concept i, and is stored as the row vector C of the report extraction concept library E where is the normal vector of the classification boundary of the support vector machine.
[0025] In a preferred embodiment of the present invention, the multi-modal concept score is expressed as:
[0026]
[0027] where scores represent the concept scores of the three modalities;
[0028] The calculation formula for the concept score of each modality is:
[0029]
[0030] where E C represents the report extraction concept library, and φ m (X m ) represents encoding the m-th radiological image X m .
[0031] In a preferred embodiment of the present invention, the linear classifier is a learnable weight matrix W, and the method for calculating the classification result data includes: applying the Sigmoid activation function to the linear classifier, and then multiplying element-wise with the multi-modal concept scores to generate an attention matrix where n represents the number of disease categories, N represents the number of concepts in the concept library, and finally sum along dimension N, and obtain the final disease prediction through Softmax that is, the classification result data, and the formula is:
[0032]
[0033] where ⊙ represents element-wise multiplication.
[0034] Compared with the prior art, the beneficial effects of the present invention are:
[0035] 1. By adopting the method based on the Concept Bottleneck Model (CBM), the system of the present invention significantly enhances the visualization and understanding ability of the model decision-making process, which is particularly important in the medical field because doctors and clinical workers can clearly understand how the model makes judgments based on specific features in the image, which not only enhances users' trust in the model, but also makes medical decisions more transparent and reliable.
[0036] 2. Traditional deep learning models require a large amount of labeled data for training, and the labeling process of these data is time-consuming and costly. The system of the present invention significantly reduces the dependence on professional medical knowledge labeling by automatically extracting key concepts from clinical diagnostic reports using large language models, thereby reducing the time and cost in the data preparation stage.
[0037] 3. When processing multi-modal medical image data, the system of the present invention effectively fuses concept scores of different modalities through multi-modal max pooling, fully leveraging the advantages of each modality data, improving the accuracy and reliability of classification. This fusion method provides a more comprehensive perspective for the model to understand disease characteristics, thereby improving the prediction accuracy. In the single-modal scenario, the structure of the system of the present invention is still applicable.
[0038] 4. The system of the present invention also sets an image-concept pair verification window and a concept score correction window, allowing user doctors to adjust the concept scores according to their professional knowledge and clinical experience to "correct" the model's prediction, thereby improving the accuracy of data classification.
[0039] To make the above objects, features and advantages of the present invention more obvious and understandable, the following specifically gives embodiments of the present invention and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0041] Figure 1 It is a schematic diagram of a multi-modal disease data processing system based on an interpretability model. Among them, in a, a large language model (LLM) is used to extract image-concept pairs from clinical medical reports. At the same time, as a comparative experiment, experts are invited to help check the matching of the image-concept pairs. In b, a report extraction concept library is constructed by learning concept activation vectors. In particular, the concept library modified by experts is named the expert-verified concept library. In c, it is the output stage of the model. The input is a series of multi-modal images. The pre-trained image encoder φ is used to extract feature vectors from these images. Subsequently, multi-modal concept scores are calculated. Then, an interpretable prediction and a judgment basis are obtained through a linear classifier. In addition, an explanatory report is generated using the LLM to improve the transparency of the data processing process.
[0042] Figure 2 It is a structural diagram of the system components. Among them, a is a schematic diagram of the pre-trained image encoder. Images of different modalities are converted into features through dedicated encoders and modal feature fusion is performed through the attention pooling module. The classifier obtains the final classification prediction. The encoders of the three modalities form the pre-trained image encoder φ. b is the detailed process of feature fusion by the attention pooling module. c refers to a linear layer mechanism, which is a matrix whose weights can be adjusted through training. It can predict the type of tumor based on the concept scores. During this process, the attention scores calculated through element-wise multiplication can clearly reveal the relevance between the current image (i.e., representing the situation of a specific patient) and each concept, thus providing strong support for the prediction.
[0043] Figure 3 It is an interface demonstration of a multi-modal disease data processing system based on an interpretability model. Specific embodiments
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0045] The present invention provides a multimodal disease data processing system based on an interpretable model, including a report extraction concept library, a pre-trained image encoder, a multimodal concept score calculation module, a multimodal concept score correction module, a linear classifier, and a report generation module.
[0046] The process of constructing this system is divided into a data preprocessing stage, a concept library construction stage, and an interpretable prediction stage.
[0047] 1. Data preprocessing stage
[0048] The dataset used to construct this system is an ophthalmic disease dataset, which includes three categories (choroidal hemangioma, choroidal metastatic carcinoma, and choroidal melanoma) and three modalities (fluorescein angiography (FA), indocyanine green angiography (ICGA), and Doppler ultrasound images (US)).
[0049] 1) Data cleaning
[0050] There are various problems in the original data of the ophthalmic disease dataset, such as missing values, duplicate data, noise interference, and inconsistent formats, etc., which will have an adverse impact on the training results of the model and thus reduce the accuracy of the model. To address these issues, we need to identify and correct incorrect or abnormal data, and the methods adopted include deletion, replacement, or correction. In addition, we also need to perform consistency checks on the data to ensure that each field in the data is logically consistent and avoid logical contradictions or errors.
[0051] 2) Pre-trained image encoder
[0052] In order to better capture the imaging features in the images, the present invention first performs pre-training of the multimodal encoder.
[0053] a. The architectural details of the pre-trained image encoder are as Figure 2 shown and divided into two parts:
[0054] (i) Modal feature extraction, Figure 2 In a of, three encoders are used corresponding to the three modalities, and features are extracted from different data modalities respectively to obtain image embeddings of the three modalities; these three encoders correspond to the three modalities of fluorescein angiography (FA), indocyanine green angiography (ICGA), and Doppler ultrasound images (US). Each encoder adopts a convolutional neural network structure based on EfficientNet-B0 to efficiently extract the image features of its respective modality and generate image embeddings of the three modalities.
[0055] (ii) Feature fusion and label prediction, as Figure 2As shown in a of, we designed a simplified pre-trained auxiliary classifier for three-class disease prediction. The auxiliary classifier consists of three modules: attention pooling, self-attention mechanism, and classifier. Among them, attention pooling effectively fuses the embeddings from different modality images and maps them into a shared feature space, enabling the model to also handle single-modality image cases; the self-attention mechanism and classifier perform refined processing on the fused features, extract the most crucial information for disease prediction, and give the classification results. The cross-entropy is used as the loss function in the three-class pre-training process, and the image encoder φ after pre-training will be used in subsequent steps.
[0056] b. Pre-training of the Image Encoder
[0057] The dataset for pre-training the image encoder includes the training set, validation set, and test set from the fundus disease dataset to ensure the generalization ability of the model on different data. The pre-trained encoder uses the EfficientNet-B0 structure combined with the pre-trained auxiliary classifier for three-class disease prediction. The Adam optimizer is used during the training process, with the initial learning rate set to 1e-4 and the batch size set to 8. The training is carried out for 200 epochs, and the early stopping strategy is adopted to prevent overfitting. Training stops when the performance on the validation set has not improved for 10 consecutive epochs. To increase the robustness of the model, data augmentation techniques such as random rotation, translation, scaling, and flipping are applied during the training process. After each epoch, the accuracy, recall, and F1 score of the model are evaluated using the validation set, and the model parameters that perform best on the validation set are selected as the final pre-trained image encoder. The entire training process is carried out on an NVIDIA 3090 GPU, based on the Pytorch framework, using its high-level API to simplify the model construction and training process. After training, the performance of the pre-trained image encoder on the test set reaches an accuracy of 92.8%, a recall of 89.2%, and an F1 score of 89.2%, demonstrating its effectiveness and reliability in multi-modal image feature extraction.
[0058] 3) Concept Extraction
[0059] The traditional Concept Bottleneck Model (CBM) requires constructing a large-scale and high-quality image-concept pair dataset for training. Traditional concept annotation methods mainly rely on manual operations, which not only require the operator to have professional medical knowledge but also are time-consuming, laborious, and costly. To reduce the cost of manual annotation and improve efficiency, the present invention innovatively utilizes the diagnostic reports in the clinical diagnosis process. As Figure 1As shown in Figure a, by designing a series of prompts, the present invention prompts a large language model (LLM) to automatically extract concepts closely related to diseases from diagnostic reports. With the powerful natural language processing ability of the large language model, it is possible to efficiently construct a dataset of corresponding relationships between images and concepts (image-concept pairs) without manual annotation by experts. Considering the training dataset where respectively represent multi-modal images, disease categories, and diagnostic reports issued by doctors. The image-concept pairs extracted using the large language model are represented as where are the extracted concepts. Then we prompt the large language model to merge concepts with the same semantic meaning, thereby obtaining a compressed concept representation as the final concept dataset after preprocessing.
[0060] 2. Concept Library Construction Phase
[0061] The core part of the multi-modal disease data processing system contains various concepts related to diseases. The construction of the report extraction concept library is mainly based on the prior knowledge about diseases contained in the diagnostic reports of the fundus disease dataset, corresponding concepts to image features to achieve concept embedding.
[0062] 1) Image Encoding
[0063] In the data preprocessing stage, the pre-trained image encoder φ has been obtained, which can map images of different modalities into a shared feature space. We first encode all radiological images to obtain image features
[0064] 2) Generation of the Report Extraction Concept Library
[0065] The features obtained from the pre-trained image encoder φ and the concepts extracted by the large language model from clinical diagnostic reports form image feature-concept pairs. Subsequently, we use these image feature-concept pairs to generate bottleneck embeddings, thereby constructing a concept library, whose mathematical expression is where N represents the number of concepts, and d represents the specific dimension of the embedding space. As Figure 1 shown in Figure b on the left, each row of the concept library E C is a concept learned through the concept activation vector (CAV) technique Embedding. This process ensures that each concept in the concept library has a unique representation, thus supporting subsequent analysis and applications. To learn the Concept Activation Vector (CAV), we first need to divide the positive and negative samples. For concept i, we take the images containing it as positive samples and the images not containing this concept as negative samples. Then we use the pre-trained image encoder φ to encode and obtain the positive sample embedding vectors
[0066]
[0067] and negative sample embedding vectors
[0068]
[0069] Then we train a Support Vector Machine (SVM) to perform binary classification on the embedding vectors P i and N i The normal vector of the SVM classification boundary is used as the embedding of concept i and is stored as a row vector in the concept library E C where where is the normal vector of the SVM classification boundary.
[0070] 3. Interpretable Prediction Phase
[0071] The ultimate goal of the multi-modal disease data processing system is to predict and interpret diseases based on the constructed report extraction concept library.
[0072] 1) Calculate the multi-modal concept scores based on the image embedding and the concept embeddings in the report extraction concept library.
[0073] As shown in c of Figure 1 we use the pre-trained image encoder φ that has been obtained to encode the corresponding modal images into feature vectors where m represents the corresponding modality. To measure the similarity between the image features and the concept representations in the concept library, we perform a dot product between the image features and the concept library E C to obtain the concept score of this image The calculation formula is:
[0074]
[0075] To meet the requirements of multi-modal tasks, we need to fuse the concept scores of each specific modality to obtain the multi-modal concept score Although there are multiple feasible fusion functions to choose from, in our scenario, the fusion function must be able to handle the scores from different modalities properly. As a simple method, we consider implementing Figure 1The multi-modal max pooling operation shown in c is used to determine the concept most likely associated with each modality. The multi-modal concept score can be expressed as:
[0076]
[0077] It should be noted that when calculating the concept score, the FA and ICGA modalities share a concept library, which stems from their similarity in imaging features. In contrast, the US modality uses an independent concept library to reflect its uniqueness in imaging features.
[0078] 2) A linear classifier is used to calculate the classification result data based on the multi-modal concept score.
[0079] Based on the calculated multi-modal concept score, the prediction result of the disease is obtained, and the concepts on which the prediction itself depends are given. We use a linear classifier ( Figure 1 shown in c) as the predictor g in the interpretable prediction stage. The details of the linear classifier are as Figure 2 shown in c, which is a learnable weight matrix that can learn the preferences of specific diseases for certain concepts. Then, we use the multi-modal concept score as the only input to the weight matrix to calculate the final prediction result. To measure the attention of the predictor to the image concept score, we first apply the Sigmoid activation function to the weight matrix W of the classifier, and then multiply it element-wise with the multi-modal concept score to generate the attention matrix where n represents the number of disease categories and N represents the number of concepts in the concept library. Finally, sum along dimension N and pass through Softmax to obtain the final disease prediction i.e., the classification result data, and the formula is:
[0080]
[0081] where ⊙ represents element-wise multiplication (Element Product).
[0082] During the training process of the linear classifier g, we use the same dataset settings as the pre-trained encoder, that is, the fundus disease dataset is divided into a training set, a validation set, and a test set. The fundus images are encoded into feature vectors by the pre-trained encoder and then dot-multiplied with the concept library E C to obtain the concept score of the image and the disease category y corresponding to the image. Cross Entropy is used as the loss function during the training process:
[0083]
[0084] The training was conducted for 200 epochs, and the early stopping strategy was adopted to prevent overfitting. Training was stopped when the performance on the validation set did not improve for 10 consecutive epochs. To increase the robustness of the model, data augmentation techniques such as random rotation, translation, scaling, and flipping were applied during the training process. After each epoch, the accuracy, recall, and F1-score of the model were evaluated using the validation set, and the model parameters that performed best on the validation set were selected as the final pre-trained image encoder. The entire training process was carried out on an NVIDIA 3090 GPU, based on the Pytorch framework, using its high-level API to simplify the model construction and training process. After training, the performance of the pre-trained image encoder on the test set reached an accuracy of 94.5%, a recall of 90.0%, and an F1-score of 91.0%, demonstrating the high effectiveness and reliability of the present invention in the multi-modal fundus disease classification task.
[0085] 3) Generate a medical data processing classification report by extracting the concept embeddings of the concept library based on the classification result data and report.
[0086] As Figure 1 shown in c, the task of automatically generating a data processing classification report refers to automatically generating a descriptive narrative report in the medical field based on the input radiological images. With the interpretable predictor and its corresponding concept scores proposed by the present invention, we can sort the concepts in descending order of concept scores, and the top k concepts are used as the final classification basis. It should be particularly noted that clinical diagnosis reports usually need to follow a specific structured format, including parts such as patient basic information, medical process details, diagnosis results, treatment suggestions, and other necessary relevant information. Therefore, by combining the large language model (LLM) with the top k core concepts refined by the interpretable predictor, the present invention can generate a relatively comprehensive and structurally complete clinical medical data processing classification report.
[0087] 4) Multi-modal concept score correction
[0088] Another significant difference between the CBM model and traditional black-box models lies in the unique interaction ability of the CBM model. When humans use the CBM model, they can interact with it by intervening in the concept scores. This intervention method is particularly useful in medical practice. We provide a window for modifying multi-modal concept scores, as shown in Figure 3 shown. Doctors can observe the concept scores output by the model, and then modify the j-th concept value to the true value c j to "correct" the model, and then use a linear classifier to predict the disease category based on the modified concept scores and observe the changes in the corresponding prediction outputs to qualitatively see the contribution of each concept.
[0089] As Figure 3 shown, the following is the final demonstration of the medical assistance interface implemented by the present invention, which shows a highly interactive, intuitive and easy-to-use medical assistance interface. This interface makes full use of the advantages of modern Web technology and provides doctors with fast, accurate and reliable assistance tools. Doctors can, through simple operations, input multi-modal images of patients. The system will then, based on this information and in combination with the built-in MMCBM model, automatically analyze and give the data processing classification results. These results are presented on the interface in a clear and intuitive manner, including concept scores, possible disease types, and reports generated by invoking a large language model (LLM). Doctors can also, as needed, further intervene and optimize the concept scores to ensure the accuracy of the final classification results.
[0090] The system of the present invention has the following advantages:
[0091] 1) The present invention proposes an interpretable multi-modal medical concept bottleneck model (MMCBM). MMCBM is designed based on the classical interpretable model - the concept bottleneck model (CBM). By encoding the input images and then calculating a series of concept scores, and then making predictions based on these concept scores, it realizes the visualization of the internal decision-making process of the model. On the other hand, the MMCBM model allows human users to interact with it more richly through manual intervention. This interaction ability not only enhances the interpretability of the model, but also brings unique value to fields such as medical practice. By intervening in these concept scores, we can affect the decision-making process of the model, thereby achieving the purpose of optimizing the model performance. This intervention can be an adjustment of a certain concept score, or a recombination and optimization of some concept scores. In medical practice, the ability to manually intervene in the CBM model is particularly important. For example, in the process of medical data processing, doctors may find that the model has deviations in the judgment of some key concepts. At this time, they can "correct" the decision of the model by adjusting the scores of relevant concepts, thereby improving the accuracy of the final classification. In addition, doctors can also observe the changes in the prediction output by removing some concepts to understand the contribution of each concept to the final decision. This qualitative analysis method helps doctors better understand the decision-making process of the model, and thus better trust and rely on the model in practical applications.
[0092] 2) The training process of the traditional CBM model requires manual annotation of images to identify the concepts present therein. However, due to the professionalism and complexity of medical images, this annotation process relies on professional medical knowledge. To address this issue, the present invention makes full use of the powerful capabilities of large language models in text understanding. Specifically, the present invention designs a series of prompts and uses a large language model (LLM) to automatically extract disease-related concepts from the patient's clinical diagnosis report, thereby constructing a dataset of image-concept pairs without the need for additional manual annotation by experts. This innovative method not only improves the efficiency and accuracy of data annotation but also provides new possibilities for the development of the field of medical image analysis.
[0093] 3) The present invention first constructs image-concept pairs using a large language model (LLM). Subsequently, by pre-training an image encoder, images of different modalities are mapped to a shared feature space. On this basis, combined with the concept activation vector (CAV) technology, the embedding of concepts is achieved. These embedded concepts together form a concept library, through which we can calculate the concept scores of each concept in the image. These concepts play a crucial role in the decision-making process of the model, providing us with a window to deeply understand the decision-making logic of the model and enabling us to clearly understand how the model makes decisions based on these concepts.
[0094] 4) Multimodal images enable the model to extract information from complex and diverse medical data, further improving the practicality and generalization ability of the model. The MMCBM of the present invention fuses medical image data of different modalities through multimodal max pooling, giving full play to the advantages of various types of data and improving the accuracy and reliability of medical data processing and classification.
[0095] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multimodal disease data processing system based on an interpretable model, characterized in that: include: Report extraction concept library, which consists of several concept embeddings, generated based on large language models, pre-trained image encoders, and image embedding learning concept activation vectors; Pre-trained image encoder to extract feature vectors of multi-modal radiological images to obtain image embeddings; A multimodal concept score calculation module is used to calculate the multimodal concept score based on the image embedding and the concept embedding in the report extraction concept library; A multimodal concept score correction module, used for providing a window for modifying the multimodal concept score; Linear classifier, used to obtain classification result data based on multimodal concept score calculation; The report generation module is used to generate a data processing classification report based on the classification result data and the concept embedding of the report extraction concept library.
2. The multimodal disease data processing system based on an interpretable model according to claim 1, characterized in that: The method for generating the report extraction concept library includes the following steps: 1) Obtain several clinical diagnosis reports; 2) For each clinical diagnosis report, a large language model is used to extract concepts from it, and the radiological image corresponding to the clinical diagnosis report is encoded using a pre-trained image encoder to obtain image features. The obtained concepts and image features are combined to form image feature-concept pairs, and the image feature-concept pairs are processed based on image embedding learning concept activation vectors to generate concept embeddings; 3) Utilize all the concept embeddings to compose the report and extract the concept library.
3. The multimodal disease data processing system based on the explainability model according to claim 2, characterized in that: The clinical diagnosis report in step 1) comes from a fundus disease dataset that has undergone data cleaning; data cleaning includes identifying and correcting erroneous or abnormal data, and checking the consistency of the data.
4. The multimodal disease data processing system based on an interpretable model according to claim 3, characterized in that: The training dataset of the pre-trained image encoder includes a training set and a test set. Both the training set and the test set include several multimodal radiological images from the fundus disease dataset. The training set uses a five-fold cross-validation method to train the pre-trained image encoder, and the test set is used to evaluate the final performance of the pre-trained image encoder.
5. The multimodal disease data processing system based on an interpretable model according to claim 4, characterized in that: The pre-trained image encoder includes three encoders that correspond one-to-one to the three modalities and are used to obtain image embeddings. During the training process of the pre-trained image encoder, a pre-trained auxiliary classifier consisting of an attention pooling module, a self-attention mechanism module and a classifier module is used. The attention pooling module is used to fuse image embeddings from different modalities and map them to a shared feature space. The self-attention mechanism module and the classifier module are used to refine the fused features, extract the most critical information for disease prediction, and give classification results.
6. The multimodal disease data processing system based on an interpretable model according to claim 5, characterized in that: Cross entropy is used as the loss function during the training of the pre-trained image encoder.
7. The multimodal disease data processing system based on an interpretable model according to claim 2, characterized in that: Each row of the report extraction concept library is the embedding of the concept C learned by the concept activation vector method. In the process of building the report extraction concept library, learning the concept activation vector first requires dividing the positive and negative samples. For concept i, the radiological images containing concept i are taken as positive samples, and the radiological images not containing concept i are taken as negative samples. Then, the pre-trained image encoder φ is used to encode the radiological images to obtain the positive sample embedding vector P i and negative sample embedding vector N i , and then train a support vector machine to classify the positive sample embedding vector and the negative sample embedding vector. The normal vector of the support vector machine classification boundary is used as the embedding of concept i and stored as the report extraction concept library E C The row vector of in It is the normal vector of the support vector machine classification boundary.
8. The multimodal disease data processing system based on an interpretable model according to claim 7, characterized in that: The multimodal concept score is expressed as: in, The points represent the concept scores of the three modalities; The concept score calculation formula for each modality is: Among them, E C Represents the report extraction concept library, Indicates the mth radiological image Encode the image.
9. The multimodal disease data processing system based on an interpretable model according to claim 8, characterized in that: The linear classifier is a learnable weight matrix W. The method for calculating the classification result data includes: applying the Sigmoid activation function to the linear classifier, and then summing it with the multimodal concept score Perform element-wise multiplication to generate the attention matrix Where n represents the number of disease categories, N represents the number of concepts in the concept library, and finally the sum is taken along dimension N, and the final disease prediction is obtained through Softmax. That is, the classification result data, the formula is: Among them, ⊙ represents element-by-element multiplication.
Citation Information
Cited By
Explanatable text classification method based on large model concept generation
CN120929605A