Explainable medical image classification system based on multi-modal prototype network

CN117392473BActive Publication Date: 2026-09-25QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311426940.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2026-09-25
Estimated Expiration
2043-10-30

AI Technical Summary

Technical Problem

除此之外,传统的基于原型的解决方案只关注原型在像素特征的相似度,忽略了原型的位置信息,在医学图像中,一些疾病往往发生在相似位置,这些信息并没有被利用

Benefits of technology

[0012]本发明中设计了一种包括特征提取层、多模态注意力层,位置嵌入层,原型层和分类层的可解释医学图像分类模型;多模态注意力层解决了长期困扰于其他基于原型模型中普遍存在的局限性问题,即存在不明显的原型含义,可以帮助训练更为准确的原型。位置嵌入层证明了在原型中嵌入其他信息的可能性,提升了模型分类的精度。设计的原型激活限制性损失,抑制了疾病原型在非文本关联区域的激活,促进了原型学习,远离了可能出现在不是原型指定类中的任何特性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117392473B_ABST
    Figure CN117392473B_ABST
Patent Text Reader

Abstract

The application discloses an interpretable medical image classification system based on a multimodal prototype network, which comprises the following steps: taking a medical image to be classified, inputting the medical image to be classified into a trained multimodal prototype network, and outputting an interpretable medical image classification result; wherein the trained multimodal prototype network extracts image features from the medical image to be classified to obtain an image feature map, embeds position features in the image feature map, divides the feature map with embedded position information into a plurality of potential patches, calculates the distance between each potential patch and a known disease prototype, finds the potential patch closest to the prototype, visualizes the original medical image region in the same position as the closest potential patch, converts the distance into a similarity score, converts the similarity score into a prediction score, and obtains a medical image classification result; wherein the known disease prototype is a feature map corresponding to a known lesion image region in a training set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an interpretable medical image classification system based on a multimodal prototype network. Background Technology

[0002] The statements in this section merely refer to the background art related to this invention and do not necessarily constitute prior art.

[0003] Medical image processing refers to the analysis and processing of medical images using computer image processing technology. It can assist doctors in performing qualitative and quantitative analysis of lesions and other regions of interest, thereby greatly improving the accuracy and reliability of medical diagnosis. In recent decades, automated diagnostic methods using deep neural networks have achieved high performance; however, the lack of interpretability of these models has led to their limited clinical use. This is because medical decisions are life-threatening, requiring models not only to have high accuracy but also to provide the rationale behind their reasoning. Explainable Artificial Intelligence (XAI) aims to study interpretable models while maintaining high levels of learning performance and predictive accuracy.

[0004] The inventors discovered that prototype networks have attracted considerable attention from researchers in recent years. Prototypes provide classification criteria by comparing the similarity between region patches in a test image and feature prototypes. Existing methods inevitably generate prototypes in recurring, similar medical background regions, potentially resulting in prototypes with disease-irrelevant features. In chest X-ray images, most areas are repetitive healthy regions, while lesion areas are very small and sparse, creating obstacles to generating accurate disease prototypes. Recent research indicates that machine learning models are prone to learning spurious associations between medically irrelevant features (e.g., patterns of healthy tissue) and prediction targets (e.g., types of tumor margins). Furthermore, traditional prototype-based solutions only focus on the similarity of pixel features, ignoring the location information of the prototypes. In medical images, some diseases often occur in similar locations, and this information is not utilized. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides an interpretable medical image classification system based on a multimodal prototype network. It leverages multimodal data to introduce expert knowledge. Notably, the model uses medical images and their corresponding medical text reports only during the training phase, while during the testing phase, it uses only the medical images to be tested. The network provides semantic support for prototype training by utilizing medical text reports. Compared to other methods, the model ensures that prototypes are generated in regions with dense medical semantics rather than useless medical background regions, providing expert guidance for prototype training and improving the model's interpretability. This invention designs a location embedding layer, allowing the generated prototypes to carry location information. Furthermore, this invention proposes a multi-factor similarity calculation method, which enables the model to integrate pixel and location information for classification decisions.

[0006] An interpretable medical image classification system based on multimodal prototype networks includes:

[0007] The acquisition module is configured to acquire a training set, which consists of medical images and medical diagnostic reports of known healthy and diseased regions.

[0008] The training module is configured to: input the training set into the multimodal prototype network, train the network to obtain the trained multimodal prototype network;

[0009] The output module is configured to: acquire the medical image to be classified, input the medical image to be classified into the trained multimodal prototype network, and output an interpretable medical image classification result;

[0010] The trained multimodal prototype network extracts image features from the medical images to be classified, obtaining image feature maps. Location features are embedded into these feature maps, which are then divided into several potential patches. The distance between each potential patch and a known disease prototype is calculated. The nearest potential patch to the prototype is found, and the original medical image region located at the same position as the nearest potential patch is visualized. Simultaneously, the distance is converted into a similarity score, and the similarity score is converted into a prediction score to obtain the medical image classification result. The known disease prototypes are the feature maps corresponding to known lesion image regions in the training set.

[0011] The above technical solution has the following advantages or beneficial effects:

[0012] This invention designs an interpretable medical image classification model comprising a feature extraction layer, a multimodal attention layer, a location embedding layer, a prototype layer, and a classification layer. The multimodal attention layer addresses a long-standing limitation common to other prototype-based models, namely the presence of unclear prototype meaning, thus aiding in training more accurate prototypes. The location embedding layer demonstrates the possibility of embedding other information into the prototype, improving the model's classification accuracy. The designed prototype activation-restricted loss suppresses the activation of disease prototypes in non-textually associated regions, promoting prototype learning and avoiding any features that might appear in classes not specified by the prototype. Attached Figure Description

[0013] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0014] Figure 1 This is a flowchart of the model according to Embodiment 1 of the present invention;

[0015] Figure 2 This is a schematic diagram of the multimodal attention module in Embodiment 1 of the present invention;

[0016] Figure 3 This is a schematic diagram of the position embedding layer in Embodiment 1 of the present invention;

[0017] Figure 4 This is a prototype visualization of Embodiment 1 of the present invention. Detailed Implementation

[0018] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0019] Example 1

[0020] This embodiment provides an interpretable medical image classification system based on a multimodal prototype network;

[0021] An interpretable medical image classification system based on multimodal prototype networks includes:

[0022] The acquisition module is configured to acquire a training set, which consists of medical images and medical diagnostic reports of known healthy and diseased regions.

[0023] The training module is configured to: input the training set into the multimodal prototype network, train the network to obtain the trained multimodal prototype network;

[0024] The output module is configured to: acquire the medical image to be classified, input the medical image to be classified into the trained multimodal prototype network, and output an interpretable medical image classification result;

[0025] The trained multimodal prototype network extracts image features from the medical images to be classified, obtaining image feature maps. Location features are embedded into these feature maps, which are then divided into several potential patches. The distance between each potential patch and a known disease prototype is calculated. The nearest potential patch to the prototype is found, and the original medical image region located at the same position as the nearest potential patch is visualized. Simultaneously, the distance is converted into a similarity score, and the similarity score is converted into a prediction score to obtain the medical image classification result. The known disease prototypes are the feature maps corresponding to known lesion image regions in the training set.

[0026] Furthermore, the training set is input into the multimodal prototype network to train the network and obtain the trained multimodal prototype network, wherein the multimodal prototype network structure includes:

[0027] Image feature extraction layer and text feature extraction layer;

[0028] The input of the image feature extraction layer is used to input medical images, and the input of the text feature extraction layer is used to input medical diagnostic reports.

[0029] The output of the image feature extraction layer is connected to the input of the location embedding layer and the input of the multimodal attention layer, respectively.

[0030] The output of the text feature extraction layer is connected to the input of the multimodal attention layer;

[0031] The output of the position embedding layer is connected to the input of the prototype layer;

[0032] The output of the multimodal attention layer is connected to the input of the prototype layer;

[0033] The output of the prototype layer is connected to the input of the classification layer, and the output of the classification layer is used to output the classification result.

[0034] like Figure 1 As shown, the medical image classification model in this embodiment can be explained by including a feature extraction layer, a multimodal attention layer, a position embedding layer, a prototype layer, and a classification layer.

[0035] Furthermore, the image feature extraction layer is implemented using a ResNet-50 network.

[0036] z = p z (E z (x z ))

[0037] Where z represents the encoded image feature, p z E represents a nonlinear image projector. z Indicates the ResNet50 encoder, x z This represents the original input image;

[0038] For example, Reset-50, which is pre-trained on ImageNet, is used as the image encoder.

[0039] t = p t (E t (x t ))

[0040] Where t represents the encoded text feature, p t E represents a text nonlinear projector. t Indicates the BERT encoder, x t This indicates the input medical text report;

[0041] Furthermore, the text feature extraction layer is implemented using a BERT network. To better extract text feature reports from medical sources, BERT is used as the text feature extractor. A non-linear projection function is used to project image features and text features into the joint embedding space, respectively.

[0042] Furthermore, such as Figure 3 As shown, the location embedding layer is used to embed location information;

[0043]

[0044]

[0045]

[0046]

[0047] Where x and y represent the horizontal and vertical position indices, respectively, and i, j ∈ [0, D / 4] represent the dimensions. Position features PE(x, y, 2i), PE(x, y, 2i+1), PE(x, y, 2j+D / 2), and PE(x, y, 2j+D / 2) are embedded into the feature map; PE(x, y, 2i) represents the embedded horizontal position feature, PE(x, y, 2i+1) represents the embedded horizontal position feature, PE(x, y, 2j+D / 2) represents the embedded vertical position feature, and PE(x, y, 2j+D / 2) represents the embedded vertical position feature.

[0048] To avoid affecting prototype projection and visualization, a feature concatenation method is used to combine the position embedding with the feature map, resulting in a new representation containing position encoding. for:

[0049]

[0050] Here, Concat() represents element concatenation, where z represents the image pixel features and PE represents the embedded position feature vector.

[0051] Understandably, 2D perceptual position embedding is used to concatenate horizontal and vertical positional information with feature maps to generate vector representations with 2D positional information. The generated prototype also carries positional information. The positional encoding has the same size and dimension as the image feature map. Specifically, a sine or cosine signal is generated in the horizontal or vertical direction, and all sine or cosine signals are concatenated into a D-dimensional array. The first D / 2 dimensions describe the horizontal position, and the last D / 2 dimensions describe the vertical position. The advantage of this positional encoding technique is that it does not add new trainable parameters to the neural network.

[0052] Furthermore, such as Figure 2 As shown, the multimodal attention layer is used to calculate similarity using image features and text features to generate a multimodal attention matrix:

[0053] First, calculate the dot product similarity between text features and all representational sub-regions of the image:

[0054] S i =t·z i

[0055] Among them, S i The similarity between text features and the feature map of the i-th sub-region of the image is represented by z, where t represents the encoded text features, and z represents the similarity between the text features and the feature map of the i-th sub-region of the image. i This represents the i-th feature vector of the encoded image feature map;

[0056] Apply the ReLU activation function to the attention map to make the attention weights between dissimilar image text regions zero;

[0057] H i =max(0, S) i )

[0058] Among them, H i S represents the similarity score processed by the Reu function. i This represents the similarity between the text features and the feature map of the i-th sub-region of the image, with max indicating the maximum value.

[0059] Calculate multimodal attention for a sub-region of an image, with attention weights a. i It is the normalized similarity of text features across all image regions:

[0060]

[0061] in, It is a temperature parameter, H i H represents the similarity score after processing by the ReLU function. j The similarity score is represented by the ReLU function, N represents the number of feature vector patches in the feature map, and j represents the j-th patch.

[0062] The multimodal attention weight matrix is ​​calculated based on image and text features. The attention learned by the model weighs the meaning of a given statement according to different image sub-regions.

[0063] Furthermore, the prototype layer is composed of C groups of prototype units, where C is the number of diseases, and each group of prototype units g contains K disease prototypes; the function of each prototype unit g is to calculate the Euclidean distance between the disease prototype of the unit and each patch z of the feature map.

[0064] The prototype layer refers to the layer that uses a multi-factor similarity mechanism to calculate the Euclidean distance between the potential patch of the feature map and the disease prototype, and converts the distance into a similarity score.

[0065]

[0066] Among them, z 1 , These are the visual features of potential image patches and disease prototypes, respectively. 2 , α and β are the location embedding features of the potential image patch and the disease prototype, respectively; α and β are the hyperparameters of visual feature similarity and location feature similarity, respectively. This represents the calculation process of the prototype unit. These represent the patches of the feature map after embedding location information (the feature map is 7x7, with a total of 49 patches, each of which is called a patch). here (Referring to each of the 49 patches);

[0067] Image latent patches are obtained by meshing the image feature map; the size of each patch is the same as the size of the prototype.

[0068] The visual features of the disease prototype are the visual features extracted from lesion images of known disease types in the training set using ResNet50.

[0069] The location embedding feature of the disease prototype is the location feature extracted from lesion images of known disease types in the training set.

[0070] It should be understood that both the prototype vector and image features consist of two parts: pixel features and positional features. Therefore, a multi-factor similarity method is used, which calculates image information and positional information separately. Through this process, the classification result can be obtained by comprehensively considering the similarity between the image features and positional features between the corresponding prototype and the input image.

[0071] The prototype layer uses prototype activation constraint loss and leverages a multimodal attention matrix to assist prototype training and activation; classification results are derived based on the patch closest to the prototype. Multiple disease prototypes are learned for each disease on the feature maps of medical chest X-rays in the training set. Using prototype activation constraint loss effectively limits the activation regions of non-text-related disease prototypes.

[0072] The feature map is input to the prototype layer, which can find the patch closest to the disease prototype, providing a basis for image classification. The disease prototype is defined as a trainable tensor of shape (H1*W1*D) for a known disease type, where H1 < H and W1 < W. Simultaneously, an unbiased generalized convolutional form can be used, where the k-th prototype of disease type c is set. Acting as a kernel, through a receptive field of shape (H*W*D) Slide up and calculate the prototype. Its current receiving domain Euclidean distance between them, receptive field It is called a patch;

[0073] Apply minimum pooling to select the closest prototype in the receptive field z. The shape is (H1*W) l *D) patches; recent potential patches and prototypes The distance between them determines the extent to which the prototype exists in the input image;

[0074] After selecting prototypes, prototype visualization is performed. To facilitate the visualization of decision interpretability, a projection operation is performed on the prototypes, employing the same strategy as the interpretable network ProtoPNet. The patch closest to the prototype is selected as the prototype projection to approximate the prototype representation, thus achieving the purpose of prototype visualization. Each prototype is projected onto the nearest latent feature block in the same class as that prototype, allowing us to view any prototype that contributes to image classification decisions. H1*W1*D is the dimension of the prototype, and H*W*D is the dimension of the feature map (7x7xD).

[0075] Furthermore, the classification layer is implemented using a fully connected layer. The similarity score is converted into a prediction score using a grouped fully connected layer. In each group of classification layers, the similarity is calculated by considering only the prototype corresponding to one disease category.

[0076] Predicted score p(y) c |x):

[0077]

[0078] Where σ represents the sigmoid activation function, Representing the prototype The weights corresponding to the generated similarity scores represent the importance of each prototype to the classification.

[0079] Multi-label classification is a binary classification problem for each category. In the classification layer, this invention uses a grouped fully connected layer to achieve this. In each group classification layer, this invention only considers the prototypes of the set 13 disease types c to calculate the similarity score.

[0080] Furthermore, the training set is input into the multimodal prototype network to train the network and obtain the trained multimodal prototype network. The network training process includes:

[0081] The medical images in the training set are input into the image feature extraction layer, and the extracted image feature maps are output.

[0082] The medical diagnostic reports from the training set are input into the text feature extraction layer, which outputs the extracted text features.

[0083] By embedding location information into the image feature map, an image feature map with the embedded location is obtained.

[0084] Both image feature maps and text features are input into the multimodal attention layer. The feature similarity between image features and text features is calculated to generate a multimodal attention matrix. Based on the multimodal attention matrix, a prototype activation restriction loss function is constructed.

[0085] The image feature map of the embedding location is input into the prototype layer. The prototype layer divides the image feature map of the embedding location into a grid to obtain several potential feature map patches. The Euclidean distance between each potential feature map patch and the known disease prototype is calculated, and the Euclidean distance is converted into a similarity score. The fully connected layer converts the similarity score into a prediction score and gives the image classification result. Here, the known disease prototype refers to the image feature map corresponding to the lesion region of a known disease type. The size of the known disease type is the same as the size of the potential feature map patch.

[0086] When the total loss function value of the network no longer decreases, training is stopped, and the trained multimodal prototype network is obtained.

[0087] Training the entire network requires learning the image encoder E. z The parameters are used for image feature mapping, text encoder E t Used for text feature mapping. A nonlinear projector p for the joint embedding semantic space of images and text. z p t Learning Prototypes and fully connected layer parameters

[0088] Classification Loss. The model struggles to learn positive instances (images with pathological findings) during training, likely because the image labels are very sparse, with far more "0"s than "1"s. To address the class label imbalance problem, a weighted balance loss is used to enhance the learning of positive instances.

[0089]

[0090] in, Represents the classification loss function; and These represent the number of samples labeled "0" and "1" for disease c, respectively. It is the i-th sample x i The predicted score; γ is the balance parameter. It is sample x i The true label in category c;

[0091] Prototype activation constraint loss:

[0092]

[0093] Among them, L res This represents the prototype activation-restricted loss function;

[0094] Representing the prototype It does not belong to the category Y to which this image belongs; Representing the prototype It belongs to the category Y of this image; Representing the prototype and Euclidean distance between them; M i represents the multimodal attention matrix of the i-th image; ⊙ represents the Hadamard product;

[0095] Clustering loss and separation loss. Through clustering loss, this invention encourages each positive sample to have some potential patches that are at least close to a prototype of its own type. Through separation loss, this invention encourages each negative sample's patches to be far removed from these prototypes.

[0096]

[0097]

[0098] in, Represents the clustering loss function; Represents the separation loss function;

[0099] y c This represents the true label of the image on disease c; A feature map containing embedded positional features; This represents the k disease prototypes of the c-th disease category.

[0100] To align representations and learn joint embeddings, a training objective for multimodal associations is needed. Here, the present invention sets a contrastive objective for learning multimodal representations.

[0101] For batches of size N, a symmetric contrast loss for global alignment of image and text projections helps the model learn shared latent semantics. Medical reports contain detailed descriptions of medical images, so it is expected that paired images and reports have similar semantic information in a multimodal semantic space.

[0102] The model uses a contrastive loss function to minimize the negative log-posterior probability:

[0103]

[0104] Where τ2 is the scaling parameter, <z i , t i > represents the cosine similarity between image representations and text features.

[0105] Furthermore, the total loss function of the network is expressed as:

[0106] L = L cls +λ cont L cont +λ res Lres +λ clst L clst +λ sep L sep

[0107] Where, λ cont , λ res , λ clst , λ sep The hyperparameters are used to balance the loss.

[0108] Prototype Visualization: The learned latent prototypes need to be projected onto the training image for interpretability. Specifically, this invention replaces the prototype with the patch closest to it in the training image. These latent patches are naturally the parts of the corresponding prototype with the strongest activation. This is achieved by upsampling the activation map generated by the prototype unit to the size of image x, where the strongest activation patch of x is indicated by the high activation region in the (upsampled) activation map. Since the prototype and latent features of this invention include both image features and location features, this invention only uses the image feature portion for prototype projection.

[0109] Each prototype Expressed using a formula:

[0110]

[0111]

[0112] in, The pixel features representing the disease prototype, z 1 Z represents the pixel features of a potential patch in a feature map, where Z represents the feature map and z represents each potential patch in the feature map. Some examples of prototype visualizations are as follows: Figure 4 As shown.

[0113] In our experiments, we applied this method to two authoritative multi-label datasets, MIMIC-CXR and OpenI. In subsequent experiments, we divided these three datasets into three subsets: training, testing, and validation sets, and compared them with a series of baseline models to validate the model's effectiveness.

[0114] The model in this embodiment was compared with other baseline models on the MIMIC-CXR and OpenI datasets, and the experimental results showed a significant improvement. The image to be classified and the paired medical report obtained image and text features through a feature extraction layer; similarity was calculated using the image and text features to generate a multimodal attention matrix. 2D perceptual location embedding was used to embed location information into the image features; the prototype layer used prototype activation loss to effectively limit prototype activation in non-text-related regions, calculated the Euclidean distance between potential patches in the feature map and disease prototypes, generated a similarity score, and converted the similarity score into a predicted score to achieve image classification decision and obtain the classification result. This approach solves the problems of inaccurate prototype generation, lack of basis for prototype generation, and loss of location information in traditional models, which easily lead to classification errors.

[0115] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An interpretable medical image classification system based on multimodal prototype networks, characterized by: include: The acquisition module is configured to acquire a training set, which consists of medical images and medical diagnostic reports of known healthy and diseased regions. The training module is configured to: input the training set into the multimodal prototype network, train the network to obtain the trained multimodal prototype network; The multimodal prototype network structure includes a prototype layer, which consists of C groups of prototype units, where C is the number of diseases. Each group of prototype units g contains K disease prototypes. The function of each prototype unit g is to calculate the Euclidean distance between the disease prototype of that unit and each patch z of the feature map. The prototype layer uses a multi-factor similarity mechanism to calculate the Euclidean distance between potential patches of the feature map and disease prototypes, and converts the distance into a similarity score. in, , These are the visual features of potential image patches and disease prototypes, respectively. , These are the location embedding features of potential image patches and disease prototypes, respectively; , These are the hyperparameters for visual feature similarity and positional feature similarity, respectively. This represents the calculation process of the prototype unit. These represent the individual patches of the feature map after embedding location information. The output module is configured to: acquire the medical image to be classified, input the medical image to be classified into the trained multimodal prototype network, and output an interpretable medical image classification result; The trained multimodal prototype network extracts image features from the medical images to be classified, obtaining image feature maps. Location features are embedded into these feature maps, which are then divided into several potential patches. The distance between each potential patch and a known disease prototype is calculated. The nearest potential patch to the prototype is found, and the original medical image region located at the same position as the nearest potential patch is visualized. Simultaneously, the distance is converted into a similarity score, and the similarity score is converted into a prediction score to obtain the medical image classification result. The known disease prototypes are the feature maps corresponding to known lesion image regions in the training set.

2. The interpretable medical image classification system based on multimodal prototype networks as described in claim 1, characterized in that, The training set is input into the multimodal prototype network to train the network, resulting in a trained multimodal prototype network. The multimodal prototype network structure includes: Image feature extraction layer and text feature extraction layer; The input of the image feature extraction layer is used to input medical images, and the input of the text feature extraction layer is used to input medical diagnostic reports. The output of the image feature extraction layer is connected to the input of the location embedding layer and the input of the multimodal attention layer, respectively. The output of the text feature extraction layer is connected to the input of the multimodal attention layer; The output of the position embedding layer is connected to the input of the prototype layer; A prototype activation constraint loss function is constructed between the multimodal attention layer and the prototype layer; The output of the prototype layer is connected to the input of the classification layer, and the output of the classification layer is used to output the classification result.

3. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The image feature extraction layer is implemented using a ResNet-50 network. = in, This represents the encoded image features. This represents a non-linear image projector. Indicates a ResNet50 encoder. This represents the original input image.

4. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The text feature extraction layer is implemented using the BERT network.

5. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The location embedding layer is used to embed location information; in, These represent the horizontal and vertical position indices, respectively. ∈[0,D / 4] represents the dimension; location features Embedded into the feature map; This represents the horizontal positional features of the embedding. This represents the embedded horizontal position features. This represents the embedded vertical position feature. Represents the vertical positional features of the embedding; The location embedding is concatenated with the feature map using feature concatenation, resulting in a new representation containing the location encoding. for: (10) in, This indicates element concatenation, where z represents the image pixel features and PE represents the embedded position feature vector.

6. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The multimodal attention layer is used to calculate similarity using image features and text features, and generate a multimodal attention matrix. First, calculate the dot product similarity between text features and all representational sub-regions of the image: in, Representing text features and image features i Similarity between feature maps of individual sub-regions Represents the features of the encoded text. The first part represents the encoded image feature map. i 1 eigenvector; Apply the ReLU activation function to the attention map to make the attention weights between dissimilar image text regions zero; in, This represents the similarity score after processing with the ReLU function. This represents the similarity between text features and the feature map of the i-th sub-region of the image. This indicates taking the maximum value; Calculate multimodal attention for image sub-regions, and the attention weights. It is the normalized similarity of text features across all image regions: in, It's a temperature parameter. This represents the similarity score after processing with the ReLU function. This represents the similarity score after processing with the ReLU function. This indicates the number of feature vector patches in the feature map. Indicates the first One patch.

7. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The classification layer is implemented using a fully connected layer. The similarity score is converted into a predicted score using a grouped fully connected layer. In each group of classification layers, only the prototype corresponding to one disease category is considered to calculate the similarity. Predicted score : (12) in, This represents the sigmoid activation function. Representing the prototype The weights corresponding to the generated similarity scores represent the importance of each prototype to the classification.

8. The interpretable medical image classification system based on a multimodal prototype network as described in claim 2, characterized in that, The training set is input into the multimodal prototype network, and the network is trained to obtain the trained multimodal prototype network. The network training process includes: The medical images in the training set are input into the image feature extraction layer, and the extracted image feature maps are output. The medical diagnostic reports from the training set are input into the text feature extraction layer, which outputs the extracted text features. By embedding location information into the image feature map, an image feature map with the embedded location is obtained. Both image feature maps and text features are input into the multimodal attention layer. The feature similarity between image features and text features is calculated to generate a multimodal attention matrix. Based on the multimodal attention matrix, a prototype activation restriction loss function is constructed. The image feature map of the embedding location is input into the prototype layer. The prototype layer divides the image feature map of the embedding location into a grid to obtain several potential feature map patches. The Euclidean distance between each potential feature map patch and the known disease prototype is calculated, and the Euclidean distance is converted into a similarity score. The fully connected layer converts the similarity score into a prediction score and gives the image classification result. Here, the known disease prototype refers to the image feature map corresponding to the lesion region of a known disease type. The size of the known disease type is the same as the size of the potential feature map patch. When the total loss function value of the network no longer decreases, training is stopped, and the trained multimodal prototype network is obtained.

9. The interpretable medical image classification system based on a multimodal prototype network as described in claim 8, characterized in that, The total loss function of the network is expressed as: in, , , Hyperparameters for balancing losses; in, Represents the classification loss function; | and | These represent the number of samples labeled "0" and "1" for disease c, respectively. It is the i-th sample The predicted score; For balancing parameters, It is a sample The true label in category c; Prototype activation constraint loss: in, This represents the prototype activation-restricted loss function; Representing the prototype This image does not belong to any category. ; Representing the prototype This image belongs to the category ; Representing the prototype and The Euclidean distance between them; This represents the multimodal attention matrix for the i-th image; It represents the Hadamardi (or Hadama) stack; in, Represents the clustering loss function; Represents the separation loss function; This represents the true label of the image on disease c; A feature map containing embedded positional features; This represents the k disease prototypes of the c-th disease category; Each prototype Expressed using a formula: in, Pixel features representing disease prototypes The pixel features representing potential patches in the feature map Representing feature maps, This represents each potential patch of the feature map; The model uses a contrastive loss function to minimize the negative log-posterior probability: in, It's a scaling parameter. It represents the cosine similarity between image representations and text features.