Explanatable Transform medical diagnosis method based on prototype learning
By introducing prototype learning mechanism and CvT architecture into the Transformer model, the self-attention layer is optimized, and the problems of large amount of computation and poor interpretation of traditional Transformer in image data processing are solved, and efficient and interpretable medical image diagnosis is achieved.
Patent Information
- Application Number
- CN202510323104.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional convolutional neural networks have limitations in the extraction of local feature of image data and modeling of global context information. The self-attention mechanism is computationally expensive and difficult to explain during image processing. The existing Transformer architecture lacks flexibility in modeling dependencies at different scales or regions.
The Convolutional Vision Transformer (CvT) architecture is adopted to combine convolutional neural networks and Transformer, and a prototype learning mechanism is introduced to optimize the self-attention layer. A new attention mechanism is generated through prototype features and attention weights, combined with cross-entropy loss, focus loss and L2 regularization optimization model, and the model is trained using the Adam algorithm.
It improves the diagnostic accuracy and efficiency of medical image data, reduces the computational complexity, provides model interpretability, enhances the ability to locate key information, and improves the accuracy and credibility of medical diagnosis.
Smart Images

Figure CN120280122A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence and medical diagnosis, and particularly relates to an interpretable Transformer medical diagnosis method based on prototype learning. Background Art
[0002] Traditional Convolutional Neural Networks (CNNs) excel at extracting local features of images, but they often have limited ability to capture global context information. In contrast, the Transformer architecture can well model the global dependencies of sequence data through the self-attention mechanism. However, when dealing with image data, its computational complexity and the number of parameters may increase sharply.
[0003] To overcome these challenges, researchers have begun to explore methods of combining the advantages of Transformers and CNNs. The Convolutional Vision Transformer (CvT) architecture is one of the important achievements of this exploration. By combining convolutional operations with the self-attention mechanism of Transformers, CvT not only retains the advantages of convolutional operations in extracting local features but also strengthens the ability to model global context information through the self-attention mechanism.
[0004] However, although architectures such as CvT have achieved good performance in tasks such as image recognition, they still have some limitations. For example, the traditional self-attention mechanism usually performs pixel-level operations when dealing with images, which leads to huge computational amounts and is difficult to interpret. In addition, existing Transformer architectures usually lack flexibility in modeling dependencies at different scales or regions. Summary of the Invention
[0005] Aiming at the technical problem that the traditional self-attention mechanism usually performs pixel-level operations when dealing with images, which leads to huge computational amounts and is difficult to interpret, the present invention provides an interpretable Transformer medical diagnosis method based on prototype learning.
[0006] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0007] An interpretable Transformer medical diagnosis method based on prototype learning, comprising the following steps:
[0008] S1. Data preprocessing: Preprocess the medical image data, including normalization, denoising, and enhancement operations, to improve the data quality and the stability of model training;
[0009] S2. Prototype Feature Extraction: Use unsupervised learning methods to extract representative prototype features from the preprocessed medical image data. These prototype features can reflect the typical features of different categories and provide strong support for subsequent classification tasks;
[0010] S3. Model Architecture: Adopt Convolutional Vision Transformer (CvT) as the basic architecture, combine the advantages of convolutional neural network CNN and Transformer, and achieve effective fusion of local features and global context information;
[0011] S4. Attention Mechanism Optimization: Introduce a prototype learning mechanism in the self-attention layer of Transformer, and generate new attention weights by calculating the similarity between the input image and the prototype features;
[0012] S5. Key-Value Pair Storage: Store key-value pairs as parameterized prototypes, which do not depend on input features, enabling the model to use a new region-to-region attention mechanism;
[0013] S6. Loss Function Design: Adopt cross-entropy loss and focal loss to optimize the classification performance of the model. At the same time, introduce a regularization term to prevent overfitting;
[0014] S7. Select the Adam algorithm to train the model, and adjust the hyperparameters of the learning rate and batch size according to the actual situation;
[0015] S8. Model Evaluation and Tuning: Evaluate the performance of the model on the validation set, including accuracy, recall, and F1 value metrics, and tune the model according to the evaluation results.
[0016] The method of data preprocessing in S1 is as follows:
[0017] S1.1. Scale the pixel values of the image data to a unified range [0, 1] to reduce the influence of illumination changes and accelerate the model training process;
[0018] S1.2. Smooth the image through a convolutional kernel to remove random noise; combine spatial proximity and pixel value similarity to effectively remove noise while retaining edge features;
[0019] S1.3. Process missing data by interpolation, mean filling, or removing outliers; identify and correct outliers to ensure data consistency and accuracy.
[0020] The method for prototype feature extraction in S2 is as follows: The contrastive learning method is used to extract high-dimensional feature representations from unsupervised data, and more informative representations are generated by comparing the correlations between instances; The most representative samples are selected from each category as the prototypes of that category, and the Euclidean distance or other distance metric methods are used to calculate the distances between the samples to be classified and the prototypes of each category. The prototypes are updated iteratively to more accurately reflect the typical features of different categories.
[0021] The method for optimizing the attention mechanism in S4 is as follows:
[0022] S4.1: Embed the prototype feature vector as an additional input into the self-attention layer of the Transformer;
[0023] S4.2: Concatenate the input feature vector with the prototype feature vector to generate new Query, Key, and Value matrices;
[0024] S4.3: Calculate the dot product between Query and Key to obtain the attention scores;
[0025] S4.4: Normalize the attention scores using the Softmax function to generate a new attention weight matrix;
[0026] S4.5: Multiply the attention weights by the Value matrix to obtain the output after weighted summation.
[0027] The method for designing the loss function in S6 is as follows:
[0028] Cross-entropy loss is a standard loss function for measuring the difference between the predicted probability distribution of the model and the true label distribution, and is applicable to multi-classification problems. Its formula is:
[0029]
[0030] where y i is the true label, is the probability value predicted by the model. This loss function can measure the classification ability of the model for each category. By minimizing this loss value, the classification performance of the model can be improved;
[0031] Focal loss is an improved cross-entropy loss, applicable to the problem of class imbalance. It reduces the weight of easy-to-classify samples by introducing a modulation factor γ and enhances the attention to difficult-to-classify samples. Its formula is:
[0032] FL(p t ) = -(1 - p t ) γ log(p t )
[0033] Among them, p t is the probability that the sample belongs to the positive class, γ is a tuning parameter with a value of 2 or higher; by adjusting the size of γ, the contribution to simple samples can be dynamically reduced, thereby enhancing the model's recognition ability for complex samples;
[0034] L2 regularization constrains the model complexity by adding the sum of squares of the weight parameters to the loss function, thereby avoiding overfitting. Its formula is:
[0035]
[0036] Among them, λ is the regularization coefficient, and ω i is the model weight;
[0037] The final combined loss function is expressed as:
[0038] L total = L + αFL(p t ) + βL reg
[0039] Among them, α and β are the weight coefficients of the two, respectively, used to balance the contributions of different loss terms.
[0040] The method of selecting the Adam and RMSprop algorithms to train the model in S7 is: use the Adam optimizer for training. Adam combines the advantages of Momentum and RMSprop. By calculating the first-order moment estimate and second-order moment estimate of the gradient, it can adaptively adjust the learning rate of each parameter, thereby improving the training efficiency and stability.
[0041] Compared with the prior art, the present invention has the following beneficial effects:
[0042] 1. The present invention introduces a prototype learning mechanism to optimize the attention mechanism of the Transformer model. This innovative approach overcomes the shortcoming of insufficient model interpretability in the prior art. By combining prototype features and attention weights, the model can more accurately locate key information and improve the accuracy of medical diagnosis.
[0043] 2. The present invention uses the Convolutional Vision Transformer (CvT) architecture as the basis, combining the advantages of convolutional neural networks and Transformers. It not only retains the ability to extract local features but also strengthens the modeling of global context information. This enables the model to more efficiently extract useful information when processing medical image data, improving performance while reducing computational complexity. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are merely exemplary, and for those of ordinary skill in the art, without creative efforts, other implementation drawings can also be derived based on the provided drawings.
[0045] The structures, ratios, sizes, etc. illustrated in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limiting conditions for the implementation of the present invention. Therefore, they do not have technical substance. Any modification of the structure, change in the proportional relationship, or adjustment of the size, without affecting the effects that the present invention can produce and the purposes that can be achieved, should still fall within the scope covered by the technical content disclosed in the present invention.
[0046] Figure 1 It is a flowchart of the method of the present invention. Specific Embodiments
[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. These descriptions are only to further illustrate the features and advantages of the present invention, rather than a limitation on the claims of the present invention; based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by the present application.
[0048] The following will further describe in detail the specific embodiments of the present invention in combination with the drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0049] An interpretable Transformer medical diagnosis method based on prototype learning, as Figure 1 shown, includes the following steps:
[0050] Step 1, Data preprocessing: Preprocess the medical image data, including normalization, denoising, and enhancement operations to improve the data quality and the stability of model training.
[0051] Step 1.1, Scale the pixel values of the image data to a unified range [0,1] to reduce the influence of illumination changes and accelerate the model training process.
[0052] Step 1.2, Smooth the image through a convolution kernel to remove random noise. Combine spatial proximity and pixel value similarity to effectively remove noise while retaining edge features.
[0053] Step 1.3: Process the missing data by means of interpolation, mean filling, or outlier removal. Identify and correct the outliers to ensure data consistency and accuracy.
[0054] Step 2: Prototype Feature Extraction: Use unsupervised learning methods to extract representative prototype features from the preprocessed medical image data. These prototype features can reflect the typical features of different classes and provide strong support for subsequent classification tasks. Adopt contrastive learning methods to extract high-dimensional feature representations from unsupervised data, and generate more informative representations by comparing the correlations between instances. Select the most representative samples from each class as the prototypes of that class, calculate the distances between the samples to be classified and the prototypes of each class using Euclidean distance or other distance metrics, and update the prototypes iteratively to make them more accurately reflect the typical features of different classes.
[0055] Step 3: Model Architecture: Adopt Convolutional Vision Transformer (CvT) as the basic architecture, combine the advantages of convolutional neural network CNN and Transformer, and achieve effective fusion of local features and global context information.
[0056] Step 4: Attention Mechanism Optimization: Introduce a prototype learning mechanism into the self-attention layer of Transformer, and generate new attention weights by calculating the similarity between the input image and the prototype features.
[0057] Step 4.1: Embed the prototype feature vector as an additional input into the self-attention layer of Transformer.
[0058] Step 4.2: Concatenate the input feature vector with the prototype feature vector to generate new Query, Key, and Value matrices.
[0059] Step 4.3: Calculate the dot product between Query and Key to obtain the attention scores.
[0060] Step 4.4: Normalize the attention scores using the Softmax function to generate a new attention weight matrix.
[0061] Step 4.5: Multiply the attention weights with the Value matrix to obtain the output after weighted summation.
[0062] Step 5: Key-Value Pair Storage: Store the key-value pairs as parameterized prototypes that do not depend on the input features, enabling the model to use a new region-to-region attention mechanism.
[0063] Step 6, Loss Function Design: Cross-entropy loss and focal loss are used to optimize the classification performance of the model. At the same time, a regularization term is introduced to prevent overfitting.
[0064] Cross-entropy loss is a standard loss function that measures the difference between the predicted probability distribution of the model and the true label distribution. It is applicable to multi-classification problems, and its formula is:
[0065]
[0066] where y i is the true label, is the probability value predicted by the model. This loss function can measure the classification ability of the model for each category. By minimizing this loss value, the classification performance of the model can be improved.
[0067] Focal loss is an improved cross-entropy loss, which is applicable to the problem of class imbalance. It reduces the weight of easy-to-classify samples and enhances the attention to difficult-to-classify samples by introducing a modulation factor γ. Its formula is:
[0068] FL(p t ) = -(1 - p t ) γ log(p t )
[0069] where p t is the probability that the sample belongs to the positive class, and γ is the modulation parameter, with a value of 2 or higher. By adjusting the size of γ, the contribution to simple samples can be dynamically reduced, thereby improving the model's recognition ability for complex samples.
[0070] L2 regularization constrains the model complexity by adding the sum of the squares of the weight parameters to the loss function, thereby avoiding overfitting. Its formula is:
[0071]
[0072] where λ is the regularization coefficient and ω i is the model weight.
[0073] The final combined loss function is expressed as:
[0074] L total = L + αFL(p t ) + βL reg
[0075] where α and β are the weight coefficients of the two, respectively, used to balance the contributions of different loss terms.
[0076] Step 7: Select the Adam algorithm to train the model and adjust the hyperparameters of the learning rate and batch size according to the actual situation. Use the Adam optimizer for training. Adam combines the advantages of Momentum and RMSprop. By calculating the first-order moment estimate and second-order moment estimate of the gradient, it can adaptively adjust the learning rate of each parameter, thus improving the training efficiency and stability.
[0077] Step 8: Model evaluation and tuning: Evaluate the performance of the model on the validation set, including accuracy, recall, and F1-score metrics, and tune the model according to the evaluation results.
[0078] Example 1: A medical diagnosis system integrating prototype learning and self-attention mechanism can assist doctors in more refined disease classification and prediction
[0079] Application in medical diagnosis: Through prototype learning, the system can learn the representative features of different diseases, while the self-attention mechanism can capture the complex correlations between these features. This combination may enable the system to more precisely identify the subtle differences of diseases, thereby classifying diseases more accurately. In addition, by learning from historical data, the system may also be able to predict the development trend of diseases, providing a basis for early warning and intervention for doctors, thus improving the treatment effect of patients.
[0080] Example 2: Application of the Transformer model based on prototype learning in medical image diagnosis
[0081] Considering the wide application and importance of medical image data in medical diagnosis, we can envision applying the Transformer model based on prototype learning to medical image diagnosis. By leveraging the deep feature extraction ability of the Transformer model for image data and combining the prototype learning algorithm, the system can learn the representative image features of different diseases. In this way, when doctors make a diagnosis, the system can display prototype images similar to the case and provide corresponding diagnostic suggestions. This not only improves the accuracy and efficiency of diagnosis but also provides doctors with intuitive and interpretable diagnostic basis, helping to improve the quality of medical care and patient satisfaction.
[0082] Consider a practical medical application scenario, namely the diagnosis of chest X-rays. In this scenario, doctors need to identify lung abnormalities in X-rays, such as lung nodules, pneumonia, etc. Traditional diagnostic methods mainly rely on doctors' experience and visual observation, but this method is often limited by doctors' personal abilities and experience levels and is easily affected by factors such as fatigue and emotions.
[0083] Now, we introduce an interpretable Transformer medical diagnosis system based on prototype learning. This system first learns a set of representative prototypes from a large amount of lung X-ray data through the prototype learning algorithm. These prototypes contain different types of lung abnormality features and can be regarded as "templates" for various lesions.
[0084] Next, when the system receives a new X-ray, it uses the self-attention mechanism in the Transformer architecture to globally analyze the image data and find the parts that are most similar to the known prototypes. By comparing the similarity between the new data points and the known prototypes, the system can accurately identify the abnormal areas in the X-ray and give corresponding diagnostic results.
[0085] More importantly, since this system combines the interpretable design of prototype learning and the Transformer architecture, it can not only provide high-precision diagnostic results but also generate visual explanations for the decision-making process. For example, the system can highlight the areas in the X-ray that best match the abnormal prototypes and give corresponding explanations. In this way, doctors can intuitively understand the decision-making basis of the system, thereby enhancing their trust in machine decisions.
[0086] In addition, by collecting and analyzing the diagnostic results and explanations of the system in different cases, doctors can also discover new lesion features and diagnostic rules, further promoting the development of medical research and education.
[0087] The above only elaborates on the preferred embodiments of the present invention in detail. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the purpose of the present invention, and all such changes should be included within the protection scope of the present invention.
Claims
1. An interpretable Transformer medical diagnosis method based on prototype learning, characterized in that, It includes the following steps: S1. Data preprocessing: Preprocess the medical image data, including normalization, denoising, and enhancement operations to improve the data quality and the stability of model training; S2. Prototype feature extraction: Use unsupervised learning methods to extract representative prototype features from the preprocessed medical image data. These prototype features can reflect the typical features of different classes and provide strong support for subsequent classification tasks; S3. Model architecture: Adopt Convolutional Vision Transformer (CvT) as the basic architecture, combine the advantages of convolutional neural network CNN and Transformer to effectively fuse local features and global context information; S4. Attention mechanism optimization: Introduce a prototype learning mechanism in the self-attention layer of Transformer, and generate new attention weights by calculating the similarity between the input image and the prototype features; S5. Key-value pair storage: Store the key-value pairs as parameterized prototypes. These prototypes do not depend on the input features, enabling the model to use a new region-to-region attention mechanism; S6. Loss function design: Adopt cross-entropy loss and focal loss to optimize the classification performance of the model. At the same time, introduce a regularization term to prevent overfitting; S7. Select the Adam algorithm to train the model and adjust the hyperparameters of the learning rate and batch size according to the actual situation; S8. Model evaluation and tuning: Evaluate the performance of the model on the validation set, including accuracy, recall, and F1 value metrics, and tune the model according to the evaluation results.
2. The interpretable Transformer medical diagnosis method based on prototype learning according to claim 1, wherein The method of data preprocessing in S1 is as follows: S1.
1. Scale the pixel values of the image data to a unified range [0, 1] to reduce the influence of light changes and accelerate the model training process; S1.
2. Smooth the image through a convolutional kernel to remove random noise; Combine spatial proximity and pixel value similarity to effectively remove noise while retaining edge features; S1.
3. Process missing data by interpolation, mean filling, or removing outliers; Identify and correct outliers to ensure data consistency and accuracy.
3. The interpretable Transformer medical diagnosis method based on prototype learning according to claim 1, wherein, The method of prototype feature extraction in S2 is as follows: Adopt a contrastive learning method to extract high-dimensional feature representations from unsupervised data, generate more informative representations by comparing the correlations between instances; Select the most representative samples from each category as the prototypes of that category, calculate the distances between the samples to be classified and the prototypes of each category using Euclidean distance or other distance metrics, and iteratively update the prototypes to make them more accurately reflect the typical features of different classes.
4. An interpretable Transformer medical diagnosis method based on prototype learning according to claim 1, characterized in that, The method of attention mechanism optimization in S4 is as follows: S4.
1. Embed the prototype feature vector as an additional input into the self-attention layer of Transformer; S4.
2. Concatenate the input feature vector with the prototype feature vector to generate new Query, Key, and Value matrices; S4.
3. Calculate the dot product between Query and Key to obtain the attention scores; S4.
4. Normalize the attention scores using the Softmax function to generate a new attention weight matrix; S4.
5. Multiply the attention weights by the Value matrix to obtain the output after weighted summation.
5. The interpretable Transformer medical diagnosis method based on prototype learning according to claim 1, characterized in that The method for designing the loss function in S6 is as follows: Cross-entropy loss is a standard loss function for measuring the difference between the predicted probability distribution of the model and the true label distribution, which is applicable to multi-classification problems. Its formula is: where y i is the true label, is the probability value predicted by the model. This loss function can measure the classification ability of the model for each category. By minimizing this loss value, the classification performance of the model can be improved; Focal loss is a modified cross-entropy loss applicable to the problem of class imbalance. It reduces the weight of easy-to-classify samples and enhances the attention to difficult-to-classify samples by introducing a modulation factor γ. Its formula is: FL(p t ) = -(1 - p t ) γ log(p t ) where p t is the probability that the sample belongs to the positive class, γ is a tuning parameter with a value of 2 or higher; by adjusting the magnitude of γ, the contribution of simple samples can be dynamically reduced, thereby enhancing the model's ability to identify complex samples; L2 regularization constrains the model complexity by adding the sum of the squares of the weight parameters to the loss function, thereby avoiding overfitting. Its formula is: where λ is the regularization coefficient and ω i are the model weights; The final comprehensive loss function is expressed as: L total = L + αFL(p t ) + βL reg Among them, α and β are the weight coefficients of the two respectively, used to balance the contributions of different loss terms.
6. The interpretable Transformer medical diagnosis method based on prototype learning according to claim 1, wherein The method for selecting the Adam and RMSprop algorithms to train the model in S7 is as follows: Use the Adam optimizer for training. Adam combines the advantages of Momentum and RMSprop and can adaptively adjust the learning rate of each parameter by calculating the first-order moment estimate and second-order moment estimate of the gradient, thereby improving the training efficiency and stability.
Citation Information
Cited By
Sandy soil liquefaction discrimination method based on CPTU and weighted nonlinear similarity
CN120974283A
A sand liquefaction discrimination method based on CPTU and weighted nonlinear similarity degree
CN120974283B
Acrylic plate performance evaluation method based on double-spectrum in-situ detection
CN121207957A