A medical image classification method based on conditional class-specific prompts
By using a method based on conditional category-specific prompts in the medical image classification task, each category-specific prompt is generated to enhance the perception of medical image details by text prompts, the problem of insufficient generalization ability in the traditional Chinese medicine image classification task is solved, and more efficient medical image recognition and reasoning is achieved.
Patent Information
- Application Number
- CN202411907934.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-12-24
AI Technical Summary
The prior art has the problem of insufficient generalization ability in medical image classification tasks, especially since different categories of medical images may have similar properties, making it difficult for pre-trained models of contrast text-image pairs to distinguish differences between images.
A medical image classification method based on conditional category-specific hints is proposed. By generating each category-specific hints, it enhances the perception of medical image details by text cues. Specific steps include obtaining medical image datasets, building a pre-trained model for contrast text-image pairs, training with specific prompts, and generating specific prompts through a lightweight network.
By generating category-specific prompts, subtle feature information in medical images can be better captured, text encoder perceived image features, enhanced model understanding and classification capabilities of medical images, and improved the generalization ability of pre-trained models for contrasting text-image pairs.
Smart Images

Figure CN119723205B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence deep learning medical image classification, and particularly relates to a medical image classification method based on conditional class-specific prompts. Background Art
[0002] Large-scale vision-language models, such as the contrastive text-image pre-training model CLIP, have excellent performance in downstream tasks. Their powerful vision and language understanding capabilities enable them to process complex multi-modal information. Their large training data enables the model to acquire diverse and comprehensive knowledge, thus having a powerful zero-shot ability. Due to the scarcity of medical data, contrastive language-image pre-training models have attracted attention in the field of medical images.
[0003] The pre-training model CLIP for contrastive text-image pairs contains a vision branch and a language branch. After obtaining the visual and text features of the two branches, the two types of features are matched to predict the image category. A common text prompt is "a photo of a [category name]". After filling in the category name and inputting it into the text encoder, the text feature representation is obtained. Existing methods show that adding the attribute names describing medical images, such as color, texture, and shape, etc., to the text prompt can better drive the CLIP model to focus on the specific features of medical images, thereby improving the model's understanding ability of medical images. The attribute information contained in these medical prompts is obtained from pre-trained language models and visual question-answering models in the medical field according to the category name.
[0004] However, relying solely on the text description of attributes to enhance the model's perception ability of medical images has certain limitations. On the one hand, different categories of medical images may have similar attributes. For example, both "ependymoma" and "medulloblastoma" of brain tumors are located in the posterior cranial fossa and may both present irregular shapes and lobulated edges. On the other hand, the similarity between medical images is relatively high. The pre-training model CLIP for contrastive text-image pairs only focuses on the attribute features obtained according to the category prior and does not necessarily distinguish the differences between images, and also misses the possibility of extracting other discriminative features. These problems significantly affect the generalization ability of the pre-training model for contrastive text-image pairs in medical image classification tasks. Summary of the Invention
[0005] The present invention is to solve the above-mentioned deficiencies existing in the prior art, and proposes a medical image classification method based on conditional class-specific prompts, in order to generate class-specific prompts for each category according to the features of all medical images, so as to characterize the feature attributes of different categories, thereby enhancing the perception of the details of medical images by text prompts and realizing more efficient medical image recognition and reasoning.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A medical image classification method based on conditional class-specific prompts according to the present invention is characterized by comprising the following steps:
[0008] Step S1, obtaining a medical image data set including K categories , where represents the th medical image, represents the total number of medical images; let the true class label of be , and
[0009] The th medical image is divided into fixed-size image patches, and after each image patch is flattened and then projected onto a fixed dimension, the projected one-dimensional vectors are concatenated into an image patch embedding ;
[0010] Construct the learnable prompt , where represents the th learnable prompt; represents the total number of prompts;
[0011] Step S2, constructing a pre-trained model for contrasting text-image pairs, including: an image encoder and a text encoder , wherein the image encoder is used to process the prompt and the image patch embedding to obtain the visual feature vector of the th medical image ;
[0012] Step S3, obtaining the average image feature vector of the kth category by using Equation (1):
[0013] (1)
[0014] In Equation (1), represents the total number of categories, is the visual feature vector of the th medical image belonging to the kth category, is the total number of medical images in the kth category;
[0015] Step S4: Construct a lightweight network and process to obtain specific cues for the k-th category ;
[0016] Step S5: Concatenate the specific cues { | k = 1, 2, …, K} for K categories with the name of each category | k = 1, 2, …, K}, and then input them into the text encoder for processing to obtain text features { | k = 1, 2, …, K}; where, represents the name of the k-th category, and represents the text feature of the k-th category;
[0017] Step S6: Match with { | k = 1, 2, …, K}, and select the category corresponding to the maximum probability as the -th medical image 's predicted category;
[0018] Step S7: Construct a loss function , and use the gradient descent method to train the pre-trained model of the contrastive text-image pair and the lightweight network. Update the cues in the image encoder and the parameter weights of the lightweight network through gradient backpropagation until the loss function converges, so as to obtain an optimal medical image classification model for classifying medical images and outputting predicted categories.
[0019] The feature of the medical image classification method based on conditional category-specific cues according to the present invention also lies in that the image encoder in Step S2 includes: category embedding , position embedding , layers of image self-attention models , where each layer of the image self-attention model consists of a multi-head image self-attention layer and a multi-layer image perceptron block;
[0020] Step S21: Initialize = 1; Define and initialize the visual feature vector output by the ( -1)-th layer of the image self-attention model ;
[0021] Initialize the cue input to the self-attention layer = ; Embed the image patch into the self-attention layer of the i-th layer as the image patch embedding ; Initialize the class embedding input to the i-th layer self-attention layer = = ;
[0022] Step S22, Prompt , Image patch embedding and class embedding are concatenated, then added to the position embedding and input to the self-attention layer of the i-th layer for processing, so as to output the image patch embedding of the i-th layer and class embedding using Equation (2);
[0023] (2)
[0024] In Equation (2), is the self-attention layer of the i-th layer of the visual encoder, and are the class embedding and image patch embedding output by the self-attention layer of the i-th layer;
[0025] Step S221, Input the prompt of the i-th layer , Image patch embedding and class embedding are concatenated, then layer normalization processing is performed, so as to obtain the embedding vector sequence using Equation (3);
[0026] (3)
[0027] Step S222, The h self-attention heads of the self-attention layer of the i-th layer respectively process using Equation (4) to obtain the h multi-head attention representations of the self-attention layer of the i-th layer , so as to obtain the comprehensive feature representation output by the self-attention layer of the i-th layer using Equation (5):
[0028] (4)
[0029] (5)
[0030] In Equation (4) - Equation (5), , , is the self - attention layer of the i - th layer of the projection parameter matrix of the k - th self - attention head, is also the projection parameter matrix, is the activation function, represents the transpose, represents the concatenation operation, represents the self - attention layer of the i - th layer of the multi - head attention representation output by the k - th attention head , is the number of parallel self - attention heads;
[0031] Step S223, obtain the residual representation in the self - attention layer of the i - th layer using Equation (6):
[0032] (6)
[0033] In Equation (6), represents the residual connection;
[0034] Step S224, obtain the image patch embedding output by the self - attention layer of the l - th layer and the class embedding using Equation (7):
[0035] (7)
[0036] In Equation (7), and are the two transformation matrices of the two linear layers, and are the two biases of the two linear layers, is the activation function; represents the layer normalization operation; represents the hint output by the self - attention layer of the l - th layer ;
[0037] Step S225, after copying i + 1 to i, return to Step 221 and execute sequentially until i > J, so as to obtain the image patch embedding{ |i = 1, 2, …, J} output by the self - attention layer{ |i = 1, 2, …, J} and the class embeddings { |i = 1, 2, …, J};
[0038] Step S23, obtain the self - attention layers from the k - th layer to the l - th layer of the self - attention layers using Equation (8), and for any k - th layer of the self - attention layers in the self - attention layers from the k - th layer to the l - th layer, the output image patch embeddings , class embeddings and cues are: :
[0039] (8)
[0040] Step S24, project the class embeddings output by the k - th layer of the self - attention layers to obtain the visual features of the medical image . .
[0041] Furthermore, the said Step S4 includes:
[0042] Step S4.1, obtain the intermediate features of the k - th class using Equation (9):
[0043] (9)
[0044] wherein, is the transformation matrix of the third linear layer, is the bias of the third linear layer;
[0045] Step S4.2, obtain the specific cue of the k - th class using Equation (10):
[0046] (10)
[0047] In Equation (10), is the transformation matrix of the fourth linear layer, is the bias of the fourth linear layer, is the activation function.
[0048] Furthermore, the said Step S5 includes:
[0049] Step S5.1, establish a text encoder including: a multi - layer text self - attention model, where each layer of the text self - attention model consists of a multi - head text self - attention layer and a multi - layer text perceptron block;
[0050] Step S5.2. Concatenate with the corresponding category name to obtain the k-th text expression, thereby obtaining K text expressions;
[0051] Step S5.3. After tokenizing the K text expressions, project the tokenization results onto a fixed dimension to obtain text embeddings ;
[0052] Step S5.4. Initialize = 1; Define and initialize the text self-attention model of the -1-th layer output text embedding = ;
[0053] Step S5.5. Input to the -th layer text self-attention model for processing, and output the text embedding feature vector ;
[0054] Step S5.5.1. The h-th head text self-attention layer of the multi-head text self-attention layer in the -th layer text self-attention model uses Equation (11) to obtain the h-th head text attention representation of the -th layer text self-attention model , and then uses Equation (12) to obtain the comprehensive feature representation output by the multi-head text self-attention layer in the -th layer text self-attention model :
[0055] (11)
[0056] (12)
[0057] In Equations (11) and (12), , , are the three projection parameter matrices of the h-th head text self-attention layer of the -th layer text self-attention model , is the global projection parameter matrix, is the activation function, represents transpose, represents concatenation operation, is the number of parallel text self-attention heads;
[0058] Step S5.5.2: Obtain the residual representation of the layer text self-attention model using Equation (13): :
[0059] (13)
[0060] In Equation (13), represents the residual connection;
[0061] Step S5.5.3: Obtain the text embedding output by the multi-layer text perceptron block of the layer text self-attention model using Equation (14): :
[0062] (14)
[0063] In Equation (14), and are the text transformation matrices of two linear layers respectively, and are the text biases of two linear layers respectively, is the activation function;
[0064] Step S5.5.4: After assigning +1 to , return to Step S5.5.2 and execute sequentially until >L, so as to obtain the text embedding output by the L-th layer text self-attention layer ;
[0065] Step S5.6: Project the last token of the text embedding onto a vector space of a fixed dimension to obtain the text feature vector { |k = 1, 2, …, K}.
[0066] Furthermore, in Step S6, the probability that the th medical image belongs to the k-th category is obtained using Equation (15):
[0067] (15)
[0068] In Equation (15), represents the cosine similarity, represents the exponential function, represents the temperature parameter.
[0069] Further, in step S7, a loss function is constructed using Equation (16). :
[0070] (16)
[0071] In Equation (16), represents the logarithmic function.
[0072] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the medical image classification method, and the processor is configured to execute the program stored in the memory.
[0073] A computer-readable storage medium according to the present invention, characterized in that a computer program is stored on the computer-readable storage medium, and the computer program executes the steps of the medical image classification method when run by a processor.
[0074] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0075] 1. The method of category-specific prompting in the present invention uses a pre-trained model for comparing text-image pairs to extract subtle feature information in the image features that may be able to distinguish medical images, thereby strengthening the text encoder's perception of the image features and improving the generalization ability of the pre-trained model for comparing text-image pairs to medical images.
[0076] 2. The image encoder in the present invention captures the detailed features in the medical images, and these detailed features are stored in the memory bank. The features in the memory bank are transformed into prompts that match the text description by category to better align the semantic information and context of the images.
[0077] 3. A small number of learnable parameters are inserted into the image encoder in the present invention, which better enhances the understanding ability of the pre-trained model CLIP for medical images, so that the model can more accurately identify and distinguish different medical image categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] Figure 1 is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0079] In this embodiment, a medical image classification method based on conditional category-specific prompts mainly captures detailed features in medical images through the image encoder of a pre-trained model for comparing text-image pairs. These detailed features are stored in a memory bank, and the features in the memory bank are transformed into prompts that match text descriptions by category to better align the semantic information and context of the images. At the same time, a small number of learnable parameters are inserted into the image encoder to better enhance the CLIP model's ability to understand medical images. As Figure 1 shown, it includes the following steps:
[0080] Step S1: Obtain a medical image dataset including K categories , where represents the th medical image, represents the total number of medical images; let 's true category label be , and ∈{1, 2, …, K};
[0081] Divide the th medical image into fixed-size image patches, and after flattening each image patch, project it onto a fixed dimension, so as to splice projected one-dimensional vectors into an image patch embedding ;
[0082] Construct the learnable prompt , where represents the th learnable prompt; represents the total number of prompts.
[0083] Step S2: Construct the image encoder of the pre-trained model for comparing text-image pairs , including: class embedding , position embedding , layers of image self-attention models , where each layer of the image self-attention model consists of a multi-head image self-attention layer and a multi-layer image perceptron block; the image encoder is used to process the prompt and the image patch embedding to obtain the visual feature vector of the th medical image ;
[0084] Step S21: Initialize = 1; Define and initialize the -1 layer of the image self-attention model Output visual feature vector ;
[0085] Initialize the self-attention layer of the input i-th layer Tips = ; Embed the image block As the i-th layer of self-attention Image patch embedding ; Initialize input to the first Self-attention layer Category embedding = ;
[0086] Step S22: Prompt , Image Patch Embedding With category embedding After concatenation, embed with position Add and enter Self-attention layer In order to process it, the formula (1) is used to output the Layer image patch embedding With category embedding ;
[0087] (1)
[0088] In formula (2), is the first self-attention layer, and It is Category embedding and image patch embedding output by the self-attention layer;
[0089] Step S221: Input the prompt of the i-th layer , Image Patch Embedding With category embedding After concatenation, layer normalization is performed to obtain the embedding vector sequence using formula (2): ;
[0090] (2)
[0091] Step S222, self-attention layer of the i-th layer The h self-attention heads of Processing is performed to obtain the self-attention layer of the i-th layer h multi-head attention representations , and then use formula (4) to get the self-attention layer of the i-th layer Comprehensive Feature Representation of the Output :
[0092] (3)
[0093] (4)
[0094] In Equation (4) - Equation (5), , , is the self - attention layer of the i - th layer of the th projection parameter matrix of the self - attention head, is also the projection parameter matrix, is the activation function, represents transpose, represents the concatenation operation, represents the self - attention layer of the i - th layer of the th multi - head attention representation output by the self - attention head, , is the number of parallel self - attention heads.
[0095] Step S223: Obtain the residual representation in the self - attention layer of the i - th layer using Equation (5): :
[0096] (5)
[0097] In Equation (6), represents the residual connection;
[0098] Step S224: Obtain the image patch embedding output by the self - attention layer of the th layer and the class embedding : :
[0099] (6)
[0100] In Equation (6), and are the two transformation matrices of the two linear layers, and are the two biases of the two linear layers, is the activation function; represents the layer normalization operation; represents the prompt output by the self - attention layer of the th layer ;
[0101] Step S225: After copying i + 1 to i, return to Step 221 and execute sequentially until i > J, so as to obtain the image patch embeddings output by the self-attention layers of layer J { | i = 1, 2, …, J} and the class embeddings { | i = 1, 2, …, J}. | i = 1, 2, …, J}.
[0102] Step S23: Use Equation (7) to obtain the self-attention layer of the th layer to the self-attention layer of the th layer, and the image patch embeddings, class embeddings, and cues output by any kth self-attention layer in the self-attention layer of the th layer to the self-attention layer of the th layer: (7) (7) and cues :
[0103] (7)
[0104] Step S24: After projecting the class embeddings output by the self-attention layer of the th layer, obtain the visual features of the medical image . of the medical image .
[0105] Step S3: Use Equation (8) to obtain the average image feature vector of the kth class :
[0106] (8)
[0107] In Equation (8), represents the total number of classes, is the visual feature vector of the th medical image belonging to the kth class , and is the total number of medical images in the kth class.
[0108] Step S4: Construct a lightweight network and process to obtain the specific cue of the kth class;
[0109] Step S4.1: Use Equation (9) to obtain the intermediate feature of the kth class:
[0110] (9)
[0111] wherein, is the transformation matrix of the third linear layer, is the bias of the third linear layer;
[0112] Step S4.2: Use formula (10) to obtain the specific prompt of the kth category :
[0113] (10)
[0114] In formula (10), is the transformation matrix of the fourth linear layer, is the bias of the fourth linear layer, is the activation function.
[0115] Step S5: Category-specific tips |k=1,2,…,K} and the name of each category |k=1,2,…,K} after corresponding concatenation, input text encoder Processed in, get the text features { |k=1,2,…,K}; where represents the name of the k-th category, Represents the text features of the k-th category;
[0116] Step S5.1: Create a text encoder include: Layer text self-attention model, where each layer of text self-attention model consists of a multi-head text self-attention layer and a multi-layer text perception machine block;
[0117] Step S5.2: The corresponding category name After concatenation, the k-th sentence text representation is obtained, thereby obtaining K sentence text representations.
[0118] Step S5.3: After segmenting the K sentences, project the segmentation results onto a fixed dimension to obtain text embedding. ;
[0119] Step S5.4: Initialization =1; define and initialize the -1 layer text self-attention model Output text embedding = ;
[0120] Step S5.5: Input to Layer Text Self-Attention Model Processed in, output text embedding feature vector 。
[0121] Step S5.5.1, the h-th head text self-attention layer of the multi-head text self-attention layer in the layer text self-attention model obtains the h-th layer text self-attention model head text attention representation using Equation (11), and thus obtains the comprehensive feature representation output by the multi-head text self-attention layer in the layer text self-attention model using Equation (12): :
[0122] (11)
[0123] (12)
[0124] In Equations (11) and (12), , , are the three projection parameter matrices of the h-th head text self-attention layer of the layer text self-attention model, is the global projection parameter matrix, is the activation function, represents transpose, represents concatenation operation, is the number of parallel text self-attention heads.
[0125] Step S5.5.2, obtain the residual representation of the layer text self-attention model
[0126] using Equation (13):
[0127] (13) In Equation (13),
[0128] represents residual connection; Step S5.5.3, obtain the text embedding output by the multi-layer text perceptron block of the layer text self-attention model
[0129] using Equation (14):
[0130] (14) In Equation (14), and They are the text transformation matrices of two linear layers respectively, and are the text biases of two linear layers respectively, is the activation function.
[0131] Step S5.5.4: After assigning +1 to , return to Step S5.5.2 and execute sequentially until >L, so as to obtain the text embedding output by the L-th layer text self-attention layer ;
[0132] Step S5.6: Project the last token of the text embedding onto a vector space of a fixed dimension to obtain the text feature vectors { |k = 1, 2, …, K}.
[0133] Step S6: Use Equation (15) to obtain the probability that the k-th medical image belongs to the k-th category, and select the category corresponding to the maximum probability as the predicted category of the k-th medical image medical image .
[0134] (15)
[0135] In Equation (15), represents the cosine similarity, represents the exponential function, represents the temperature parameter.
[0136] Step S7: Use Equation (16) to construct the loss function :
[0137] (16)
[0138] In Equation (16), represents the logarithmic function. Use the gradient descent method to train the pre-trained model of the contrast text-image pair and the lightweight network, and update the hint in the image encoder and the parameter weights of the lightweight network until the loss function converges, so as to obtain the optimal medical image classification model for classifying medical images and outputting the predicted category.
[0139] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the medical image classification method, and the processor is configured to execute the program stored in the memory.
[0140] In this embodiment, a computer-readable storage medium stores a computer program. When the computer program is run by a processor, it executes the steps of the medical image classification method.
[0141] To verify the effectiveness of the present invention, five advanced prompt learning classification methods, namely COOP (Context Optimization), COCOOP (Conditional Context Optimization), VPT (Visual Prompt Tuning), MAPLE (Multi-modal Prompt Learning), and COOPLVT (Context Optimization with Learnable Visual Tokens), are selected in this embodiment. They are trained on three medical image datasets to evaluate the performance of different method models, namely the brain MRI dataset, the melanoma dataset, and the tuberculosis detection dataset. The accuracy is used as the evaluation index for the brain MRI dataset and the tuberculosis detection dataset. Since there is a long-tailed distribution problem in the melanoma dataset, the average class accuracy is used as the evaluation index on this dataset to ensure that the final result will not be biased due to the difference in the number of samples in different classes. The experimental results are shown in Table 1.
[0142] Table 1 Experimental results of the method of the present invention and the selected comparison methods on three medical image datasets
[0143]
[0144] It can be seen from the experimental results that after the model is processed by the method (Ours) of the present invention, the pre-trained model of the contrast text-image pair can capture the possible subtle differences between different classes in the medical image, enhance the perception of the text prompt for the image details, improve the understanding of the medical image by the contrast language-image pre-trained model, so that the model can more accurately identify and distinguish different medical image classes.
Claims
1. A method for medical image classification based on conditional category specific prompts, characterized in that: The following steps are involved: Step S1: Obtain a medical image dataset including K categories ,in, Indicates Medical images, represents the total number of medical images; let The true category label is ,and ∈{1,2,…,K}; The first Medical images Divide A fixed-size image block is formed, and each image block is flattened and then projected onto a fixed dimension, thereby The projected one-dimensional vectors are concatenated into image blocks for embedding ; Constructing Hints to Learn ,in, Indicates Tips to learn; Indicates the total number of prompts; Step S2: Construct a pre-trained model for comparing text-image pairs, including: image encoder , Text Encoder , wherein the image encoder is used to prompt and image patch embedding Process it and get Medical images The visual feature vector ; Step S3: Use formula (1) to obtain the average image feature vector of the kth category: : (1) In formula (1), represents the total number of categories, is the first Medical images The visual feature vector of is the total number of medical images in the kth category; Step S4: Build a lightweight network and Processing to get specific tips for the kth category ; Step S5: Category-specific tips |k=1,2,…,K} and the name of each category |k=1,2,…,K} after corresponding concatenation, input text encoder Processed in, get the text features { |k=1,2,…,K}; where represents the name of the k-th category, Represents the text features of the kth category; Step S6: and{ |k=1,2,…,K} to match, and select the category corresponding to the maximum probability as the first Medical images The prediction category of Step S7: Construct loss function , and train the pre-trained model and lightweight network of the contrasting text-image pair using the gradient descent method, and update the image encoder by gradient backpropagation Tips in and the parameter weights of the lightweight network until the loss function Until convergence, the optimal medical image classification model is obtained, which is used to classify medical images and output the predicted category.
2. The medical image classification method based on conditional category specific prompts according to claim 1, characterized in that: Image encoder in step S2 Includes: Category Embedding , Position Embedding , Layer Image Self-Attention Model , where each layer of the image self-attention model consists of a multi-head image self-attention layer and a multi-layer image perception machine block; Step S21: Initialization =1; define and initialize the -1 layer image self-attention model Output visual feature vector ; Initialize the self-attention layer of the input i-th layer Tips = ; Embed the image block As the i-th layer of self-attention Image patch embedding ; Initialize input to the first Self-attention layer Category embedding = ; Step S22: Prompt , Image Patch Embedding With category embedding After concatenation, embed with position Add and enter Self-attention layer In order to process it, the formula (2) is used to output the Layer image patch embedding With category embedding ; (2) In formula (2), is the first self-attention layer, and It is Category embedding and image patch embedding output by the self-attention layer; Step S221: Input the prompt of the i-th layer , Image Patch Embedding With category embedding After concatenation, layer normalization is performed to obtain the embedding vector sequence using formula (3): ; (3) Step S222, self-attention layer of the i-th layer The h self-attention heads of Processing is performed to obtain the self-attention layer of the i-th layer h multi-head attention representations , and then use formula (5) to get the self-attention layer of the i-th layer Comprehensive characterization of output : (4) (5) In formula (4)-formula (5), , , is the self-attention layer of the i-th layer No. The projection parameter matrix of the self-attention head, is also the projection parameter matrix, is the activation function, represents transpose, Represents a serial operation, Represents the self-attention layer of the i-th layer No. Multi-head attention representation of the output of the attention heads , is the number of parallel self-attention heads; Step S223: Use formula (6) to obtain the self-attention layer of the i-th layer Residual representation in : (6) In formula (6), represents residual connection; Step S224: Use formula (7) to obtain Self-attention layer Output image patch embedding With category embedding : (7) In formula (7), and are the two transformation matrices of the two linear layers, and are the two biases of the two linear layers, is the activation function; Representation layer normalization operation; Indicates Self-attention layer Output prompts; Step S225, after copying i+1 to i, return to step 221 and execute sequentially until i>J, thereby obtaining the J-layer self-attention layer { |i=1,2,…,J} The output image patch embedding { |i=1,2,…,J} and category embedding { |i=1,2,…,J}; Step S23: Use formula (8) to obtain Self-attention layer To Self-attention layer The self-attention layer of any kth layer in Output image patch embedding , Category Embedding and tips : (8) Step S24: Self-attention layer Output category embedding After projection, the medical image is obtained Visual features .
3. The medical image classification method based on conditional category specific prompts according to claim 1, characterized in that: The step S4 comprises: Step S4.1: Use formula (9) to obtain the intermediate features of the kth category : (9) In the formula, is the transformation matrix of the third linear layer, is the bias of the third linear layer; Step S4.2: Use formula (10) to obtain the specific prompt of the kth category : (10) In formula (10), is the transformation matrix of the fourth linear layer, is the bias of the fourth linear layer, is the activation function.
4. The medical image classification method based on conditional category specific prompts according to claim 1, characterized in that: The step S5 comprises: Step S5.1: Create a text encoder include: Layer text self-attention model, where each layer of text self-attention model consists of a multi-head text self-attention layer and a multi-layer text perception machine block; Step S5.2: The corresponding category name After concatenation, the k-th sentence text representation is obtained, thereby obtaining the K-sentence text representation; Step S5.3: After segmenting the K sentences, project the segmentation results onto a fixed dimension to obtain text embedding. ; Step S5.4: Initialization =1; define and initialize the -1 layer text self-attention model Output text embedding = ; Step S5.5: Input to Layer Text Self-Attention Model Processed in, output text embedding feature vector ; Step S5.5.1, Layer Text Self-Attention Model The h-th head text self-attention layer in the multi-head text self-attention layer uses formula (11) to obtain the h-th head text self-attention layer Layer Text Self-Attention Model The h-th head text attention representation , and then use formula (12) to get the Layer Text Self-Attention Model Comprehensive feature representation of the output of multi-head text self-attention layer : (11) (12) In formulas (11) and (12), , , It is Layer Text Self-Attention Model The three projection parameter matrices of the h-th head text self-attention layer, is the global projection parameter matrix, is the activation function, represents transpose, Represents a serial operation, is the number of parallel text self-attention heads; Step S5.5.2: Use formula (13) to get Layer Text Self-Attention Model The residual representation of : (13) In formula (13), represents residual connection; Step S5.5.3: Use formula (14) to obtain Layer Text Self-Attention Model The text embedding output by the multi-layer text perceptron block : (14) In formula (14), and They are the text transformation matrices of the two linear layers, and They are the text biases of the two linear layers, is the activation function; Step S5.5.4: +1 assigned to Then, return to step S5.5.2 and execute sequentially until >L, thus obtaining the Lth text self-attention layer Output text embedding ; Step S5.6: Embed text The last token of is projected onto a fixed-dimensional vector space to obtain the text feature vector { |k=1,2,…,K}.
5. The medical image classification method based on conditional category specific prompts according to claim 1, characterized in that: In step S6, equation (15) is used to obtain Medical images The probability of belonging to the kth category : (15) In formula (15), represents the cosine similarity, represents the exponential function, Represents the temperature parameter.
6. The medical image classification method based on conditional category specific prompts according to claim 1, characterized in that: In step S7, the loss function is constructed using formula (16): : (16) In formula (16), Represents a logarithmic function.
7. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports the processor to execute the medical image classification method according to any one of claims 1 to 6, and the processor is configured to execute the program stored in the memory.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the medical image classification method according to any one of claims 1 to 6 are performed.
Citation Information
Patent Citations
Image classification method and device, electronic equipment and storage medium
CN112749737A
Remote sensing few-sample target detection method based on condition prompt and causal learning
CN118334519A