A Few-Shot Semantic Segmentation Method Based on Image-Text Fusion
By introducing graphic and text fusion technology and attention mechanism into the semantic segmentation method of few samples, the problem of poor segmentation effect of existing methods in small samples scenarios is solved, and higher segmentation accuracy and performance are achieved.
Patent Information
- Application Number
- CN202411284272.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The existing semantic segmentation method with few samples is difficult to effectively capture the details and semantic information of the target category when facing a small amount of labeled data, and the spatial differences between image visual features and semantic information lead to low feature fusion efficiency, limiting the segmentation performance of the model.
By introducing graphic and text fusion technology into the semantic segmentation method of few samples, the pre-trained deep neural network extracts multi-scale image features, and the text features of preset text prompts are extracted through the text encoder, and the two are fused. At the same time, attention mechanism is used to mine the correlation between supporting image features and query image features to generate segmentation prediction results.
It significantly improves the accuracy and performance of semantic segmentation of few samples, can accurately identify target categories in the image with limited data, and improves the segmentation effect and generalization ability of the model.
Smart Images

Figure CN119273914B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and specifically to a few-shot semantic segmentation method based on text-image fusion. Background Art
[0002] In the field of computer vision, few-shot semantic segmentation has always been a challenging research problem. Traditional semantic segmentation methods rely on a large amount of labeled data to train models for accurate recognition and segmentation of target regions in images. However, in practical applications, it is often very difficult and time-consuming to obtain a large amount of labeled data, especially in some specific fields or emerging scenarios where the labeled samples are extremely limited. This makes the existing segmentation methods trained based on a large amount of data perform poorly in the face of few-shot scenarios and are difficult to effectively promote and apply.
[0003] To solve this problem, researchers have proposed few-shot semantic segmentation methods, attempting to maintain high segmentation performance under the condition of a small number of labeled samples. However, these existing methods mainly rely on extracting visual features from limited support images. Due to the small amount of data and lack of diversity, when the model segments query images, it often fails to fully capture the details and semantic information of the target class. In addition, due to the spatial differences between image visual features and semantic information, it is difficult for existing methods to achieve efficient feature fusion, further limiting the segmentation performance of the model. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a few-shot semantic segmentation method based on text-image fusion, which supplements semantic information through text-image fusion and text prompts in few-shot scenarios, improving the accuracy and performance of semantic segmentation.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A few-shot semantic segmentation method based on text-image fusion, comprising the following steps:
[0006] Collect and preprocess a data set, the data set is grouped by category, and each group contains a number of support images and query images;
[0007] Use a pre-trained deep neural network to extract multi-scale image features;
[0008] Use a text encoder to extract text features of preset text prompts;
[0009] Fuse the extracted text features with the image features;
[0010] Use an attention mechanism to mine the correlation between support image features and query image features;
[0011] Upsample the fused image features and decode to generate the segmentation prediction results.
[0012] Preferably, the step of collecting and preprocessing the dataset includes:
[0013] Group the dataset according to the total number of image categories in the dataset, divide the dataset into four groups, and each group is called a fold;
[0014] The data in each fold consists of support images and query images. Among them, the support images are denoted as The query images are denoted as D q ={X, Y}, where X i (X) represents the original image, Y i (Y) represents the mask corresponding to the image, and K represents the shot value set in the few-shot scenario, indicating how many support images there are for each category;
[0015] Evenly distribute the image categories in the dataset into the four folds to ensure that the categories contained in each fold do not intersect.
[0016] Preferably, the step of using the pre-trained deep neural network to extract multi-scale image features includes:
[0017] Input the input image X i into the pre-trained deep neural network to extract multi-scale image features;
[0018] Through different layers of the pre-trained deep neural network, extract four scales of image features in sequence, which are respectively denoted as X 1 , X 2 , X 3 , X 4 , where X 1 , X 2 , X 3 , X 4 represent the image features from shallow to deep in sequence;
[0019] Perform global average pooling on the deep feature map to obtain the global feature of the image The formula is:
[0020]
[0021] Input the global feature and the deep feature Figure X 4 into the multi-head self-attention mechanism together to enhance the feature expression ability of the image. The formula is:
[0022]
[0023] Among them, E represents the enhanced global feature, and Z represents the enhanced image feature.
[0024] Preferably, the step of using a text encoder to extract the text feature of a preset text prompt includes:
[0025] A preset text prompt, which is related to the target category in the image and is used to indicate the target category to be segmented;
[0026] Input the text prompt into the text encoder Transformer to extract the text feature related to this prompt;
[0027] The text encoder Transformer processes the input text prompt, captures the context information in the text through the self-attention mechanism, and generates a feature vector t representing the text semantics. The formula is:
[0028] t = Transformer(text prompt)
[0029] where text prompt represents the preset text prompt, and t represents the text feature vector extracted from the text prompt.
[0030] Preferably, the step of fusing the extracted text feature with the image feature includes:
[0031] Through a Transformer decoder composed of three layers of Transformers, fuse the image feature Z with the text feature t to generate a text feature V containing visual information. The formula is:
[0032]
[0033] where Z represents the multi-scale feature extracted from the image, t represents the text feature generated by the text prompt, and V represents the text feature containing visual information after fusion;
[0034] Use a learnable scaling parameter α to perform a weighted combination of the fused visual information V and the original text feature t to generate an updated text feature t'. The formula is:
[0035] t′ = t + αV
[0036] where α is a learnable parameter that controls the influence of visual information on the text feature, and t' is the updated text feature, which contains semantic information from the image;
[0037] Calculate the correlation between the updated text feature t' and the image feature Z to generate a score map S. The formula is:
[0038] S = Zt′
[0039] Concatenate the score map S with the original deep image feature X 4 to form the fused image feature X'. 4 For subsequent decoding and segmentation, the formula is:
[0040] X′ 4 = [X 4 , S]
[0041] where X 4 is the original deep image feature, S is the score map, and X' 4 is the fused image feature, which combines multi-modal information of images and text.
[0042] Preferably, the step of using the attention mechanism to mine the correlation between the support image feature and the query image feature includes:
[0043] Flatten the feature F s extracted from the support image and the feature F q extracted from the query image. The flattening operation converts the multi-dimensional feature map into a one-dimensional vector. The formula is:
[0044] F′ s = flatten(F s )
[0045] F′ q = flatten(F q )
[0046] Perform linear mapping on the flattened support image feature F' s and the query image feature F' q respectively to generate the query vector Q, the key vector K, and the value vector V. The formula is:
[0047] Q = linear(flatten(F' q ))
[0048] K, V = linear(flatten(F' s ))
[0049] Use the self-attention mechanism to calculate the correlation between the query vector Q and the key vector K. The correlation is calculated by the dot product of the query vector Q and the key vector K, and the attention weight is generated through normalization operation, and then the value vector V is weighted. The formula is:
[0050]
[0051] where softmax(·) represents the normalization operation, Let \(\alpha\) be the scaling factor, \(d\) be the dimension of the vector, and Atten(Q, K, V) be the weighted feature representation;
[0052] Through the above attention mechanism, the support image feature \(F\) is mined s and the query image feature \(F\) q The correlation between them is obtained to get the enhanced query image feature The formula is:
[0053]
[0054] where \(\hat{F}\) is the query image feature enhanced by the attention mechanism, which contains semantic information related to the support image feature and is further used to improve the accuracy of few-shot semantic segmentation.
[0055] Preferably, the step of upsampling the fused image feature and decoding to generate a segmentation prediction result includes:
[0056] Perform an upsampling operation on the fused image feature \(X'\) 4 to restore the feature map to the same resolution as the input image. The upsampling can be achieved by bilinear interpolation or transposed convolution. The formula is:
[0057] \(X\) up = Upsample(X' 4 )
[0058] where \(X'\) 4 represents the fused image feature, and \(X\) up is the upsampled feature map;
[0059] During the upsampling process, the upsampled feature Figure X up is fused with the image feature \(X\) of the corresponding scale extracted in the encoding stage i to retain the global semantic information and detail information. The formula is:
[0060] \(X\) fused = Concat(X up , X i )
[0061] where Concat(·) represents the feature concatenation operation, \(X\) i represents the multi-scale features extracted in the encoding stage, and \(X\) fused represents the fused feature map;
[0062] The fused feature Figure X fusedInput to the decoder, which consists of multiple convolutional layers or transposed convolutional layers to gradually restore and generate the final segmentation prediction results The specific expression is:
[0063]
[0064] Among them, Decoder(·) represents the decoder, represents the segmentation prediction mask output by the decoder;
[0065] The final segmentation prediction mask It has the same resolution as the original input image, and the value of each pixel represents the probability or classification label of the pixel belonging to a certain category, which is used for the prediction output of the few-shot semantic segmentation task.
[0066] The present invention also provides a few-sample semantic segmentation device based on image-text fusion, comprising:
[0067] Data preprocessing module, used to collect and preprocess data sets;
[0068] Image feature extraction module, which uses pre-trained deep neural network to extract multi-scale image features;
[0069] A text feature extraction module, which uses a text encoder to extract text features of a preset text prompt;
[0070] Feature fusion module, which fuses the extracted text features with image features;
[0071] The correlation mining module uses the attention mechanism to mine the correlation between the supporting image features and the query image features;
[0072] The decoding module upsamples the fused image features and decodes them to generate segmentation prediction results.
[0073] The present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the above method is implemented.
[0074] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.
[0075] The present invention provides a few-sample semantic segmentation method based on image-text fusion, which has the following beneficial effects:
[0076] 1. The present invention proposes a few-sample semantic segmentation method based on image-text fusion, which supplements additional semantic information by adding text prompts and improves the existing few-sample semantic segmentation method by using the image-text fusion method, thereby significantly improving the segmentation performance. This method effectively overcomes the problem of poor segmentation effect caused by insufficient data in the case of few samples.
[0077] 2. When solving the problem of few-sample semantic segmentation, the present invention introduces textual hints related to image categories as a supplementary source of semantic information. Textual hints provide additional semantic guidance for the model, helping the model to better understand and identify the target categories in the image, so that semantic segmentation can still be performed accurately when data is limited.
[0078] 3. The present invention effectively alleviates the fusion problem of visual and semantic space due to spatial differences by aligning the image visual space and the text semantic space. Through the effective fusion of image and text features, the model can establish a closer connection between the semantic level and the visual level, improve the expressive power of the fused features, and thus enhance the segmentation effect.
[0079] 4. The proposed method shows good segmentation performance in the case of few samples and can significantly improve the accuracy of segmentation. Whether in the case of limited training data or in the face of new categories, the method can effectively generalize and provide high-precision segmentation results in various challenging tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Figure 1 It is a schematic diagram of the method flow of the present invention;
[0081] Figure 2 A flowchart of a new method for few-sample semantic segmentation based on prototype perception is provided for an embodiment of the present invention;
[0082] Figure 3 A schematic diagram of a support image and a query image in a training and testing process provided by an embodiment of the present invention;
[0083] Figure 4 An example diagram of extracting multi-scale image features using the ResNet50 network provided in an embodiment of the present invention;
[0084] Figure 5 A schematic diagram of text features updated using visual information provided by an embodiment of the present invention;
[0085] Figure 6 A schematic diagram of image features that introduce semantic information provided by an embodiment of the present invention;
[0086] Figure 7Schematic diagram for mining the relationship between the support image and the query image provided by the embodiment of the present invention;
[0087] Figure 8 Segmentation performance of the model provided by the embodiment of the present invention on the COCO dataset;
[0088] Figure 9 Schematic diagram of the device structure of the present invention;
[0089] Figure 10 Schematic diagram of the computer device structure of the present invention.
[0090] Among them, 100, data preprocessing module; 200, image feature extraction module; 300, text feature extraction module; 400, feature fusion module; 500, correlation mining module; 600, decoding module; 40, computer device; 41, processor; 42, memory; 43, storage medium. Detailed implementation manners
[0091] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0092] Please refer to the attached Figure 1 , the embodiment of the present invention provides a few-shot semantic segmentation method based on text-image fusion, including the following steps:
[0093] S1. Collect and preprocess the dataset
[0094] First, perform the operations of collecting and preprocessing the dataset. The task of few-shot semantic segmentation depends on the effective grouping and processing of the dataset. In order to ensure that the model can accurately perform semantic segmentation under few-shot conditions, it is necessary to reasonably group the dataset by category and perform standardized preprocessing on the image data to provide high-quality input data for the subsequent steps.
[0095] In this embodiment, the specific implementation steps are as follows:
[0096] 1. Dataset grouping:
[0097] The dataset is grouped according to the total number of image categories in the dataset, and the dataset is divided into multiple groups (folds). For example, for the COCO dataset, which contains 80 image categories, it is evenly divided into four groups, with each group containing 20 categories, and the categories between different groups do not overlap. This grouping method ensures that the categories within different groups are independent of each other and can provide diversity for model training in few-shot scenarios.
[0098] 2. Definition and division of support images and query images:
[0099] The data of each group is further divided into a support image set D s and a query image D q :
[0100] The support image set D s , which is used for model training and provides a small amount of annotation information to guide the learning of the model:
[0101]
[0102] The query image D q , which is used for model testing and verification to evaluate the segmentation performance of the model in actual applications:
[0103] D q ={X,Y}
[0104] where X i (X) represents the original image, Y i (Y) represents the mask corresponding to the image, and K represents the shot value (usually 1 or 5) set in the few-shot scenario, indicating how many support images there are for each category.
[0105] 3. Data preprocessing:
[0106] Before model training, necessary preprocessing is performed on the image data. The specific operations include the following:
[0107] Size adjustment: Resize the images to the input size required by the network to ensure that the images input into the neural network have a unified resolution.
[0108] Data augmentation: Apply data augmentation techniques such as random cropping, rotation, flipping, and color jittering to increase the diversity of the data and improve the robustness of the model.
[0109] Normalization processing: Normalize the pixel values of the images, scaling the pixel values to the range of [0,1] or [-1,1], which helps to accelerate the model training process and improve the training effect.
[0110] Through the above steps, the effective grouping and preprocessing of the dataset are achieved in this embodiment, laying a solid foundation for the few-shot semantic segmentation task of the model.
[0111] S2. Extract multi-scale image features using a pre-trained deep neural network
[0112] After completing the collection and preprocessing of the dataset, next, multi-scale image features are extracted using a pre-trained deep neural network. The goal of this step is to extract feature information at different levels from the image to provide rich visual feature support in the subsequent text-image fusion and semantic segmentation processes.
[0113] In this embodiment, the specific implementation steps are as follows:
[0114] 1. Multi-scale feature extraction:
[0115] Input image X i Process it through a pre-trained ResNet50 network. ResNet50 is a commonly used convolutional neural network that has been pre-trained on large-scale datasets (such as ImageNet) and has good feature extraction capabilities.
[0116] In the ResNet50 network, different levels of convolutional layers will extract feature maps of different scales. Specifically, in this embodiment, four scales of feature maps are extracted, denoted as X 1 , X 2 , X 3 , X 4 . Among them:
[0117] X 1 represents the low-level features extracted by the shallow convolutional layer, containing more edge and texture information;
[0118] X 2 and X 3 represent the intermediate-level features, containing gradually abstracted semantic information;
[0119] X 4 represents the high-level features extracted by the deep convolutional layer, mainly containing global semantic information.
[0120] 2. Global feature extraction:
[0121] Further process the deep features Figure X 4 Use global average pooling to compress the spatial dimension of the feature map into a vector to obtain the global feature The specific operation of global average pooling is to calculate the average of all pixels in each channel of the feature map:
[0122]
[0123] This global feature represents the global information of the entire image and can provide semantic support for multi-modal feature fusion in subsequent steps.
[0124] 3. Enhancement of multi-head self-attention mechanism:
[0125] To further improve the expression ability of image features, the global feature and the deep feature Figure X 4 are input into the multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA) for processing. The multi-head self-attention mechanism captures the mutual relationships at different positions in the feature map through multiple parallel attention heads, generating the enhanced global feature and the deep feature Z:
[0126]
[0127] Among them, represents the global feature enhanced by the attention mechanism, and Z represents the enhanced deep image feature.
[0128] Through the above steps, in this embodiment, the extraction and enhancement of multi-scale image features are completed, providing rich and diverse visual information for subsequent text-image fusion and segmentation prediction.
[0129] S3. Use a text encoder to extract the text features of the preset text prompt
[0130] After the extraction of multi-scale image features is completed, to further enhance the model's semantic understanding ability of the target category, the present invention uses a text encoder to extract the text features related to the preset text prompt. By combining the image features and the text features, the semantic information of the model can be enriched, thereby improving the accuracy of few-shot semantic segmentation.
[0131] In this embodiment, the specific implementation steps are as follows:
[0132] 1. Text prompt design:
[0133] First, design a text prompt related to the target category. This prompt is used to clarify the target category that the model should focus on. For example, for the target category "cat", the text prompt can be set as "a photo of a cat". This prompt provides clear semantic guidance for the model, enabling the subsequent extracted text features to more accurately reflect the semantic information of the target category.
[0134] 2. Use of the text encoder:
[0135] Input the preset text prompt into the Transformer of the text encoder. Transformer is a model based on the self-attention mechanism and is widely used in natural language processing tasks. Its self-attention mechanism can effectively capture the context information in the text, enabling the extracted text features to better represent the semantic content of the text prompt.
[0136] 3. Extraction of text features:
[0137] Specifically, the text encoder Transformer processes the input text prompt to generate a high-dimensional text feature vector t. This vector encodes the semantic information of the text prompt and can be effectively combined with the image features in the subsequent feature fusion process. The process of extracting text features can be expressed as:
[0138] t = Transformer(text prompt)
[0139] where text prompt is the input text prompt and t is the text feature vector extracted from the text prompt.
[0140] Through the above steps, the text features related to the target category are successfully extracted in this embodiment, providing the necessary semantic information support for the fusion of image and text features in the subsequent steps.
[0141] S4. Fuse the extracted text features with the image features
[0142] After the extraction of image features and text features is completed, in order to make full use of the information of these two different modalities, the present invention fuses the extracted text features with the image features. Through this fusion method, the model can combine visual and semantic information, thereby enhancing the ability to understand and segment the target category.
[0143] In this embodiment, the specific implementation steps are as follows:
[0144] 1. Incorporate visual information into text features:
[0145] First, fuse the image features Z extracted in step S2 with the text features t extracted in step S3. Specifically, take the image features Z and the text features t as inputs and send them into the Transformer decoder composed of three layers of Transformer. This decoder incorporates the visual information into the text features through multiple layers of self-attention mechanisms and feed-forward neural networks to generate a new feature representation V. This process can be expressed as:
[0146]
[0147] Among them, Z represents the multi-scale features extracted from the image, t represents the text features generated by the text prompt, and V represents the text features fused with visual information.
[0148] 2. Update text features:
[0149] To enable the fused visual information to better affect the text features, a learnable scaling parameter α is used to perform a weighted combination of the visual information V and the original text features t to generate the updated text features t'. This update operation adjusts the weight of the visual information to make the text features more adaptable to the few-shot segmentation task. The specific operation is as follows:
[0150] t′ = t + αV
[0151] Among them, α is a learnable parameter that controls the degree of influence of the visual information, and t' is the updated text features, which combines the semantic information of the image.
[0152] 3. Incorporate semantic information into image features:
[0153] Next, the updated text features t' are further fused with the image features Z. Specifically, the correlation between the text features t' and the image features Z is calculated to generate a score map S, which reflects the semantic matching degree between the text and the image features. The fusion operation can be expressed as:
[0154] S = Zt′
[0155] Finally, the generated score map S is concatenated with the deep image features X 4 to form the new fused image features X' 4 , which are used for subsequent upsampling and decoding steps. The concatenation operation can be expressed as:
[0156] X′ 4 = [X 4 , S]
[0157] Among them, X 4 is the original deep image features, S is the semantic score map, and X' 4 is the image features combined with semantic information.
[0158] Through the above steps, in this embodiment, an effective fusion of the image features and the text features is achieved. The generated fusion features contain both the details of the visual information and the text semantic information, providing richer feature support for subsequent segmentation prediction.
[0159] S5. Use the attention mechanism to mine the correlation between the support image features and the query image features
[0160] After the fusion of image features and text features, in order to further improve the accuracy of few-shot semantic segmentation, the present invention uses an attention mechanism to mine the correlation between the support image features and the query image features. In this way, the semantic relationship between the support image and the query image can be better captured, so as to more accurately identify the target category during the segmentation process.
[0161] In this embodiment, the specific implementation steps are as follows:
[0162] 1. Feature flattening and linear mapping:
[0163] First, the feature F s extracted from the support image q and the feature F
[0164] F′ s =flatten(F s )
[0165] F′ q =flatten(F q )
[0166] extracted from the query image are flattened. The flatten operation converts the originally multi-dimensional feature map into a one-dimensional vector for subsequent linear mapping processing. The expression of the flatten operation is: s The flattened support image feature F' q and the query image feature F' are respectively used to generate a query vector Q, a key vector K, and a value vector V through linear mapping. The purpose of the linear mapping operation is to convert the feature vector into a suitable dimension for self-attention mechanism calculation. The specific expression is:
[0167] Q=linear(flatten(F' q ))
[0168] K,V=linear(flatten(F' s ))
[0169] where Q represents the query vector of the query image feature, and K and V represent the key vector and value vector of the support image feature.
[0170] 2. Self-attention mechanism calculation:
[0171] The self-attention mechanism is used to calculate the correlation between the query vector Q and the key vector K. The specific operation is to calculate the dot product of the query vector Q and the key vector K to obtain the similarity score between them. To avoid the values being too large or too small, a scaling process is performed during the calculation, and the softmax function is used to normalize the similarity score to generate the attention weights. The calculation formula of the self-attention mechanism is:
[0172]
[0173] Among them, is the scaling factor, d is the dimension of the vector, and Atten(Q, K, V) represents the feature representation after attention weighting.
[0174] 3. Feature enhancement:
[0175] The generated attention weights are used to weight the value vector V to obtain the weighted support image features. Then, the weighted support image features are further combined with the query image feature F q to form the enhanced query image feature This combination method helps the model to better utilize the semantic information in the support image when segmenting the query image. The expression of the enhanced query image feature is:
[0176]
[0177] Among them, is the query image feature processed by the self-attention mechanism, which combines the relevant semantic information from the support image.
[0178] Through the above steps, in this embodiment, the effective mining of the correlation between the support image features and the query image features is realized. The generated enhanced features combine the semantic information of both, which helps to improve the accuracy and robustness of the few-shot semantic segmentation task.
[0179] S6. Upsample the fused image features and decode to generate the segmentation prediction result
[0180] After completing the mining of the correlation between the support image features and the query image features, in order to obtain the final segmentation result, the present invention upsamples the fused image features and generates the final segmentation prediction mask through the decoder. The upsampling and decoding processes can restore the spatial resolution of the image, while retaining and utilizing the feature information at different scales to achieve high-precision semantic segmentation.
[0181] In this embodiment, the specific implementation steps are as follows:
[0182] 1. Upsample to restore the spatial resolution:
[0183] First, perform an upsampling operation on the fused image features X' generated in the previous step. 4 The purpose of upsampling is to restore the low-resolution feature map to the spatial resolution of the original image, so that the output segmentation mask can accurately match the size of the input image. Upsampling can usually be achieved by methods such as bilinear interpolation or transposed convolution. The specific operation expression is:
[0184] X up = Upsample(X' 4 )
[0185] where X' 4 is the fused image feature, and X up is the upsampled feature map with a higher spatial resolution.
[0186] 2. Multi-scale feature fusion:
[0187] During the upsampling process, fuse the upsampled feature Figure X up with the corresponding scale feature extracted in the encoding stage Figure X i By combining these multi-scale features, both global semantic information and detailed information can be retained simultaneously, thereby enhancing the accuracy of the segmentation result. The specific expression of the fusion operation is:
[0188] X fused = Concat(X up , X i )
[0189] where X i represents the multi-scale features extracted in the encoding stage, and X fused is the fused feature map containing multi-scale visual information.
[0190] 3. Decoding to generate a segmentation mask:
[0191] Input the fused feature Figure X fused into the decoder for layer-by-layer decoding. The decoder usually consists of multiple convolutional layers or transposed convolutional layers, and generates the final segmentation prediction result by gradually restoring the spatial resolution of the feature map The decoding process restores the feature map to a segmentation mask of the same size as the original input image, and the value of each pixel represents the probability or classification label that the position belongs to a certain category. The expression of the decoding process is:
[0192]
[0193] Among them, Decoder(·) represents the decoder, which represents the segmentation prediction mask output by the decoder.
[0194] 4. Precision verification (mIoU):
[0195] To verify the precision of the segmentation results, the generated segmentation masks are evaluated using the mIoU (Mean Intersection over Union) metric. mIoU is a commonly used semantic segmentation evaluation metric that measures the accuracy of segmentation by calculating the intersection over union between the prediction results and the ground truth labels.
[0196] Specifically, for each class i, its IoU value is calculated:
[0197]
[0198] where TP i represents the number of true positives for class i, FP i represents the number of false positives, and FN i represents the number of false negatives.
[0199] Then, the IoU values of all classes are averaged to obtain the mIoU value:
[0200]
[0201] where K is the total number of classes, and mIoU is the average segmentation precision over all classes.
[0202] Through the above steps, in this embodiment, not only the upsampling and decoding processes of the fused features are realized, generating high-precision segmentation prediction results, but also the segmentation results are quantitatively evaluated using the mIoU metric, ensuring the precision and effectiveness of the model in the few-shot semantic segmentation task.
[0203] Embodiment:
[0204] The embodiment of the present invention provides a few-shot semantic segmentation method based on prototype perception, which realizes high-precision semantic segmentation in the few-shot scenario by combining image features and text features and using the multi-head self-attention mechanism and cross-attention mechanism. The embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0205] Figure 2 The following shows the overall flowchart of a new few-shot semantic segmentation method based on prototype perception provided by the embodiment of the present invention. The process includes the following main steps:
[0206] S1. Collect and download the datasets required for the experiment and perform preprocessing.
[0207] S2. Extract multi-scale image features using a deep neural network.
[0208] S3. Extract text features using a text encoder.
[0209] S4. Perform image and text feature fusion.
[0210] S5. Mine the correlation between the support image features and the query image features.
[0211] S6. Decode the fused image features to obtain the prediction results.
[0212] S7. Input the test set to evaluate the overall segmentation accuracy of the model.
[0213] Figure 3 The following is a schematic diagram of the support image and the query image in the training and testing processes of the embodiment of the present invention. The support image set contains N*K images, where N represents the number of categories (as shown in the figure, the categories include Dog, Goose, Bus, etc.), and K represents the number of support images corresponding to each category (as shown in the figure, K is 1, that is, each category has one support image). The query image consists of the original image and the corresponding image mask. During the training process, the model learns the features of each category through the support images and applies these features to the segmentation task of the query image.
[0214] Figure 4 The following is an example diagram of extracting multi-scale image features using the ResNet50 network provided by the embodiment of the present invention. The ResNet50 network is composed of stacked convolutional layers, pooling layers, activation layers, etc., and is divided into four blocks. Different blocks contain different numbers of convolutional layers and can extract image features of four different scales. The shallow feature maps retain more spatial detail information, while the deep feature maps contain higher-level semantic information.
[0215] Figure 5 The following is a schematic diagram of the text feature update process provided by the embodiment of the present invention. First, preset the text prompt "a photo of [CLS]", where CLS represents the target category to be segmented. Then, use the Transformer encoder to encode this text prompt to obtain the initial text features. Next, perform residual update on the text features by combining the visual information extracted from the Transformer Decoder to achieve the fusion of image and text information, thereby enhancing the semantic expression ability of the text features.
[0216] Figure 6 The following is a schematic diagram of the image feature and text feature fusion process in the embodiment of the present invention. First, use the multi-head self-attention mechanism (MHSA) to obtain the global features And the feature map Z. Then, multiply the feature map Z by the text feature t to obtain the correspondence between the image feature and the text feature, and perform preliminary fusion of the image and text features through this relationship, thereby introducing the semantic information of the text into the image feature.
[0217] Figure 7 The figure shows a schematic diagram of using the cross-attention mechanism to mine the relationship between the support image and the query image provided by the embodiment of the present invention. The feature of the query image is used as Query (Q), and the features of the support image are used as Key (K) and Value (V). Through the cross-attention mechanism, the class-related information in the query image is mapped to the support image, thereby further emphasizing the target category in the query image. This process effectively enhances the segmentation ability of the model in the few-shot scenario.
[0218] Figure 8 The figure shows a comparison chart of the segmentation performance of the embodiment of the present invention on the COCO dataset. The figure shows that the segmentation accuracy of the model of the present invention on fold2 and fold3 is better than the other three methods, and the overall segmentation performance is better than the other two methods, second only to CWTd. Through the evaluation results of mIoU, it can be seen that the model of the present invention has excellent performance in the few-shot semantic segmentation task.
[0219] The present invention realizes high-precision semantic segmentation in the few-shot scenario by combining the multi-modal features of images and texts and using the multi-head self-attention and cross-attention mechanisms. The multi-scale image features are extracted by ResNet50, the text features are extracted by the Transformer encoder, and the semantic understanding ability of the model is effectively enhanced in the process of feature fusion and relationship mining. Finally, the precise segmentation result is generated through upsampling and decoding, and the performance is evaluated by mIoU, verifying the superiority of the present invention in the few-shot semantic segmentation task.
[0220] The few-shot semantic segmentation device based on text-image fusion described below can be correspondingly referred to the few-shot semantic segmentation method based on text-image fusion described above.
[0221] Please refer to the appendix Figure 9 , the present invention also provides a few-shot semantic segmentation device based on text-image fusion, including:
[0222] The data preprocessing module 100 is used to collect and preprocess the dataset;
[0223] The image feature extraction module 200 extracts multi-scale image features by using a pre-trained deep neural network;
[0224] The text feature extraction module 300 extracts the text features of the preset text prompt by using a text encoder;
[0225] The feature fusion module 400 fuses the extracted text features and image features;
[0226] The correlation mining module 500 uses the attention mechanism to mine the correlation between the support image features and the query image features;
[0227] The decoding module 600 upsamples the fused image features and decodes them to generate the segmentation prediction results.
[0228] The device of this embodiment can be used to execute the above method embodiment, and its principle and technical effect are similar, which will not be elaborated here.
[0229] Please refer to the appendix Figure 10 , the present invention also provides a computer device 40, including: a processor 41 and a memory 42, the memory 42 stores a computer program executable by the processor, and when the computer program is executed by the processor, it executes the above method.
[0230] The present invention also provides a storage medium 43, on which a computer program is stored, and when the computer program is run by the processor 41, it executes the above method.
[0231] Among them, the storage medium 43 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (abbreviated as SRAM), electrically erasable programmable read-only memory (abbreviated as EEPROM), erasable programmable read-only memory (abbreviated as EPROM), programmable read-only memory (abbreviated as PROM), read-only memory (abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0232] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A few-sample semantic segmentation method based on image-text fusion, characterized in that: The following steps are involved: Collect and preprocess a dataset, wherein the dataset is grouped by category, each group containing a number of support images and query images; Extract multi-scale image features using pre-trained deep neural networks; Using a text encoder to extract text features of a preset text prompt; Fusion of extracted text features with image features; Use the attention mechanism to mine the correlation between the supporting image features and the query image features; Upsample the fused image features and decode them to generate segmentation prediction results; The step of extracting multi-scale image features using a pre-trained deep neural network includes: The input image X i Input into the pre-trained deep neural network to extract multi-scale image features Z1; Through different levels of the pre-trained deep neural network, four scales of image features are extracted in turn, which are denoted as X1, X2, X3, and X4, respectively, where X1, X2, X3, and X4 represent image features from shallow to deep layers in turn; Perform global average pooling on the deep feature map X4 to obtain the global features of the image The formula is: Global Features Together with the deep feature map X4, it is input into the multi-head self-attention mechanism to enhance the feature expression ability of the image. The formula is: in, represents the enhanced global features, and Z represents the enhanced image features; The step of using a text encoder to extract text features of a preset text prompt comprises: Preset text prompts, where the text prompts are related to the target category in the image and are used to indicate the target category to be segmented; Inputting the text prompt into a text encoder Transformer to extract text features related to the prompt; The text encoder Transformer processes the input text prompts, captures the contextual information in the text through the self-attention mechanism, and generates a text feature vector t representing the semantics of the text. The formula is: t=Transformer(text prompt) Wherein, text prompt represents the preset text prompt, and t represents the text feature vector extracted from the text prompt; The step of fusing the extracted text features with the image features comprises: Through a Transformer decoder consisting of three layers of Transformers, the enhanced image feature Z is fused with the text feature vector t to generate a text feature V containing visual information. The formula is: Among them, Z represents the enhanced image features, t represents the text feature vector extracted from the text prompt, and V represents the fused text features containing visual information; Using the learnable scaling parameter α, the fused text feature V containing visual information is weightedly combined with the text feature vector t extracted from the text prompt to generate the updated text feature t′, which is: t′=t+αV Among them, α is a learnable scaling parameter that controls the influence of visual information on text features, t′ is the updated text feature, which contains semantic information from the image; Calculate the correlation between the updated text feature t′ and the multi-scale image feature Z1 to generate the score map S. The formula is: S=Z1t′ The score map S is concatenated with the original deep feature map X4 to form the fused image feature X′4 for subsequent decoding and segmentation. The formula is: X′4=[X4,S] Among them, X4 is the original deep feature map, S is the score map, and X′4 is the fused image feature, which combines the multimodal information of image and text; The step of using the attention mechanism to mine the correlation between the supporting image features and the query image features comprises: The features F extracted from the support image s and the features F extracted from the query image q Flattening is performed to convert the multidimensional feature map into a one-dimensional vector. The formula is: F′ s =flatten(F s ) F′ q =flatten(F q ) For the flattened support image feature F′ s and query image features F′ q Perform linear mapping respectively to generate query vector Q, key vector K and value vector V. The formula is: Q=linear(flatten(F′ q )) K,V=linear(flatten(F′ s )) The self-attention mechanism is used to calculate the correlation between the query vector Q and the key vector K. The correlation is calculated by the dot product of the query vector Q and the key vector K, and the attention weight is generated by the normalization operation, and then the value vector V is weighted. The formula is: Among them, softmax(·) represents the normalization operation, is the scaling factor, d is the dimension of the vector, and Atten(Q, K, V) is the weighted feature representation; Through the above attention mechanism, we can mine the supporting image features F s and query image features F q The correlation between them is used to obtain enhanced query image features. The formula is: in, The query image features are enhanced by the attention mechanism and contain semantic information related to the supporting image features, which are further used to improve the accuracy of few-shot semantic segmentation.
2. The method for few-sample semantic segmentation based on image-text fusion according to claim 1, characterized in that: The steps of collecting and preprocessing the data set include: The dataset is grouped according to the total image category of the dataset, and the dataset is divided into four groups, each group is called a fold; The data in each fold consists of support images and query images, where the support images are denoted as The query image is denoted as D q ={X i , Y i }, where X i represents the original input image, Y i represents the mask corresponding to the image, K represents the shot value set in the few-sample scenario, and represents the number of supporting images for each category; The image categories in the dataset are evenly distributed into four folds to ensure that the categories contained in each fold do not overlap with each other.
3. The few-sample semantic segmentation method based on image-text fusion according to claim 1, characterized in that: The steps of upsampling the fused image features and decoding to generate segmentation prediction results include: The fused image feature X′4 is upsampled to restore the feature map to the same resolution as the input image. Upsampling can be achieved through bilinear interpolation or transposed convolution. The formula is: X up =Upsample(X′4) Among them, X′4 represents the fused image features, X up is the feature map after upsampling; In the upsampling process, the upsampled feature map X up Image features of the corresponding scale extracted in the encoding stage Fusion is performed to retain global semantic information and detail information. The formula is: Among them, Concat(·) represents the feature concatenation operation, represents the multi-scale features extracted in the encoding stage, X fused Represents the fused feature map; The fused feature map X fused Input to the decoder, which consists of multiple convolutional layers or transposed convolutional layers to gradually restore and generate the final segmentation prediction results The specific expression is: Among them, Decoder(·) represents the decoder, represents the segmentation prediction mask output by the decoder; The final segmentation prediction mask It has the same resolution as the original input image, and the value of each pixel represents the probability or classification label of the pixel belonging to a certain category, which is used for the prediction output of the few-shot semantic segmentation task.
4. A few-sample semantic segmentation device based on image-text fusion, used to implement the few-sample semantic segmentation method based on image-text fusion as claimed in any one of claims 1 to 3, characterized in that: include: Data preprocessing module, used to collect and preprocess data sets; Image feature extraction module, which uses pre-trained deep neural network to extract multi-scale image features; A text feature extraction module, which uses a text encoder to extract text features of a preset text prompt; Feature fusion module, which fuses the extracted text features with image features; The correlation mining module uses the attention mechanism to mine the correlation between the supporting image features and the query image features; The decoding module upsamples the fused image features and decodes them to generate segmentation prediction results.
5. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 3 is implemented.
6. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
Model training method and device, target detection method and device and computer storage medium
CN117809021A
Small sample medical image segmentation method based on text semantic guidance
CN118314161A