Small sample semantic segmentation method based on object query

The object prototype extraction module and feature fusion module are used to enhance feature correlation, and the bidirectional Transformer module is used to perform foreground-background separation, which solves the problems of detail loss and noise sensitivity in small-sample semantic segmentation and improves the accuracy and generalization ability of segmentation.

CN120689613APending Publication Date: 2025-09-23WUXI RES INST OF NANJING UNIV OF INFORMATION ENG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510752510.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Among the existing small-sample semantic segmentation technologies, the prototype matching method has problems of loss of detail information and insufficient generalization ability, and the pixel comparison method has noise sensitivity and computational efficiency bottlenecks. Especially in the small-sample scenario, the model is prone to overfitting, making it difficult to accurately segment complex or detailed objects and distinguish semantically similar areas.

Method used

The object prototype extraction module is used to generate foreground and background prototypes, the feature fusion module is used to enhance feature correlation, and the bidirectional Transformer module is used to perform the mask attention mechanism for foreground-background separation, combining position embedding and cross entropy loss to optimize the segmentation mask.

Benefits of technology

It effectively solves the problems of detail loss and noise sensitivity, improves the accuracy and generalization ability of segmentation, and can better segment complex objects and distinguish semantically similar areas in small sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689613A_ABST
    Figure CN120689613A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample semantic segmentation method based on object query, and the method comprises the following steps: generating N object prototypes through mask pooling operation, the object prototypes comprising first N / 2 foreground prototypes and last N / 2 background prototypes; based on the object prototype, performing class correlation enhancement on the support set features and the query image features through a feature fusion module to generate enhanced features; inputting the enhanced features and the object prototype into a bidirectional Transform module, carrying out pixel-level and object-level bidirectional interaction through a foreground-background separated mask attention mechanism, and outputting segmentation features; generating a final semantic segmentation mask according to the segmentation features; according to the invention, the background interference problem is effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semantic segmentation technology, and in particular to a small sample semantic segmentation method based on object query. Background Art

[0002] Few-shot semantic segmentation is a computer vision task that aims to accurately segment new categories using limited labeled examples, aiming to improve the model's generalization ability to unseen categories. FSS methods follow the meta-learning paradigm and are generally categorized into prototype-based comparison methods and pixel-based matching methods.

[0003] Problems with existing prototype matching-based solutions: The core flaw of this solution lies in the loss of detailed information and insufficient generalization capability caused by its global feature compression mechanism. Because global average pooling is used to compress the support set foreground features into a single prototype vector, the spatial structural information of the target (such as edge contours and local texture differences) is over-smoothed during the compression process, making it difficult to accurately depict objects with complex morphology or rich details (such as branch branches and animal hair). Especially in few-sample scenarios, when the number of support set samples is limited and there may be differences in perspective, occlusion, or illumination, a single prototype cannot fully cover the diversity within the class and is easily affected by abnormal samples. For example, if the support set targets are partially occluded or have atypical postures, the prototype vector will deviate from the true class distribution, leading to missegmentation of similar regions in the query image. In addition, prototype matching relies heavily on high-level semantic features and lacks sensitivity to local details. When segmenting targets with significant intra-class differences (such as different dog breeds), it has difficulty distinguishing subtle feature differences, resulting in blurred segmentation boundaries or local missed detections.

[0004] Problems with existing technical solutions based on pixel comparison: The main problems of this solution are noise sensitivity and computational efficiency bottlenecks. Due to direct pixel feature comparison, any local similarity (such as the similarity between background texture and target surface pattern) may cause mismatching, resulting in artifacts or breaks in the segmentation results. Although this method retains fine-grained features, it lacks high-level semantic guidance and is difficult to distinguish between semantically similar but category-different regions (such as the local feature similarity between cat ears and dog ears), which easily leads to category confusion. In addition, the dense pixel pair similarity calculation causes the computational complexity to increase quadratically with the image resolution. For example, when processing a 512×512 image, more than 2.6 billion pixel pair similarities need to be calculated, which seriously restricts real-time performance. In the few-sample scenario, the limited training data further amplifies the noise sensitivity, and the model is prone to overfitting to the local features of the support set samples, and the generalization ability is significantly reduced. For example, when the support set target contains rare background elements, the same background in the query image will be mistakenly activated, interfering with the accuracy of foreground segmentation. Summary of the Invention

[0005] Purpose of the invention: The purpose of the present invention is to provide a small-sample semantic segmentation method based on object query, which solves the problems of detail loss and insufficient intra-class generalization ability caused by global feature compression in the prototype matching method; eliminates the inherent defects of the pixel-level comparison method in terms of noise sensitivity and computational redundancy; and alleviates the problem of model overfitting caused by data distribution deviation in small-sample scenarios.

[0006] Technical solution: The small sample semantic segmentation method based on object query described in the present invention includes the following steps:

[0007] (1) Generate N object prototypes through mask pooling operation, which include the first N / 2 foreground prototypes and the last N / 2 background prototypes;

[0008] (2) Based on the object prototype, the feature fusion module is used to enhance the class relevance of the support set features and the query image features to generate enhanced features;

[0009] (3) The enhanced features and object prototypes are input into the bidirectional Transformer module, and bidirectional interaction between pixel level and object level is performed through the mask attention mechanism of foreground-background separation to output segmentation features;

[0010] (4) Generate the final semantic segmentation mask based on the segmentation features.

[0011] Furthermore, in step (1), the mask pooling operation includes the following steps:

[0012] (11) Extract the mid-level feature map F∈H×W×C of the support set image, where H and W are spatial dimensions and C is the number of channels;

[0013] (12) Generate the category correlation matrix A(i) based on the support set mask to distinguish the foreground and background areas;

[0014] (13) Calculate the weight mask W through a two-layer MLP network q (i) and combined with the two-dimensional sinusoidal position code R sin (i) Generate dynamic weights:

[0015]

[0016] Among them, W q (i) represents the weight of the i-th pixel of the q-th mask; A(i) represents the similarity matrix between the support sample foreground and the prototype; σ is the sigmoid function; f poolingweight It is a multi-layer perceptron.

[0017] (14) Perform weighted average pooling on each weight mask to obtain N object prototypes S∈N×C.

[0018] Furthermore, the feature fusion module in step (2) includes the following steps:

[0019] (21) High-level features of the support set and query image high-level features Perform convolution respectively to generate query vector Q, key vector K and value vector V;

[0020] (22) Calculate the correlation matrix A between the support feature and the query feature through spatial attention ji :

[0021]

[0022] Among them, K i and Q j Indicates the query and key obtained by convolution of support features and query features.

[0023] (23) Perform matrix multiplication on the correlation matrix and the value vector V to generate enhanced features

[0024] (24) will support prototype P s and After splicing, the enhanced features R are output through 1×1 convolution fusion.

[0025] Furthermore, in step (3), the bidirectional Transformer module includes the following steps:

[0026] (31) Through the masked cross-attention mechanism, the first N / 2 object queries focus only on the foreground area, and the last N / 2 queries focus only on the background area;

[0027] (32) In the self-attention layer, global contextual interactions are performed between object queries;

[0028] (33) Injecting spatial information through position embedding.

[0029] Furthermore, position embedding includes: position embedding of object query Generated by linear projection of dynamic prototype S; position embedding of pixel features By a fixed sine code R sin Generated by fusion with the initial feature r0.

[0030] 6. The object query-based small-sample semantic segmentation method according to claim 5, wherein the masked cross-attention mechanism constrains the attention range by the following formula:

[0031]

[0032] Among them, M q(i) indicates whether the qth object query focuses on the i-th pixel; A(i) represents the similarity matrix between the support sample foreground and the prototype.

[0033] By setting the attention weights of irrelevant pixels to negative infinity, they are suppressed in the Softmax calculation.

[0034] Furthermore, the method adopts a meta-learning paradigm for training, with the backbone network being a pre-trained VGG-16 or ResNet-50. During the training process, the backbone network parameters are frozen, and only the parameters of the object prototype extraction module, the feature fusion module, and the bidirectional Transformer module are updated.

[0035] Furthermore, the generation of segmentation masks is optimized by jointly using cross entropy loss and Dice loss:

[0036]

[0037] Among them, Mi represents the prediction result of pixel i; Mq,i represents the ground truth.

[0038] An electronic device according to the present invention includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods when executing the program.

[0039] The computer-readable storage medium of the present invention stores a computer program, which implements the steps of any one of the methods when executed by a processor.

[0040] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: the present invention effectively solves the matching granularity problem and the data deviation caused by insufficient samples and the resulting intra-class differences and inter-class similarities through the object prototype extraction module (OPE), feature fusion module (FFM), and object transformer block (OTB), and effectively reduces the background interference problem. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 is a flow chart of the present invention;

[0042] Figure 2 is the fusion module of the present invention;

[0043] Figure 3 These are the training curves of the method of the present invention and the baseline method. DETAILED DESCRIPTION

[0044] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0045] like Figure 1-2As shown, an embodiment of the present invention provides a method for controlling current limiting of a grid-connected doubly-fed wind turbine generator system with a timing fault, which is compatible with grid-connected guidelines, comprising the following steps:

[0046] Step 1: The present invention uses object prototype To fully extract the semantic information of the supporting image. Specifically, using N weight masks To perform mask pooling on the support features (middle-level). Given the target feature Mask W, the prototype of the qth object is:

[0047]

[0048] Step 2: To obtain more accurate and representative semantic information, we use a foreground-background separation strategy to calculate the mask W and a 2D sinusoidal position embedding to enhance features. Each pixel is assigned a category-related weight and some noise is removed to make the extracted information cleaner. The i-th pixel of the q-th weight mask is calculated by:

[0049]

[0050] There are N object prototypes in total, of which the first N / 2 are used to represent the foreground and the last N / 2 are used to represent the background. σ is the sigmoid function, f poolingweight It is a 2-layer N-dimensional multilayer perceptron. It calculates the category correlation matrix A for the foreground and background separately. Specifically, it performs GAP on the mid-level features of the foreground in the support set using the support mask to obtain the prototype. Then, it performs pixel-level cosine similarity calculation to obtain A(i).

[0051] Step 3: The present invention uses a feature fusion module to enhance feature representation (such as Figure 2 As shown in Figure 2, due to the diversity within the category, there is low correlation between regions of the same category. The accuracy of feature matching is improved by information interaction between the query image and the support image. First, the high-level support features after encoding are and query features Convolution is respectively obtained Notice The mask is used to remove the background and only include the foreground, C k Indicates the number of low-dimensional channels after convolution. Then Q and K are reshaped into C k ×N size, N=H h ×W h . Perform matrix multiplication on the transpose of Q and K and perform softmax to obtain the final feature attention matrix A ji ; A jiIt represents the correlation between the i-th position in the support feature and the j-th position in the query feature. The more similar the features of the two positions are, the higher the correlation is.

[0052]

[0053] Step 4: Simultaneously reshape the V into a C k ×N size, N=H h ×W h , the present invention is in A ji Perform matrix multiplication with V and reshape into H h ×W h ×C k The size of the feature map is obtained by emphasizing the important information.

[0054] Step 5: To further alleviate intra-class differences, the present invention will support prototype P s and The final enhanced feature R is obtained by fusion, which can highlight the category-related parts of the query features and suppress the irrelevant parts, thereby improving the accuracy of segmentation. Specifically, the prototype P s Refactor and After fusion, it is fused according to the dimension through 1×1 convolution.

[0055]

[0056] The final enhanced feature R and object prototype X will be input into the Object Transformer module for subsequent information interaction.

[0057] Step 6: The object transformer acts as a bidirectional cross-attention and self-attention based module to extract the object query from the support set and pixel features enhanced by class correlation As input, H and W are the feature dimensions output after feature fusion. This module iterates in a bidirectional manner of pixel-to-query and query-to-pixel. The present invention has 4 object transformer blocks, R0 is the output feature of the feature fusion module. The output X in the lth block is l and R l As the input of the next block. The output of the last block R L is the final output feature.

[0058] Step 7: Object Query X l First, read the pixel-level feature R through the foreground / background mask attention mechanism l-1 , focus on the foreground and background, and extract semantic information respectively. Then Xl Standard self-attention and feedforward networks are used to perform object-level reasoning and feature fusion, so that object queries can summarize foreground and background prior knowledge. Although the foreground and background are separated by masks, the query vectors can still interact with each other in the Self-Attention layer, ensuring the integration of global context information. Then a reverse cross-attention mechanism is used to combine X l The object semantic information in R is rewritten l-1 , and then get the final R through pixel-level FFN l-1 .

[0059] Step 8: This paper uses masked cross attention to forcibly separate the foreground and background regions of interest. The first half of the object query is only allowed to focus on pixels with high category relevance, and the second half of the object query is only allowed to focus on pixels with low category relevance. By setting the attention weight of irrelevant pixels to negative infinity, noise is completely discarded in the Softmax calculation:

[0060]

[0061] The mask matrix M controls the attention range and determines whether the qth object query focuses on the i-th pixel. It is defined as follows:

[0062]

[0063] Step 9: Since conventional attention mechanisms such as permutation are insensitive to input order (for example, swapping the positions of two input tokens does not change the output), this paper injects position information through position embedding to enable the model to perceive spatial structure. This paper adds position embedding to the query and key in the attention layer instead of the value, because the attention weight is calculated based on the similarity between the query and key, while the value is responsible for conveying content information.

[0064] For object queries, position embeddings are used It is achieved by linear projection f ObjEmbed Map the dynamic object prototype S to the position embedding space and combine it with the learnable embedding E X Add together, where is an end-to-end learnable embedding, f ObjEmbed is a trainable linear projection layer.

[0065] P X =E X +f ObjEmbed (S)

[0066] Step 10: For pixel features, the present invention uses position embedding It is encoded by a fixed sinusoidal code R sinIt is fused with the initial pixel feature R0, where R sin Provides absolute position information, which is generated based on normalized coordinates and scaled to the input image size during testing, f PixEmbed is another trainable linear projection layer.

[0067] P R =R sin +f PixEmbed (R0)

[0068] During the training process, cross entropy loss and Dice loss are used jointly to balance the classification accuracy and region overlap rate.

[0069]

[0070] Table 1 shows the mIoU comparison between the proposed method and the SOTA method on the Pascal dataset. The proposed method uses vgg16 and resnet50 as the backbone to demonstrate the mIoU in 1-shot and 5-shot contexts, respectively. It can be clearly seen from Table 1 that when the backbone is vgg16, the proposed method outperforms the SOTA method by 1.9% and 2.6% in 1-shot and 5-shot contexts, respectively. When the backbone is resnet50, the proposed method outperforms the SOTA method by 2.2% and 2.9%, respectively. This shows that the proposed method has superior performance and can obtain more accurate segmentation masks in the query set.

[0071] Table 1 Performance of various methods on the Pascal-5i dataset

[0072]

[0073] Table 2 shows a comparison of mIoU between our method and the state-of-the-art method on the COCO dataset. The COCO dataset has more complex scenes and is therefore more challenging. We also demonstrate mIoU using VGG16 and ResNet50 as the backbone for 1-shot and 5-shot scenarios, respectively. As can be seen from Table 2, our method, when using VGG16 as the backbone, outperforms the state-of-the-art method by 3.0% and 2.3% for 1-shot and 5-shot scenarios, respectively. This demonstrates that the object query and object transformer modules proposed in this paper can achieve good segmentation accuracy.

[0074] Table 2 Performance of various methods on the COCO-20i dataset

[0075]

[0076] The method of the present invention mainly includes three modules: object prototype extraction module (OPE), feature fusion module (FFM), and object transformer block (OTB). Table 3 shows the verification of the effectiveness of each component of the method of the present invention. The present invention first uses OPE to extract the object query and the query feature extracted by the backbone to obtain the segmentation mask using the simplest cosine similarity. Then the FFM and OTB modules are added in sequence. As can be seen from Table 3, after adding the FFM and OTB modules, the results are improved by 2.4% and 3.8% respectively. The present invention also conducted an experiment to remove the OPE and FFM modules in the method and test the effect of the OBT module. Here, the present invention uses ordinary support prototypes and does not use query feature enhancement. The results show that better results can also be achieved. After the addition of FFM, the results exceeded the baseline by 5.9%, indicating that the method proposed by the present invention effectively solves the matching granularity problem and the data deviation caused by insufficient samples and the resulting intra-class differences and inter-class similarities, and effectively reduces the background interference problem.

[0077] Table 3. Validity experimental results of the core modules of the method of the present invention

[0078]

[0079]

[0080] Since a single prototype cannot correctly summarize the true data distribution of the category, and pixel-level feature comparison lacks the guidance of high-level semantic information, the present invention designs an object prototype extraction module to extract object queries (OPE) to mine information. Here, the present invention is compared with the traditional single prototype (Proto) and Super pixel (SP) methods to prove its effectiveness. As shown in Table 4, the present invention first tested three feature extraction methods respectively. Among them, the multi-prototype method based on SP achieved good results, but the object query-based method of the present invention exceeded SP by 2.3%. In addition, the present invention adds the OBT module to the three methods. It can be seen that the object query can better extract supporting features, exceeding the baseline by 5.2%.

[0081] Table 4 Experimental results on the effectiveness of object prototype extraction module

[0082] S-Pro Spix OPE OTB mIoU (%) √ 50.3 √ 57.2 √ 59.5 √ √ 59.9 √ √ 62.7 √ √ 67.9

[0083] Foreground-background cross attention (MCA) is a core component of the object transformer block. In order to solve background interference and noise interference, the present invention changes the traditional transformer's cross attention part into a foreground-background mask attention mechanism to help the model better distinguish the foreground and background while removing some of the interference noise. Table 5 shows the effect of MCA in combination with different transformer models. The present invention first uses the original transformer (VA) as the baseline. It can be seen from the table that the OT method with MCA exceeds the baseline by 1.2%, and after combining with OTB, it exceeds the baseline by 3.8%. This shows that the foreground-background mask attention mechanism plays a key role in this. It also shows that the bidirectional transformer module OTB designed by the present invention can better separate the foreground and background of the query image, alleviate background noise interference, and achieve better segmentation effect.

[0084] Table 5 Masked attention mechanism effectiveness experiment

[0085] MCA VA OTB mIoU (%) √ 64.1 √ 65.3 √ √ 66.1 √ √ 67.9

[0086] like Figure 3 As shown in the figure, we compare the changes in the mIoU metric during the training phase between our method and a baseline method (IPMT, which also uses an iterative transformer architecture). The experimental results show that during training, the baseline model's mIoU metric slightly outperforms our method (+0.5%); however, during inference, our method achieves a higher mIoU value (+1.5%) than the baseline model, demonstrating its superior generalization capabilities.

Claims

1. A small sample semantic segmentation method based on object query, characterized by: The following steps are involved: (1) Generate N object prototypes through mask pooling operation, which include the first N / 2 foreground prototypes and the last N / 2 background prototypes; (2) Based on the object prototype, the feature fusion module is used to enhance the class relevance of the support set features and the query image features to generate enhanced features; (3) The enhanced features and object prototypes are input into the bidirectional Transformer module, and bidirectional interaction between pixel level and object level is performed through the mask attention mechanism of foreground-background separation to output segmentation features; (4) Generate the final semantic segmentation mask based on the segmentation features.

2. The small sample semantic segmentation method based on object query according to claim 1, characterized in that: In step (1), the mask pooling operation includes the following steps: (11) Extract the mid-level feature map F∈H×W×C of the support set image, where H and W are spatial dimensions and C is the number of channels; (12) Generate the category correlation matrix A(i) based on the support set mask to distinguish the foreground and background areas; (13) Calculate the weight mask W through a two-layer MLP network q (i) and combined with the two-dimensional sinusoidal position code R sin (i) Generate dynamic weights: Among them, W q (i) represents the weight of the i-th pixel of the q-th mask; A(i) represents the similarity matrix between the support sample foreground and the prototype; σ is the sigmoid function; f poolingWeight It is a multi-layer perceptron. (14) Perform weighted average pooling on each weight mask to obtain N object prototypes S∈N×C.

3. The small sample semantic segmentation method based on object query according to claim 1, characterized in that: The feature fusion module in step (2) includes the following steps: (21) High-level features of the support set and query image high-level features Perform convolution respectively to generate query vector Q, key vector K and value vector V; (22) Calculate the correlation matrix A between the support feature and the query feature through spatial attention ji : Among them, K i and Q j Indicates the query and key obtained by convolution of support features and query features. (23) Perform matrix multiplication on the correlation matrix and the value vector V to generate enhanced features (24) will support prototype P s and After splicing, the enhanced features R are output through 1×1 convolution fusion.

4. The small sample semantic segmentation method based on object query according to claim 1, characterized in that: In step (3), the bidirectional Transformer module includes the following steps: (31) Through the masked cross-attention mechanism, the first N / 2 object queries focus only on the foreground area, and the last N / 2 queries focus only on the background area; (32) In the self-attention layer, global contextual interactions are performed between object queries; (33) Injecting spatial information through position embedding.

5. The small sample semantic segmentation method based on object query according to claim 4 is characterized in that: Position embedding includes: Position embedding of object queries Generated by linear projection of dynamic prototype S; position embedding of pixel features By a fixed sine code R sin Generated by fusion with the initial feature R0.

6. The small sample semantic segmentation method based on object query according to claim 5, characterized in that: The masked crisscross attention mechanism constrains the attention range through the following formula: Among them, M q (i) indicates whether the qth object query focuses on the i-th pixel; A(i) represents the similarity matrix between the support sample foreground and the prototype. By setting the attention weights of irrelevant pixels to negative infinity, they are suppressed in the Softmax calculation.

7. The small sample semantic segmentation method based on object query according to claim 1, characterized in that: The method adopts a meta-learning paradigm for training. The backbone network is a pre-trained VGG-16 or ResNet-50. During the training process, the backbone network parameters are frozen, and only the parameters of the object prototype extraction module, feature fusion module and bidirectional Transformer module are updated.

8. The small sample semantic segmentation method based on object query according to claim 1, characterized in that: The generation of segmentation masks is optimized by jointly cross entropy loss and Dice loss: Among them, Mi represents the prediction result of pixel i; Mq,i represents the ground truth.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium, characterized in that A computer program is stored, and when the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.