Cross-task universal foreground segmentation method

By adopting a cross-task common method in the foreground segmentation task, combining multi-modal comparison learning of multi-scale strategies, mask attention mechanism, edge enhancement module and CLIP model, the problem of lack of uniformity and poor segmentation effect in the existing technology is solved, and the segmentation effect of high accuracy and detailed expression is achieved.

CN119991724AActive Publication Date: 2025-05-13FUDAN UNIVERSITY

Patent Information

Application Number
CN202411889134.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-13
Estimated Expiration
2044-12-20

AI Technical Summary

Technical Problem

The prior art lacks uniformity in the foreground segmentation task, making it difficult to effectively distinguish the foreground from the background, resulting in poor segmentation effect.

Method used

A common foreground segmentation method across tasks is adopted, combining multi-modal comparison learning of multi-scale strategies, mask attention mechanism, edge enhancement module and CLIP model, and a variety of foreground segmentation tasks are uniformly processed, significantly improving the accuracy and detailed expression of segmentation boundaries.

Benefits of technology

The unified processing of multiple prospect segmentation tasks is achieved, which significantly improves the accuracy and detailed expression of segmentation boundaries, and improves the accuracy and robustness of segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991724A_ABST
    Figure CN119991724A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer vision, and discloses a cross-task universal foreground segmentation method, which comprises the following steps of: firstly, constructing a unified foreground segmentation framework based on a multi-scale strategy and a mask attention mechanism, introducing binary query, and representing foreground and background features in an image through foreground query and background query; through an edge enhancement module, edge information of the image is extracted by adopting a convolutional neural network, and edge features and multi-scale features of the image are fused in combination with a multi-scale deformation attention mechanism; the multi-scale features and binary query are input into a Transform decoder together, a mask attention mechanism is applied, the binary query is updated, and accurate foreground and background segmentation masks are obtained; the segmentation results of the foreground and the background are refined by using a multi-modal contrast learning strategy, and the accuracy of the segmentation boundary and the detail retention effect are improved. The method can be widely applied to different types of foreground segmentation tasks, and high-precision segmentation results can be realized in a plurality of complex scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and in particular to a cross-task universal foreground segmentation method based on multimodal learning. Background Art

[0002] Foreground segmentation is an important task in computer vision, which aims to distinguish salient objects (foreground) from the rest (background) in an image. This task covers multiple subfields, such as salient object detection, camouflaged object detection, shadow detection, focus blur detection, and forgery detection. Traditionally, these tasks often use specially designed architectures that lack versatility and are therefore difficult to share and extend across different tasks. In addition, these methods mainly focus on the identification of the foreground, while often ignoring the background and the relationship between the background and the foreground, resulting in poor segmentation results.

[0003] In the field of computer vision, some general segmentation tasks such as instance segmentation, semantic segmentation, and panoptic segmentation have made significant progress, and many advanced models have performed well in these tasks. However, these models usually lack specialized optimization for specific foreground segmentation tasks. For example, in camouflaged object detection, existing general models are insufficient in distinguishing camouflaged objects from the background. At the same time, traditional segmentation methods often generate multiple redundant masks for the same image, while in practical applications users often only need specific foreground masks, such as image background removal, which makes these methods perform poorly in practical scenarios.

[0004] Although some works in recent years have attempted to achieve universality between salient object detection and camouflaged object detection, the scope of application of these methods is still limited, especially when faced with different types of foreground segmentation tasks, they are still not comparable to task-specific models. In addition, most foreground segmentation models focus mainly on foreground objects, while not making full use of background information and its relationship with the foreground. In traditional segmentation tasks, the background is usually a key category. Therefore, in foreground segmentation tasks, how to effectively use background information to optimize segmentation results is an urgent problem to be solved.

[0005] Therefore, based on the challenges faced by existing technologies, it is urgent to develop a framework that can uniformly handle multiple foreground segmentation tasks and make full use of background information to improve the accuracy and robustness of segmentation. Summary of the invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned prior art and provide a cross-task general foreground segmentation method. By combining multi-scale strategy, mask attention mechanism, edge enhancement module and multimodal contrast learning of CLIP model, it effectively realizes the unified processing of multiple foreground segmentation tasks and significantly improves the accuracy of segmentation boundaries and detail expression.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A cross-task universal foreground segmentation method, the method steps comprising:

[0009] Construct a unified foreground segmentation model, introduce binary query, and represent the foreground and background features in the image through foreground query and background query;

[0010] Extract multi-scale image features through visual backbone network;

[0011] The edge features of the image are extracted through the edge enhancement module, and combined with the multi-scale deformation attention mechanism, the edge features are fused with the multi-scale image features to obtain the edge-enhanced multi-scale image features;

[0012] The edge-enhanced multi-scale features are input into the Transformer decoder together with the binary query. The Transformer decoder applies the masked attention mechanism to update the binary query so that the binary query focuses on the features related to the foreground and background.

[0013] Decodes a binary query and produces foreground and background segmentation masks and class predictions.

[0014] As a preferred technical solution, the edge enhancement module includes multiple pairs of edge feature injection modules and backbone feature extraction modules, and each hierarchical block of the backbone network corresponds to an edge injection module and a backbone feature extraction module;

[0015] The edge feature injection module extracts edge features of the image through a convolutional neural network. The convolutional neural network adopts ResNet50, including STEM and multiple convolution blocks, wherein STEM is used for preliminary feature extraction; the feature map output by the convolution block is flattened and 1×1 convolutionally projected to the same dimension, and connected to form a feature pyramid of edge features. The image features extracted by the backbone network DINOv2 are fused with the edge features using the cross-attention based edge feature injection structure;

[0016] The backbone feature extraction module applies multi-scale deformation attention and further processes the features fused by the edge feature injection module through multiple convolution blocks to obtain enhanced foreground and background feature expressions.

[0017] As a preferred technical solution, the edge feature injection is expressed as:

[0018]

[0019] Among them, MSDA represents multi-scale deformation attention, with the normalized output features of the i-th layer of the backbone network As the query, take the normalized edge features as keys and values; γ i It is a learnable parameter used to balance the backbone features and fusion features.

[0020] As a preferred technical solution, the backbone feature extraction module is expressed as:

[0021]

[0022] Among them, MSDA stands for multi-scale deformation attention, with normalized edge features As a query, take the output features of the next layer of the backbone network As keys and values; the ConvFFN structure consists of two fully connected layers and a depth-separable convolutional layer, the output It will be used as the input of the edge feature injection module of the next layer.

[0023] As a preferred technical solution, the edge enhancement module performs multi-scale feature fusion and output for edge features and image features, as follows:

[0024] The features output by different blocks of the backbone network are upsampled to resolutions of 1 / 4, 1 / 8, 1 / 16 and 1 / 32, the output of the backbone feature extraction module is split and restored to the original size, and then the upsampled backbone features are added to the corresponding split output and STEM output from the backbone feature extraction module to obtain edge-enhanced multi-scale image features.

[0025] As a preferred technical solution, the Transformer decoder applies a masked attention mechanism to update the binary query, as follows:

[0026] The multi-scale image features extracted by the edge enhancement module are input into the pixel decoder based on multi-scale deformation attention. The pixel decoder combines the feature map generated by the backbone network to perform dense pixel-level prediction and generate pixel-level output features.

[0027] The pixel-level output features are input into the Transformer decoder together with the binary query, and the binary query is updated through the mask attention mechanism. The process is expressed as:

[0028]

[0029] Among them, K l and V l is the image feature obtained by linear transformation from the l-th layer block of the pixel decoder, all of which belong to space; represents the query features extracted from the l-th layer Transformer decoder block, and X0 is initialized by the input query features of the Transformer decoder; is a binary query at level l, is the attention mask of the previous layer, defined as follows:

[0030]

[0031] in, By decoding the binary query GQ l-1 And perform binarization to obtain, and the dimension size is the same as K l Be consistent.

[0032] As a preferred technical solution, the method initializes the attention mask by principal component analysis and binarization of the feature map of the last block of the backbone network.

[0033]

[0034] Among them, F DINOv2 Represents the binarized feature map from the last backbone network block of DINOv2, which is resized to the same resolution as K1.

[0035] As a preferred technical solution, the foreground segmentation model further uses a multimodal refinement module to refine the segmentation results of the foreground and background, as follows:

[0036] Decode the binary query to obtain foreground and background masks, resize the foreground and background masks and overlay them on the original image to form a mask-fused image;

[0037] Use text descriptions to represent foreground and background separately;

[0038] Use the image encoder and text encoder of the CLIP model to encode the mask fusion image and text description respectively;

[0039] Compute the contrastive loss between mask-fused image and text description in multimodal feature space The foreground and background segmentation results are refined by minimizing the contrast loss.

[0040] As a preferred technical solution, the contrast loss between the mask fusion image and the text feature The details are as follows:

[0041]

[0042]

[0043]

[0044] in, is the contrast loss from image to text; is the contrast loss from text to image; I f ,I b ,T f ,T b ∈R N×C They represent the image features and text features of the foreground and background obtained by CLIP respectively; τ is the temperature parameter used to control the smoothness of the softmax function.

[0045] As a preferred technical solution, during the training phase, the multimodal refinement module optimizes the boundaries of the foreground and background masks through multiple iterations to achieve the best effect; during the reasoning phase, the multimodal refinement module is removed and only the optimized foreground segmentation model is retained.

[0046] Compared with the prior art, the present invention has the following beneficial effects:

[0047] 1) This paper proposes a unified foreground segmentation model architecture, which represents the features of the foreground and background as foreground queries and background queries respectively, and inputs them into the Transformer decoder together with the visual features extracted from the image by the visual feature extraction module through a multi-scale strategy, and applies a mask attention mechanism to make the foreground query and background query focus on the foreground and background information in the image respectively. Then, a multi-layer perceptron is used to decode these query features to generate the final foreground and background masks, as well as the corresponding category predictions. The binary query is a universal feature with strong generalization ability, and can learn the corresponding foreground and background information according to the context of different tasks.

[0048] 2) The present invention constructs an edge enhancement module to perform multi-scale edge enhancement features, extracts edge information of the image through a convolutional neural network, combines the multi-scale deformation attention mechanism, and fuses the edge features with the multi-scale features of the image, thereby enhancing the expression of the foreground and background in the image. It can effectively fuse the global features and edge information of the image, thereby significantly improving the clarity of the boundary and the accuracy of segmentation in the foreground segmentation task.

[0049] 3) The present invention also proposes a multimodal refinement module, which iteratively refines the generated mask edges to ensure that only appropriate pixels are included in the foreground and background. This process makes the mask fusion image closer to the corresponding text in the feature space, while keeping a distance from the unmatched text features. This not only makes the mask edge more accurate, but also further increases the gap between the foreground and background. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1It is a flowchart of the cross-task universal foreground segmentation method of the present invention;

[0051] Figure 2 It is a schematic diagram of the structure of an edge enhancement module for extracting and fusing image edge features according to the present invention;

[0052] Figure 3 Schematic diagram of the workflow of the multimodal refinement module of the present invention. DETAILED DESCRIPTION

[0053] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0054] Example 1

[0055] The present invention aims to solve the problem that the foreground segmentation method in the prior art lacks uniformity and is difficult to effectively distinguish the foreground from the background. The general multimodal foreground segmentation method is proposed. The overall process is as follows: Figure 1 As shown:

[0056] Firstly, a unified foreground segmentation framework is constructed based on multi-scale strategy and mask attention mechanism. By introducing the concept of binary query (GQ), the features of foreground and background are represented as foreground query and background query respectively. The binary query is a pair of learnable tensors, which are input into the Transformer decoder together with the visual features extracted from the image by the visual feature extraction module through the multi-scale strategy. The mask attention mechanism is applied so that the foreground query and background query focus on the foreground and background information in the image respectively. The binary query is randomly initialized, and the binary query will learn the corresponding features during the forward propagation process. Then, the multi-layer perceptron is used to decode these query features to generate the final foreground and background masks and the corresponding category predictions. The binary query is a universal feature with strong generalization ability, which can learn the corresponding foreground and background information according to the context of different tasks. The constructed unified foreground segmentation model can handle different foreground segmentation tasks including salient object detection, camouflaged object detection, shadow detection, focus blur detection and forgery detection.

[0057] Secondly, an edge enhancement module is proposed, which extracts the edge information of the image through a convolutional neural network (CNN) and combines the multi-scale deformation attention mechanism to fuse the edge features with the multi-scale features of the image to enhance the edge features of the image. This fusion can improve the accuracy of the segmentation boundary.

[0058] Finally, the present invention also proposes a multimodal refinement module, which further refines the generated foreground and background segmentation results through multimodal contrastive learning of the CLIP model. The foreground and background masks are fused with the image to generate a mask fusion image, and the segmentation boundaries of the foreground and background are optimized through a multimodal contrastive learning strategy. This refinement process significantly improves the detail expression and boundary clarity of the segmentation results, improves the accuracy of the segmentation boundaries and the detail retention effect, and can be selectively removed in the inference stage to improve segmentation efficiency.

[0059] The method of the present invention can be widely applied to different types of foreground segmentation tasks and achieve high-precision segmentation results in multiple complex scenarios.

[0060] The present invention proposes a FOCUS architecture based on a multi-scale strategy and a mask attention mechanism, which can handle a variety of foreground segmentation tasks and is suitable for tasks such as salient object detection, camouflaged object detection, shadow detection, focus blur detection, and forgery detection. The concept of binary query (GQ) is introduced. The binary query consists of two independent vectors, namely the foreground query and the background query, which are used to embed and represent the foreground and background in the image in a specific task context.

[0061] (S1.1) Extracting image features by a backbone network based on DINOv2;

[0062] (S1.2) using the edge information of the foreground object to correct the image features extracted by the backbone network through the edge enhancement module. The edge enhancement module extracts the edge information of the image through the convolutional neural network, and combines the multi-scale deformation attention mechanism to fuse the edge features with the image features extracted by the backbone network to obtain multi-scale features.

[0063] (S1.3) The multi-scale features are input into the Transformer decoder, and a mask attention mechanism is applied in the Transformer decoder to make the binary query focus on the features related to the foreground and background, thereby obtaining accurate foreground and background segmentation masks;

[0064] (S1.4) Decode the binary query through a multi-layer perceptron to generate segmentation masks and category predictions for foreground and background;

[0065] (S1.5) Compare the generated segmentation mask with the true label, and optimize the segmentation mask using the binary cross entropy loss function and the Dice loss function.

[0066] The steps of optimizing image features through the edge enhancement module are as follows:

[0067] (S2.1) using a ResNet50 convolutional neural network to extract edge information of the image, wherein the convolutional neural network is composed of a STEM and a plurality of convolutional blocks, wherein the STEM is used for preliminary feature extraction, and the convolutional blocks are used for extracting multi-scale edge features;

[0068] (S2.2) fusing the extracted edge features with the image features generated by the DINOv2 backbone network through a multi-scale deformation attention mechanism, wherein the deformation attention mechanism enhances the expression capability of foreground and background features through multi-scale spatial transformation operations;

[0069] (S2.3) The fused features are further processed through multiple convolution blocks to obtain enhanced foreground and background feature expressions, thereby improving the accuracy of foreground segmentation.

[0070] (S2.4) In the fused features, the local spatial information and global semantic information of the image features are further optimized through iterative fusion operations of multi-scale spatial dimensions and channel dimensions, making the segmentation boundary between the foreground and the background clearer and more precise.

[0071] The present invention also refines the segmentation results of the foreground and background through a multimodal refinement module, and the steps are as follows:

[0072] (S3.1) fusing the foreground and background masks generated by the binary query with the original image to form a mask-fused image;

[0073] (S3.2) using the image encoder and text encoder of the CLIP model to encode the mask fusion image and the preset text description respectively, and the text description can be replaced according to different processing tasks;

[0074] (S3.3) Calculate the contrast loss of the mask fusion image and the text description in the multimodal feature space, and refine the segmentation results of the foreground and background by minimizing the contrast loss, thereby improving the accuracy of the segmentation boundary and the detail expression.

[0075] During the training phase, the multimodal refinement module optimizes the boundaries of the foreground and background masks through multiple iterations to achieve the best effect; during the reasoning phase, the multimodal refinement module can be removed, and only the optimized foreground segmentation model is retained to improve reasoning efficiency.

[0076] The implementation of the unified foreground segmentation model proposed in the present invention is specifically as follows:

[0077] The present invention selects DINOv2 as the backbone network of the FOCUS architecture to extract image features. DINOv2 is a self-supervised learning model based on Transformer. It uses a multi-head self-attention mechanism to analyze global dependencies in images, thereby effectively capturing rich semantic information. The visualization of its feature map shows that the model is already able to focus on salient objects in images without supervision. In the present invention, we use the features of each level block of DINOv2 to extract multi-scale image features, and initialize the mask attention mechanism of the first layer by using the binary feature map of the last backbone network block. This processing strategy effectively enhances the model's ability to recognize key visual elements.

[0078] like Figure 2 As shown, the edge enhancement module of the present invention is intended to use the edge information of the foreground object to correct the image features extracted by the backbone network, thereby improving the accuracy of foreground segmentation. The edge enhancement module includes two key parts: an edge feature injection module and a backbone feature extraction module. Each DINOv2 backbone network block corresponds to an edge injection module and a backbone feature extraction module.

[0079] 2.1 Edge feature injection module

[0080] First, the edge features of the image are extracted through the ResNet50 convolutional neural network. In order to reduce the confusion caused by color, the image is converted to a grayscale image, and Gaussian smoothing is applied to reduce noise. Then, the gradient map is extracted using the Canny edge detector and superimposed on the original image. The ResNet50 network is divided into STEM and the rest, where STEM is used for preliminary feature extraction, including a series of convolution, batch normalization, and ReLU activation layers. The feature maps output by the remaining convolution blocks are flattened and projected to the same dimension D by 1×1 convolution, and connected to form a feature pyramid of edge features.

[0081] Next, the edge feature injection structure based on cross attention is used to fuse the image features extracted by the backbone network DINOv2 with the edge features. The formula for the edge feature injection part is:

[0082]

[0083] Among them, MSDA (Multi-Scale Deformable Attention) uses the normalized output features of the i-th layer of the backbone network As the query, the normalized edge features As key and value. iIt is a learnable parameter used to balance the backbone features and fusion features.

[0084] 2.2 Backbone feature extraction module

[0085] After completing the edge feature injection module, the backbone feature extraction module is used to further process the fused features. The formula of the backbone feature extraction module is:

[0086]

[0087] Here, MSDA is applied again, however, this time with the normalized edge features As a query, the output features of DINOv2 of the next layer As keys and values. The ConvFFN structure consists of two fully connected layers and a depth-wise separable convolutional layer. The output It will be used as the input of the edge feature injection module of the next layer.

[0088] 2.3 Multi-scale feature fusion and output

[0089] Finally, the features output by different blocks of the backbone network are upsampled to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 resolutions. The output of the backbone feature extraction module is split and restored to its original size, and then the upsampled backbone features are added to the corresponding split outputs from the backbone feature extraction module and the STEM output to obtain edge-enhanced multi-scale image features. These features will be input to the pixel decoder based on multi-scale deformation attention for dense pixel-level prediction.

[0090] In this way, the edge enhancement module can effectively fuse the global features and edge information of the image, thereby significantly improving the clarity of the boundaries and the accuracy of segmentation in the foreground segmentation task.

[0091] After obtaining the multi-scale edge enhancement features extracted by the backbone network and the edge enhancement module, these pixel-level features will be input into the pixel decoder to generate pixel-level outputs. These output features will be input into the Transformer decoder together with the binary query, which will be updated through the masked attention mechanism. The process can be expressed as:

[0092]

[0093] Among them, K l and V l is the image feature obtained by linear transformation from the l-th layer block of the pixel decoder, all of which belong to space; represents the query features extracted from the l-th layer Transformer decoder block, and X0 is initialized by the input query features of the Transformer decoder; is a binary query at level l, is the attention mask of the previous layer, defined as follows:

[0094]

[0095] in, is obtained by decoding the binary query GQ l-1 And it is obtained by binarization, and its dimension size is the same as K l Be consistent.

[0096] The present invention initializes the attention mask by principal component analysis and binarization of the feature map of the last block of the backbone network The formula is:

[0097]

[0098] Among them, F DINOv2 Represents the binarized feature map from the last backbone network block of DINOv2, which is adjusted to the same resolution as K0. This new initialization method can utilize the localization prior knowledge learned by DINOv2 on large-scale data, thereby further improving the computational efficiency of the mask attention mechanism.

[0099] The implementation of the multimodal refinement module in the present invention is as follows Figure 3 As shown in Figure 1, we first decode the binary query through a multi-layer perceptron to obtain the foreground and background masks, which are resized and superimposed on the original image. Then, we use text descriptions to represent the foreground and background respectively, such as "This is a salient object image without background" and "This is a background image after removing the salient object". It should be noted that these text descriptions can be adjusted according to specific tasks. For example, in the shadow detection task, the text can be replaced with "This is a shadow image without background" and "This is a background image after removing the shadow", thereby extending the multimodal refinement module to other foreground segmentation tasks.

[0100] In the encoding process of images and texts, CLIP's image encoder and text encoder are used to encode the fused image and text respectively. Then, the contrast loss between the mask fused image and the text features is calculated. The loss function consists of the following two parts:

[0101] 3.1. Image-to-text contrast loss

[0102]

[0103] 3.2. Contrastive loss for text-to-image

[0104]

[0105] The final contrastive loss is the average of the above two contrastive losses:

[0106]

[0107] Among them, I f ,I b ,T f ,T b ∈R N×C They represent the image features and text features of the foreground and background obtained by CLIP respectively, and τ is the temperature parameter used to control the smoothness of the softmax function.

[0108] The multimodal refinement module iteratively refines the mask edges generated by the previous module to ensure that only appropriate pixels are included in the foreground and background. This process makes the mask-fused image closer to the corresponding text in the feature space while keeping a distance from the mismatched text features. This not only makes the mask edges more accurate, but also further increases the gap between the foreground and background.

[0109] It is worth noting that the multimodal refinement module is only used to distill knowledge from CLIP and will be discarded during the inference phase. In addition, to fully exploit the multimodal capabilities of CLIP, we keep the parameters of the image and text encoders completely frozen to avoid potential performance degradation caused by fine-tuning.

[0110] The preferred specific embodiments of the present invention are described in detail above. It should be understood that a person skilled in the art can make many modifications and changes based on the concept of the present invention without creative work. Therefore, any technical solution that can be obtained by a person skilled in the art through logical analysis, reasoning or limited experiments based on the concept of the present invention on the basis of the prior art should be within the scope of protection determined by the claims.

Claims

1. A cross-task general foreground segmentation method, characterized in that: The method steps include: Construct a unified foreground segmentation model, introduce binary query, and represent the foreground and background features in the image through foreground query and background query; Extract multi-scale image features through visual backbone network; The edge features of the image are extracted through the edge enhancement module, and combined with the multi-scale deformation attention mechanism, the edge features are fused with the multi-scale image features to obtain the edge-enhanced multi-scale image features; The edge-enhanced multi-scale features are input into the Transformer decoder together with the binary query. The Transformer decoder applies the masked attention mechanism to update the binary query so that the binary query focuses on the features related to the foreground and background. Decodes a binary query and produces foreground and background segmentation masks and class predictions.

2. A cross-task general foreground segmentation method according to claim 1, characterized in that: The edge enhancement module includes multiple pairs of edge feature injection modules and backbone feature extraction modules, and each hierarchical block of the backbone network corresponds to an edge injection module and a backbone feature extraction module; The edge feature injection module extracts edge features of the image through a convolutional neural network. The convolutional neural network adopts ResNet50, including STEM and multiple convolution blocks, wherein STEM is used for preliminary feature extraction; the feature map output by the convolution block is flattened and 1×1 convolutionally projected to the same dimension, and connected to form a feature pyramid of edge features. The image features extracted by the backbone network DINOv2 are fused with the edge features using the cross-attention based edge feature injection structure; The backbone feature extraction module applies multi-scale deformation attention and further processes the features fused by the edge feature injection module through multiple convolution blocks to obtain enhanced foreground and background feature expressions.

3. A cross-task general foreground segmentation method according to claim 2, characterized in that: The edge feature injection is expressed as: Among them, MSDA represents multi-scale deformation attention, with the normalized output features of the i-th layer of the backbone network As the query, take the normalized edge features as keys and values; γ i It is a learnable parameter used to balance the backbone features and fusion features.

4. The cross-task universal foreground segmentation method according to claim 2, characterized in that: The backbone feature extraction module is expressed as: Among them, MSDA stands for multi-scale deformation attention, with normalized edge features As a query, take the output features of the next layer of the backbone network As keys and values; the ConvFFN structure consists of two fully connected layers and a depth-separable convolutional layer, the output It will be used as the input of the edge feature injection module of the next layer.

5. The cross-task universal foreground segmentation method according to claim 2, characterized in that: The edge enhancement module performs multi-scale feature fusion and output for edge features and image features, as follows: The features output by different blocks of the backbone network are upsampled to resolutions of 1 / 4, 1 / 8, 1 / 16 and 1 / 32, the output of the backbone feature extraction module is split and restored to the original size, and then the upsampled backbone features are added to the corresponding split output and STEM output from the backbone feature extraction module to obtain edge-enhanced multi-scale image features.

6. The cross-task universal foreground segmentation method according to claim 1, characterized in that: The Transformer decoder applies the masked attention mechanism to update the binary query as follows: The multi-scale image features extracted by the edge enhancement module are input into the pixel decoder based on multi-scale deformation attention. The pixel decoder combines the feature map generated by the backbone network to perform dense pixel-level prediction and generate pixel-level output features. The pixel-level output features are input into the Transformer decoder together with the binary query, and the binary query is updated through the mask attention mechanism. The process is expressed as: Among them, K l and V l is the image feature obtained by linear transformation from the l-th layer block of the pixel decoder, all of which belong to space; represents the query features extracted from the l-th layer Transformer decoder block, and X0 is initialized by the input query features of the Transformer decoder; is a binary query at level l, is the attention mask of the previous layer, defined as follows: in, By decoding the binary query GQ l-1 And perform binarization to obtain, and the dimension size is the same as K l Be consistent.

7. A cross-task universal foreground segmentation method according to claim 6, characterized in that: The method initializes the attention mask by principal component analysis and binarization of the feature map of the last block of the backbone network. Among them, F DINOv2 Represents the binarized feature map from the last backbone network block of DINOv2, which is resized to the same resolution as K1.

8. The cross-task universal foreground segmentation method according to claim 1, characterized in that: The foreground segmentation model also uses a multimodal refinement module to refine the segmentation results of the foreground and background, as follows: Decode the binary query to obtain foreground and background masks, resize the foreground and background masks and overlay them on the original image to form a mask-fused image; Use text descriptions to represent foreground and background separately; Use the image encoder and text encoder of the CLIP model to encode the mask fusion image and text description respectively; Compute the contrastive loss between mask-fused image and text description in multimodal feature space The foreground and background segmentation results are refined by minimizing the contrast loss.

9. The cross-task universal foreground segmentation method according to claim 8, characterized in that: The contrast loss between the mask fusion image and text features The details are as follows: in, is the contrast loss from image to text; is the contrast loss from text to image; I f ,I b ,T f ,T b ∈R N ×C They represent the image features and text features of the foreground and background obtained by CLIP respectively; τ is the temperature parameter used to control the smoothness of the softmax function.

10. The cross-task universal foreground segmentation method according to claim 8, characterized in that: During the training phase, the multimodal refinement module optimizes the boundaries of the foreground and background masks through multiple iterations to achieve the best effect; during the inference phase, the multimodal refinement module is removed and only the optimized foreground segmentation model is retained.

Citation Information

Patent Citations

  • Digital camera with non-uniform image resolution

    CN101156434A

  • Text refinement network

    CN114529903A

  • Scene text recognition method based on character distance perception

    CN115116066A

  • Small sample medical image segmentation method based on double attention mechanism and multi-scale fusion

    CN118334045A

  • Multisource remote sensing image semantic segmentation method based on Transform, Mama and diffusion model

    CN119152205A

Cited By

  • Contrast learning and foreground extraction-based dense scene target detection method and system, and medium

    CN120783019A

  • A dense scene target detection method, system and medium based on contrast learning and foreground extraction

    CN120783019B