Zero sample image anaphora segmentation method based on global and local mixed representation

By using a global and local mixed representation method in zero-sample image reference segmentation, combined with the improved CLIP model and segmentation mask, mixed feature extraction and spatial relationship correction are performed, and the problems of model redundancy and spatial relationship understanding in the existing methods are solved, achieving efficient reference image segmentation and comprehensive utilization of the model.

CN120032124APending Publication Date: 2025-05-23NORTHWESTERN POLYTECHNICAL UNIV

Patent Information

Application Number
CN202510092856.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing zero-sample image segmentation method has problems such as complex model, high computing resource consumption, and difficulty in understanding spatial relationships when utilizing the basic model, and it is difficult to fully explore the full understanding ability of the basic model.

Method used

Using a method based on global and local mixed representation, the improved CLIP visual encoder and text encoder combine the segmentation mask generated by segmentation all models to perform mixed global-local feature extraction, and a space guidance enhancement module is added to perform spatial relationship correction to achieve efficient referential image segmentation.

Benefits of technology

It realizes accurate reference image segmentation without training, which can be efficient and relational comprehension capabilities, unlocks the full potential of SAM and CLIP in reference segmentation tasks, and improves the adaptability of the model and the understanding of segmentation masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032124A_ABST
    Figure CN120032124A_ABST
Patent Text Reader

Abstract

The invention discloses a zero sample image anaphora segmentation method and device based on global and local mixed representation, a medium and equipment. A series of high-quality label-free segmentation masks are generated by using SAM; a modified CLIP image encoder is combined to carry out global-local feature mixing on each segmentation mask to understand context information of each segmentation mask, a space guiding enhancement module is added to carry out space relation correction, and finally anaphora segmentation is realized through similarity scoring. The method realizes the release of all potential of SAM and CLIP in the anaphora segmentation task, and can have the ability of spatial relationship understanding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image segmentation, and in particular to a zero-sample image reference segmentation method, apparatus, medium and device based on global and local mixed representation. Background Art

[0002] In the field of computer vision, image semantic segmentation is a key technology that involves assigning each pixel in an image to a category label. This process has important promoting effects in many fields such as medical imaging, image editing, and environmental monitoring. Referring Image Segmentation (RIS) is an important task in the field of computer vision. Its core goal is to accurately segment the specific objects or regions mentioned in the image based on natural language expressions. This capability is extremely important for a variety of application scenarios, including but not limited to visual search, robot perception, and human-computer interaction. In visual search, RIS can help users find specific elements in an image by describing them; in robot perception, it enables robots to better understand and operate their environment, especially when performing tasks that require precise identification and positioning; and in the field of human-computer interaction, RIS can enhance the user experience of the system, allowing the machine to respond to user instructions more naturally and intuitively. Therefore, RIS is not only a hot topic in computer vision research, but also one of the key factors in promoting the practical application of related technologies.

[0003] Currently, fully supervised learning methods have achieved remarkable results in RIS tasks, achieving accurate segmentation of specific objects by effectively integrating image and text description information. These methods rely on large-scale, exhaustively annotated datasets, in which each object mentioned in the text is accurately segmented in the corresponding image. However, the construction of such datasets is expensive and time-consuming, which limits the scalability and application scope of supervised learning methods. Therefore, some weakly supervised learning methods have also been proposed, which can learn and perform RIS tasks based on a small amount of annotated data, but still face certain challenges.

[0004] Recent studies have shown that pre-trained foundation models (FMs), such as CLIP and Stable Diffusion, can also demonstrate segmentation capabilities without specific task training. In addition, the Segment Anything Model (SAM) has demonstrated strong segmentation capabilities in multiple works, further proving the value of FMs in zero-shot segmentation and related tasks. There are also some works exploring the referential segmentation capabilities of FMs. For example, the Global-Local method first uses pre-trained models such as FreeSOLO and CLIP to achieve zero-shot RIS through simple token masking and feature addition. On its basis, Ref-Diff introduces a diffusion model to enhance target localization capabilities; TAS uses BLIP2 to generate descriptive captions to provide additional context for images, thereby enhancing segmentation results. Pseudo-RIS improves segmentation accuracy by modifying the Global-Local process and combining the CoCa image caption generator with unsupervised training. Although existing reference image segmentation methods have achieved remarkable results in specific tasks, these methods either fail to fully tap the full understanding capabilities of the underlying models (such as CLIP, SAM) through simple operations, or introduce too many models, making the system redundant and even requiring additional training, consuming a lot of computing resources. At the same time, these models also have the problem of difficulty in understanding spatial relationships. Summary of the invention

[0005] The main purpose of this application is to provide a zero-sample image referential segmentation method, device, medium and equipment based on global and local mixed representation, aiming to achieve the goal of accurate referential image segmentation with high efficiency and relational understanding capability without training using a training-free method with as few basic models as possible.

[0006] To achieve the above-mentioned objectives, the first aspect of the present application provides a zero-sample image reference segmentation method based on global and local hybrid representation, comprising: obtaining an image to be segmented; performing masking and blurring processing on the image to be segmented and a set of candidate segmentation masks respectively to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on a segmentation model; using an improved CLIP visual encoder to process the masked image and the blurred image respectively to obtain global and local hybrid encoding feature vectors of each segmentation mask; obtaining a reference text, processing the reference text based on the CLIP text encoder to obtain a text encoding feature vector, and determining a semantic alignment score based on the text encoding feature vector and the global and local hybrid encoding feature vector; using the CLIP text encoder to extract a non-subject text encoding feature vector in the reference text, and according to the non-subject text encoding feature vector, determining a semantic alignment score. The feature vector and the global and local mixed encoding feature vectors are used to determine the non-subject similarity score of each segmentation mask; the position vocabulary of the referential text is obtained, and the spatial relationship correction of the non-subject similarity score and the semantic alignment score is performed based on the position vocabulary to obtain the corrected semantic alignment score of each segmentation mask, and the spatial position matrix is ​​determined based on the position vocabulary; a spatial positioning map is obtained based on the modified CLIP and the subject text encoding feature vector, and a spatial position guidance map is determined according to the spatial positioning map and the spatial position matrix; the spatial guidance score of each segmentation mask is obtained based on the spatial position guidance map and each segmentation mask, and the spatial guidance score and the corrected semantic alignment score of each segmentation mask are weightedly added to obtain the candidate segmentation mask corresponding to the maximum score, and the candidate segmentation mask is used as the segmentation corresponding to the referential text.

[0007] Optionally, the improved CLIP visual encoder is used to process the masked image and the blurred image respectively to obtain global and local mixed coding feature vectors of each segmentation mask, including: processing the masked image and the blurred image through local branches and global branches respectively to obtain global and local mixed coding feature vectors corresponding to each segmentation mask, wherein the local branch is obtained based on the CLIP visual encoder, and the global branch is obtained after performing an attention masking operation on each layer of the CLIP-based visual encoder.

[0008] Optionally, the visual encoder has a total of K layers, and the masked image and the blurred image are processed by local branches and global branches respectively to obtain global and local mixed coding feature vectors corresponding to each segmentation mask, including: in the first L-1 layers of the visual encoder, the masked image is encoded by the local branch to obtain a first image code, and the blurred image is encoded by the global branch to obtain a second image code; in the Lth to Kth layers of the visual encoder, the global branch performs an attention masking operation on the second image code to obtain the second image code of the Lth to Kth layers, and after token masking the second image code by the Lth layer of the global branch, in the Lth layer of the local branch, it is added to the first image code to obtain a third image code, and the third image code is propagated to the L+1th layer of the local branch, and so on, and the global and local mixed coding feature vectors of each segmentation mask are output in the Kth layer of the visual encoder.

[0009] Optionally, the processing of the referent text based on the CLIP text encoder to obtain a text encoding feature vector, and determining a semantic alignment score based on the text encoding feature vector and the global and local mixed encoding feature vector, includes: using a word segmenter to obtain a main phrase of the referent text, using the CLIP text encoder to perform feature extraction on the referent text and the main phrase to obtain the text encoding feature vector; and obtaining a semantic alignment score of each segmentation mask relative to the referent text based on the text encoding feature vector and the global and local mixed encoding feature vector of each segmentation mask.

[0010] Optionally, the determining the non-subject similarity score of each segmentation mask based on the non-subject text encoding feature vector and each of the global and local mixed encoding feature vectors includes: performing softmax normalization processing on the non-subject similarity score and the semantic alignment score, respectively, to obtain a first probability that the segmentation mask is a target and a second probability that the segmentation mask is a non-subject noun; obtaining a predefined spatial relationship function; and obtaining the semantic similarity score after spatial relationship correction based on the first probability, the second probability and the spatial relationship function, to obtain the non-subject similarity score after spatial relationship correction.

[0011] Optionally, the spatial positioning map is obtained based on the modified CLIP and the subject text encoding feature vector, and the spatial position guidance map is determined according to the spatial positioning map and the spatial position matrix, including: modifying the self-attention module of the CLIP based on the grounded everything model GEM, and processing the image to be segmented in sequence based on the visual encoder and the self-attention module to generate a spatial positioning map; using the word segmenter to obtain the spatial relationship words of the reference text, and determining the spatial position relationship matrix based on the spatial relationship words; and obtaining the spatial position guidance map according to the product of the spatial position relationship matrix and the spatial positioning map.

[0012] Optionally, obtaining a spatial guidance score based on the spatial position guidance map and each of the segmentation masks includes: calculating an inner and outer mean difference between the spatial position guidance map and each of the segmentation masks to obtain the spatial guidance score.

[0013] In addition, to achieve the above-mentioned purpose, the present application also provides a zero-sample image reference segmentation device based on global and local mixed representation, including: an acquisition module, used to acquire an image to be segmented; an image processing module, used to mask and blur the image to be segmented and a set of candidate segmentation masks respectively, to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on a segmentation model; a feature extraction module, used to use an improved CLIP visual encoder to process the masked image and the blurred image respectively, to obtain global and local mixed encoding feature vectors of each segmentation mask; a first scoring module, used to acquire a reference text, process the reference text based on the CLIP text encoder, to obtain a text encoding feature vector, and determine a semantic alignment score based on the text encoding feature vector and the global and local mixed encoding feature vector; a second scoring module, used to use the CLIP text encoder to extract the non-subject text encoding feature vector in the reference text, and The non-subject similarity score of each segmentation mask is determined according to the non-subject text encoding feature vector and each of the global and local mixed encoding feature vectors; a spatial guidance module is used to obtain the position vocabulary of the referential text, perform spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary, obtain the corrected semantic alignment score of each segmentation mask, and determine the spatial position matrix based on the position vocabulary, obtain the spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determine the spatial position guidance map according to the spatial positioning map and the spatial position matrix; a segmentation determination module is used to obtain the spatial guidance score of each segmentation mask based on the spatial position guidance map and each segmentation mask, and weightedly add the spatial guidance score and the corrected semantic alignment score of each segmentation mask to obtain the candidate segmentation mask corresponding to the maximum score, and use the candidate segmentation mask as the segmentation corresponding to the referential text.

[0014] To achieve the above-mentioned purpose, the second aspect of the present application further provides a computer-readable storage medium, which includes instructions, which, when executed on a computer, enables the computer to execute the zero-sample image reference segmentation method based on global and local mixed representation provided in the first aspect.

[0015] To achieve the above-mentioned purpose, the third aspect of the present application also provides an electronic device, which includes: at least one processor, a memory and an input and output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the zero-sample image reference segmentation method based on global and local mixed representation provided by the first aspect.

[0016] The embodiments of the present application propose a method, apparatus, medium and device for zero-sample image representative segmentation based on global and local hybrid representation. The method utilizes SAM to generate a series of high-quality unlabeled segmentation masks, combines the modified CLIP image encoder to mix global-local features for each segmentation mask to understand the contextual information of each segmentation mask, and then adds a spatial guidance enhancement module to perform spatial relationship correction. Finally, similarity scoring is used to achieve representative segmentation, thereby releasing the full potential of SAM and CLIP in the representative segmentation task and having the ability to understand spatial relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 A flowchart diagram of an embodiment of a zero-sample image reference segmentation method based on global and local mixed representation provided by the present application;

[0018] Figure 2 A schematic diagram of an embodiment of a zero-sample image reference segmentation method based on global and local mixed representation provided by the present application;

[0019] Figure 3 This is a structural block diagram of a zero-sample image reference segmentation device based on global and local mixed representation in this application.

[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0021] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] This application aims to solve the problems existing in the related art to a certain extent, and proposes a zero-sample representative segmentation method based on hybrid global-local feature extraction and spatial guidance enhancement. This application aims to unleash the full potential of SAM and CLIP in the representative segmentation task, and to have the ability to understand spatial relationships. This application generates a series of high-quality unlabeled segmentation masks by using SAM, combines the modified CLIP image encoder to mix global-local features for each segmentation mask to understand the contextual information of each segmentation mask, and then adds a spatial guidance enhancement module for spatial relationship correction, and finally achieves representative segmentation through similarity scoring.

[0023] The hybrid global-local feature extraction module (Hybrid Global-Local Feature Extraction) designed in this application calculates an image feature including global context information and local target attribute information for each segmentation mask through a set of unlabeled segmentation masks of a given image, which is used to calculate the similarity with the text feature for reference segmentation. This process only modifies the CLIP pipeline and does not require any training or fine-tuning. In addition, unlike the simple addition of global-local features in previous methods, this module mixes global-local features in the forward transmission process of the CLIP image encoder, which can not only obtain the global-local information of a segmentation mask, but also alleviate the problem of insufficient alignment between global and local features, resulting in information loss after addition. Therefore, this application has strong adaptability and segmentation mask understanding capabilities, and is also characterized by a streamlined and efficient model.

[0024] The spatial guidance enhancement module designed in this application identifies the spatial relationship words existing in the reference text, and performs spatial relationship correction on the similarity score obtained by the hybrid global-local feature extraction module, so that the model has the ability to understand spatial relationships. This process only uses a simple word segmenter for semantic analysis, and the modified CLIP, which also does not require any training and fine-tuning. This module uses three parts, spatial relationship guidance, spatial positioning guidance, and spatial position guidance, to correct the similarity score obtained by the hybrid global-local feature extraction module, so that the model can fully analyze the spatial relationship between different targets in the image and complete more accurate positioning and segmentation.

[0025] Reference Figure 1 The first embodiment of the present application provides a zero-sample image reference segmentation method based on global and local mixed representation. The zero-sample image reference segmentation method based on global and local mixed representation may include:

[0026] S10, obtaining an image to be segmented;

[0027] S20, performing masking and blurring processing on the image to be segmented and a set of candidate segmentation masks respectively to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on the segmentation model;

[0028] Among them, the goal of the zero-shot referential semantic segmentation task is to segment a region in an image based on a referential expression (usually a natural language description) without explicit training on specific objects or scenes. Given an input image I∈R H×W×3 , where H and W are the height and width of the image, respectively, and refers to the text T, describing the target object or region. The goal of this application is to predict the segmentation mask m corresponding to the region in I described by T. * ∈{0,1} H×WTo this end, this application first uses the Segment Anything Model (SAM) to generate a set of segmentation mask proposals M = {m 1 ,m 2 ,…,m n}, each m i Each is a binary segmentation mask, highlighting a candidate segmentation region in I.

[0029] S30, using the improved CLIP visual encoder to process the masked image and the blurred image respectively, and obtain the global and local mixed encoding feature vectors of each segmentation mask;

[0030] Specifically, for each segmentation mask m i ∈M, this application uses the CLIP visual encoder to extract the proposed hybrid global-local feature x i ∈R d .

[0031] S40, obtain the reference text, process the reference text t based on the CLIP text encoder φtext(·), and obtain the text encoding feature vector f t , and determining a semantic alignment score based on the text encoding feature vector and the global and local mixed encoding feature vector;

[0032] Then, based on the text feature f t With each feature x i Get each segmentation mask m i Semantic alignment score

[0033] S50, using the CLIP text encoder to extract the non-subject text encoding feature vector in the reference text, and determining the non-subject similarity score of each segmentation mask according to the non-subject text encoding feature vector and each global and local mixed encoding feature vector;

[0034] S60, obtaining a position vocabulary referring to the text, performing spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary, obtaining the corrected semantic alignment score of each segmentation mask, and determining a spatial position matrix based on the position vocabulary, obtaining a spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determining a spatial position guide map based on the spatial positioning map and the spatial position matrix;

[0035] This application introduces a spatial guidance enhancement module that prioritizes semantic and spatially aligned segmentation masks and integrates multiple spatial cues, especially spatial relationship guidance, spatial positioning guidance, and spatial location guidance, to enhance the segmentation mask scoring process.

[0036] S70. Based on the spatial position guidance map and each segmentation mask, a spatial guidance score of each segmentation mask is obtained, and the spatial guidance score of each segmentation mask and the corrected semantic alignment score are weightedly added to obtain a candidate segmentation mask corresponding to the maximum score, and the candidate segmentation mask is used as the segmentation corresponding to the reference text.

[0037] After spatial relationship enhancement, this application takes the highest corrected semantic alignment score The corresponding segmentation mask m i as the final segmentation result.

[0038] In summary, in this embodiment, the present application utilizes SAM to generate a series of high-quality unlabeled segmentation masks, combines the modified CLIP image encoder to mix the global-local features of each segmentation mask to understand the contextual information of each segmentation mask, and then adds a spatial guidance enhancement module to perform spatial relationship correction. Finally, similarity scoring is used to achieve reference segmentation, thereby realizing the full potential of SAM and CLIP in the reference segmentation task and being able to have the ability to understand spatial relationships.

[0039] In an embodiment of the present application, step S30 may include the following execution process:

[0040] S301. Process the masked image and the blurred image through the local branch and the global branch respectively to obtain the global and local mixed coding feature vectors corresponding to each segmentation mask, wherein the local branch is obtained based on the CLIP original visual encoder, the first L-1 layers of the global branch are obtained based on the CLIP original visual encoder, and the Lth layer and subsequent layers are obtained based on the attention masking operation on the output of the transformer layer.

[0041] Specifically, step S301 may include the following execution process:

[0042] S3011, in the first L-1 layers of the visual encoder, using the local branch to encode the masked image to obtain a first image code, and using the global branch to encode the blurred image to obtain a second image code;

[0043] S3012. In the Lth to Kth layers of the visual encoder, the global branch performs an attention masking operation on the second image code to obtain the second image code of the Lth to Kth layers. After token masking is performed on the second image code using the Lth layer of the global branch, the third image code is obtained by adding the first image code to the Lth layer of the local branch. The third image code is propagated to the L+1th layer of the local branch, and so on. The global and local mixed coding feature vectors of each segmentation mask are output at the Kth layer of the visual encoder.

[0044] By executing step S3011 to step S3012, global and local mixed coding feature vectors are extracted. Figure 2 Steps S3011 to S3012 can be performed by a hybrid global-local feature extraction module. The process of the hybrid global-local feature extraction module is as follows: Figure 2 As shown in the left part, the image to be segmented is first compared with a series of segmentation mask proposals M = {m 1 ,m 2 ,…,m n} performs masking and blurring operations, and inputs the two sets of images into the CLIP image encoder as local branches and global branches respectively. The local branch is normally encoded in the first L-1 layers of the CLIP image encoder ViT, and the global branch performs attention masking operations at each layer to focus on the segmentation mask target. After the Lth layer, the encoding of the global branch is merged into the local branch, and finally the hybrid global-local image encoding X corresponding to each segmentation mask output by the local branch is obtained. hybrid ={x 1 ,x 2 ,…,x n}.

[0045] Exemplarily, the hybrid global-local feature extraction module may perform steps S3011 to S3012 including the following execution process:

[0046] Step 1: Set the reference image I∈R H × W × 3 Input SAM, use SAM to obtain a series of candidate segmentation masks M = {m 1 ,m 2 ,…,m n}, where m i ∈{0,1} H×W ;

[0047] Step 2: The image to be segmented I∈R H × W × 3 Respectively with each candidate segmentation mask m i Perform masking and blurring operations to obtain I mask ∈R n×H×W×3 and I blur ∈R n×H×W×3 , use CLIP to analyze each group Extract and mix features to get X hybrid ={x 1 ,x 2 ,…,x n}, where x i ∈R d The specific process is as follows:

[0048] Step 2-1: Compare the image to be segmented I with each candidate segmentation mask m i Perform masking operation get The image to be segmented I is respectively compared with each candidate segmentation mask m i Perform blur operation Gaussian blur is performed on the segmentation mask area to obtain

[0049] Step 2-2: First, the two sets of images I mask and I blur Preprocessing to obtain local and global features and Then the two features are input into the image encoder φ of CLIP respectively. visual As local branches and global branches, the image encoder here is composed of 12 layers of transformer encoders with the same structure and different parameters in series. The lth layer encoder is In the first k-1 layers, the local branches are encoded normally, i.e. The global branch is also encoded normally, i.e.

[0050] Step 2-2: From the kth layer to the 12th layer of the image encoder, the global branch performs an attention masking operation when performing the self-attention operation of the encoding, and masks the cls token of the self-attention according to the segmentation mask, that is, The features of the global branch Perform token masking and local branching Perform additive blending as a new blending feature Perform forward pass on local branch.

[0051]

[0052] Where β is a hyperparameter that adjusts the ratio of the two features. is the segmentation mask m i The token segmentation mask obtained by deformation. Finally, the cls token output of the last layer of the local branch is obtained Project it Get the final mixed global-local feature X hybrid ={x 1 ,x 2 ,…,x n}, x i ∈R d In this implementation, k is set to 10 and β is set to 2.

[0053] The hybrid global-local feature extraction module (Hybrid Global-Local Feature Extraction) designed in this application calculates an image feature including global context information and local target attribute information for each segmentation mask through a set of unlabeled segmentation masks of a given image, which is used to calculate the similarity with the text feature for reference segmentation. This process only modifies the CLIP pipeline and does not require any training or fine-tuning. In addition, unlike the simple addition of global-local features in previous methods, this module mixes global-local features in the forward transmission process of the CLIP image encoder, which can not only obtain the global-local information of a segmentation mask, but also alleviate the problem of insufficient alignment between global and local features, resulting in information loss after addition. Therefore, this application has strong adaptability and segmentation mask understanding capabilities, and is also characterized by a streamlined and efficient model.

[0054] In an embodiment of the present application, step S40 may include the following execution process:

[0055] S401, using a word segmenter to obtain a main phrase of the reference text, and using a CLIP text encoder to extract features of the reference text and the main phrase to obtain a text encoding feature vector;

[0056] S402: Obtain a semantic alignment score of each segmentation mask relative to the referenced text according to the text encoding feature vector and the global and local mixed encoding feature vectors of each segmentation mask.

[0057] Specifically, after obtaining the mixed global-local features, it is necessary to obtain each mask m i Semantic alignment score To obtain the alignment between semantics and images.

[0058] Exemplarily, the above process may include the following execution process:

[0059] First, use the tokenizer to convert the main phrase t referring to the text t 0 Then, we use the text encoder φtext(·) of CLIP to decode the referential text f and its main phrase t 0 Encode and obtain text features

[0060] Calculate text feature f t With each feature x i The cosine similarity between:

[0061]

[0062] in That is, for each mask m iThe semantic alignment score of the hybrid global-local feature extraction module is

[0063] In an embodiment of the present application, step S50 may include the following execution process:

[0064] S501, performing softmax normalization processing on the non-subject similarity score and the semantic alignment score, respectively, to obtain a first probability that the segmentation mask is a target and a second probability that the segmentation mask is a non-subject noun;

[0065] S502, obtaining a predefined spatial relationship function;

[0066] S503: Based on the first probability, the second probability and the spatial relationship function, obtain a semantic similarity score after spatial relationship correction, and obtain a non-subject similarity score after spatial relationship correction.

[0067] It can be understood that in order to enable the model to have the ability to distinguish spatial relationships (up and down, left and right, big and small, etc.), the present application uses a spatial guidance enhancement module to correct the score obtained by the hybrid global-local feature extraction module using three parts: spatial relationship guidance, spatial positioning guidance, and spatial position guidance. This process can be performed by the spatial guidance enhancement module.

[0068] Continue to refer Figure 2 ,The process of the space guidance enhancement module is as follows Figure 2 As shown in the right part, firstly, the cosine similarity of the output of the hybrid global-local feature extraction module is obtained s Conduct spatial relationship guidance and generate new S with spatial relationship information s .

[0069] Exemplarily, the present application firstly performs a semantic score S of the hybrid global-local feature extraction module. s Perform spatial relationship correction. Here, a word segmenter is used to obtain the non-subject noun t that refers to the text -1 , and then use the text encoder φ of CLIP text (·) Get other noun text features f′ t , and other noun similarity scores The two groups of similarity scores are normalized by softmax respectively, and P(m i ) is the segmentation mask m i is the probability of the target, Q(m i ) is the segmentation mask m i For the probability of other nouns, define the spatial relationship function R(m i ,m j ) to quantize the two masks m i ,m jThe spatial relationship between:

[0070]

[0071] Then the similarity score after spatial relationship correction is:

[0072]

[0073] If the referent text only has one noun, the subject, then the similarity score after spatial relationship correction is:

[0074]

[0075] In an embodiment of the present application, step S60 may include the following execution process:

[0076] S601, modifying the self-attention module of CLIP based on the grounded everything model GEM, and processing the image to be segmented in sequence based on the visual encoder and the self-attention module to generate a spatial positioning map;

[0077] S602, using a word segmenter to obtain spatial relationship words referring to the text, and determining a spatial position relationship matrix based on the spatial relationship words;

[0078] S603: Obtain a spatial position guidance map according to the product of the spatial position relationship matrix and the spatial positioning map.

[0079] The spatial guidance enhancement module performs steps S601 to S603, inputs the image into the modified CLIP image encoder for self-attention operation to obtain the spatial positioning map G co , and then the spatial positioning guidance graph G cp And the spatial position relationship matrix G pos Multiply them to get the spatial position guidance map G, and calculate the internal and external mean difference between the spatial position guidance map G and the segmentation mask proposal M to get the spatial guidance score Add the weighted scores of the two modules to get S, and take the highest score The corresponding segmentation mask m i as the final segmentation result.

[0080] Specifically, to enhance spatial positioning capabilities, the algorithm of the Grounding Everything Model is used to modify the self-attention operation of CLIP, and the main phrase t 0 , generate the spatial positioning graph G co ∈[0,1] H×W In order to make the spatial positioning guidance also have the ability to distinguish spatial relationships, this application introduces the spatial position relationship matrix G pos :

[0081]

[0082] Among them, pos is a relational word of one of up, down, left, and right. is a matrix that changes linearly from 0 to 1 according to the relation word pos, then the final spatial position guidance graph G is calculated as:

[0083] G=G co ⊙G pos

[0084] The spatial position guidance map G is combined with each segmentation mask proposal m i Calculate the difference between the inside and outside means to get the spatial bootstrap score:

[0085]

[0086] Where Sum(·) is the matrix element summation function, λ is a hyperparameter that controls the target size, which is set to 9 by default. When the relational word is related to “big” or “small”, λ is set to 3 and 14 respectively.

[0087] In an embodiment of the present application, step S70 may include the following execution process:

[0088] The spatial guidance score is obtained by calculating the difference between the spatial position guidance map and the inside and outside mean of each segmentation mask.

[0089] Finally, this application scores the semantic score S s and spatial guidance score S g First perform softmax normalization and then weighted addition:

[0090]

[0091] Among them, α is a hyperparameter that controls the ratio of the two, and is set to 0.6 in implementation. The corresponding segmentation mask m * This is the segmentation corresponding to the desired reference text.

[0092] Based on the above embodiments, the present application further provides a zero-sample image reference segmentation device based on global and local mixed representation, including an acquisition module 101, an image processing module 102, a feature extraction module 103, a first scoring module 104, a second scoring module 105, a space guidance module 106, and a segmentation determination module 107:

[0093] The acquisition module 101 is used to acquire the image to be segmented;

[0094] The image processing module 102 is used to perform masking and blurring processing on the image to be segmented and a set of candidate segmentation masks, respectively, to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on the segmentation model;

[0095] The feature extraction module 103 is used to use the improved CLIP visual encoder to process the masked image and the blurred image respectively, and obtain the global and local mixed encoding feature vectors of each segmentation mask;

[0096] The first scoring module 104 is used to obtain the reference text, process the reference text based on the CLIP text encoder, obtain a text encoding feature vector, and determine a semantic alignment score based on the text encoding feature vector and the global and local mixed encoding feature vectors;

[0097] The second scoring module 105 is used to extract the non-subject text encoding feature vector in the reference text by using the CLIP text encoder, and determine the non-subject similarity score of each segmentation mask according to the non-subject text encoding feature vector and each global and local mixed encoding feature vector;

[0098] The spatial guidance module 106 is used to obtain a position vocabulary referring to the text, perform spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary, obtain the corrected semantic alignment score of each segmentation mask, determine the spatial position matrix based on the position vocabulary, obtain the spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determine the spatial position guidance map according to the spatial positioning map and the spatial position matrix;

[0099] The segmentation determination module 107 is used to obtain the spatial guidance score of each segmentation mask based on the spatial position guidance map and each segmentation mask, and weightedly add the spatial guidance score of each segmentation mask and the corrected semantic alignment score to obtain the candidate segmentation mask corresponding to the maximum score, and use the candidate segmentation mask as the segmentation corresponding to the reference text.

[0100] To achieve the above objectives, the present application also provides a computer-readable storage medium, which includes instructions, which, when executed on a computer, enables the computer to execute the zero-sample image reference segmentation method based on global and local mixed representation provided by any of the aforementioned embodiments.

[0101] To achieve the above-mentioned purpose, the present application also provides an electronic device, which includes: at least one processor, a memory and an input-output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the zero-sample image reference segmentation method based on global and local mixed representation provided in the above-mentioned embodiment.

[0102] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A zero-shot image referential segmentation method based on global and local hybrid representation, characterized in that: include: Obtain the image to be segmented; The image to be segmented and a set of candidate segmentation masks are respectively subjected to masking processing and blurring processing to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on a segmentation model; Using the improved CLIP visual encoder to process the masked image and the blurred image respectively, to obtain global and local mixed encoding feature vectors of each segmentation mask; Acquire a referent text, process the referent text based on the CLIP text encoder to obtain a text encoding feature vector, and determine a semantic alignment score based on the text encoding feature vector and the global and local hybrid encoding feature vector; Extracting non-subject text encoding feature vectors and subject text encoding feature vectors in the reference text using the CLIP text encoder, and determining non-subject similarity scores of each segmentation mask according to the non-subject text encoding feature vectors and each of the global and local mixed encoding feature vectors; Obtaining a position vocabulary of the reference text, performing spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary, obtaining a corrected semantic alignment score for each segmentation mask, and determining a spatial position matrix based on the position vocabulary, obtaining a spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determining a spatial position guide map according to the spatial positioning map and the spatial position matrix; Based on the spatial position guidance map and each of the segmentation masks, a spatial guidance score of each of the segmentation masks is obtained, and the spatial guidance score of each of the segmentation masks and the corrected semantic alignment score are weightedly added to obtain the candidate segmentation mask corresponding to the maximum score, and the candidate segmentation mask is used as the segmentation corresponding to the reference text.

2. The zero-sample image referential segmentation method based on global and local hybrid representation as claimed in claim 1, characterized in that: The method of using the improved CLIP visual encoder to process the masked image and the blurred image respectively to obtain global and local mixed coding feature vectors of each segmentation mask includes: The masked image and the blurred image are processed by local branches and global branches respectively to obtain global and local mixed coding feature vectors corresponding to each segmentation mask, wherein the local branch is obtained based on the CLIP original visual encoder, the first L-1 layers of the global branch are obtained based on the CLIP original visual encoder, and the Lth layer and subsequent layers are obtained based on the attention masking operation on the output of the transformer layer.

3. The zero-sample image referential segmentation method based on global and local hybrid representation as claimed in claim 2, characterized in that: The visual encoder has a total of K layers, and the masked image and the blurred image are processed by local branches and global branches respectively to obtain global and local mixed coding feature vectors corresponding to each segmentation mask, including: In the first L-1 layers of the visual encoder, the masked image is encoded using the local branch to obtain a first image encoding, and the blurred image is encoded using the global branch to obtain a second image encoding; In the Lth to Kth layers of the visual encoder, the global branch performs an attention masking operation on the second image code to obtain the second image code of the Lth to Kth layers, and after token masking the second image code using the Lth layer of the global branch, the third image code is obtained by adding it to the first image code in the Lth layer of the local branch, and the third image code is propagated to the L+1th layer of the local branch, and so on, and the global and local mixed coding feature vectors of each segmentation mask are output in the Kth layer of the visual encoder.

4. The zero-sample image referential segmentation method based on global and local hybrid representation as claimed in claim 1, characterized in that: The step of processing the reference text based on the CLIP text encoder to obtain a text encoding feature vector, and determining a semantic alignment score based on the text encoding feature vector and the global and local hybrid encoding feature vectors, comprises: Using a word segmenter to obtain a main phrase of the reference text, and using the CLIP text encoder to extract features of the reference text and the main phrase to obtain the text encoding feature vector; According to the text encoding feature vector and the global and local mixed encoding feature vectors of each segmentation mask, a semantic alignment score of each segmentation mask relative to the reference text is obtained.

5. The zero-sample image referential segmentation method based on global and local hybrid representation as claimed in claim 1, characterized in that: The performing spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary includes: Performing softmax normalization processing on the non-subject similarity score and the semantic alignment score respectively, and obtaining a first probability that the segmentation mask is a target and a second probability that the segmentation mask is a non-subject noun; Get predefined spatial relationship functions; Based on the first probability, the second probability and the spatial relationship function, the semantic similarity score after spatial relationship correction is obtained, and the non-subject similarity score after spatial relationship correction is obtained.

6. The zero-sample image referential segmentation method based on global and local hybrid representation according to claim 1, characterized in that: The step of obtaining a spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determining a spatial position guide map according to the spatial positioning map and the spatial position matrix includes: Modifying the self-attention module of the CLIP based on the grounded everything model GEM, and sequentially processing the image to be segmented based on the visual encoder and the self-attention module to generate a spatial positioning map; Using the word segmenter to obtain the spatial relationship words of the reference text, and determining the spatial position relationship matrix based on the spatial relationship words; The spatial position guidance map is obtained according to the product of the spatial position relationship matrix and the spatial positioning map.

7. The zero-sample image referential segmentation method based on global and local hybrid representation according to claim 1, characterized in that: The obtaining of a spatial guidance score based on the spatial position guidance map and each of the segmentation masks comprises: The spatial position guidance map and the inside and outside mean difference of each segmentation mask are calculated to obtain the spatial guidance score.

8. A zero-sample image referential segmentation device based on global and local mixed representation, characterized in that: include: An acquisition module, used for acquiring an image to be segmented; An image processing module, used for performing masking and blurring processing on the image to be segmented and a set of candidate segmentation masks respectively, to obtain a masked image and a blurred image, wherein the candidate segmentation masks are generated based on a segmentation model; A feature extraction module, used for processing the masked image and the blurred image respectively using the improved CLIP visual encoder to obtain global and local mixed encoding feature vectors of each segmentation mask; A first scoring module is used to obtain a referent text, process the referent text based on the CLIP text encoder to obtain a text encoding feature vector, and determine a semantic alignment score based on the text encoding feature vector and the global and local mixed encoding feature vector; A second scoring module is used to extract a non-subject text encoding feature vector in the reference text using the CLIP text encoder, and determine a non-subject similarity score of each of the segmentation masks according to the non-subject text encoding feature vector and each of the global and local mixed encoding feature vectors; A spatial guidance module is used to obtain a position vocabulary of the reference text, perform spatial relationship correction on the non-subject similarity score and the semantic alignment score based on the position vocabulary, obtain the corrected semantic alignment score of each segmentation mask, determine a spatial position matrix based on the position vocabulary, obtain a spatial positioning map based on the modified CLIP and the subject text encoding feature vector, and determine a spatial position guidance map according to the spatial positioning map and the spatial position matrix; A segmentation determination module is used to obtain the spatial guidance score of each segmentation mask based on the spatial position guidance map and each segmentation mask, and to weightedly add the spatial guidance score of each segmentation mask and the corrected semantic alignment score to obtain the candidate segmentation mask corresponding to the maximum score, and use the candidate segmentation mask as the segmentation corresponding to the reference text.

9. A computer-readable storage medium, characterized in that: The invention comprises instructions, which, when executed on a computer, enable the computer to execute the zero-sample image reference segmentation method based on global and local mixed representation as claimed in any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device comprises: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the zero-sample image reference segmentation method based on global and local mixed representation according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and system for training zero sample classification model, electronic equipment and storage medium

    CN113298096A

  • Text segmentation recognition method, system and equipment in artificial intelligence field and medium

    CN116229584A

  • Underwater multi-modal target segmentation method based on mask complementary cross-layer fusion

    CN117975000A

  • CLIP-based single-stage zero-sample semantic segmentation method and apparatus

    CN119693649A

Cited By

  • Image anaphora segmentation method based on multilevel feature fusion and dual-channel information enhancement

    CN120997225A

  • Zero sample indication image segmentation method based on recursive semantic optimization

    CN121505278A

  • Zero-shot image segmentation method based on recursive semantic optimization

    CN121505278B

  • Zero sample anaphora image segmentation method based on text perception adapter

    CN121725007A