Weak supervision directivity segmentation method based on semantic and detail cooperation

By introducing semantic perception module, detail cognitive module and collaborative learning strategy into the weakly supervised directive segmentation method, combined with the negative text generation method, the problem of low segmentation accuracy when dealing with complex natural language and visually similar backgrounds is solved, and higher image segmentation accuracy and alignment comprehensiveness are achieved.

CN119992096AActive Publication Date: 2025-05-13DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510154787.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-13
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

Existing weakly supervised directive segmentation methods are difficult to utilize useful knowledge in image text pairs when dealing with rich vocabulary and complex grammar in natural language, resulting in low segmentation accuracy, especially when objects look similar to background visual appearance or vague reference expressions.

Method used

Using a method based on semantics and detail collaboration, we integrate image features and text features through semantic perception modules and detail cognitive modules, and combine collaborative learning strategies and negative text generation methods to enhance the comprehensiveness and fine-grainedness of image text alignment.

Benefits of technology

Improved image segmentation accuracy, especially in dealing with complex natural language descriptions and visually similar backgrounds, significantly outperforms the most advanced weak supervision methods at present.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992096A_ABST
    Figure CN119992096A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and discloses a weak supervision directivity segmentation method based on semantic and detail collaboration. The method comprises the following steps of: firstly, respectively extracting features of an input image and a text by utilizing two encoders based on Transform, and then inputting the two features into a semantic perception module and a detail cognition module; the two modules respectively pay attention to high-level semantics and low-level details and are used for a collaborative learning strategy, and the strategy integrates cross-modal internal loss, matching invariance loss and regional contrast loss to achieve sufficient visual language alignment. The semantic perception module and the detail cognition module respectively generate an activation graph, and then the activation graphs are integrated by a cooperation module to output accurate segmentation masks. According to the method, only the image text is used for training the supervision model, the method does not depend on dense pixel-level labels, and good image segmentation performance indicated by natural languages is achieved while the manual labeling cost is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a weakly supervised directed segmentation method based on semantics and detail collaboration. Background Art

[0002] The task of directional segmentation is to segment the pixel regions related to an image according to the input natural language description given an image and a language description of the target object. This task is different from the semantic segmentation of images, which combines image segmentation with natural language. This is a promising task due to the increasing demand for visual language understanding in practical applications, such as text-based image editing, human-computer interaction, and visual navigation. Compared with the unimodal visual segmentation task, directional segmentation is more difficult because it has to deal with richer vocabulary and complex grammar in natural language, while understanding the relationship between text and corresponding instances. The difficulty of this task lies in locating the relevant target objects and segmenting them using the given language conditions. Recently, fully supervised directional segmentation techniques have made significant progress, but the high acquisition cost of pixel-level annotations for each target object limits the applicability and scalability of this technique. Therefore, weakly supervised directional segmentation has gradually attracted the attention of researchers, which uses the useful knowledge hidden in image-text pairs to effectively identify the alignment between language and referring image regions.

[0003] Existing weakly supervised directed segmentation methods can be divided into two categories: attention-based methods, such as "Dongwon Kim, Namyup Kim, Cuiling Lan, and Suha Kwak. Shatter and Gather: Learning Referring Image Segmentation with Text Supervision [C]. In International Conference on Computer Vision, IEEE, 2023", and contrastive learning-based methods, such as "Fang Liu, Yuhao Liu, Yuqiu Kong, et al. Referring Image Segmentation Using Text Supervision [C]. In International Conference on Computer Vision, IEEE, 2023". The former captures the long-range dependencies between image-text modalities and generates cross-attention maps rich in high-level semantic information. However, the attention maps often only focus on a small number of objects because they do not explicitly exploit the underlying detail information of the scene. To alleviate this problem, researchers use the attention maps to select candidate masks generated by unsupervised object discovery techniques as the final result. However, the performance of such methods is significantly affected by the quality of candidate masks. Contrastive learning-based methods mainly constrain the alignment between image features and text features by calculating the affinity of cross-modal features and produce activation maps. Since such methods implicitly enforce the consistency between image and text features, they do not introduce high-level semantics into cross-modal embeddings, causing activation maps to focus more on low-level details and incorrectly respond to irrelevant areas. Such methods perform poorly when the referent exhibits a similar visual appearance to the background or when the referential expression is ambiguous.

[0004] The method based on semantics and details collaboration explicitly utilizes and combines high-level semantics and low-level details in the cross-modal feature learning process, which can well solve the above problems. Therefore, the present invention proposes a weakly supervised directional segmentation method based on semantics and details collaboration. Summary of the invention

[0005] The technical problem to be solved by the present invention is to make up for the shortcomings of the current image segmentation method of natural language indication under weak supervision conditions, and propose a weakly supervised directional segmentation method based on the collaboration of semantics and details. This method involves deep learning and computer vision content, and is based on a semantic perception module, a detail recognition module and a collaborative learning strategy to achieve weakly supervised natural language indication image segmentation, so as to achieve the purpose of improving image segmentation accuracy.

[0006] The technical solution of the present invention is as follows: a weakly supervised directional segmentation method based on semantics and detail collaboration, the steps are as follows:

[0007] Step 1: Build an image encoder to extract image features One of the images is split into N v non-overlapping patches, D v Represent the channel dimension of image features; build a text encoder; extract text features The input is a string containing N t words, the sentence is tokenized into a token sequence, and a [CLS] token is added to the beginning of the sequence; t Channel dimension representing text features;

[0008] Step 2: Build a semantic perception module;

[0009] The semantic perception module captures the long-term dependency between image features and text features through the attention mechanism, and fuses the two features together to obtain the activation map M. sem ;

[0010] Step 3: Construct detail recognition module;

[0011] The detail recognition module integrates image features and text features by exploring low-level details of cross-modal features to obtain the activation map M det ; by M det Weight the image features to obtain the cross-modal feature F det ;

[0012] Step 4: Build a collaborative learning strategy

[0013] A collaborative learning strategy is adopted to understand the complex correlation between images and texts. A collaborative module uses activation maps that highlight high-level semantics and low-level details respectively to generate segmentation masks.

[0014] Subsequently, a collaborative learning strategy leverages cross-modal information from three perspectives to ensure comprehensive image-text alignment: cross-modal internal loss, matching invariance loss, and region contrast loss. In addition, a total variation regularization loss is introduced to improve segmentation performance. A unified loss function simultaneously constrains the semantic perception module and the detail recognition module to ensure cross-modal features are aligned at two different levels.

[0015] Step 5: After the training of steps 1 to 4 is completed, an image and a short sentence in natural language are input, the image is sent to the image encoder to extract image features, and the short sentence in natural language is sent to the text encoder to extract text features; the image features and text features are sent to the semantic perception module and the detail recognition module to generate cross-modal features F focusing on high-level semantics.sem and cross-modal features F that focus on low-level details det , and generate two activation maps at the same time; the two activation maps are sent to the collaboration module to output the segmentation results.

[0016] Furthermore, a new negative text generation method is added to ensure accurate understanding of referential expressions;

[0017] The new negative text generation method refers to introducing more difficult to recognize negative text by replacing keywords in the paired text; first, the text is parsed and a part-of-speech tag is assigned to each word using a text parser; then two strategies are adopted to generate challenging negative text; the first is to randomly mask an adjective or noun and fill it with RoBERTa; if there is no adjective in the original sentence, a color word or a position word is randomly added to the sentence; the second is to swap the positions of the two nouns in the text; if there is only one noun, a category name is randomly selected from the COCO dataset and added to the text, and the positions of the two nouns are swapped; by modifying the keywords of the sentence to make negative samples, the segmentation model perceives the subtle differences in semantic information in natural language when learning the matching between images and text, thereby effectively improving its ability to distinguish between real targets and image backgrounds.

[0018] The semantic perception module adopts the gradient weighted class activation mapping technology, which uses the gradient of image-text similarity to determine the area in the image that has the greatest impact on the final classification result, and then identifies the target object from an overall perspective;

[0019] The semantic perception module consists of 6 stacked network layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a feedforward network, with image features and text features as input; for the cross-attention layer, the text features are used as queries and the image features are used as keys and values; the semantic perception module outputs cross-modal features represented as Where D h Represents the channel dimension of the cross-modal feature; the cross-modal feature is constrained by the matching invariance loss to ensure the alignment between the image-text pairs; the gradient weighted class activation mapping technique is used to sem Generate activation map This activation map indicates the image regions that are most semantically relevant to the reference expression.

[0020]

[0021] in, is the affinity matrix of the cross-attention layer of the semantic perception module, Represents the matrix A cross The nth column element taken out from Reflects the affinity matrix to the similarity score S sem contribution; the similarity score is used to evaluate the matching degree between image features and text features, and ⊙ represents an element-by-element multiplication operation.

[0022] The detail recognition module first uses a convolutional neural network to extract local features from the image, and uses upsampling technology to improve the resolution of image features to obtain a feature with higher resolution. Where N up Represents the number of non-overlapping patches; the detail recognition module generates an activation map by calculating the affinity between normalized image features and text features

[0023]

[0024] in and It maps the input image features and text features to dimension D h Linear projection layer of is the [CLS] embedding of text features, norm(·) represents L2 normalization; using M det Weight the image features to obtain cross-modal features

[0025]

[0026] in It maps the input image features to dimension D h Linear projection layer.

[0027] The collaborative modules jointly integrate M sem The high-level semantic information and M det low-level details; first, considering the M generated by the semantic perception module sem Focusing on the limited object area, a block similarity propagation method is used to expand the effective activation area. sem The activation score of each block in is extended to its neighboring blocks according to their similarity;

[0028] First, calculate the similarity matrix between image patches

[0029]

[0030] in and is a linear projection layer; then, the corrected activation map is calculated

[0031]

[0032] Where N(n) is the set of eight blocks spatially adjacent to the nth image block on the 2D plane, σ(·) is the softmax function; rec With M det The final activation map obtained by merging is expressed as:

[0033] M final =Up(M rec )⊙M det

[0034] Where Up(·) means to increase M rec Upsample to M det Same size.

[0035] The cross-modal internal loss enables the method to directly use cross-modal feature information to distinguish whether image-text pairs match; this loss is used to learn the [CLS] embedding of cross-modal features and obtain the true value label by judging the match of image-text pairs:

[0036] L cmi =CE(MLP(F sem [CLS]), y label )+CE(MLP(F det ), y label )

[0037] Among them, MLP is a multi-layer perceptron, which is used to reduce the number of channels of multimodal features to 2; when the cross-modal feature F sem and F det is obtained from a correctly aligned image-text pair, y label The value of is 1, otherwise it is 0; CE(·,·) represents the cross entropy loss;

[0038] The matching invariance loss is used to bring the cross-modal features generated by the real image-text pairs closer to the text features, while pushing the cross-modal features from the unmatched image-text pairs away from the text features; if the i-th image and the j-th text match, the cross-modal features and will be combined with the text feature T j closely aligned; on the contrary, the distance between them will be enlarged; first calculate With T j The similarity between

[0039]

[0040] Among them, m and n represent the m-th word and the n-th word respectively; the matching invariance loss is represented by the InfoNCE loss:

[0041] L miv =InfoNCE(Ssem )+InfoNCE(S det )

[0042]

[0043] where τ denotes a learnable temperature parameter and B denotes the size of the training batch; the alignment between images and text is discriminated at various levels by applying a matching invariance loss to cross-modal features that contain both high-level semantics and low-level details;

[0044] To further distinguish highly similar negative image-text pairs and increase the influence of negative image-text pairs showing high matching similarity during training, similarity scores are used to select image-text pairs; then, negative texts for each image are sampled from the batch based on the similarity scores, and unmatched images of each original text are sampled based on the similarity scores; the matching invariance loss is further modified as:

[0045] L′ miv =L miv +InfoNCE(S sem′ )+InfoNCE(S det′ )

[0046] Among them, S sem′ Is not from S sem The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000, S det′ Is not from S det The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000;

[0047] The region contrast loss is used to align image features and text features at a fine-grained level; the segmentation mask generated in the training phase is used to promote region-level contrast learning; the loss formula is:

[0048] L rct =InfoNCE(S mask )

[0049] Where S mask Represents the image feature V of the mask region in the i-th image in the same batch mask Cosine similarity graph with feature T of the j-th text;

[0050]

[0051] in and It maps the input features to dimension D h Linear projection layer of M; using segmentation mask M binaryAnd the original image I obtains the mask image features Expressed as

[0052] V mask =E v (I⊙UP(M binary ))

[0053] UP(·) means to convert M binary Upsample to size I,

[0054] M binary =Gumbel-Max(reshape 2d (M final ))

[0055] in Gumbel-Max(·) is a hard assignment method that ensures that the binary mask remains differentiable and the entire process is end-to-end trainable; reshape 2d (·) indicates that M final From 1D to 2D;

[0056] Applying total variation regularization loss makes V up and V are smoother in the feature space:

[0057] L reg =ψ(V up )+ψ(V)

[0058] where ψ(·) is the anisotropic total variation norm.

[0059] The image encoder uses the Swin-Transformer network structure as the backbone network; the text encoder uses a 6-layer Transformer network structure as the backbone network.

[0060] Beneficial effects of the invention: The invention solves the directed segmentation task only through text supervision; it uses a semantic perception module and a detail recognition module to integrate image features and text features, and explores high-level semantics and low-level details respectively; the collaborative learning strategy aims to enhance the interaction between the two modules and achieve comprehensive visual language alignment; in addition, the invention introduces a new negative text generation method to ensure accurate understanding of referring expressions; extensive experiments on three benchmarks show that the invention outperforms the current state-of-the-art weakly supervised methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 Schematic diagram of the training framework for the weakly supervised directional segmentation method based on the collaboration of semantics and details.

[0062] Figure 2 This is the internal structure diagram of the semantic perception module.

[0063] Figure 3 It is the internal structure of the detail cognition module.

[0064] Figure 4 Schematic diagram of the new negative text generation method. DETAILED DESCRIPTION

[0065] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0066] This method uses a semantic perception module and a detail cognition module, and explores the potential coordination of semantics and details through a collaborative learning strategy; the semantic perception module adopts an attention-based cross-modal interaction method and uses a gradient weighting technique to generate an attention map indicating the image area; this top-down approach encourages accurate localization of the target object, but cannot detect the entire object area; the detail cognition module explores potential related areas in a bottom-up manner by calculating the similarity between image and text features, thereby producing an activation map; this two-pronged approach ensures comprehensive information extraction and lays a solid foundation for accurate segmentation guided by text descriptions; the collaborative learning strategy uses collaborative modules to fuse the advantages of the above two modules; in addition, unlike traditional contrastive learning-based methods that directly align visual and text features, our strategy restricts the potential consistency of cross-modal features from three perspectives: internal alignment of cross-modal features, alignment of cross-modal features with text features, and alignment of mask area image features with text features; to further encourage fine-grained alignment between cross-modal features, we propose a negative text generation method, which is applied in the collaborative learning stage; it guides the model to focus on more subtle basic elements, such as attributes and relations, improving the performance of the model in distinguishing and understanding natural language.

[0067] Figure 1 Schematic diagram of the network training framework based on semantics and detail collaboration; the network includes an image encoder, a text encoder, a semantic perception module, a detail recognition module, and a collaboration module; V i represents the features of the i-th image, T i Represents the text features that match the image, T j It does not match the image; the semantic perception module and detail recognition module are used for cross-modal feature fusion, which integrates image features and text features by exploring high-level semantics and low-level detail information, and generates two complementary activation maps at the same time; the collaborative learning strategy uses collaborative modules to combine activation maps and generate the final segmentation mask; it mainly uses three losses to jointly train the entire network, namely, the cross-modal internal loss L cmi , matching invariance loss L miv and region contrast loss L rct, for comprehensive visual language alignment.

[0068] Figure 2 is the specific structure of the semantic perception module, where Represents cross-modal features containing high-level semantics generated by paired image features and text features. on the contrary.

[0069] Figure 3 is the specific structure of the detail recognition module, where represents the image features after upsampling, Refers to the matrix multiplication operation; represents the cross-modal features containing low-level details generated by paired image features and text features, on the contrary.

[0070] Figure 4 Introduced a negative text generation method, where "girl under a green umbrella" is an example.

[0071] A weakly supervised directional segmentation method based on semantics and detail collaboration includes the following steps:

[0072] Step 1: Build an image encoder and a text encoder;

[0073] The Swin-Transformer network structure is used as the backbone network to establish an image encoder for extracting image features. One of the images is split into N v non-overlapping patches, D v Represents the channel dimension of the visual features; the input of the image encoder is a three-channel RGB image, and the number of channels corresponding to the output image features is 1024.

[0074] A 6-layer Transformer network structure is used as the backbone network to build a text encoder for extracting text features. Its input is a string containing N t words, the sentence is tokenized into a token sequence, and a [CLS] token is added to the beginning of the sequence; t Represents the channel dimension of text features, and its value is 768.

[0075] Step 2: Build a semantic perception module;

[0076] The semantic perception module captures the long-term dependency between image features and text features through the attention mechanism and fuses the two features together. It adopts the gradient weighted class activation mapping technology and uses the gradient of image-text similarity to determine which areas in the image have the greatest impact on the final classification results, thereby identifying the target object from a holistic perspective.

[0077] The semantic perception module consists of 6 stacked network layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer, and a feedforward network. The module takes image features and text features as input, such as Figure 2 As shown in the figure, for the cross attention layer, text features are used as queries and image features are used as keys and values. The output fusion feature of the semantic perception module is represented as This feature is subject to a matching invariance loss to ensure alignment between image-text pairs; it is then weighted using a gradient-weighted class activation mapping technique based on F sem Generate activation map This activation map indicates the visual regions that are most semantically relevant to the reference expression.

[0078]

[0079] in is the affinity matrix of the cross-attention layer of the semantic perception module; Reflects the affinity matrix to the similarity score S sem The score evaluates the matching degree between image features and text features; ⊙ represents the element-wise multiplication operation.

[0080] Activation Map M sem The image regions most relevant to the referring expression are highlighted; the global perspective helps capture a broader contextual framework and overall semantic content; however, when determining the matches of cross-modal features, this top-down localization approach mainly focuses on the most prominent regions, resulting in M sem Subtle details are often overlooked, especially at object boundaries; therefore, more detailed information about the referenced objects is needed.

[0081] Step 3: Construct detail recognition module;

[0082] The detail recognition module integrates image features and text features by exploring low-level details of cross-modal features, such as Figure 3 As shown in the figure; first, a convolutional neural network is used to extract local features from the image, and an upsampling technique is used to increase the resolution of the visual features; it obtains a feature with higher resolution Where N uprepresents the number of non-overlapping patches; the upsampled image features contain more detailed information such as color and texture, thus facilitating subsequent low-level perception; then, the module generates an activation map by calculating the affinity between the normalized image features and the text features

[0083]

[0084] in and is to map the input to dimension D h Linear projection layer of is the [CLS] embedding of text features, and norm(·) denotes L2 normalization. This method is able to leverage the alignment of low-level image features with text details to recognize the entire object from a low-level perspective. The activation map generated by the shallow interaction between the two modalities mainly highlights the low-level information and performs a broader recognition of the entire object.

[0085] Then, use M det Weight the image features to obtain cross-modal features

[0086]

[0087] in is to map the input to dimension D h Projection.

[0088] Step 4: Build a collaborative learning strategy

[0089] The present invention adopts a collaborative learning strategy to understand the complex correlations between images and texts, such as Figure 1 As shown in the right half; it uses a collaborative module to generate segmentation masks using two activation maps that highlight high-level semantics and low-level details respectively.

[0090] The collaboration module jointly integrates M sem The high-level semantic information and M det low-level details; first, considering the M generated by the semantic perception module sem Focusing on the limited object area, the present invention uses a block similarity propagation method to expand the effective activation area; this involves sem The activation scores of each block in are extended to its neighboring blocks according to their similarity.

[0091] First, calculate the similarity matrix between image patches

[0092]

[0093] in and is a linear projection layer; then, the corrected activation map is calculated

[0094]

[0095] where N(n) is the set of eight patches that are spatially adjacent to the nth patch in the 2D plane, and σ(·) is the softmax function; compared to previous methods, this approach eliminates interference from patches far away from the nth patch, since distant patches are not always highly correlated with the nth patch; nevertheless, M rec Still, background regions will be activated incorrectly; by changing M rec With M det Merging can solve this problem; the final activation map is represented as:

[0096] M final =Up(M rec )⊙M det

[0097] Where Up(·) means to increase M rec Upsample to M det Same size;

[0098] Subsequently, a collaborative learning strategy exploits cross-modal information from three perspectives to ensure comprehensive visual-language alignment: cross-modal internal loss, matching invariance loss, and region contrast loss; in addition, a total variation regularization loss is introduced to improve segmentation performance; a unified loss function simultaneously constrains the semantic perception module and the detail cognition module to ensure that cross-modal features are aligned at two different levels; the semantic perception module will have a more significant impact on the loss components related to high-level semantic information, while the detail cognition module tends to have a more obvious impact on the alignment of low-level features.

[0099] The cross-modal internal loss enables the model to directly use cross-modal feature information to distinguish whether the image-text pair matches; this loss is used to learn the [CLS] embedding of cross-modal features and obtain the true value label by judging the match of the image-text pair:

[0100] L cmi =CE(MLP(F sem [CLS]), y label )+CE(MLP(F det ), y label )

[0101] Among them, MLP is a multi-layer perceptron, which reduces the number of channels of multimodal features to 2; when the cross-modal feature F sem and F det is obtained from a correctly aligned image-text pair, y labelThe value of is 1, otherwise it is 0; CE(·,·) represents the cross entropy loss.

[0102] The present invention also uses matching invariance loss to bring the cross-modal features generated by real image-text pairs closer to the text features, while pushing the cross-modal features from mismatched image-text pairs away from the text features; intuitively, cross-modal features can be seen as language features enriched with visual context; therefore, if the i-th image and the j-th text match, then the cross-modal features and will be combined with the text feature T j Closely aligned; on the contrary, the distance between them will be enlarged; first calculate the similarity between the cross-modal features and the text features,

[0103]

[0104] where m and n represent the m-th and n-th words respectively; then, the matching invariance loss can be expressed by the InfoNCE loss:

[0105] L miv =InfoNCE(S sem )+InfoNCE(S det )

[0106]

[0107] Where τ represents a learnable temperature parameter and B represents the size of the training batch; by applying matching invariance loss to cross-modal features that contain high-level semantics and low-level details, the model can discern the alignment between images and texts at all levels, thereby improving the performance of the model;

[0108] In order to further distinguish highly similar but unmatched image-text pairs, the present invention adds the influence of negative image-text pairs showing high matching similarity during the training process, that is, the similarity score is used to select image-text pairs; then, the negative text of each image is sampled from the batch according to the similarity score, and the unmatched image of each original text is sampled according to the similarity score; therefore, the matching invariance loss is further modified as:

[0109] L′ miv =L miv +InfoNCE(S sem′ )+InfoNCE(S det′ )

[0110] Among them, S sem′ Is not from S sem The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000, S det′ Is not from Sdet The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000;

[0111] The third loss used in the present invention is the region contrast loss; this loss is used to align image features and text features at a fine-grained level; compared to traditional contrastive learning methods that perform feature alignment at the image level, the present invention utilizes the segmentation masks generated during the training phase to promote region-level contrastive learning; this aims to make the image features of the masked region closer to the corresponding text features while ensuring a clear distinction from the unmatched text features; the loss formula is:

[0112] L rct =InfoNCE(S mask )

[0113] Where S mask Represents the image feature V of the mask area in the i-th image in the same batch mask Cosine similarity graph between the feature T of the j-th text.

[0114]

[0115] in and is to map the input to dimension D h Linear projection of binary And the original image I obtains the mask image features Expressed as

[0116] V mask =E v (I⊙UP(M binary ))

[0117] UP(·) means to convert M binary Upsample to the size of the original image I, M binary represents the segmentation mask,

[0118] M binary =Gumbel-Max(reshape 2d (M final ))

[0119] in Gumbel-Max(·) is a hard assignment method that ensures that the binary mask remains differentiable and the entire process is end-to-end trainable; reshape 2d (·) indicates that M final From 1D to 2D.

[0120] In addition, the embedding similarity between adjacent blocks of the image should be stronger; the present invention applies total variation regularization loss to make Vup and V are smoother in the embedding space:

[0121] L reg =ψ(V up )+ψ(V)

[0122] where ψ(·) is the anisotropic total variation norm.

[0123] Furthermore, the negative text generation method is introduced

[0124] Directly aligning the entire reference text with the image may not be able to effectively identify the subtleties of basic elements such as attributes and relationships; in order to avoid learning bag-of-words representation, the present invention adopts negative text generation technology in collaborative learning, which helps to understand the reference expression in a fine-grained manner, such as Figure 4 shown.

[0125] In addition to the negative text in the batch that does not match the image, the present invention also introduces more difficult to identify negative text by replacing the keywords in the paired text; first parse the text, that is, use the text parser Spacy to assign a part-of-speech tag to each word; then adopt two strategies to generate challenging negative text; the first is to randomly mask an adjective or noun and fill it with RoBERTa; if there is no adjective in the original sentence, randomly add a color word or a position word to the sentence; the second is to swap the positions of the two nouns in the text; if there is only one noun, randomly select a category name from the COCO dataset and add it to the text, and swap the positions of the two nouns; by modifying the keywords of the sentence to make negative samples, the model can perceive the subtle differences in semantic information in natural language when learning the matching between images and text, thereby effectively improving its ability to distinguish between real targets and image backgrounds.

[0126] Construct the overall network structure; input an image and a short sentence in natural language, send the image to the image encoder to extract image features, and send the short sentence in natural language to the text encoder to extract text features; send the image features and text features to the semantic perception module and the detail recognition module to generate cross-modal features focusing on high-level semantics and low-level details respectively, and generate two activation maps at the same time; send the two activation maps to the collaboration module to output the segmentation result.

[0127] Training phase: using Swin-Transformer-Base eAs the image encoder, use a 6-layer Transformer as the text encoder, load the pre-trained weights at the beginning of training, and load the weights of X-VLM at the beginning of training, such as "Yan Zeng, Xinsong Zhang, Hang Li. Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts [C]. In International Conference on Machine Learning, ACM, 2022"; use the overall loss L = L reg +λ1L′ miv +λ2L cmi +λ3L rct The model was trained, where λ1=1.0, λ2=2.0, and λ3=0.2 are hyperparameters for balancing the importance of loss. The model was trained on the RefCOCO, RefCOCO+, and RefCOCOg datasets, and finally validated on the corresponding test sets. The training was repeated for 15 rounds with a total of 56,000 iterations, a batch size of 32, and an input image resolution of 224×224. The network was optimized by the AdamW optimizer with a weight decay of 1e -2 ; In the first round, the learning rate is 1e -8 Heating to 1e -5 , the subsequent rounds decay linearly to 1e -8 .

[0128] The model performance is improved by 6.21%, 7.25%, 6.11% and 5.88% compared with the existing technology on the validation sets of RefCOCO, RefCOCO+, G-Ref (Google partition) and G-Ref (UMD partition), respectively.

Claims

1. A weakly supervised directional segmentation method based on semantics and detail collaboration, characterized in that: Here are the steps: Step 1: Build an image encoder to extract image features One of the images is split into N v non-overlapping patches, D v Represent the channel dimension of image features; build a text encoder; extract text features The input is a string containing N t words, the sentence is tokenized into a token sequence, and a [CLS] token is added to the beginning of the sequence; t Channel dimension representing text features; Step 2: Build a semantic perception module; The semantic perception module captures the long-term dependency between image features and text features through the attention mechanism, and fuses the two features together to obtain the activation map M. sem ; Step 3: Construct detail recognition module; The detail recognition module integrates image features and text features by exploring low-level details of cross-modal features to obtain the activation map M det ; by M det Weight the image features to obtain the cross-modal feature F det ; Step 4: Build a collaborative learning strategy A collaborative learning strategy is adopted to understand the complex correlation between images and texts. A collaborative module uses activation maps that highlight high-level semantics and low-level details respectively to generate segmentation masks. Subsequently, a collaborative learning strategy leverages cross-modal information from three perspectives to ensure comprehensive image-text alignment: cross-modal internal loss, matching invariance loss, and region contrast loss. In addition, a total variation regularization loss is introduced to improve segmentation performance. A unified loss function simultaneously constrains the semantic perception module and the detail recognition module to ensure cross-modal features are aligned at two different levels. Step 5: After the training of steps 1 to 4 is completed, an image and a short sentence in natural language are input, the image is sent to the image encoder to extract image features, and the short sentence in natural language is sent to the text encoder to extract text features; the image features and text features are sent to the semantic perception module and the detail recognition module to generate cross-modal features F focusing on high-level semantics. sem and cross-modal features F that focus on low-level details det , and generate two activation maps at the same time; the two activation maps are sent to the collaboration module to output the segmentation results.

2. The method for directivity segmentation based on semantics and detail collaboration according to claim 1, characterized in that: A new method for generating negative text is also added to ensure accurate understanding of referring expressions; The new negative text generation method refers to introducing more difficult to recognize negative text by replacing keywords in the paired text; first, the text is parsed and a part-of-speech tag is assigned to each word using a text parser; then two strategies are adopted to generate challenging negative text; the first is to randomly mask an adjective or noun and fill it with RoBERTa; if there is no adjective in the original sentence, a color word or a position word is randomly added to the sentence; the second is to swap the positions of the two nouns in the text; if there is only one noun, a category name is randomly selected from the COCO dataset and added to the text, and the positions of the two nouns are swapped; by modifying the keywords of the sentence to make negative samples, the segmentation model perceives the subtle differences in semantic information in natural language when learning the matching between images and text, thereby effectively improving its ability to distinguish between real targets and image backgrounds.

3. The method for directivity segmentation based on semantics and detail collaboration according to claim 1 or 2, characterized in that: The semantic perception module adopts the gradient weighted class activation mapping technology, which uses the gradient of image-text similarity to determine the area in the image that has the greatest impact on the final classification result, and then identifies the target object from an overall perspective; The semantic perception module consists of 6 stacked network layers, each of which contains a multi-head self-attention layer, a multi-head cross-attention layer and a feedforward network, with image features and text features as input; for the cross-attention layer, the text features are used as queries and the image features are used as keys and values; the semantic perception module outputs cross-modal features represented as Where D h Represents the channel dimension of the cross-modal feature; the cross-modal feature is constrained by the matching invariance loss to ensure the alignment between the image-text pairs; the gradient weighted class activation mapping technique is used to sem Generate activation map This activation map indicates the image regions that are most semantically relevant to the reference expression. in, is the affinity matrix of the cross-attention layer of the semantic perception module, Represents the matrix A cross The nth column element taken out from Reflects the affinity matrix to the similarity score S sem contribution; the similarity score is used to evaluate the matching degree between image features and text features, and ⊙ represents an element-by-element multiplication operation.

4. The method for directivity segmentation based on semantics and detail collaboration according to claim 1 or 2, characterized in that: The detail recognition module first uses a convolutional neural network to extract local features from the image, and uses upsampling technology to improve the resolution of image features to obtain a feature with higher resolution. Where N up Represents the number of non-overlapping patches; the detail recognition module generates an activation map by calculating the affinity between normalized image features and text features in and It maps the input image features and text features to dimension D h Linear projection layer of is the [CLS] embedding of text features, norm(·) represents L2 normalization; using M det Weight the image features to obtain cross-modal features in It maps the input image features to dimension D h Linear projection layer.

5. The method for directivity segmentation based on semantics and detail collaboration according to claim 1 or 2, characterized in that: The collaborative modules jointly integrate M sem The high-level semantic information and M det low-level details; first, considering the M generated by the semantic perception module sem Focusing on the limited object area, a block similarity propagation method is used to expand the effective activation area. sem The activation score of each block in is extended to its neighboring blocks according to their similarity; First, calculate the similarity matrix between image patches in and is a linear projection layer; Then, calculate the corrected activation map Where N(n) is the set of eight blocks spatially adjacent to the nth image block on the 2D plane, σ(·) is the softmax function; rec With M det The final activation map obtained by merging is expressed as: M final =Up(M rec )⊙M det Where Up(·) means to increase M rec Upsample to M det Same size.

6. The method for directional segmentation based on semantics and detail collaboration according to claim 1 or 2, characterized in that: The cross-modal internal loss enables the method to directly use cross-modal feature information to distinguish whether image-text pairs match; this loss is used to learn the [CLS] embedding of cross-modal features and obtain the true value label by judging the match of image-text pairs: L cmi =CE(MLP(F sem [CLS]),y label )+CE(MLP(F det ),y label ) Among them, MLP is a multi-layer perceptron, which is used to reduce the number of channels of multimodal features to 2; when the cross-modal feature F sem and F det is obtained from a correctly aligned image-text pair, y label The value of is 1, otherwise it is 0; CE(·,·) represents the cross entropy loss; The matching invariance loss is used to bring the cross-modal features generated by the real image-text pairs closer to the text features, while pushing the cross-modal features from the unmatched image-text pairs away from the text features; If the i-th image and the j-th text match, then the cross-modal feature and will be combined with the text feature T j closely aligned; on the contrary, the distance between them will be enlarged; first calculate With T j The similarity between Among them, m and n represent the m-th word and the n-th word respectively; the matching invariance loss is represented by the InfoNCE loss: L miv =InfoNCE(S sem )+InfoNCE(S det ) where τ denotes a learnable temperature parameter and B denotes the size of the training batch; the alignment between images and text is discriminated at various levels by applying a matching invariance loss to cross-modal features that contain both high-level semantics and low-level details; To further distinguish highly similar negative image-text pairs and increase the influence of negative image-text pairs showing high matching similarity during training, similarity scores are used to select image-text pairs; then, negative texts for each image are sampled from the batch based on the similarity scores, and unmatched images of each original text are sampled based on the similarity scores; the matching invariance loss is further modified as: L′ miv =L miv +InfoNCE(S sem′ )+InfoNCE(S det′ ) Among them, S sem′ Is not from S sem The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000, S det′ Is not from S det The similarity graph obtained by setting the elements corresponding to the sampled mismatched image-text pairs to -10000; The region contrast loss is used to align image features and text features at a fine-grained level; the segmentation mask generated in the training phase is used to promote region-level contrast learning; the loss formula is: L rct =InfoNCE(S mask ) Where S mask Represents the image feature V of the mask region in the i-th image in the same batch mask Cosine similarity graph with feature T of the j-th text; in and It maps the input features to dimension D h Linear projection layer of M; using segmentation mask M binary And the original image I obtains the mask image features Expressed as In mask =E v (I⊙UP(M binary )) UP(·) means to convert M binary Upsample to size I, M binary =Gumbel-Max(reshape 2d (M final )) in Gumbel-Max(·) is a hard assignment method that ensures that the binary mask remains differentiable and the entire process is end-to-end trainable; reshape 2d (·) indicates that M final From 1D to 2D; Applying total variation regularization loss makes V up and V are smoother in the feature space: L reg =ψ(V up )+ψ(V) where ψ(·) is the anisotropic total variation norm.

7. The method for directional segmentation based on semantics and detail collaboration according to claim 1 or 2, characterized in that: The image encoder uses the Swin-Transformer network structure as the backbone network; the text encoder uses a 6-layer Transformer network structure as the backbone network.

Citation Information

Patent Citations

  • Efficient weak supervision semantic segmentation method and device based on text driving

    CN115937852A

  • Semantic relationship mining and reasoning-based reference image segmentation method

    CN117078939A

  • Novel image generation method jointly driven by text and semantic segmentation map

    CN117557683A

  • Transform weak supervision semantic segmentation method combined with context attention

    CN118411522A

  • Segmentation of objects in an image

    WO2024233818A1

Cited By

  • Power equipment image comparison retrieval method and system based on semantic object relationship

    CN121030032A

  • Cross-modal image-text analysis method for machine vision

    CN121210958A