Feature modulation and cross-modal alignment-based weak supervision image semantic segmentation method
By using beta distribution modulation features and cross-modal alignment strategies in weakly supervised semantic segmentation, the problem of poor quality of category activation graph generation and indistinguishable categories is solved, and higher segmentation accuracy and more accurate pseudo-label generation are achieved.
Patent Information
- Application Number
- CN202510431011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-24
AI Technical Summary
The existing weakly supervised semantic segmentation method is difficult to pay full attention to all areas of the target when generating category activation maps, resulting in incomplete and inaccurate pseudo-labels, affecting segmentation accuracy; at the same time, the semantic similarity between categories makes it difficult for the model to distinguish between true and false positive characteristics.
The probability density function of beta distribution is used to modulate the feature processing of image tokens and text tags, and the recognition ability of the model is enhanced through cross-modal alignment strategies to generate more fine and accurate pseudo-labels.
It effectively improves the accuracy and segmentation accuracy of target activation, reduces the occurrence of misjudgments, and enhances the model's ability to distinguish different categories.
Smart Images

Figure CN120198672A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image semantic segmentation, and particularly relates to a weakly supervised image semantic segmentation method, system, device and storage medium based on feature modulation and cross-modal alignment. Background Art
[0002] With the development of computer vision, weakly supervised semantic segmentation (WSSS) has become an important research direction in the field of computer vision in recent years. Different from fully supervised learning methods, weakly supervised semantic segmentation uses more easily obtained weak labels such as image-level labels, bounding boxes, points, and scribbles to train the segmentation network, greatly reducing the cost of obtaining high-quality pixel-level semantic segmentation labels. This advantage makes weakly supervised semantic segmentation methods have broad development prospects in practical application scenarios such as smart cities, intelligent healthcare, and autonomous driving.
[0003] The image-level label-driven weakly supervised semantic segmentation method mainly involves two stages. In the first stage, a classification network is trained and a class activation map (CAM) is generated as a rough localization map of the target object. Then, the class activation map after post-processing is usually called a pseudo-mask. In the second stage, the pseudo-mask is used as pixel-level supervision to end-to-end train a fully supervised semantic segmentation network under the guidance of the cross-entropy loss and generate the final segmentation result.
[0004] In the prior art, there are mainly two problems with weakly supervised semantic segmentation methods. One is that due to insufficient label information, the class activation map often fails to fully focus on all regions of the target and only focuses on the regions with the most class discriminability of the target. For example, when classifying the category of "person", the model often focuses on the face and ignores the torso, which will directly lead to the generated pseudo-mask being often incomplete and inaccurate, ultimately affecting the accuracy of segmentation and resulting in a decrease in precision. The other is that there are often semantic similarities between categories in the dataset, such as sofas and chairs, which makes it difficult for the model to effectively distinguish true positive features and false positive features, resulting in misjudgment. In fact, the existence of false positive predictions is also one of the important factors affecting the performance of WSSS.
[0005] Therefore, there is an urgent need for a weakly supervised image semantic segmentation method that can effectively solve the problems of poor generation quality of class activation maps and difficulty in distinguishing categories in the dataset, and effectively improve the segmentation performance. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the above prior art and provide a weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment. This method modulates the features of the image tokens and text tokens by using the probability density function of the beta distribution during the feature extraction process of images and texts, effectively guiding the model to automatically activate the foreground and background, thereby generating accurate pseudo-labels with clearer boundaries and effectively improving the segmentation accuracy.
[0007] Meanwhile, the second purpose of the present invention is to provide a weakly supervised image semantic segmentation system based on feature modulation and cross-modal alignment.
[0008] Meanwhile, the third purpose of the present invention is to provide a computer device.
[0009] Meanwhile, the fourth purpose of the present invention is to provide a storage medium.
[0010] The purpose of the present invention is achieved through the following technical solutions:
[0011] A weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment, comprising the following steps:
[0012] Construct an image dataset, preprocess the image dataset, and divide the preprocessed image dataset into a training set and a test set. The image dataset includes picture information, category information, and text prompt information;
[0013] S2. Use the text encoder of the CLIP model to extract the text prompt information to obtain text tokens, and use the VIT model to extract the image data and category information to obtain image tokens and category tokens; then perform linear transformation and position encoding processing on the generated image tokens, category tokens, and text tokens respectively to obtain processed image tokens, category tokens, and text tokens;
[0014] S3. Construct a classification network model based on feature modulation and cross-modal alignment, and train the classification network model using the beta modulation mechanism and cross-modal alignment strategy to obtain a trained classification network model based on feature modulation and cross-modal alignment; the classification model includes a Transformer encoder module improved based on the beta modulation mechanism, a contrast learning module, and a classification module;
[0015] S4. Process the processed image tokens, class tokens, and text tokens using the trained classification network model based on feature modulation and cross-modal alignment to obtain a class activation map. Extract the mapping information between the image tokens and text tokens in the Transformer encoder module improved based on the beta modulation mechanism and add it to the class activation map to obtain a more refined class activation map, and convert the more refined class activation map into corresponding pseudo-labels;
[0016] S5. Construct and train an image semantic segmentation network, and use the pseudo-labels obtained in step S4 to evaluate the performance of the image semantic segmentation network during the training process of the image semantic segmentation network. Take the optimal image semantic segmentation network model obtained during training as the trained image semantic segmentation network model, and use the trained image semantic segmentation network model to complete the semantic segmentation of the target image.
[0017] Preferably, the training process of the classification network model based on feature modulation and cross-modal alignment in step S3 is as follows:
[0018] S31. Construct a classification network model based on feature modulation and cross-modal alignment. The classification model includes a Transformer encoder module improved based on the beta modulation mechanism, a contrastive learning module, and a classification module; the improved Transformer encoder is composed of multiple processing sub-layers stacked, and the processing sub-layer is composed of a parallel block and a beta modulation unit; the classification module includes a class activation map generation unit, a top classification head unit, an auxiliary classification head unit, a text classification loss unit, and an image classification loss unit;
[0019] S32. After processing the training set obtained in step S1 using step S2, obtain the preprocessed image tokens, class tokens, and text tokens, and input the preprocessed image tokens, class tokens, and text tokens into the classification network model based on feature modulation and cross-modal alignment;
[0020] S33. The Transformer encoder module improved based on the beta modulation mechanism performs feature stacking and beta feature modulation operations on the preprocessed image tokens, class tokens, and text tokens to obtain modulated image tokens, class tokens, and text tokens;
[0021] S34. Use the contrastive learning module to perform cross-modal alignment on the modulated image features and text features in step S33. During the alignment process, use the InfoNCE contrastive loss as the cross-modal alignment loss function to maximize the similarity between positive samples;
[0022] S35. Use the category activation map generation unit in the classification module to perform convolution and shape reshaping processing on the modulated image labels, and convert them into corresponding category activation maps;
[0023] S36. The top classification head unit obtains corresponding classification prediction values according to the modulated category labels, and calculates the top classification loss. The auxiliary classification head unit uses the category labels obtained from the penultimate layer processing sublayer of the Transformer encoder module improved based on the beta modulation mechanism in step S33 as input to calculate the auxiliary classification loss; the image classification loss unit calculates the image classification loss by applying global weighted sorting pooling to the modulated image labels; the text classification loss module obtains the prediction values based on the text labels through global average pooling and obtains the text classification loss;
[0024] S37. Construct a total loss function based on the top classification loss, auxiliary classification loss, image classification loss, text classification loss, and cross-modal alignment loss, and use minimizing the total loss function as the training objective to jointly train the network model to obtain a trained classification network model.
[0025] Preferably, the specific process of step S33 is as follows:
[0026] S331. Input the preprocessed image labels, category labels, and text labels into the Transformer encoder module improved based on the beta modulation mechanism; the improved Transformer encoder is composed of multiple stacked processing sublayers, and each processing sublayer is composed of a parallel block and a beta modulation unit;
[0027] S332. The preprocessed image labels, category labels, and text labels are further feature-extracted through each processing sublayer in the Transformer encoder module improved based on the beta modulation mechanism, and the output of each layer of processing sublayer is used as the input of the next layer of processing sublayer; the block in the processing sublayer performs feature extraction on the preprocessed image labels, category labels, and text labels, and the beta modulation unit in the processing sublayer performs beta feature modulation operations on the outputs of the blocks in the same processing sublayer to obtain the modulated image labels, category labels, and text labels;
[0028] Specifically, the working process of the beta modulation unit is as follows:
[0029] Taking the preprocessed image label as an example, first, normalize the image label output by the block on the D-dimensional channel so that the value range of the normalized image label meets the input conditions of the beta distribution probability density function. Then, modulate the normalized image label using the probability density function of the beta distribution to obtain the modulated image label.
[0030] Preferably, the working process of the contrast learning module in step S34 is as follows:
[0031] S341. First, define the set of categories included in the images in the training set; the specific categories are as follows:
[0032]
[0033] Among them, C p is the set of all categories in the training set, and C is the category actually present in the current image;
[0034] S342. Extract the modulated image features through convolution operations and compress them into a vector Q img through global average pooling. The specific representation of the vector Q img is as follows:
[0035] Q img = GAP(Conv(F pat ))
[0036] Among them, Conu(·) represents the convolutional layer, and GAP(·) represents the global average pooling operation;
[0037] S343. Project the modulated text features onto a new matrix U txt through a linear layer;
[0038] S344. Select the positive sample set and the negative sample set from the new matrix U txt . The text features of the categories that appear in the image dataset are used as the positive sample set, while the text features of the categories that do not appear in the image dataset are used as the negative sample set;
[0039] S345. Align the vector Q img and the matrix U txt . During the alignment process, use the InfoNCE contrast loss as the cross-modal alignment loss function to maximize the similarity between positive samples, minimize the similarity between negative samples, and calculate the corresponding cross-modal alignment loss. The specific representation of the cross-modal alignment loss function is as follows:
[0040]
[0041] Among them, LCA is the cross-modal alignment loss, τ is the temperature coefficient, τ is used to adjust the smoothness of the loss function, U pos is the positive sample set, U neg is the negative sample set.
[0042] Preferably, the more refined class activation map described in step S4 is represented as follows:
[0043] CAM = ReLU(Upsample(M)
[0044] M = A p2p ·(F′ pat oA t2p )
[0045] where CAM is the generated more refined class activation map, the size of CAM is H×W×C, H and W are the length and width of the original image respectively, F′ pat is the modulated image feature after convolution and shape reshaping transformation, and o is element-wise multiplication.
[0046] Preferably, the total loss function is represented as follows:
[0047] L = L top + L aux + L pat + L txt + λ CA L CA
[0048] where L top is the top classification loss, L aux is the auxiliary classification loss, L pat is the image label classification loss, L txt is the text label classification loss, L CA is the cross-modal alignment loss, and λ is an empirical parameter.
[0049] Preferably, the image semantic segmentation network adopts the DeeplabV3 network model.
[0050] A weakly supervised image semantic segmentation system based on feature modulation and cross-modal alignment, used to implement the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment, includes:
[0051] An image data module, used to preprocess the image data set and divide the preprocessed image data set into a training set and a test set, and the image data set includes image information, category information, and text prompt information;
[0052] A text marking module, used to extract the text prompt information of the image data set by using the text encoder of the CLIP model to obtain text marks;
[0053] An image marking module, configured to extract image information of the image dataset using a VIT model to obtain an image mark;
[0054] A category marking module, configured to extract category information of the image dataset using a VIT model to obtain a category mark;
[0055] A linear transformation and position encoding module, configured to perform linear transformation and position encoding processing on the image mark, category mark, and text mark respectively to obtain the processed image mark, category mark, and text mark;
[0056] A classification network model construction module, configured to construct a classification network model based on feature modulation and cross-modal alignment;
[0057] A classification network model training module, configured to train the classification network model using a beta modulation mechanism and a cross-modal alignment strategy to obtain a trained classification network model based on feature modulation and cross-modal alignment;
[0058] A pseudo-label generation module, which processes the processed image mark, category mark, and text mark using the trained classification network model based on feature modulation and cross-modal alignment to obtain a class activation map, and extracts the mapping information between the image mark and the text mark in the Transformer encoder module improved based on the beta modulation mechanism and adds it to the class activation map to obtain a more refined class activation map, and converts the more refined class activation map into a corresponding pseudo-label;
[0059] An image semantic segmentation network training module, which constructs and trains an image semantic segmentation network, and evaluates the effect of the image semantic segmentation network using the pseudo-label of the pseudo-label generation module during the training process of the image semantic segmentation network, and uses the optimal image semantic segmentation network model obtained during the training as the trained image semantic segmentation network model;
[0060] An image semantic segmentation module, configured to perform semantic segmentation of a target image using the trained image semantic segmentation network model.
[0061] A computer device, including a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, it implements the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment.
[0062] A storage medium stores a program, which when executed by a processor, implements the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment.
[0063] The present invention has the following advantages and beneficial effects compared with the prior art:
[0064] (1) The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment of the present invention effectively guides the model to automatically activate the foreground and background and capture local targets with less typical features by using the beta modulation scheme to modulate the features of image tokens and text tokens. Moreover, it is not necessary to obtain the modulation function parameters by fitting the data. The modulation shape is flexible and the computational cost is low, which can effectively improve the accuracy of target activation, so that the model generates clearer and more accurate pseudo-labels and effectively improves the segmentation accuracy.
[0065] (2) The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment of the present invention not only enhances the model's recognition ability for positive categories but also effectively distinguishes the subtle differences between different categories through the cross-modal alignment strategy, thus reducing the occurrence of misjudgments. The selection of negative samples is dynamically adjusted according to the training process, focusing only on the most confusing categories, thereby guiding the model to learn key features and improving the accuracy of target recognition, thus effectively improving the segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 It is a schematic diagram of the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment provided in Embodiment 1 of the present invention;
[0067] Figure 2 It is a partial flow schematic diagram of the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment provided in Embodiment 1 of the present invention;
[0068] Figure 3 It is a flow schematic diagram of cross-modal alignment of the optimal image features and text features provided in Embodiment 1 of the present invention;
[0069] Figure 4 It is the beta distribution form of the beta modulation unit provided in Embodiment 1 of the present invention.
[0070] Figure 5 It is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention.
[0071] Figure 6 It is a schematic diagram of the structure of a storage medium provided in Embodiment 4 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The present invention will be further described below with reference to the drawings and embodiments.
[0073] To make the object, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0074] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and through specific embodiments.
[0075] Embodiment 1
[0076] A weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment, characterized by comprising the following steps:
[0077] S1. Construct an image data set, preprocess the image data set, and divide the preprocessed image data set into a training set and a test set. The image data set includes graphic information, category information, and text prompt information;
[0078] S2. Use the text encoder of the CLIP model to extract the text prompt information to obtain text tokens, and use the VIT model to extract the image data and category information to obtain image tokens and category tokens; then perform linear transformation and position encoding processing on the generated image tokens, category tokens, and text tokens respectively to obtain processed image tokens, category tokens, and text tokens;
[0079] The specific process of step S2 is as follows;
[0080] The text encoder of the CLIP model extracts the text prompt information to obtain a text token T of size C*D txt , where C is the category and D is the dimension of the feature space; this setting is to provide additional semantic information for the model through the text token and effectively improve the discrimination ability of the model between similar categories. Then use the VIT model to extract the image tokens and category tokens in the image data set to obtain the corresponding image token T pat , category token T cls , and perform linear transformation and position encoding on the image token T pat , category token T cls and text token T txt respectively to obtain processed image tokens, category tokens, and text tokens; this setting is to construct an effective multi-modal information fusion mechanism and retain the category token T that aggregates global features of the VIT cls, this adopts a mixed method of multiple markers, which can not only retain the original image features, but also introduce and not completely rely on text supervision, which is beneficial to capturing and retaining richer and more diverse semantic information.
[0081] S3. Construct a classification network model based on feature modulation and cross-modal alignment, and use the beta modulation mechanism and cross-modal alignment strategy to train the classification network model to obtain a trained classification network model based on feature modulation and cross-modal alignment; the classification model includes a Transformer encoder module improved based on the beta modulation mechanism, a contrastive learning module, and a classification module;
[0082] Specifically, the training process of the classification network model based on feature modulation and cross-modal alignment is as follows:
[0083] S31. Construct a classification network model based on feature modulation and cross-modal alignment. The classification model includes a Transformer encoder module improved based on the beta modulation mechanism, a contrastive learning module, and a classification module; the improved Transformer encoder is composed of multiple processing sub-layers stacked, and the processing sub-layer is composed of a parallel block and a beta modulation unit; the classification module includes a class activation map generation unit, a top classification head unit, an auxiliary classification head unit, a text classification loss unit, and an image classification loss unit;
[0084] S32. After the training set obtained in step S1 is processed by step S2, obtain the preprocessed image markers, class markers, and text markers, and input the preprocessed image markers, class markers, and text markers into the classification network model based on feature modulation and cross-modal alignment;
[0085] S33. The Transformer encoder module improved based on the beta modulation mechanism performs feature superposition and beta feature modulation operations on the preprocessed image markers, class markers, and text markers to obtain the modulated image markers, class markers, and text markers; the specific process is as follows:
[0086] S331. Input the preprocessed image markers, class markers, and text markers into the Transformer encoder module improved based on the beta modulation mechanism; the improved Transformer encoder is composed of multiple processing sub-layers stacked, and the processing sub-layer is composed of a parallel block and a beta modulation unit;
[0087] S332. The pre-processed image tokens, class tokens, and text tokens are further feature-extracted by each processing sub-layer in the Transformer encoder module improved based on the beta modulation mechanism, and the output of each processing sub-layer is used as the input of the next processing sub-layer; in the processing sub-layer, the block extracts features from the pre-processed image tokens, class tokens, and text tokens, and the beta modulation unit in the processing sub-layer performs beta feature modulation operations on the outputs of the blocks in the same processing sub-layer to obtain the modulated image tokens, class tokens, and text tokens;
[0088] Specifically, the working process of the beta modulation unit is as follows:
[0089] Taking the image features as an example, first, the input feature embedding of the i-th block is represented as x i , x i with a size of (1 + P + C) * D. Normalize x i on the D-dimensional channel so that the value range of the normalized x i meets the input conditions of the probability density function of the beta distribution. Then, modulate the normalized image tokens using the probability density function of the beta distribution. After modulation, x i is de-normalized so that the modulated feature values remain within the original size range;
[0090] Specifically, the probability density function of the beta distribution is shown as follows:
[0091]
[0092] where Γ(·) represents the gamma function.
[0093] Specifically, the modulated feature values are represented as follows:
[0094]
[0095] where UnNorm is the de-normalization operation and Norm is the normalization operation.
[0096] For example Figure 4As shown, in this embodiment, in order to more comprehensively guide network attention, the beta modulation unit has three different shapes of beta distributions, which are respectively used to guide the network to focus on high-activation, medium-activation, and low-activation regions, corresponding to different levels of feature activation. Therefore, the network can effectively capture different features of the target and enhance the attention to details and backgrounds. At this time, the features of the three modulated shapes are stacked together, with a size of (3, 1 + P + C, D), and then the modulated features are reduced in dimension to a size of (3, 1 + P + C, D / 3) through a convolution operation, followed by batch normalization, and then the features are fused through a channel shuffling mechanism shuffle(·) to obtain the modulated image features.
[0097] The representation of the modulated image features is as follows:
[0098]
[0099] Specifically, the modulated features As a supplementary branch of the i-th block of the Transformer encoder, it participates in subsequent network training;
[0100]
[0101] Among them, blk() is the operation of the block, and λ BM is an empirical hyperparameter.
[0102] S34. Use the contrast learning module to perform cross-modal alignment on the modulated image features and text features in step S33. During the alignment process, use the InfoNCE contrast loss as the cross-modal alignment loss function to maximize the similarity between positive samples; the text features of the categories in the training set are the positive sample set, and the text features of the categories not present in the image dataset are the negative sample set;
[0103] The working process of the contrast learning module is as follows:
[0104] S341. First, define the set of categories included in the images in the training set; the specific categories are as follows:
[0105]
[0106] Among them, C p is the set of all categories in the training set, and C is the category actually present in the current image;
[0107] S342. Extract the modulated image features through a convolution operation and compress them into a vector Q img , and the specific representation of the vector Q img is as follows:
[0108] Q img = GAP(Conv(F pat ))
[0109] Among them, Conu(·) represents the convolutional layer, and GAP(·) represents the global average pooling operation. This process can extract the semantic information in the image and map it into a low-dimensional space for easy alignment with text features.
[0110] S343. Project the modulated text features into a new matrix U through a linear layer txt ;
[0111] S344. Select the positive sample set and the negative sample set from the new matrix U txt The text features of the categories that appear in the image dataset are used as the positive sample set, while the text features of the categories that do not appear in the image dataset are used as the negative sample set;
[0112] Specifically, the positive sample set and the negative sample set are represented as follows:
[0113]
[0114] Among them, U pos is the positive sample set, U neg is the negative sample set, P(·) is the classification probability of category j obtained by the top classification head, and topK(P(·)) refers to selecting the highest K probabilities.
[0115] S345. Align the vector Q img and the matrix U txt During the alignment process, the InfoNCE contrast loss is used as the cross-modal alignment loss function to maximize the similarity between positive samples, minimize the similarity between negative samples, and calculate the corresponding cross-modal alignment loss; The specific representation of the cross-modal alignment loss function is as follows:
[0116]
[0117] Among them, L CA is the cross-modal alignment loss, τ is the temperature coefficient, τ is used to adjust the smoothness of the loss function, U pos is the positive sample set, U neg is the negative sample set.
[0118] Setting the cross-modal alignment loss is to optimize the contrast learning process between text features and image features, ensure the discrimination between categories, reduce false positive predictions, and finally the combined loss function is jointly optimized by multiple supervision signals, enabling the model to be effectively enhanced and improving the final semantic segmentation performance.
[0119] S35. Use the class activation map generation unit in the classification module to perform convolution and shape reshaping on the modulated image labels, and convert them into corresponding class activation maps;
[0120] S36. The top classification head unit obtains corresponding classification prediction values based on the modulated class labels, and calculates the top classification loss. The auxiliary classification head unit uses the class labels obtained from the penultimate layer processing sublayer of the Transformer encoder module improved based on the beta modulation mechanism in step S33 as inputs to calculate the auxiliary classification loss; the image classification loss unit calculates the image classification loss by applying global weighted sorting pooling to the modulated image labels; the text classification loss module obtains the prediction values based on the text labels through global average pooling and obtains the text classification loss;
[0121] S37. Construct a total loss function based on the top classification loss, auxiliary classification loss, image classification loss, text classification loss, and cross-modal alignment loss, and use minimizing the total loss function as the training objective to jointly train the network model to obtain a trained classification network model.
[0122] The total loss function is expressed as follows:
[0123] L = L top + L aux + L pat + L txt + λ CA L CA
[0124] Among them, L top is the top classification loss, L aux is the auxiliary classification loss, L pat is the image label classification loss, L txt is the text label classification loss, L CA is the cross-modal alignment loss, and λ is an empirical parameter.
[0125] S4. Process the processed image labels, class labels, and text labels using the trained classification network model based on feature modulation and cross-modal alignment to obtain class activation maps, and extract the mapping information between the image labels and text labels in the Transformer encoder module improved based on the beta modulation mechanism and add it to the class activation maps to obtain more refined class activation maps, and convert the more refined class activation maps into corresponding pseudo-labels;
[0126] Specifically, in this embodiment, the mapping information is extracted from the multi-layer self-attention blocks of the Transformer encoder module improved based on the beta modulation mechanism. The mapping information includes the relationship mapping graph A from image tokens to text tokens t2p , with a size of CxP, and the relationship mapping graph A from image tokens to image tokens p2p , with a size of PxP;
[0127] Specifically, the more refined class activation map is expressed as follows:
[0128] CAM = ReLU(Upsample(M)
[0129] M = A p2p ·(F′ pat oA t2p )
[0130] where CAM is the generated more refined class activation map, with a size of H×W×C, where H and W are the length and width of the original image respectively, F′ pat is the modulated image feature after convolution and shape reshaping transformation, and o is element-wise multiplication.
[0131] Then the more refined class activation map generates corresponding pseudo-labels. The pseudo-label P is specifically expressed as follows:
[0132] P = arg max(BG∪CAM)
[0133] where ε is the probability that a pixel is the background, ε is set to 0.5, P is the pseudo-label, BG is a matrix with all element values being ε, and its size is H*W*1.
[0134] Specifically, for natural images, pseudo-label optimization methods (the seed propagation module of CRF or AffinityNet or FMA based on the Segment Anything Model) can be used to optimize the pseudo-labels for subsequent evaluation of the image semantic segmentation network.
[0135] S5. Construct an image semantic segmentation network and train it, and use the pseudo-labels obtained in step S4 to evaluate the effect of the image semantic segmentation network during the training process of the image semantic segmentation network. Take the optimal image semantic segmentation network model obtained during training as the trained image semantic segmentation network model, and use the trained image semantic segmentation network model to complete the semantic segmentation of the target image.
[0136] Specifically, in this embodiment, the image semantic segmentation network uses the DeeplabV3 network model.
[0137] Embodiment 2
[0138] A weakly supervised image semantic segmentation system based on feature modulation and cross-modal alignment, which is used to implement the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment as described in Embodiment 1. The system includes:
[0139] An image data module, which is used to preprocess the image data set and divide the preprocessed image data set into a training set and a test set. The image data set includes image information, category information, and text prompt information;
[0140] A text marking module, which is used to extract the text prompt information of the image data set by using the text encoder of the CLIP model to obtain text marks;
[0141] An image marking module, which is used to extract the image information of the image data set by using the VIT model to obtain image marks;
[0142] A category marking module, which is used to extract the category information of the image data set by using the VIT model to obtain category marks;
[0143] A linear transformation and position encoding module, which is used to perform linear transformation and position encoding processing on the image marks, category marks, and text marks respectively to obtain the processed image marks, category marks, and text marks;
[0144] A classification network model construction module, which is used to construct a classification network model based on feature modulation and cross-modal alignment;
[0145] A classification network model training module, which is used to train the classification network model by using the beta modulation mechanism and cross-modal alignment strategy to obtain a trained classification network model based on feature modulation and cross-modal alignment;
[0146] A pseudo-label generation module, which processes the processed image marks, category marks, and text marks by using the trained classification network model based on feature modulation and cross-modal alignment to obtain a class activation map, and extracts the mapping information between the image marks and the text marks in the Transformer encoder module improved based on the beta modulation mechanism and adds it to the class activation map to obtain a more refined class activation map, and converts the more refined class activation map into corresponding pseudo-labels;
[0147] An image semantic segmentation network training module, which constructs and trains an image semantic segmentation network, and evaluates the effect of the image semantic segmentation network by using the pseudo-labels of the pseudo-label generation module during the training process of the image semantic segmentation network, and uses the optimal image semantic segmentation network model obtained during the training as the trained image semantic segmentation network model;
[0148] An image semantic segmentation module for performing semantic segmentation of a target image using the trained image semantic segmentation network model.
[0149] Embodiment 3
[0150] As Figure 5 shown, this embodiment provides a computer device, which includes a processor 102, a memory, an input device 103, a display 104, and a network interface 105 connected through a system bus 101. Among them, the processor 102 is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 106 and an internal memory 107. The non-volatile storage medium 106 stores an operating system, a computer program, and a database. The internal memory 107 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium 106. When the computer program is executed by the processor 102, it implements the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment described in Embodiment 1.
[0151] Embodiment 4
[0152] As Figure 6 shown, this embodiment provides a storage medium storing a program, which when executed by a processor, implements the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment described in Embodiment 1.
[0153] It should be noted that the computer-readable storage medium of this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0154] In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In this embodiment, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable storage medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0155] The above computer-readable storage medium can be written in one or more programming languages or combinations thereof for executing the computer program of this embodiment. The above programming languages include object-oriented programming languages - such as Java, Python, C++, and also include conventional procedural programming languages - such as C language or similar programming languages. The program can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0156] In summary, the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment of the present invention effectively guides the model to automatically activate the foreground and background and capture local targets with less typical features by using the beta modulation scheme to perform modulation feature processing on image tokens and text tokens. Moreover, it does not need to obtain the modulation function parameters by fitting data, has a flexible modulation shape and low computational cost, can effectively improve the accuracy of target activation, so that the model generates clearer and more accurate pseudo-labels, and effectively improves the segmentation accuracy. In addition, through the cross-modal alignment strategy of the present invention, not only the model's recognition ability for positive categories is enhanced, but also the subtle differences between different categories can be effectively distinguished, thus reducing the occurrence of misjudgments. The selection of negative samples is dynamically adjusted according to the training process, only focusing on the most confusing categories, thus guiding the model to learn key features, improving the accuracy of target recognition, and effectively improving the segmentation performance and the average intersection over union of segmentation.
[0157] The above specific implementation manners are the preferred embodiments of the present invention and cannot limit the present invention. Any other changes or other equivalent replacement manners made without departing from the technical solutions of the present invention are included within the protection scope of the present invention.
Claims
1. A weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment, characterized in that: The following steps are involved: S1, constructing an image data set, preprocessing the image data set, and dividing the preprocessed image data set into a training set and a test set, wherein the image data set includes graphic information, category information, and text prompt information; S2, using the text encoder of the CLIP model to extract the text prompt information to obtain a text tag, and using the VIT model to extract the image data and category information to obtain an image tag and a category tag; then performing linear transformation and position encoding processing on the generated image tag, category tag and text tag respectively to obtain processed image tags, category tags and text tags; S3, constructing a classification network model based on feature modulation and cross-modal alignment, and training the classification network model using a beta modulation mechanism and a cross-modal alignment strategy to obtain a trained classification network model based on feature modulation and cross-modal alignment; the classification model includes a Transformer encoder module improved based on the beta modulation mechanism, a contrastive learning module, and a classification module; S4, processing the processed image tags, category tags and text tags using the trained classification network model based on feature modulation and cross-modal alignment to obtain a category activation map, extracting mapping information between image tags and text tags in the Transformer encoder module improved based on the beta modulation mechanism and adding it to the category activation map to obtain a more refined category activation map, and converting the more refined category activation map into corresponding pseudo labels; S5. Construct an image semantic segmentation network and train it. During the training process of the image semantic segmentation network, use the pseudo labels obtained in step S4 to evaluate the effect of the image semantic segmentation network. Use the optimal image semantic segmentation network model obtained in the training as the trained image semantic segmentation network model, and use the trained image semantic segmentation network model to complete the semantic segmentation of the target image.
2. The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 1, characterized in that: The classification network model training process based on feature modulation and cross-modal alignment described in step S3 is as follows: S31, constructing a classification network model based on feature modulation and cross-modal alignment, wherein the classification model includes a Transformer encoder module improved based on a beta modulation mechanism, a contrastive learning module, and a classification module; the improved Transformer encoder is composed of a plurality of processing sublayers stacked together, and the processing sublayers are composed of parallel blocks and beta modulation units; the classification module includes a category activation map generation unit, a top classification head unit, an auxiliary classification head unit, a text classification loss unit, and an image classification loss unit; S32, processing the training set obtained in step S1 by using step S2 to obtain preprocessed image tags, category tags and text tags, and inputting the preprocessed image tags, category tags and text tags into the classification network model based on feature modulation and cross-modal alignment; S33, the Transformer encoder module improved based on the beta modulation mechanism performs feature superposition and beta feature modulation operations on the preprocessed image label, category label and text label to obtain modulated image label, category label and text label; S34, using a contrastive learning module to perform cross-modal alignment on the modulated image features and text features in step S33, and using InfoNCE contrast loss as a cross-modal alignment loss function in the alignment process to maximize the similarity between positive samples; S35, using the category activation map generation unit in the classification module to perform convolution and reshape processing on the modulated image mark, and convert it into a corresponding category activation map; S36, the top classification head unit obtains the corresponding classification prediction value according to the modulated category mark, and calculates the top classification loss, the auxiliary classification head unit uses the category mark obtained by the third-to-last processing sublayer of the Transformer encoder module improved based on the beta modulation mechanism in step S33 as input to calculate the auxiliary classification loss; the image classification loss unit calculates the modulated image mark by applying global weighted sorting pooling to obtain the image classification loss; the text classification loss module obtains the prediction value based on the text mark by global average pooling, and obtains the text classification loss; S37. Construct a total loss function based on the top classification loss, auxiliary classification loss, image classification loss, text classification loss and cross-modal alignment loss, and jointly train the network model with minimizing the total loss function as the training goal to obtain a trained classification network model.
3. The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 2, characterized in that: The specific process of step S33 is as follows: S331, inputting the pre-processed image tag, category tag and text tag into the Transformer encoder module improved based on the beta modulation mechanism; the improved Transformer encoder is composed of a plurality of processing sub-layers, and the processing sub-layers are composed of parallel blocks and beta modulation units; S332, the pre-processed image mark, category mark and text mark are further feature extracted through each processing sub-layer in the Transformer encoder module improved based on the beta modulation mechanism, and the output of each processing sub-layer is used as the input of the next processing sub-layer; the blocks in the processing sub-layer extract features of the pre-processed image mark, category mark and text mark, and the beta modulation unit in the processing sub-layer performs beta feature modulation operation on the output of the block in the same processing sub-layer to obtain the modulated image mark, category mark and text mark; Specifically, the working process of the beta modulation unit is as follows: Taking the preprocessed image label as an example, the image label output by the block is first normalized on the D-dimensional channel so that the value range of the normalized image label meets the input condition of the Beta distribution probability density function, and then the Beta distribution probability density function is used to modulate the normalized image label to obtain the modulated image label.
4. According to the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 2, the working process of the contrastive learning module in step S34 is as follows: S341, firstly, a category set included in the images in the training set is defined; the categories are specifically as follows: in, C p is the set of all categories in the training set, and C is the category actually existing in the current image; S342, extracting the modulated image features through a convolution operation, and compressing them into a vector Q through global average pooling img , the vector Q img The specific representation is as follows: Q i mg =GAP(Conv(F pat )) Among them, Conu(·) represents the convolution layer, GAP(·) represents the global average pooling operation; S343, projecting the modulated image features to a new matrix U through a linear layer txt ; S344, from the new matrix U txt Selecting a positive sample set and a negative sample set, wherein the text features of the categories that appear in the image data set are used as the positive sample set, and the text features of the categories that do not appear in the image data set are used as the negative sample set; S345, the vector Q img and the matrix U txt Alignment is performed, and in the alignment process, InfoNCE contrast loss is used as the cross-modal alignment loss function to maximize the similarity between positive samples, while minimizing the similarity between negative samples, and calculate the corresponding cross-modal alignment loss; the specific expression of the cross-modal alignment loss function is as follows: Among them, L CA is the cross-modal alignment loss, τ is the temperature coefficient, τ is used to adjust the smoothness of the loss function, U pos is the positive sample set, U neg is the negative sample set.
5. The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 1, characterized in that: The more refined category activation map described in step S4 is expressed as follows: CAM=ReLU(Upsample(M) M=A p2p ·(F′ pat ofA t2p ) Among them, CAM is a more refined category activation map generated, and the size of CAM is H×W×C, where H and W are the length and width of the original image, respectively, and F′ pat is the modulated image feature after convolution and reshape transformation, o is element-wise multiplication.
6. The weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 1, characterized in that: The total loss function is expressed as follows: L=L top +L aux +L pat +L txt +λ CA L CA Among them, L top is the top classification loss, L aux is the auxiliary classification loss, L pat is the image label classification loss, L txt is the text label classification loss, L CA is the cross-modal alignment loss and λ is an empirical parameter.
7. According to the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to claim 1, the image semantic segmentation network adopts the DeeplabV3 network model.
8. A weakly supervised image semantic segmentation system based on feature modulation and cross-modal alignment, used to implement the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment according to any one of claims 1 to 7, characterized in that: include: An image data module, used to preprocess the image data set and divide the preprocessed image data set into a training set and a test set, wherein the image data set includes image information, category information and text prompt information; A text labeling module, used for extracting text prompt information from the image data set using a text encoder of a CLIP model to obtain a text label; An image labeling module, used to extract image information of the image data set using a VIT model to obtain an image label; A category labeling module, used to extract category information of the image data set using a VIT model to obtain a category label; A linear transformation and position encoding module is used to perform linear transformation and position encoding processing on the image label, the category label and the text label respectively to obtain the processed image label, the category label and the text label; Classification network model building module, used to build a classification network model based on feature modulation and cross-modal alignment; A classification network model training module, used to train the classification network model using a beta modulation mechanism and a cross-modal alignment strategy to obtain a trained classification network model based on feature modulation and cross-modal alignment; A pseudo-label generation module processes the processed image labels, category labels and text labels using the trained classification network model based on feature modulation and cross-modal alignment to obtain a category activation map, extracts mapping information between image labels and text labels in the Transformer encoder module improved based on the beta modulation mechanism and adds it to the category activation map to obtain a more refined category activation map, and converts the more refined category activation map into a corresponding pseudo-label; An image semantic segmentation network training module is used to construct and train an image semantic segmentation network, and during the training process of the image semantic segmentation network, pseudo labels of a pseudo label generation module are used to evaluate the effect of the image semantic segmentation network, and the optimal image semantic segmentation network model obtained during the training is used as the trained image semantic segmentation network model; The image semantic segmentation module is used to perform semantic segmentation on the target image using the trained image semantic segmentation network model.
9. A computer device comprising a processor and a memory for storing a program executable by the processor, characterized in that: When the processor executes the program stored in the memory, it implements the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment described in any one of claims 1-7.
10. A storage medium storing a program, characterized in that: When the program is executed by a processor, the weakly supervised image semantic segmentation method based on feature modulation and cross-modal alignment described in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Nerve radiation field plant rendering method and device fused with large language model
CN120411325A
Multi-mode weak supervision small sample semantic segmentation method based on semantic anchoring and double-branch coupling
CN121883855A
A Multimodal Weakly Supervised Small Sample Semantic Segmentation Method Based on Semantic Anchoring and Dual-Branch Coupling
CN121883855B