Weak supervision semantic segmentation method based on channel specific prototype
By adopting channel-specific prototypes and global information fusion branches in weakly supervised semantic segmentation, the problem of error activation of CAMs background areas is solved, and segmentation performance is improved, especially when foreground classes and background classes co-occur frequently.
Patent Information
- Application Number
- CN202510055718.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-23
AI Technical Summary
In the existing multi-stage weakly supervised semantic segmentation methods, initial class activation maps (CAMs) are susceptible to the inaccurate activation of background areas, resulting in inaccurate pseudomasks and degrade final segmentation performance, especially when foreground classes and background classes often co-occur.
Weak supervised semantic segmentation method based on channel-specific prototypes is adopted to generate channel-specific foreground and background prototypes through channel-specific feature extraction modules, and force network to distinguish foreground and background through pixel-to-prototype comparison learning. At the same time, a channel-aware global information fusion branch is introduced, and global information is obtained and prototype-based channel information fusion module is used to obtain and refine the prototype.
It effectively alleviates the problem of error activation of CAMs background areas in weakly supervised semantic segmentation, and improves the quality and final segmentation performance of CAMs, especially when the foreground and background classes co-occur frequently.
Smart Images

Figure CN120032123A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic segmentation, and in particular to a weakly supervised semantic segmentation method based on channel-specific prototypes. Background Art
[0002] Although existing semantic segmentation methods have a variety of model designs and have achieved impressive results, training deep neural networks for semantic segmentation requires a large amount of pixel-level annotations, which requires a lot of manpower and resources. In order to reduce the reliance on dense annotations, researchers have explored the use of low-cost labels for semantic segmentation, such as bounding boxes, lines, points, and image-level labels. Among them, weakly supervised semantic segmentation based on image-level labels is the most cost-effective but also the most difficult.
[0003] Mainstream weakly supervised semantic segmentation methods based on convolutional neural networks (CNNs) follow a multi-stage framework. The first stage trains a classification model on a given dataset to generate initial class activation maps (CAMs), the second stage refines and expands the CAMs to generate pseudo masks, and the third stage uses the generated pseudo masks as labels to train the segmentation network. However, the final segmentation performance of these multi-stage methods is usually highly dependent on the quality of the CAMs generated in the first stage, which are often affected by false activations of background regions. This problem is particularly evident for categories that often co-occur with background elements, such as trains and railways or ships and water. The noise introduced by these false activations tends to accumulate during the pseudo-mask refinement process, resulting in inaccurate pseudo masks and degrading the final segmentation performance. Summary of the invention
[0004] Purpose of the invention: The purpose of the present invention is to provide a weakly supervised semantic segmentation method based on channel-specific prototypes to solve the problem of erroneous activation of background areas of CAMs in multi-stage weakly supervised semantic segmentation, especially for the problem of foreground and background classes that often co-occur.
[0005] Technical solution: The weakly supervised semantic segmentation method based on channel-specific prototypes described in the present invention comprises the following steps: (1) Input the image into the mainstream classification network to extract deep features and generate initial CAMs; (2) The deep features and the initial CAMs are fed into a channel-specific feature extraction module to generate channel-specific foreground prototypes and channel-specific background prototypes, and pixel-to-prototype comparison and prototype-to-prototype comparison are performed to separate foreground features and background features in the feature space; (3) The prototype and deep features are sent to the channel-aware global information fusion branch. After the spatial self-attention weighted module and the prototype-based channel information fusion module, the network obtains global information and uses the global information to refine the prototype. (4) Train the network based on the total loss of the backbone network and the two branches to find the optimal model, and use the trained model to generate the final CAMs; (5) Refine the CAMs to generate pseudo masks to guide the training of the segmentation network and obtain the final segmentation result.
[0006] Furthermore, step (1) includes the following steps: (11) The image is input into the ResNet 50 network to obtain deep features. The deep features are input into the classifier after global average pooling GAP to obtain the prediction score, and then the category score is generated after sigmoid activation. The multi-label classification loss is calculated with the image-level true value label. ; Among them, delete the ReLU layer; (12) The deep features are multiplied by the classifier weights to obtain the initial class activation maps CAMs.
[0007] Furthermore, in step (2), designing a channel-level local information capture branch includes: a channel-specific feature extraction module, including the following steps: (21) Inputting the initial CAMs generated by the network and the deep features extracted by the backbone network into the channel-specific feature extraction module; performing a sigmoid operation on the CAMs, normalizing the values to between 0 and 1, and generating channel-specific foreground soft masks and background soft masks; (22) Applying the channel-specific soft mask to the deep features extracted by the backbone network to obtain channel-specific foreground features and channel-specific background features; (23) A unique set of background feature prototypes is calculated for each image in the batch; assuming that the foreground features corresponding to the same class are similar, a set of foreground prototypes is calculated for the entire batch; (24) Three loss functions are introduced in the channel-level local information capture branch: including: channel-specific foreground to background prototype loss , Overall prototype loss , pixel-to-prototype contrast loss .
[0008] Furthermore, in step (3), the design of the global information fusion branch includes a spatial self-attention weighted module and a prototype-based channel information fusion module; the following steps are included: (31) The spatial self-attention weighted module first applies three parallel 1×1 convolutional layers to the deep features to obtain three feature maps; (32) Perform transposed matrix multiplication between the first two feature maps to calculate the attention score, and then perform softmax normalization to convert the obtained score into a global spatial attention map; (33) Apply the result obtained in step (32) to the third feature map so that each position aggregates spatial information, and then connect it with the original third feature map to obtain a global feature map; (34) A prototype-based channel information fusion module is introduced. First, the channel-specific foreground prototype and channel-specific background prototype of the channel-level local information capture branch are multiplied with the deep features of the backbone network and the corresponding foreground and background context attention weights are obtained using the softmax layer. Secondly, the obtained foreground and background context attention weights are applied to the corresponding prototypes to obtain the foreground and background context prototypes. Finally, the context prototype is connected to the global feature map, and then projected into the category space through a 3×3 convolutional layer to generate the final attention weighted feature map.
[0009] Furthermore, step (4) includes the following steps: (41) For the channel-level local information capture branch, according to the channel-specific foreground to background prototype loss , overall prototype loss and pixel-to-prototype comparison loss , calculate its loss ; (42) For the channel-aware global information fusion branch, the global maximum pooling (GMP) is applied to the attention-weighted feature map to obtain the global category score and calculate the global information fusion loss. ; (43) According to the classification loss , the loss of the channel local information capture branch and channel-aware global information fusion branch loss , calculate the overall loss of the network; (44) The model is trained by constraining the total loss. An evaluation is performed after each iteration during training, and the model with the highest accuracy is saved. During the inference phase, the image is fed into the trained model to generate the final CAMs.
[0010] Furthermore, step (5) includes the following steps: (51) Use IRNet to refine the CAMs generated in step (4) and generate pixel-level pseudo-truth masks; (52) Use the obtained pixel-level pseudo-truth masks and images to train the segmentation network, find the optimal model, and generate the final segmentation result.
[0011] The weakly supervised semantic segmentation system based on channel-specific prototypes described in the present invention comprises: Deep feature module: used to input images into the mainstream classification network to extract deep features and generate initial CAMs; Separation module: used to input deep features and initial CAMs into the channel-specific feature extraction module, generate channel-specific foreground prototypes and channel-specific background prototypes, and perform pixel-to-prototype comparison and prototype-to-prototype comparison to separate foreground features and background features in the feature space; Global information fusion branch module: used to send prototypes and deep features into the global information fusion branch of channel perception. After the spatial self-attention weighted module and the prototype-based channel information fusion module, the network obtains global information and uses the global information to refine the prototype. Optimal module: used to train the network to find the optimal model based on the total loss of the backbone network and the two branches, and use the trained model to generate the final CAMs; Segmentation module: used to refine CAMs to generate pseudo masks, guide the training of segmentation networks, and obtain the final segmentation results.
[0012] An electronic device described in the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, a weakly supervised semantic segmentation method based on a channel-specific prototype according to any one of the items is implemented.
[0013] A storage medium described in the present invention stores a computer program, and when the computer program is executed by a processor, it implements any one of the weakly supervised semantic segmentation methods based on channel-specific prototypes.
[0014] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages: The present invention introduces a channel-level local information capture branch and a channel-aware global information fusion branch; the channel-level local information capture branch uses initial CAMs to guide background feature modeling, generates channel-specific foreground and background prototypes, and performs prototype-to-prototype and pixel-to-prototype contrast learning, thereby forcing the network to distinguish between foreground and background. The channel-aware global information fusion branch enables the network to pay attention to more global spatial information and further refine the prototype. This framework successfully alleviates the problem of incorrect activation of the background area of CAMs generated in the first stage of weakly supervised semantic segmentation, especially for categories that often co-occur with the background, and brings significant improvements to the quality of the overall CAMs and improves the final segmentation performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 It is a schematic diagram of the structure of the present invention; Figure 2 is a network structure diagram of the CAMs generation stage of the present invention; Figure 3 It is a schematic diagram of the channel-level local information capture branch of the present invention; Figure 4It is a schematic diagram of the global information fusion branch of the present invention. DETAILED DESCRIPTION
[0016] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0017] The embodiment of the present invention provides a weakly supervised semantic segmentation method based on a channel-specific prototype, comprising the following steps: S1: Figure 2 As shown, the image is input into the mainstream classification network to extract deep features and generate initial CAMs, including the following steps: S1-1: Given an image x and the corresponding image-level label First, the image is input into a mainstream backbone network (such as ResNet 50) to obtain deep features , where C represents the number of channels, Represents spatial dimensions. Deep features After global average pooling (GAP), the prediction score is sent to the classifier, and then the category score is generated after sigmoid activation, and the multi-label classification loss is calculated with the image-level true value label. The formula is as follows: ; in, is the prediction score after the classification layer, is the image-level truth label obtained from the dataset, is the sigmoid activation function.
[0018] S1-2: Multiply the deep features extracted by the backbone network with the classifier weights to obtain the initial class activation map CAMs: ; in, represents the weight of the classifier (i.e., the fully connected layer of the ResNet network) corresponding to the kth category, Represents the CAM of the k-th category (i.e., the k-th channel of the classification layer output).
[0019] Different from the conventional generation process of obtaining CAMs through ReLU, the present invention deletes the ReLU layer and instead uses the classifier weights to multiply the initial CAMs obtained by deep features to achieve a more thorough application of the activation value.
[0020] S2: Figure 3 As shown in Figure 1, the deep features and initial CAMs are fed into the channel-level local information capture branch, which includes a channel-specific feature extraction module and three loss functions. The channel-specific feature extraction module extracts channel-specific prototypes and is constrained by the loss function at the feature space level.
[0021] S2-1: In order to alleviate the problem of false activation of the background area of CAMs, a channel-level local information capture branch is designed. In the channel-level local information capture branch, the initial CAMs generated by the network and the deep features extracted by the backbone network are first fed into the channel-specific feature extraction module.
[0022] S2-2: Since the initial CAMs of each channel reflect the network's attention to the foreground category corresponding to this channel, a sigmoid operation is performed on the CAMs to normalize the values between 0 and 1 to generate channel-specific foreground soft masks and background soft masks: ; ; in, represents the foreground mask of the kth channel, Represents the background mask of the k-th channel.
[0023] S2-3: Apply channel-specific masks to deep features extracted by the backbone network , get channel-specific foreground features and channel-specific background features : ; ; S2-4: Computing pairwise distances between all feature embeddings is very expensive, so next we compute channel-specific prototypes to guide feature constraints. Since background features tend to vary from image to image, a unique set of background feature prototypes is computed for each image in a batch. However, for foreground prototypes, it is assumed that foreground features corresponding to the same class should be similar, so a set of foreground prototypes is computed for the entire batch. The channel-specific prototypes are computed as follows: ; ; in, Represents the foreground prototype of the kth channel of the current batch, represents the background prototype of the kth channel of the bth image in the current batch, and B represents the batch size.
[0024] S2-5: In order to improve the network's ability to distinguish foreground and background, three loss functions are designed in the channel-level local information capture branch: First, in order to make full use of information from different channels, we explicitly enforce the separation of channel-specific foreground and background prototypes, resulting in a channel-specific foreground-to-background prototype loss : ; in, Represents cosine similarity. It aims to enhance background modeling by more effectively utilizing the often-neglected channel-specific feature information.
[0025] Step 2-6: In order to further combine the interaction between channel features, the foreground and background features are processed from a holistic perspective. The foreground and background should be different, and for the foreground features, different categories should also show differences. Based on this concept, the overall prototype loss is designed as follows: ; in, represents the overall foreground-to-foreground loss, represents the overall foreground to background loss, i and j represent the foreground categories, represents the overall prototype of the foreground of the current batch, Represents the background overall prototype of the bth image in the current batch. The foreground overall prototype and the background overall prototype are obtained by averaging the corresponding channel-specific prototypes in each channel.
[0026] S2-7: Finally, to further enhance the alignment of foreground and background features with their respective prototypes in feature space, a pixel-to-prototype contrast loss is calculated , the formula is as follows: ; in, represents the loss from background pixels to prototypes, represents the loss from foreground pixels to prototypes.
[0027] Under the constraints imposed by the channel-level local information capturing branch, the network exhibits a stronger ability to distinguish foreground and background pixels. This improvement is expected to reduce the false activations of background regions that are highly correlated with the foreground class in the final CAMs.
[0028] S3: The prototype and deep features are fed into the channel-aware global information fusion branch, such as Figure 4 As shown in the figure, the global information fusion branch includes a prototype-based channel information fusion module and a spatial self-attention weighted module, which enables the network to pay attention to more global information and use the global information to refine the prototype.
[0029] The spatial self-attention weighted module computes the spatial self-attention distribution based on the input feature maps. By capturing the global dependencies in these feature maps, it can dynamically adjust the feature responses of each spatial position. It includes the following steps: S3-1: Given the deep feature map extracted from the backbone network , firstly apply three parallel 1×1 convolutional layers to further reduce the computational cost, thus obtaining three feature maps F 1 、F 2 and F 3 .
[0030] S3-2: In order to capture global information, the first two feature maps F 1 and F 2 We perform transposed matrix multiplication between ∠ and ∠ to compute the attention scores, and then perform softmax normalization to transform these scores into a global spatial attention map. : .
[0031] S3-3: Then we will get Applied to the third feature map F 3 , so that F 3 Each position of F aggregates spatial information and then compares it with the original F 3 Concatenate to get the global feature map , as shown below: .
[0032] in, Indicates a connection operation. A feature map representing weighted spatial information.
[0033] To more effectively integrate the previously neglected channel information, the prototype-based channel information fusion module utilizes deep features to provide contextual information for each prototype.
[0034] S3-4: First, the channel-specific foreground prototype and channel-specific background prototype from the channel-level local information capture branch are multiplied with the deep features from the backbone network. The corresponding foreground and background contextual attention weights are obtained using a softmax layer. and : .
[0035] S3-5: These weights are then applied to the prototypes to obtain foreground and background context prototypes respectively and : ; S3-6: Final attention weighted feature map is by passing the context prototype , With global feature map Concatenate and then project it into the category space through a 3×3 convolutional layer: .
[0036] S4: Train the network based on the total loss of the backbone network and the two branches to find the optimal model, and use the trained model to generate the final CAMs. It includes the following steps: S4-1: For the branch capturing local information at the channel level, its total loss Channel-specific foreground-to-background prototype loss , overall prototype loss and pixel-to-prototype comparison loss composition: .
[0037] S4-2: For the channel-aware global information fusion branch, such as Figure 4 As shown, global maximum pooling (GMP) is applied to the attention-weighted feature map To obtain the global category score and calculate the global information fusion loss : , in, represents the standard cross entropy loss, Represents global maximum pooling.
[0038] GMP tends to highlight the regions with the strongest responses in a specific channel. Compared to GAP, GMP better preserves inter-channel differences because it only captures the most prominent responses in each channel instead of averaging all information like GAP. This is crucial for channel-aware fusion because it allows each channel to maintain its specificity and diversity.
[0039] S4-3: Figure 2 As shown, the overall loss of the network is composed of the multi-label classification loss of the backbone network , the loss from the channel local information capture branch and the channel-aware global information fusion branch loss composition: , in and are hyperparameters and are set to 1 and 0.5 respectively.
[0040] S4-4: Train the model using the total loss constraint above, perform an evaluation after each iteration during training, and save the model with the highest accuracy. In the inference phase, feed the image into the trained model to generate the final CAMs.
[0041] S5: Refine CAMs to generate pseudo masks to guide the training of the segmentation network and obtain the final segmentation result. It includes the following steps: S5-1: In the pseudo mask generation stage, IRNet is used as a refinement method to obtain high-quality pseudo ground-truth masks.
[0042] S5-2: In the segmentation stage, following the latest work of BECO, DeeplabV2 with ResNet101 backbone is adopted as the segmentation network to obtain the final segmentation result.
[0043] Experimental results: (A) Improvement of CAMs performance: The present invention can be directly integrated into existing methods. Table 1 shows the CAMs performance comparison obtained by applying the present invention to the established weakly supervised semantic segmentation baseline methods on the PASCAL VOC 2012 training set. As shown in Table 1, the present invention brings improvements in CAMs quality for all five methods.
[0044] Table 1 Quality comparison of CAMs and pseudo-truth masks evaluated on the PASCAL VOC 2012 training set .
[0045] (B) Suppression of background misactivation: Table 2 reports the mIoU when the invention is applied to different baselines, focusing on the background category (background) and two foreground categories that are highly correlated with the background, namely boat and train. The results show that the invention enhances the background category performance of all five baseline models and improves the mIoU of boat and train categories in most methods.
[0046] Table 2 Evaluation results of two specific object categories and background categories on the PASCAL VOC 2012 training set .
[0047] (C) Segmentation performance The segmentation results of the present invention on the PASCAL VOC 2012 validation set (Val) and test set (Test) are reported in Table 3. Obviously, the present invention achieves leading performance among CNN-based methods on the PASCAL VOC 2012 validation set using only image-level labels without any additional supervision, and also achieves competitive performance on the test set.
[0048] Table 3 Segmentation performance (mIoU) evaluation results on the PASCAL VOC 2012 validation set and test set .
[0049] The embodiment of the present invention further provides a weakly supervised semantic segmentation system based on a channel-specific prototype, comprising: Deep feature module: used to input images into the mainstream classification network to extract deep features and generate initial CAMs; Separation module: used to input deep features and initial CAMs into the channel-specific feature extraction module, generate channel-specific foreground prototypes and channel-specific background prototypes, and perform pixel-to-prototype comparison and prototype-to-prototype comparison to separate foreground features and background features in the feature space; Global information fusion branch module: used to send prototypes and deep features into the global information fusion branch of channel perception. After the spatial self-attention weighted module and the prototype-based channel information fusion module, the network obtains global information and uses the global information to refine the prototype. Optimal module: used to train the network to find the optimal model based on the total loss of the backbone network and the two branches, and use the trained model to generate the final CAMs; Segmentation module: used to refine CAMs to generate pseudo masks, guide the training of segmentation networks, and obtain the final segmentation results.
[0050] An embodiment of the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is loaded into the processor, the method for weakly supervised semantic segmentation based on a channel-specific prototype according to any one of the items is implemented.
[0051] An embodiment of the present invention further provides a storage medium storing a computer program, wherein the computer program, when executed by a processor, implements any one of the weakly supervised semantic segmentation methods based on channel-specific prototypes.
Claims
1. A weakly supervised semantic segmentation method based on channel-specific prototypes, characterized in that: The following steps are involved: (1) Input the image into the mainstream classification network to extract deep features and generate initial CAMs; (2) The deep features and initial CAMs are fed into the channel-level local information capture branch to generate channel-specific foreground prototypes and channel-specific background prototypes, and pixel-to-prototype comparison and prototype-to-prototype comparison are performed to separate foreground features and background features in the feature space; (3) The prototype and deep features are sent to the channel-aware global information fusion branch. After the spatial self-attention weighted module and the prototype-based channel information fusion module, the network obtains global information and uses the global information to refine the prototype. (4) Train the network based on the total loss of the backbone network and the two branches to find the optimal model, and use the trained model to generate the final CAMs; (5) Refine the CAMs to generate pseudo masks to guide the training of the segmentation network and obtain the final segmentation result.
2. The weakly supervised semantic segmentation method based on channel-specific prototypes according to claim 1, characterized in that: Step (1) includes the following steps: (11) The image is input into the ResNet 50 network to obtain deep features. The deep features are input into the classifier after global average pooling GAP to obtain the prediction score, and then the category score is generated after sigmoid activation. The multi-label classification loss is calculated with the image-level true value label. ; Among them, delete the ReLU layer; (12) The deep features are multiplied by the classifier weights to obtain the initial class activation maps CAMs.
3. The weakly supervised semantic segmentation method based on channel-specific prototypes according to claim 1, characterized in that: In step (2), the channel-level local information capture branch is designed to include: a channel-specific feature extraction module, including the following steps: (21) Inputting the initial CAMs generated by the network and the deep features extracted by the backbone network into the channel-specific feature extraction module; performing a sigmoid operation on the CAMs, normalizing the values to between 0 and 1, and generating channel-specific foreground soft masks and background soft masks; (22) Applying the channel-specific soft mask to the deep features extracted by the backbone network to obtain channel-specific foreground features and channel-specific background features; (23) A unique set of background feature prototypes is calculated for each image in the batch; assuming that the foreground features corresponding to the same class are similar, a set of foreground prototypes is calculated for the entire batch; (24) Three loss functions are introduced in the channel-level local information capture branch: including: channel-specific foreground to background prototype loss , Overall prototype loss , pixel-to-prototype contrast loss .
4. The weakly supervised semantic segmentation method based on channel-specific prototypes according to claim 1, characterized in that: In step (3), the design of the global information fusion branch includes a spatial self-attention weighted module and a prototype-based channel information fusion module; the following steps are included: (31) The spatial self-attention weighted module first applies three parallel 1×1 convolutional layers to the deep features to obtain three feature maps; (32) Perform transposed matrix multiplication between the first two feature maps to calculate the attention score, and then perform softmax normalization to convert the obtained score into a global spatial attention map; (33) Apply the result obtained in step (32) to the third feature map so that each position aggregates spatial information, and then connect it with the original third feature map to obtain a global feature map; (34) A prototype-based channel information fusion module is introduced. First, the channel-specific foreground prototype and channel-specific background prototype of the channel-level local information capture branch are multiplied with the deep features of the backbone network and the corresponding foreground and background context attention weights are obtained using the softmax layer. Secondly, the obtained foreground and background context attention weights are applied to the corresponding prototypes to obtain the foreground and background context prototypes. Finally, the context prototype is connected to the global feature map, and then projected into the category space through a 3×3 convolutional layer to generate the final attention weighted feature map.
5. The weakly supervised semantic segmentation method based on channel-specific prototypes according to claim 1, characterized in that: Step (4) includes the following steps: (41) For the channel-level local information capture branch, according to the channel-specific foreground to background prototype loss , overall prototype loss and pixel-to-prototype comparison loss , calculate its loss ; (42) For the channel-aware global information fusion branch, the global maximum pooling (GMP) is applied to the attention-weighted feature map to obtain the global category score and calculate the global information fusion loss. ; (43) According to the classification loss , the loss of the channel local information capture branch and channel-aware global information fusion branch loss , calculate the overall loss of the network; (44) The model is trained by constraining the total loss. An evaluation is performed after each iteration during training, and the model with the highest accuracy is saved. During the inference phase, the image is fed into the trained model to generate the final CAMs.
6. The weakly supervised semantic segmentation method based on channel-specific prototypes according to claim 1, characterized in that: Step (5) includes the following steps: (51) Use IRNet to refine the CAMs generated in step (4) and generate pixel-level pseudo-truth masks; (52) Use the obtained pixel-level pseudo-truth masks and images to train the segmentation network, find the optimal model, and generate the final segmentation result.
7. A weakly supervised semantic segmentation system based on channel-specific prototypes, characterized in that include: Deep feature module: used to input images into the mainstream classification network to extract deep features and generate initial CAMs; Separation module: used to input deep features and initial CAMs into the channel-specific feature extraction module, generate channel-specific foreground prototypes and channel-specific background prototypes, and perform pixel-to-prototype comparison and prototype-to-prototype comparison to separate foreground features and background features in the feature space; Global information fusion branch module: used to send prototypes and deep features into the global information fusion branch of channel perception. After the spatial self-attention weighted module and the prototype-based channel information fusion module, the network obtains global information and uses the global information to refine the prototype. Optimal module: used to train the network to find the optimal model based on the total loss of the backbone network and the two branches, and use the trained model to generate the final CAMs; Segmentation module: used to refine CAMs to generate pseudo masks, guide the training of segmentation networks, and obtain the final segmentation results.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into a processor, the method for weakly supervised semantic segmentation based on channel-specific prototypes according to any one of claims 1 to 6 is implemented.
9. A storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, a weakly supervised semantic segmentation method based on a channel-specific prototype according to any one of claims 1 to 6 is implemented.