Weakly supervised semantic segmentation method and device based on commonality-specific supervision mechanism

By combining convolutional modules and commonality-specificity supervision modules, the problems of sparse activation regions and blurred boundaries in weakly supervised semantic segmentation at the image level are solved, achieving more accurate target object segmentation, which is suitable for applications such as medical lesion segmentation.

CN117036683BActive Publication Date: 2025-11-28ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310388689.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-13
Publication Date
2025-11-28
Estimated Expiration
2043-04-13

AI Technical Summary

Technical Problem

In existing image-level weakly supervised semantic segmentation methods, the class activation map method suffers from the problem of sparse activation regions for erroneous negative and positive samples, resulting in sparse localization regions and blurred segmentation boundaries, which limits performance.

Method used

We adopt a weakly supervised semantic segmentation method based on the commonality-specificity supervision mechanism. The method enhances the internal structure distribution of images by contrastive convolutional modules, mines similar structural distributions between images by commonality-specificity supervision modules, and uses knowledge gap modules to construct contrastive generation images to overcome incomplete activation correspondence.

Benefits of technology

It improves the problems of sparse localization regions and blurred segmentation boundaries, and enhances the performance of image-level weakly supervised semantic segmentation, especially in applications such as medical lesion segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117036683B_ABST
    Figure CN117036683B_ABST
Patent Text Reader

Abstract

The application discloses a weakly supervised semantic segmentation method and device based on commonality-specificity supervision mechanism, establishes a contrast convolution module, utilizes the convolution cognitive difference of different receptive fields in an image to identify the boundary area with ambiguity in the image, and overcomes the problem of fuzzy segmentation boundary in the weakly supervised semantic segmentation task; a commonality-specificity supervision module is established, a commonality supervision mechanism is utilized to find the similar structural background distribution between different category images, a specificity supervision mechanism is utilized to identify the prominent area in the image distribution, semantic segmentation of a target object is realized, the positioning area is improved to be sparse, and the segmentation boundary is optimized; a knowledge gap module constructs a structure distribution enhanced contrast generated image, and the knowledge gap between the contrast generated image and the category image effectively overcomes the incomplete activation corresponding relationship in mainstream methods, and improves the image level weakly supervised semantic segmentation performance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image-level weakly supervised semantic segmentation, and particularly relates to a weakly supervised semantic segmentation method and device based on commonality-specificity supervision mechanism. BACKGROUND

[0002] In recent years, with the development of large-scale deep learning networks and a large number of pixel-level semantic annotations, semantic segmentation has achieved great success in various real-world applications, such as autonomous driving, robotics and medical diagnosis. However, these models rely heavily on a large number of pixel-level annotations, which require high-intensity human labor. On the contrary, some weakly supervised annotations, such as image-level labels, points, scribbles and bounding boxes, are easily available. Therefore, it is extremely attractive to explore the potential of weakly supervised annotations in the task of semantic segmentation.

[0003] It is extremely challenging to solve the problem of image-level weakly supervised semantic segmentation, because image-level annotations can only indicate whether the target object exists in an image, but lack the necessary location information. To solve this problem, the mainstream method mainly uses class activation maps to give convolutional networks the ability to locate, such as the causal intervention-based method C-CAM (Zhang, Dong, et al. "Causal intervention for weakly-supervised semantic segmentation." Advances in Neural Information Processing Systems 33 (2020): 655-666) and the region semantic-based method RCA (Zhou, Tianfei, et al. "Regional semantic contrast and aggregation for weakly supervised semantic segmentation." Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2022).

[0004] However, the above class activation map method can only identify the most discriminative region in the image, which leads to two main problems. One is the false negative example, the activation region is often sparse, and the activation of the wrong target object region as the background region; the other is the false positive example, the activation region is overflow, and the activation of the wrong background region as the target object. The incomplete activation correspondence limits the performance of the class activation map method, causing serious positioning region sparsity and segmentation boundary ambiguity. Some recent work is trying to use different network frameworks or different training strategies to solve the problem of incomplete activation correspondence.

[0005] Therefore, exploring an image-level weakly supervised semantic segmentation method to avoid the problem of incomplete activation correspondence in class activation map and improve the positioning ability and boundary segmentation ability of weakly supervised labels for semantic segmentation has become a technical problem to be solved. SUMMARY

[0006] In view of the above, the purpose of the present application is to provide a weakly supervised semantic segmentation method and device based on common-specific supervision mechanism, which overcomes the technical defects of incomplete activation correspondence in class activation map, to improve the positioning ability and accuracy of weakly supervised labels for semantic segmentation.

[0007] To achieve the above application purpose, the present application provides the following scheme:

[0008] In a first aspect, the embodiment provides a weakly supervised semantic segmentation method based on common-specific supervision mechanism, comprising the following steps:

[0009] Establishing a class 1 data set and a class 2 data set, the class 1 data set containing class 1 images and their image-level labels, and the class 2 data set containing class 2 images and their image-level labels;

[0010] Establishing a weakly supervised semantic segmentation model, the weakly supervised semantic segmentation model comprising an embedding layer, a contrastive convolution module, a common-specific supervision module, a generator, a discriminator and a knowledge gap module, the embedding layer being used for spatial mapping of the class 1 images and the class 2 images to obtain embedded representations, the contrastive convolution module being used for spatial enhancement of the embedded representations to obtain enhanced distribution representations, the common-specific supervision module being used for constructing a common supervision map based on the enhanced distribution representations using a common supervision mechanism, constructing a specific supervision map based on the common supervision map using a specific supervision, and constructing a specific class target object region based on the common supervision map and the specific supervision map, the generator being used for generating a contrastive generated image based on the target object region, the discriminator being used for discriminating the authenticity of the contrastive generated image, and the knowledge gap module being used for calculating a semantic segmentation result according to the class 1 images and their corresponding contrastive generated images;

[0011] An objective function for establishing a weakly supervised semantic segmentation model, the objective function including an adversarial loss for training of a generator and a discriminator and a consistency loss for constructing a consistency of a structure of a contrastive generated image and a class 1 image based on a semantic segmentation result;

[0012] A category 1 dataset and a category 2 dataset are adopted, and parameters of the weakly supervised semantic segmentation model are optimized by using the objective function;

[0013] The weakly supervised semantic segmentation model after the parameter optimization is used to segment a target image to be detected to obtain a semantic segmentation label of the target image at a pixel level.

[0014] Preferably, the embedding layer includes a boundary padding layer, a two-dimensional convolution layer, an instance regularization layer and a linear rectifier activation layer connected in sequence, and the embedding layer is used to spatially map the category 1 image and the category 2 image to obtain an embedded representation Embedding1 of the category 1 image and an embedded representation Embedding2 of the category 2 image.

[0015] Preferably, the contrastive convolution module includes a double-channel mode, wherein a first channel includes a two-dimensional convolution layer and a linear rectifier activation layer connected in sequence, and is used to extract a corresponding standard local representation S_Embedding1 from the embedded representation Embedding1 of the category 1 image and a corresponding standard local representation S_Embedding2 from the embedded representation Embedding2 of the category 2 image;

[0016] A second channel includes a contrastive convolution, and the contrastive convolution includes an extended convolution layer and a two-dimensional convolution layer, and is used to extract a corresponding difference representation D_Embedding1 from the embedded representation Embedding1 of the category 1 image and a corresponding difference representation D_Embedding2 from the embedded representation Embedding2 of the category 2 image;

[0017] The contrastive convolution module further includes a class activation map calculation operation and an enhanced representation calculation operation, specifically:

[0018] The difference representation D_Embedding1 and the difference representation D_Embedding2 are used to calculate a class activation map M1 corresponding to the category 1 image ca and a class activation map M2 corresponding to the category 2 image ca ;

[0019] The class activation map M1 ca is dot multiplied with the standard local representation S_Embedding1 to obtain an enhanced distribution representation E_Embedding1 of the category 1 image, and the class activation map M2 caThe enhanced distribution representation E_Embedding2 of the class 2 image is obtained by dot product with the standard local representation S_Embedding2.

[0020] Preferably, in the commonality-specific supervision module, the commonality supervision mechanism is adopted to construct a commonality supervision graph based on the enhanced distribution representation, comprising:

[0021] The enhanced distribution representation E_Embedding1 of the class 1 image is projected to a Reshape layer for size adjustment to obtain a reshaped distribution E_Embedding1 of the class 1 image. re ;

[0022] The enhanced distribution representation E_Embedding2 of the class 2 image is projected to an average buffer, and the average value of the enhanced distribution representation E_Embedding2 arranged in sequence is calculated to obtain a mean enhanced distribution representation E_Embedding2 of the class 2 image. ave ;

[0023] The enhanced distribution representation E_Embedding2 ave is projected to an SE layer to extract key structural features.

[0024] According to the E_Embedding2 ave of the class 2 image and the reshaped distribution E_Embedding1 re of the class 1 image, an element correlation matrix R of the class 1 image and the class 2 image is calculated.

[0025] Based on the element correlation matrix R and the reshaped distribution E_Embedding1 re , a commonality supervision graph M c is calculated.

[0026] Preferably, in the commonality-specific supervision module, the specificity supervision is constructed based on the commonality supervision graph, comprising:

[0027] The commonality supervision graph M c is reversely mapped to obtain a reverse mapping graph M c of the commonality supervision graph M c ′ , and the calculation process is as follows:

[0028]

[0029] Based on the reverse mapping graph M c ′ , a specificity supervision graph M s is calculated, and the calculation process is as follows:

[0030] M sReshape -1 (E_Embedding1 re )×M c ′

[0031] wherein, softmax represents a softmax activation function, and Reshape represents a reshape function.

[0032] Preferably, in the commonality-specificity supervision module, the target object region is constructed based on the commonality supervision graph and the specificity supervision graph, and the construction comprises:

[0033] The commonality supervision graph M c is added to the embedded representation Embedding1 of the class 1 image and the embedded representation Embedding2 of the class 2 image, to obtain the class 1 image general structure enhanced representation and the class 2 image general structure enhanced representation, and the specificity supervision graph M s is used to filter the specific class 1 target object region F1 in the class 1 image general structure enhanced representation cs and the specific class 2 target object region F2 in the class 2 image general structure enhanced representation cs , and the calculation process is as follows:

[0034] F1 cs =Embedding1+M c -M s

[0035] F2 cs =Embedding2+M c -M s .

[0036] Preferably, in the knowledge gap module, the semantic segmentation result is calculated according to the class 1 image and the corresponding contrast generated image, and the calculation comprises:

[0037] The difference between the class 1 image and the corresponding contrast generated image is taken as the semantic segmentation result.

[0038] The adversarial loss L adv is expressed as:

[0039]

[0040] The consistency loss is expressed as:

[0041] L cons =up(M c )·|Seg1|

[0042] wherein, G represents a generator, and D represents a discriminator. an element value in the (i, j) position of a specificity class 2 target object region F2 corresponding to a class 2 image cs an element value in the (i, j) position of a specificity class 2 target object region F2 corresponding to a class 2 image an element value in the (i, j) position of a specificity class 2 target object region F2 corresponding to a class 2 image cs an element value in the (i, j) position of a specificity class 2 target object region F2 corresponding to a class 2 image c sampling M into the size of the class 1 image c representing a common supervision graph.

[0043] In a second aspect, the embodiments further provide a weakly supervised semantic segmentation device based on a common-specificity supervision mechanism, comprising:

[0044] a data set construction module configured to construct a class 1 data set and a class 2 data set, wherein the class 1 data set contains class 1 images and their image-level labels, and the class 2 data set contains class 2 images and their image-level labels;

[0045] a model construction module configured to construct a weakly supervised semantic segmentation model, wherein the weakly supervised semantic segmentation model comprises an embedding layer, a contrastive convolution module, a common-specificity supervision module, a generator, a discriminator, and a knowledge gap module, the embedding layer is configured to spatially map the class 1 images and the class 2 images to obtain embedded representations, the contrastive convolution module is configured to spatially enhance the embedded representations to obtain enhanced distribution representations, the common-specificity supervision module is configured to construct a common supervision graph based on the enhanced distribution representations using a common supervision mechanism, construct a specificity supervision graph based on the common supervision graph using a specificity supervision, and construct a specific class target object region based on the common supervision graph and the specificity supervision graph, the generator is configured to generate a contrastive generated image based on the target object region, the discriminator is configured to determine the authenticity of the contrastive generated image, and the knowledge gap module is configured to calculate a semantic segmentation result according to the class 1 images and their corresponding contrastive generated images;

[0046] a target function construction module configured to construct a target function of the weakly supervised semantic segmentation model, wherein the target function contains an adversarial loss for training the generator and the discriminator, and a consistency loss for constructing a consistency between the contrastive generated image and the class 1 image based on the semantic segmentation result;

[0047] a parameter optimization module configured to optimize the parameters of the weakly supervised semantic segmentation model using the class 1 data set and the class 2 data set and the target function;

[0048] a detection module configured to segment a target image to be detected using the weakly supervised semantic segmentation model with the optimized parameters to obtain a semantic segmentation label of the target image at a pixel level.

[0049] In a third aspect, the embodiments further provide a computing device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the weakly supervised semantic segmentation method based on the commonality-specific supervision mechanism when executing the computer program.

[0050] In a fourth aspect, the embodiments further provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the weakly supervised semantic segmentation method based on the commonality-specific supervision mechanism.

[0051] Compared with the prior art, the present application has at least the following beneficial effects:

[0052] Firstly, the contrast convolution module is established to identify the ambiguous boundary region in the image by using the convolution cognitive difference of different receptive fields in the image, thereby overcoming the problem of fuzzy segmentation boundary in the weakly supervised semantic segmentation task; secondly, the commonality-specific supervision module is established to discover the similar structural background distribution between different categories of images by using the commonality supervision mechanism and to identify the prominent region in the image distribution by using the specificity supervision mechanism, thereby realizing the semantic segmentation of the target object, which not only improves the positioning region sparsity, but also optimizes the segmentation boundary; finally, the knowledge gap module inputs the generator of the image whose intra-image distribution and inter-image similar structural distribution are enhanced, constructs the contrast generated image with structural distribution enhancement, and effectively overcomes the incomplete activation correspondence in the mainstream method by using the knowledge gap between the contrast generated image and the category image, thereby improving the image-level weakly supervised semantic segmentation performance. BRIEF DESCRIPTION OF DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0054] Figure 1 FIG. 1 is a flowchart of a weakly supervised semantic segmentation method based on a commonality-specific supervision mechanism in an embodiment;

[0055] Figure 2 FIG. 2 is a schematic diagram of a weakly supervised semantic segmentation model framework in an embodiment;

[0056] Figure 3 FIG. 3 is a schematic diagram of a contrast convolution module of a weakly supervised semantic segmentation model in an embodiment;

[0057] Figure 4Fig. 1 is a schematic diagram of a common-specific supervision module of a weakly supervised semantic segmentation model in an embodiment.

[0058] Figure 5 Fig. 2 is a structural schematic diagram of a weakly supervised semantic segmentation device based on a common-specific supervision mechanism in an embodiment. DETAILED DESCRIPTION

[0059] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be made to the present application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the protection scope of the present application.

[0060] Current mainstream weakly supervised semantic segmentation methods mainly use image-level class labels as supervision signals, and class activation maps (CAM) as the main positioning area of the target object. However, these class activation map methods can only identify the most discriminative regions in the image, which leads to two main problems: one is the false negative examples, the activated regions are often sparse, and the activated wrong target object regions are taken as background regions; the other is the false positive examples, the activated regions are overflowed, and the activated wrong background regions are taken as target objects. The incomplete activation correspondence limits the performance of the class activation map method, causing serious positioning region sparsity and segmentation boundary ambiguity.

[0061] To solve the above problems, the embodiment of the present application provides a weakly supervised semantic segmentation method and device based on a common-specific supervision mechanism, which aims to remove ambiguous boundary regions by comparing and enhancing the internal structure distribution of the image, overcome the boundary ambiguity problem caused by the overflow of the activated region, and reduce the false positive examples; through the common-specific supervision mechanism, the similar structure distribution between images is mined, and the specific target segmentation region between images is separated, which strengthens the common-specific distribution pattern between images, overcomes the positioning problem caused by the sparsity of the activated region, avoids the false negative examples, and finally realizes the weakly supervised semantic segmentation with the help of the knowledge gap module between different categories of images. The method and device can be applied to medical lesion segmentation and other applications.

[0062] As shown in Figure 1 Fig. 1, the weakly supervised semantic segmentation method based on the common-specific supervision mechanism provided by the embodiment includes the following steps:

[0063] Step 1, a class 1 dataset with image-level annotation and a class 2 dataset with image-level annotation are established.

[0064] In an embodiment, the category 1 dataset includes category 1 images and image-level labels of the category 1 images, and the category 2 dataset includes category 2 images and image-level labels of the category 2 images. The background structure distribution of the category 1 images and the category 2 images often has similar structure distribution, but the specific classification has obvious distinction.

[0065] Step 2, establishing a weakly supervised semantic segmentation model.

[0066] As shown in Figure 2 , the weakly supervised semantic segmentation model constructed includes an embedding layer, a contrastive convolution module, a commonality-specificity supervision module, a generator, a discriminator, and a knowledge gap module.

[0067] In an embodiment, the embedding layer is used for embedding space mapping of the category 1 images and the category 2 images, and specifically maps the category 1 images and the category 2 images to obtain embedding representations. The embedding layer uses but is not limited to the following network structure, and an available embedding layer example is provided below, including a boundary padding layer (ReflectionPad2d()), a two-dimensional convolution layer (Conv2d()), an instance normalization layer (InstanceNorm2d()), and a linear rectifier activation layer (ReLU()) connected in sequence. The embedding layer projects the input category 1 images and category 2 images into an embedding representation space to obtain the embedding representation Embedding1 of the category 1 images and the embedding representation Embedding2 of the category 2 images.

[0068] The contrastive convolution module is used for feature enhancement in the representation space, and specifically enhances the embedding representation in the space to obtain an enhanced distribution representation. As shown in Figure 2 and 3 , the contrastive convolution adopts a double-channel mode, wherein the first channel includes a two-dimensional convolution layer (Conv2d()) and a linear rectifier activation layer (ReLU()) connected in sequence, which is used to extract the corresponding standard local representation S_Embedding1 from the embedding representation Embedding1 of the category 1 images and the corresponding standard local representation S_Embedding2 from the embedding representation Embedding2 of the category 2 images. The calculation process is as follows:

[0069]

[0070]

[0071] wherein k represents the side length of the two-dimensional convolution kernel, Embedding1 i,j represents the input element of Embedding1 at the (i, j) position, represents the weight parameter of the first channel at the k x k two-dimensional convolution kernel (p, q) position, represents the Embedding1 representation of all input elements Embedding1 within a k x k local range i,j with the corresponding position weight , and a weighted sum of represents the Embedding2 representation of all input elements Embedding2 within a k x k local range i,j with the corresponding position weight , and a weighted sum of max() represents the maximum value selection function.

[0072] The second channel includes a contrastive convolution (C-Conv) for calculating the internal structure distribution difference of the image, the contrastive convolution (C-Conv) includes a dilated convolution layer (D-Conv2d()) and a two-dimensional convolution layer (Conv2d()), according to the embedding representation Embedding1 of the class 1 image, the corresponding difference table D_Embedding1 is extracted, and according to the embedding representation Embedding2 of the class 2 image, the corresponding difference representation D_Embedding2 is extracted, and the calculation process is:

[0073]

[0074]

[0075] wherein, represents the weight parameter of the second channel at the k x k dimension dilated convolution kernel (s, t) position, represents the weight parameter of the second channel at the k x k dimension two-dimensional convolution kernel (p, q) position.

[0076] The contrastive convolution module further includes a class activation map calculation operation and an enhanced representation calculation operation, specifically: using the difference representation D_Embedding1 and the difference representation D_Embedding2 to calculate the class activation map M1 ca corresponding to the class 1 image and the class activation map M2 ca corresponding to the class 2 image, respectively, and the calculation process is:

[0077]

[0078]

[0079]

[0080]

[0081] wherein, LS_Embedding1 i,j represents the Embedding1 representation of all input elements Embedding1 within a k x k local rangei,j Weights at corresponding positions The weighted summation is used to represent the input element Embedding1. i,j Local representation, LS_Embedding2 i,j This represents all input elements Embedding2 within a local range of k×k in the second channel. i,j Weights at corresponding positions The weighted summation is used to represent the input element Embedding2. i,j The local representation of the class is λ, which is a hyperparameter used to measure the degree of class activation.

[0082] Activate category map M1 ca The dot product of the standard local representation S_Embedding1 and the standard local representation S_Embedding1 yields the enhanced distribution representation E_Embedding1 of the class 1 image, which in turn generates the class activation map M2. ca The enhanced distribution representation E_Embedding2 of the class 2 image is obtained by dot product of the standard local representation S_Embedding2. The calculation process is as follows:

[0083]

[0084] E_Embedding1∈R 1×C×h×w

[0085]

[0086] E_Embedding2∈R 1×C×h×w

[0087] in, S_Embedding1 represents the standard local representation at position (i,j). i,j The corresponding category activation value, S_Embedding2 represents the standard local representation at position (i,j). i,j The corresponding category activation value, R 1×C×h×w The dimension of E_Embedding1 (or E_Embedding2) is represented by C, the number of channels in the two-dimensional convolutional layer is represented by C, and h×w represents the height and width represented by E_Embedding1.

[0088] In this embodiment, the commonality-specificity supervision module is used for commonality representation and specificity representation mapping. It includes a commonality supervision mechanism and a specificity supervision mechanism. Specifically, the commonality supervision mechanism is used to construct a commonality supervision graph based on the enhanced distribution representation, and the specificity supervision is used to construct a specificity supervision graph based on the commonality supervision graph. Based on the commonality supervision graph and the specificity supervision graph, a special category target object region is constructed.

[0089] Specifically, as shown in Figure 4 , a common supervision mechanism is adopted to construct a common supervision graph based on the enhanced distribution representation, including:

[0090] (a) The enhanced distribution representation E_Embedding1 of the class 1 image is projected to the Reshape layer for size adjustment, and the E_Embedding1 representation with the dimension of 1×C×h×w is mapped to the dimension of C×hw to obtain the reshaped distribution E_Embedding1 of the class 1 image. re The calculation process is as follows:

[0091] E_Embedding1 re =Reshape(E_Embedding1)

[0092] E_Embedding1 re ∈R C×hw

[0093] Where Reshape() is a dimension mapping function, which is used to map the representation dimension to the R C×hw dimension.

[0094] (b) The enhanced distribution representation E_Embedding2 of the class 2 image is projected to the average buffer, and the average value of the enhanced distribution representation E_Embedding2 is calculated after being sequentially arranged to obtain the average enhanced distribution representation E_Embedding2 of the class 2 image. ave Specifically, it includes:

[0095] The enhanced distribution representation E_Embedding2 is sequentially arranged to form an enhanced distribution representation Butter_Embedding with the dimension of n×C×h×w. If n≤N, the average value of the enhanced distribution representation Butter_Embedding with the dimension of n×C×h×w is directly calculated to obtain the average enhanced distribution representation E_Embedding2. ave If n>N, the newly input enhanced distribution representation E_Embedding2 ′ randomly replaces one enhanced distribution representation E_Embedding2 with the dimension of 1×C×h×w in the existing enhanced distribution representation Butter_Embedding with the dimension of n×C×h×w. m Then the average value of the replaced enhanced distribution representation with the dimension of N×C×h×w is calculated to obtain the average enhanced distribution representation E_Embedding2. ave The calculation process is as follows:

[0096]

[0097] Wherein, the average buffer calculates the mean value of the enhanced distribution representation process occurs in the weakly supervised semantic segmentation method iterative training parameter model. There are new input class 2 image through the embedding layer and contrast convolution module to calculate the new class 2 image enhanced distribution representation E_Embedding2 ′ N is the representation space of the average buffer, preferably, N is set to 3. m is a random number in the range of 1 to N integer. n is the number of enhanced distribution representation Butter_Embedding arranged in sequence. E_Embedding2 l represents the lth enhanced distribution representation in the average buffer.

[0098] (c) the enhanced distribution representation E_Embedding2 ave is projected to the SE layer to extract the key Cxhw dimensional structure features E_Embedding2 struct , the calculation process is:

[0099] E_Embedding2 struct =SE(E_Embedding2 ave )

[0100] E_Embedding2 struct ∈R C×hw

[0101] (d) according to the E_Embedding2 ave of the class 2 image and the reorganization distribution E_Embedding1 re of the class 1 image, the element correlation matrix R of the class 1 image and the class 2 image is calculated, and the calculation process is:

[0102] R=softmax((E_Embedding2 struct ) T ×E_Embedding1 re )

[0103] R∈R hw×hw

[0104] Wherein, softmax represents the normalization exponential function, and the element features with high correlation between the class 1 image and the class 2 image are activated.

[0105] (e) based on the element correlation matrix R and the reorganization distribution E_Embedding1 re , the common supervision graph M c is calculated, and the calculation process is:

[0106] M c =Reshape -1(E_Embedding1 re ×R)

[0107] M c ∈R c×h×w

[0108] Among them, Reshape -1 This indicates that E_Embedding1 of dimension c×hw is used. re ×R is mapped to the c×h×w dimension.

[0109] Specifically, such as Figure 4 As shown, a specific supervision graph is constructed based on the commonality supervision graph, including:

[0110] First, the common supervision graph M c Reverse mapping yields the commonality supervision graph M. c The reverse mapping graph M c ′ The calculation process is as follows:

[0111]

[0112] Then, based on the reverse mapping graph M c ′ Calculate the specificity supervision map M s The calculation process is as follows:

[0113] M s =Reshape -1 (E_Embedding1 re )×M c ′

[0114] Where softmax represents the softmax activation function and Reshape represents the reshape function.

[0115] Specifically, regions for target objects of specific categories are constructed based on commonality supervision maps and specificity supervision maps, including:

[0116] During the training of a weakly supervised semantic segmentation model, the common supervision graph M... c The embedded representations Embedding1 and Embedding2 of the category 1 image are added to obtain the general structure-enhanced representations of the category 1 image and the general structure-enhanced representations of the category 2 image, respectively, and the specificity supervision map M. s F1 filtering was used to filter specific category 1 target object regions in the general structure enhancement representation of category 1 images. cs In the representation of general structure enhancement of category 2 images, specific category 2 target object regions F2 cs, the calculation process is:

[0117] F1 cs = Embedding1 + M c -M s

[0118] F2 cs = Embedding2 + M c -M s .

[0119] In the embodiment, the generator is configured to generate the contrast generated image, and the discriminator is configured to determine the authenticity of the contrast generated image. The generator and the discriminator form an adversarial framework, which can adopt any structure. Optionally, the generator and the discriminator adopt the basic framework of the CycleGAN network. The generator G is configured to generate the contrast generated image I1 C of the category 1 image.

[0120] I1 C = G (F1 cs )

[0121] The generated contrast generated image I1 C is a generated image for driving the structural distribution of the category 1 image to the structural distribution conversion of the category 2 image.

[0122] In the embodiment, the knowledge gap module is configured to perform object segmentation, and the semantic segmentation result Seg1 is calculated according to the category 1 image I1 and the corresponding contrast generated image, which is expressed by the formula as follows:

[0123] Seg1 = I1 - I1 C

[0124] Step 3, establishing a target function of the weakly supervised semantic segmentation model.

[0125] In the embodiment, the constructed target function includes an adversarial loss for training the generator and the discriminator and a consistency loss for constructing the consistency of the structure of the contrast generated image and the category 1 image based on the semantic segmentation result. Specifically, the adversarial loss and the consistency loss are weighted and summed to form the target function of the weakly supervised semantic segmentation model. Preferably, the weight of each loss is 1. The following will be described in detail for each loss.

[0126] The adversarial loss of the generator G and the discriminator D is denoted as L adv , and the calculation process is as follows:

[0127]

[0128] The consistency loss is denoted as L cons , and the calculation process is as follows:

[0129] L cons =up(M c )·|Seg1|

[0130] wherein D denotes a discriminator, denotes the element value of F1 cs in the (i,j) position, denotes the element value of F2 cs in the (i,j) position, and up() denotes sampling M c into the size of the category 1 image, and the consistency loss realizes keeping as much as possible unchanged for the common structure part.

[0131] Step 4: inputting the category 1 dataset and the category 2 dataset into the weakly supervised semantic segmentation model, optimizing the parameters of the weakly supervised semantic segmentation model by using the objective function, and obtaining the weakly supervised semantic segmentation model with optimized parameters.

[0132] In the embodiment, when the weakly supervised semantic segmentation model is parameter-optimized, the parameters of the discriminator model are fixed, the corresponding generator model parameter gradient of the adversarial loss L adv is calculated, the corresponding generator model parameter gradient of the consistency loss L cons is calculated, the parameters of the generator model are updated according to the parameter gradient, the parameters of the generator model are fixed, the corresponding discriminator model parameter gradient of the adversarial loss L adv is calculated, and the parameters of the discriminator model are updated according to the parameter gradient.

[0133] Step 5: segmenting the target image to be detected by using the weakly supervised semantic segmentation model with optimized parameters, and obtaining the semantic segmentation label of the target image at the pixel level.

[0134] After the training is completed, the weakly supervised semantic segmentation model with optimized parameters can be used to perform a semantic segmentation task. A selected target image to be segmented can be input into the weakly supervised semantic segmentation model, and the semantic segmentation result of the target image at the pixel level is obtained by calculation as a semantic label.

[0135] The weakly supervised semantic segmentation method based on the commonness-specificity supervision mechanism provided by the embodiment is implemented based on a class label supervision signal at an image level, and the weak supervision signal at the image level is the most easily obtained in daily applications. Among them, the embodiment proposes a contrast convolution module, which identifies ambiguous boundary regions in an image by using the convolution cognitive difference of different receptive fields in the image, and overcomes the problem of fuzzy segmentation boundary in the weakly supervised semantic segmentation task; then, a commonness-specificity supervision module is designed, which uses a commonness supervision mechanism to find similar structural background distribution between different class images, and uses a specificity supervision mechanism to identify prominent regions in the image distribution, to realize semantic segmentation of the target object, which not only improves the sparseness of the positioning region, but also optimizes the segmentation boundary; finally, the proposed knowledge gap module enhances the image intra-distribution and image inter-similarity structural distribution image input generator, constructs a structural distribution enhanced contrast generated image, and the knowledge gap between the contrast generated image and the class image effectively overcomes the incomplete activation correspondence in the mainstream method, and improves the performance of the weakly supervised semantic segmentation at the image level.

[0136] Based on the same inventive concept, the embodiment further provides a weakly supervised semantic segmentation device based on a commonness-specificity supervision mechanism, as shown in Figure 5 The unsupervised domain adaptation device 500 comprises:

[0137] A data set construction module 510 is configured to establish a class 1 data set with image level labeling and a class 2 data set with image level labeling.

[0138] A model construction module 520 is configured to establish a weakly supervised semantic segmentation model.

[0139] A target function construction module 530 is configured to establish a target function of the weakly supervised semantic segmentation model.

[0140] A parameter optimization module 540 is configured to input the class 1 data set and the class 2 data set into the weakly supervised semantic segmentation model, optimize the parameters of the weakly supervised semantic segmentation model by using the target function, and obtain a parameter-optimized weakly supervised semantic segmentation model.

[0141] A detection module 550 is configured to segment a target image to be detected by using the parameter-optimized weakly supervised semantic segmentation model, and obtain a semantic segmentation label at a pixel level of the target image.

[0142] It should be noted that the weakly supervised semantic segmentation device based on the common-specific supervision mechanism provided by the embodiment should be illustrated by the above-mentioned division of each functional module when performing the image-level weakly supervised semantic segmentation learning and application process. The above-mentioned functions can be completed by different functional modules according to the needs, that is, the internal structure of the terminal or server is divided into different functional modules to complete all or part of the above-described functions. In addition, the weakly supervised semantic segmentation device and the weakly supervised semantic segmentation method embodiment provided by the embodiment belong to the same concept, and the specific implementation process is detailed in the weakly supervised semantic segmentation method based on the common-specific supervision mechanism. Embodiment, which will not be repeated here.

[0143] Based on the same inventive concept, the embodiment further provides a weakly supervised semantic segmentation system based on a common-specific supervision mechanism, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-mentioned weakly supervised semantic segmentation method based on the common-specific supervision mechanism is realized, and specifically comprises:

[0144] Step 1, establishing a class 1 data set with image-level annotation and a class 2 data set with image-level annotation;

[0145] Step 2, establishing a weakly supervised semantic segmentation model;

[0146] Step 3, establishing a target function of the weakly supervised semantic segmentation model;

[0147] Step 4, inputting the class 1 data set and the class 2 data set into the weakly supervised semantic segmentation model, optimizing the parameters of the weakly supervised semantic segmentation model using the target function, and obtaining the weakly supervised semantic segmentation model after parameter optimization;

[0148] Step 5, using the weakly supervised semantic segmentation model after parameter optimization to segment the target image to be detected, and obtaining the pixel-level semantic segmentation annotation of the target image.

[0149] Based on the same inventive concept, the embodiment further provides a computer readable storage medium having a computer program stored thereon, characterized in that the computer program is executed by a processor to implement the steps of the above-mentioned weakly supervised semantic segmentation method based on the common-specific supervision mechanism.

[0150] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0151] The specific embodiments described above are intended to explain the technical solutions and advantages of the present application. It should be understood that the above description is only the most preferred embodiment of the present application and is not intended to limit the present application. Any modification, supplement and equivalent replacement within the principle range of the present application should be included in the protection scope of the present application.

Claims

1. A weakly supervised semantic segmentation method based on co-specificity supervision mechanism, characterized in that, The method comprises the following steps: establishing a category 1 dataset containing category 1 images and image-level labels thereof and a category 2 dataset containing category 2 images and image-level labels thereof; establishing a weakly supervised semantic segmentation model, which comprises an embedding layer, a contrastive convolution module, a commonality-specificity supervision module, a generator, a discriminator and a knowledge gap module, the embedding layer being configured to spatially map the category 1 images and the category 2 images to obtain embedded representations, the contrastive convolution module being configured to spatially enhance the embedded representations to obtain enhanced distribution representations, the commonality-specificity supervision module being configured to construct a commonality supervision graph based on the enhanced distribution representations using a commonality supervision mechanism, construct a specificity supervision graph based on the commonality supervision graph using a specificity supervision mechanism, and construct a target object region of a specific category based on the commonality supervision graph and the specificity supervision graph, the generator being configured to generate a contrastive generated image based on the target object region, the discriminator being configured to determine whether the contrastive generated image is true or false, and the knowledge gap module being configured to calculate a semantic segmentation result according to the category 1 images and the corresponding contrastive generated images thereof; Wherein, the commonality supervision mechanism is based on the enhanced distribution representation to construct a common supervision graph, including: the enhanced distribution representation E_Embedding1 of the category 1 image is projected to the reshape layer for size adjustment to obtain the reorganized distribution E_Embedding1 of the category 1 image re ; the enhanced distribution representation E_Embedding2 of the category 2 image is projected to the average buffer, and the average value of the enhanced distribution representation E_Embedding2 arranged in sequence is calculated to obtain the mean enhanced distribution representation E_Embedding2 of the category 2 image ave ; the enhanced distribution representation E_Embedding2 ave is projected to the SE layer to extract key structural features; according to the E_Embedding2 ave of the category 2 image and the reorganized distribution E_Embedding1 re of the category 1 image, the element correlation matrix R of the category 1 image and the category 2 image is calculated; based on the element correlation matrix R and the reorganized distribution E_Embedding1 re , the common supervision graph M c is calculated; constructing the specificity supervision graph based on the commonality supervision graph using the specificity supervision mechanism comprises: The common supervision graph M c The reverse mapping of the common supervision graph M c The reverse mapping graph M c ′ The calculation process is: Based on the inverse mapping map M c ′ Computing the specificity supervision map M s , the computing process being: M s = Reshape -1 (E_Embedding1 re ) x M c ′ wherein, softmax represents a softmax activation function, and Reshape represents a reshape function; constructing the target object region based on the commonality supervision graph and the specificity supervision graph comprises: M c Embedding1 and Embedding2, to obtain the class 1 image general structure enhanced representation and the class 2 image general structure enhanced representation, the specificity supervision graph M s is used to filter the specific class 1 target object region F1 in the class 1 image general structure enhanced representation cs and the specific class 2 target object region F2 in the class 2 image general structure enhanced representation cs , and the calculation process is: F1 cs = Embedding1+M c -M s F2 cs = Embedding2 + M c -M s establishing an objective function of the weakly supervised semantic segmentation model, which comprises an adversarial loss for training the generator and the discriminator and a consistency loss for constructing the contrastive generated image and the category 1 image to maintain structural consistency therebetween based on the semantic segmentation result; adopting the category 1 dataset and the category 2 dataset and optimizing the parameters of the weakly supervised semantic segmentation model using the objective function; segmenting a target image to be detected using the weakly supervised semantic segmentation model whose parameters have been optimized to obtain a semantic segmentation label of the target image at a pixel level.

2. The weakly supervised semantic segmentation method based on co-specificity supervision mechanism according to claim 1, characterized in that, The embedding layer comprises a boundary padding layer, a two-dimensional convolution layer, an instance regularization layer and a linear rectifier activation layer connected in sequence, and the embedding layer is configured to spatially map the category 1 images and the category 2 images to obtain an embedded representation Embedding1 of the category 1 images and an embedded representation Embedding2 of the category 2 images.

3. The weakly supervised semantic segmentation method based on co-specificity supervision mechanism according to claim 1, characterized in that, The contrastive convolution module comprises a double-channel mode, wherein a first channel comprises a two-dimensional convolution layer and a linear rectifier activation layer connected in sequence, and is configured to extract a corresponding standard local representation S_Embedding1 from the embedded representation Embedding1 of the category 1 images and a corresponding standard local representation S_Embedding2 from the embedded representation Embedding2 of the category 2 images; a second channel comprises a contrastive convolution, which comprises an extended convolution layer and a two-dimensional convolution layer, and is configured to extract a corresponding difference representation D_Embedding1 from the embedded representation Embedding1 of the category 1 images and a corresponding difference representation D_Embedding2 from the embedded representation Embedding2 of the category 2 images; The contrast convolution module further comprises a category activation map calculation operation and an enhanced representation calculation operation, specifically: D_Embedding1 and difference representation D_Embedding2 are calculated using the difference representation D_Embedding1 and difference representation D_Embedding2 corresponding to the class activation map M1 of the class 1 image ca and the class activation map M2 of the class 2 image ca ; M1 = M1 + M2 ca E_Embedding1 = S_Embedding1. M1 ca E_Embedding2 = S_Embedding2. M2 4. The weakly supervised semantic segmentation method based on co-specificity supervision mechanism according to claim 1, characterized in that, In the knowledge gap module, the semantic segmentation result is calculated according to the category 1 image and the corresponding contrast generated image, comprising: The difference between the category 1 image and the corresponding contrast generated image is taken as the semantic segmentation result; The adversarial loss L adv is represented as: The consistency loss is represented as: L cons = up(M c ) · |Seg1| wherein G represents a generator, and D represents a discriminator, represents an element value in the specific class 1 target object region F1 corresponding to the class 1 image at the (i, j) position cs , represents an element value in the specific class 2 target object region F2 corresponding to the class 2 image at the (i, j) position cs , Seg1 represents a semantic segmentation result, I1 represents a class 1 image, up() represents sampling M c into the size of the class 1 image, and M c represents a common supervision graph.

5. An apparatus for weakly supervised semantic segmentation based on co-specific supervision mechanism, comprising: Comprising: A dataset construction module is configured to establish a category 1 dataset and a category 2 dataset, wherein the category 1 dataset contains category 1 images and their image-level labels, and the category 2 dataset contains category 2 images and their image-level labels; A model construction module is configured to establish a weakly supervised semantic segmentation model, wherein the weakly supervised semantic segmentation model comprises an embedding layer, a contrast convolution module, a commonality-specificity supervision module, a generator, a discriminator, and a knowledge gap module, the embedding layer is configured to spatially map the category 1 images and the category 2 images to obtain embedded representations, the contrast convolution module is configured to spatially enhance the embedded representations to obtain enhanced distributed representations, the commonality-specificity supervision module is configured to construct a commonality supervision map based on the enhanced distributed representations using a commonality supervision mechanism, construct a specificity supervision map based on the commonality supervision map using a specificity supervision, and construct a target object region of a specific category based on the commonality supervision map and the specificity supervision map, the generator is configured to generate a contrast generated image based on the target object region, the discriminator is configured to determine the authenticity of the contrast generated image, and the knowledge gap module is configured to calculate a semantic segmentation result according to the category 1 image and the corresponding contrast generated image; Wherein, the commonality supervision mechanism is based on the enhanced distribution representation to construct a common supervision graph, including: the enhanced distribution representation E_Embedding1 of the category 1 image is projected to the reshape layer for size adjustment to obtain the reorganized distribution E_Embedding1 of the category 1 image re ; the enhanced distribution representation E_Embedding2 of the category 2 image is projected to the average buffer, and the average value of the enhanced distribution representation E_Embedding2 arranged in sequence is calculated to obtain the mean enhanced distribution representation E_Embedding2 of the category 2 image ave ; the enhanced distribution representation E_Embedding2 ave is projected to the SE layer to extract key structural features; according to the E_Embedding2 ave of the category 2 image and the reorganized distribution E_Embedding1 re of the category 1 image, the element correlation matrix R of the category 1 image and the category 2 image is calculated; based on the element correlation matrix R and the reorganized distribution E_Embedding1 re , the common supervision graph M c is calculated; The specificity supervision map is constructed based on the commonality supervision map using a specificity supervision, comprising: The common supervision graph M c The reverse mapping of the common supervision graph M c The reverse mapping graph M c ′ The calculation process is: Based on the inverse mapping map M c ′ Computing the specificity supervision map M s , the computing process being: M s = Reshape -1 (E_Embedding1 re ) x M c ′ Wherein, softmax represents a softmax activation function, and Reshape represents a reshape function; The target object region is constructed based on the commonality supervision map and the specificity supervision map, comprising: M c Embedding1 and Embedding2 of the category 2 image, to obtain the category 1 image general structure enhanced representation and the category 2 image general structure enhanced representation, the specific supervision graph M s is used to filter the specific category 1 target object region F1 in the category 1 image general structure enhanced representation cs and the specific category 2 target object region F2 in the category 2 image general structure enhanced representation cs , and the calculation process is: F1 cs = Embedding1+M c -M s F2 cs = Embedding2 + M c -M s A target function construction module is configured to establish a target function of the weakly supervised semantic segmentation model, wherein the target function comprises an adversarial loss for training the generator and the discriminator, and a consistency loss for constructing a structure consistency between the contrast generated image and the category 1 image based on the semantic segmentation result; A parameter optimization module is configured to optimize the parameters of the weakly supervised semantic segmentation model using the category 1 dataset and the category 2 dataset and the target function; A detection module is configured to segment a target image to be detected using the weakly supervised semantic segmentation model with optimized parameters to obtain a semantic segmentation label of the target image at a pixel level.

6. A computing device, comprising: The computer program is executed by the processor to implement the weakly supervised semantic segmentation method based on the commonality-specificity supervision mechanism according to any one of claims 1-4.

7. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the weakly supervised semantic segmentation method based on the commonality-specificity supervision mechanism according to any one of claims 1-4.