Weakly Supervised Instance Segmentation Method, System and Medium Based on Vision Transformer Network

By using the Vision Transformer network to generate category activation maps and COB candidate regions in weakly supervised instance segmentation, and combining pseudo-labels and feature generators for instance segmentation, the problems of poor instance segmentation effect and high computing resource consumption in the prior art are solved, and more efficient and accurate instance segmentation results are achieved.

CN115359254BActive Publication Date: 2025-06-17SOUTH CHINA UNIV OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210877230.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-25
Publication Date
2025-06-17
Estimated Expiration
2042-07-25

AI Technical Summary

Technical Problem

When the prior art performs instance segmentation under weak supervision conditions, there are disadvantages of the candidate mask scoring mechanism and the category activation map generated by the CNN network only focuses on the most prominent areas, resulting in the instance segmentation results only focus on the most recognizable parts, and the effect is poor; at the same time, the prior art adopts a large and deep CNN network, resulting in large computing resources consumption and long training and inference time.

Method used

The weakly supervised instance segmentation method based on the Vision Transformer network is adopted, and the global information of the image is learned through the Vision Transformer network, and the category activation map is generated. The COB candidate regions are generated in combination with the convolutional guide boundary algorithm and the hierarchical segmentation algorithm, and the fake label is constructed. Finally, the classification scores of the COB candidate regions are predicted through the ViT candidate region feature generator to obtain the instance segmentation result.

Benefits of technology

This method can generate candidate area pseudo-labels more accurately, improve the accuracy and efficiency of instance segmentation, reduce training and testing time, and save computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359254B_ABST
    Figure CN115359254B_ABST
Patent Text Reader

Abstract

The present invention discloses a weakly supervised instance segmentation method, system and medium based on the Vision Transformer network. The method includes: obtaining a labeled natural image dataset and a natural image to be segmented; constructing a weakly supervised instance segmentation model; the weakly supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator; initializing the weakly supervised instance segmentation model, constructing a loss function and performing iterative training on the labeled natural image dataset, and optimizing the loss function to obtain a trained weakly supervised instance segmentation model; inputting the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain an instance segmentation result. The present invention realizes the instance segmentation of natural images, while maintaining high performance, speeds up the inference speed and reduces the consumption of computing power.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of weakly supervised instance segmentation, and particularly relates to a weakly supervised instance segmentation method, system and medium based on a Vision Transformer network. Background Art

[0002] Instance segmentation is one of the key problems in the fields of image understanding and computer vision. It is a task of predicting the category of each single pixel in an image and providing different instance labels for separate instances belonging to the same class. In medical image analysis, instance segmentation enables extremely accurate understanding of data, greatly improving the efficiency and accuracy of diagnosis; in the fields of robotics and autonomous driving, instance segmentation provides pixel-level scene understanding for them, improving the efficiency and accuracy of recognition. However, achieving high-performance instance segmentation results requires pixel-level fine annotation. To save time and money costs, researchers are more concerned about how to train a model with image-level category annotation to approach the instance segmentation performance in the fully supervised case. As a new research direction, the core of weakly supervised instance segmentation is to locate each instance and find clear instance boundaries. However, there are a series of challenges: First, under the condition of only image-level category annotation, the multi-label classification network trained only classifies according to the most discriminative regional features of various objects in the dataset, which results in the class activation map or saliency map obtained by the CNN network often only focusing on incomplete numbers of instances and parts of instances of the same class, making the positioning information of different instances defective; Second, it is not easy to find clear instance boundaries. Without pixel-level instance boundary annotation, the CNN network will not automatically draw a clear boundary between instances.

[0003] To address the above technical problems, there are three solutions in the prior art: First, use an advanced method to generate candidate masks for images. Most likely, all instances in the picture will be included in these redundant candidate masks, and they have relatively accurate boundaries. Second, based on the class activation map or saliency map generated by the CNN network, establish a supervision signal for the instance boundary according to the change of pixel gray values, and use this to train the instance filling module. Third, design a new backpropagation method based on the CNN architecture, backpropagate the peak response points of the instance, and finally obtain the contour information of the instance in the original image; such as a PRM method proposed by Zhou et al. in the literature "Weakly Supervised Instance Segmentation using Class Peak Response"; the WS-RCNN framework proposed by Ou J R et al. in the literature "Learning to Score Proposals for Weakly Supervised Instance Segmentation". However, on the one hand, due to the drawbacks of the candidate mask scoring mechanism in the prior art and the problem that the class activation map generated by the CNN network only focuses on the most significant regions, a large number of insignificant instances are often lost, resulting in the instance segmentation result only focusing on the most recognizable part, with poor effects. On the other hand, in order to ensure high performance, an advanced large and deep CNN network is used as the basic framework to design the instance segmentation model, which consumes a large amount of computing resources during the model training process, resulting in long training and inference times and low efficiency. In particular, the prior art only solves the instance segmentation task in a single computer vision modality (CV), and cannot be directly integrated with the tasks in the natural language processing modality (NLP) to achieve more advanced tasks. Summary of the Invention

[0004] The main objective of the present invention is to overcome the drawbacks and deficiencies of the prior art, and provide a weakly supervised instance segmentation method, system, and medium based on Vision Transformer. The present invention learns the global information of the image through the Vision Transformer network to generate a class activation map, and then combines the COB candidate regions generated by the convolutional guidance boundary algorithm and the hierarchical segmentation algorithm to construct pseudo-labels; finally, the ViT candidate region feature generator predicts the classification score of the COB candidate region based on the class activation map to obtain the instance segmentation result, with fewer training parameters, shorter time, and accurate and effective results.

[0005] To achieve the above objective, the present invention adopts the following technical solutions:

[0006] On the one hand, the present invention provides a weakly supervised instance segmentation method based on the Vision Transformer network, including the following steps:

[0007] Obtain a labeled natural image dataset and a natural image to be segmented;

[0008] Construct a weakly-supervised instance segmentation model; the weakly-supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator;

[0009] The Vision Transformer network is used to obtain a multi-label classification result and generate a class activation map; the candidate region pseudo-label generator generates candidate region pseudo-labels according to the class activation map; the candidate region generator uses a convolution-guided boundary algorithm and a hierarchical segmentation algorithm to generate COB candidate regions; the ViT candidate region feature generator uses the SegAlign method to generate a feature vector of the COB candidate region and passes it through a fully-connected layer to map it to the classification score of the COB candidate region;

[0010] Initialize the weakly-supervised instance segmentation model, construct a loss function and perform iterative training on the labeled natural image dataset, and optimize the loss function to obtain a trained weakly-supervised instance segmentation model;

[0011] Input the natural image to be segmented into the trained weakly-supervised instance segmentation model to obtain an instance segmentation result.

[0012] In a preferred technical solution, the labeled natural image dataset is expressed as:

[0013]

[0014] where X i represents the i-th labeled natural image, and Y i represents the label of the i-th natural image; represents the number of images in the labeled natural image dataset, and C represents the number of labels;

[0015] Before using the labeled natural image dataset to perform iterative training on the weakly-supervised instance segmentation model, randomly crop the natural images in the labeled natural image dataset into images of a set size, perform random horizontal flipping on the images, and then perform normalization processing by channel;

[0016] The initialization of the weakly-supervised instance segmentation model means pre-training the weakly-supervised instance segmentation model on a large image dataset and using the model parameters after pre-training as initialization parameters.

[0017] Preferred technical solution, the loss function includes the Focal Loss function and the CE Loss function;

[0018] The Focal Loss function is used to train the ViT multi-label classification module, expressed as:

[0019]

[0020] where y is the true label, p t is the predicted probability, and the definition of p t is as follows:

[0021]

[0022] where p is the output value of the Vision Transformer network without any activation function processing, and the acquisition method is:

[0023] The natural image with an input size of W×H is divided into w×h image patches, each image patch contains P×P pixels, where w = W / P and h = W / P; the image patches are input into the Vision Transformer network to output a feature matrix, and then through a convolutional layer and a global average pooling layer, the feature matrix is mapped into a C-dimensional prediction score vector, that is, p is the output value of the Vision Transformer network;

[0024] The CE Loss function is used to train the ViT candidate region scoring module, expressed as:

[0025]

[0026] where y i,k represents the true label k of the i-th COB candidate region, there are a total of K label values for N COB candidate regions, and p i ′ ,k represents the probability that the i-th COB candidate region is predicted as the k-th label value.

[0027] Preferred technical solution, iteratively training on the labeled natural image dataset, specifically:

[0028] Use the Vision Transformer network to classify the labeled natural image dataset, obtain the multi-label classification result and generate a class activation map;

[0029] Input the labeled natural image dataset into the candidate region generator, and use the convolutional guided boundary algorithm and the hierarchical segmentation algorithm to generate COB candidate regions;

[0030] Activate the class activation map and the COB candidate regions, and use the candidate region pseudo-label generator to obtain the candidate region pseudo-labels;

[0031] Input the COB candidate regions into the ViT candidate region generator, and use the SegAlign method and the fully connected layer to generate the feature vectors of the COB candidate regions and pass them through the fully connected layer to map to the classification scores and classes of the COB candidate regions;

[0032] Calculate the loss value and optimize the loss function, and iteratively train until the function converges to obtain the trained weakly supervised instance segmentation model.

[0033] Preferred technical solution, the Vision Transformer network includes a convolutional layer, L cascaded transformer blocks, and a global average pooling layer; the transformer blocks include a linear transformation layer, a multi-head self-attention layer, and a multi-layer perception block;

[0034] The obtaining of the multi-label classification result and the generation of the class activation map are specifically as follows:

[0035] Input the labeled natural image dataset into the Vision Transformer network, cut each natural image with a size of W×H in the labeled natural image dataset into w×h image patches, perform convolution operations through the convolutional layer to become one-dimensional vectors, and obtain N patch tokens t; add class tokens to the patch tokens D represents the dimension of each patch token;

[0036] Send all the patch tokens with added class tokens into L cascaded transformer blocks for feature extraction to obtain the feature matrix S of the image c and L attention vectors

[0037] Input the feature matrix S of the image c into the convolutional layer and the global average pooling layer to obtain the multi-label classification result;

[0038] For the L attention vectors Take the mean and deform according to the positions of the image patches in the natural image to obtain the attention map, and the formula is:

[0039]

[0040] A′ * =Γ w×h (A * )

[0041] where Γ w×h (·) is the deformation function;

[0042] Multiply the attention map and the feature matrix of the image element - by - element to generate a class activation map TS - CAM, denoted as The element - by - element multiplication formula is:

[0043]

[0044] In a preferred technical solution, the candidate region pseudo - label obtained by using the candidate region pseudo - label generator is specifically:

[0045] In the candidate region pseudo - label generator, obtain the local peaks on the class activation map

[0046] Use each local peak and the position relationship with the COB candidate region to obtain an auxiliary mask

[0047] Take the auxiliary mask Arrange it in ascending order according to the size of the local peak Calculate the overlap degree of a certain COB candidate region R n and the auxiliary mask in sequence If the overlap degree exceeds a certain threshold λ, then mark the pseudo - label z of this COB candidate region n as class c, that is, z n = c; the overlap degree IOU represents the ratio of the overlapping part of two regions to the union part of the two regions;

[0048] If the overlap degree of a certain COB candidate region with all auxiliary masks is lower than the threshold, then mark this COB candidate region as the background class.

[0049] In a preferred technical solution, the obtaining of the local peaks on the class activation map is specifically:

[0050] Take out a certain class activation map M according to the multi - label classification result c ;

[0051] Perform a max - pooling operation on the class activation map M c with a pooling kernel size of m×m. The center of the pooling kernel traverses each position of the class activation map and records a local maximum value and the corresponding position coordinates;

[0052] When the position coordinates of the local maximum value recorded at a certain pixel on the class activation map are exactly the position coordinates of this pixel, record it as a local peak

[0053] The use of each local peak and the COB candidate region The positional relationship gives an auxiliary mask The specific operation is as follows:

[0054] For each local peak Find all COB candidate regions containing the local peak and average them, and obtain the auxiliary mask corresponding to the local peak point by taking a threshold That is:

[0055]

[0056]

[0057] Where refers to the number of COB candidate regions containing the local peak point , p ∈ [0, H] and q ∈ [0, W] are integers representing coordinate indices, and the threshold β ∈ [0, 1] is a hyperparameter

[0058] In a preferred technical solution, the mapping is the classification score and category of the COB candidate region, specifically:

[0059] In the ViT candidate region feature generator, divide the class activation map into n×n image patches, and input them into the Vision Transformer network to obtain the feature vectors of each image patch;

[0060] Concatenate the feature vectors of all image patches in order to form a feature matrix, and then reconstruct the concatenated feature matrix into a new feature matrix according to the positions of each image patch in the corresponding natural image;

[0061] Input the new feature matrix into a 1×1 convolution to fuse the features of each channel to obtain the feature layer F;

[0062] Use the SegAlign method to obtain the features of each COB candidate region on the feature layer F and align them to obtain the aligned features

[0063] Flatten the aligned feature f n into one dimension and input it into a three-layer fully connected layer. After Softmax, obtain the classification score of the COB candidate region

[0064] On the other hand, the present invention provides a weakly supervised instance segmentation system based on the Vision Transformer network, which is applied to the above-mentioned weakly supervised instance segmentation method based on the Vision Transformer network, and includes a data acquisition module, a model construction module, a model training module, and an instance segmentation module;

[0065] The data acquisition module is used to acquire a labeled natural image dataset and a natural image to be segmented;

[0066] The model construction module is used to construct a weakly supervised instance segmentation model; the weakly supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator;

[0067] The model training module is used to initialize the weakly supervised instance segmentation model, construct a loss function and perform iterative training on the labeled natural image dataset, and optimize the loss function to obtain a trained weakly supervised instance segmentation model;

[0068] The instance segmentation module is used to input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain an instance segmentation result.

[0069] On the other hand, the present invention provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned weakly supervised instance segmentation method based on the Vision Transformer network.

[0070] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0071] 1. In terms of pseudo-label generation, the class activation map generated by the Vision Transformer network can cover a larger range of the object, and the local peaks generated on this class activation map can cover a larger area of the object. The candidate region pseudo-labels constructed from the relationship between these local peaks and the candidate regions are more accurate. The ViT candidate region scoring module trained thereby can perform more correct scoring, thus obtaining a more accurate instance segmentation result.

[0072] 2. Due to the multi-head attention mechanism, the feature layer of the image generated by the Vision Transformer network can pay attention to the global information of the image. And in transfer learning, the Vision Transformer network for feature extraction greatly reduces the number of parameters to be learned, can shorten the training and testing time, and save computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0074] Figure 1 This is a flowchart of the weakly supervised instance segmentation method based on Vision Transformer according to an embodiment of the present invention;

[0075] Figure 2 This is a schematic diagram of generating pseudo-labels for candidate regions according to an embodiment of the present invention;

[0076] Figure 3 This is a schematic diagram of a candidate region feature generator according to an embodiment of the present invention;

[0077] Figure 4 This is a schematic diagram of the SegAlign method according to an embodiment of the present invention;

[0078] Figure 5 This is a block diagram of a weakly supervised instance segmentation system based on Vision Transformer according to an embodiment of the present invention;

[0079] Figure 6 This is a schematic diagram of a computer storage medium according to an embodiment of the present invention. Detailed implementation manners

[0080] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0081] Referring to "embodiments" in the present application means that the specific features, structures, or characteristics described in combination with the embodiments may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described in the present application can be combined with other embodiments.

[0082] Please refer to Figure 1 , in an embodiment of the present application, a weakly supervised instance segmentation method based on Vision Transformer is provided, including the following steps:

[0083] S1. Obtain a labeled natural image dataset and a natural image to be segmented;

[0084] In this embodiment, X is used to represent a natural image containing multiple types of objects, and Y is used to represent a multi-class label, Y = [y1,..., y c∈ {0, 1} C×1 , y c = 1 indicates that the object of class c is included in the image, and y c = 0 indicates that the object of class c is not included in the image; therefore, the obtained labeled natural image dataset is represented as:

[0085]

[0086] where X i represents the i-th labeled natural image, and Y i represents the label of the i-th natural image; represents the number of images in the labeled natural image dataset, and C represents the number of labels.

[0087] Before iteratively training the weakly supervised instance segmentation model using the labeled natural image dataset, randomly crop the natural images in the labeled natural image dataset into images of a set size, randomly horizontally flip the images, and then perform normalization processing by channel. In this embodiment, the natural images are cropped into images of size 224×224.

[0088] S2. Construct a weakly supervised instance segmentation model. As Figure 1 shown, the weakly supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator;

[0089] where the Vision Transformer network is used to obtain multi-label classification results and generate class activation maps; the candidate region pseudo-label generator generates candidate region pseudo-labels according to the class activation maps; the candidate region generator uses a convolutional orientation boundary algorithm and a hierarchical segmentation algorithm to generate COB candidate regions; the ViT candidate region feature generator uses the SegAlign method to generate feature vectors of the COB candidate regions and passes through a fully connected layer to map to the classification scores of the COB candidate regions;

[0090] S3. Initialize the weakly supervised instance segmentation model, construct a loss function, and perform iterative training on the labeled natural image dataset to optimize the loss function to obtain a trained weakly supervised instance segmentation model;

[0091] S31. Among them, initializing the weakly supervised instance segmentation model means pre-training the weakly supervised instance segmentation model on a large image dataset and using the model parameters after pre-training as the initialization parameters. In this embodiment, the weakly supervised instance segmentation model is pre-trained on the ImageNet21k dataset, and the parameters after pre-training are used as the initialization parameters. Since the ImageNet21k dataset is an image database organized according to the WordNet hierarchical structure, it has a large number of images, high resolution, many categories, and at the same time, the images also contain more irrelevant noises and variations. Therefore, the recognition difficulty is high, and it is often used for the evaluation of classification, localization, and detection tasks to prevent overfitting.

[0092] S32. Since the weakly supervised instance segmentation model constructed in the present invention contains two module branches, these two module branches need to be trained separately. For the ViT multi-label classification module, in order to solve the problem of class imbalance, the FocalLoss loss function is used to train this module branch, which is expressed as:

[0093]

[0094] Among them, y is the true label, p t is the predicted probability, and the definition of p t is as follows:

[0095]

[0096] Among them, p is the output value of the Vision Transformer network without any activation function processing, and its acquisition method is:

[0097] A natural image with an input size of W×H is sliced into w×h image patches, each image patch contains P×P pixels, where w = W / P and h = W / P; the image patches are input into the Vision Transformer network to output a feature matrix, and then through a convolutional layer and a global average pooling layer, the feature matrix is mapped into a C-dimensional prediction score vector, which is p. In this embodiment, each natural image with a size of 224×224 is sliced into 14×14 image patches, each image patch has 16×16 pixels, and a feature matrix of 768×14×14 size is obtained.

[0098] For the ViT candidate region scoring module, the CELoss loss function is used to train it, which is expressed as:

[0099]

[0100] Among them, y i,k represents the true label k of the i-th COB candidate region, there are a total of K label values for N COB candidate regions, p′i,k It represents the probability that the i-th COB candidate region is predicted as the k-th label value. By fitting the CELoss function, the inter-class distance is also increased to a certain extent.

[0101] S33. During the iterative training of the two module branches on the labeled natural image dataset, no weights are frozen. The training process is as follows:

[0102] S331. Use the Vision Transformer network to classify the labeled natural image dataset, obtain the multi-label classification result and generate the class activation map;

[0103] Specifically, the Vision Transformer network includes a convolutional layer, L cascaded transformer blocks, and a global average pooling layer; each transformer block contains a linear transformation layer, a multi-head self-attention layer, and a multi-layer perceptron block MLP;

[0104] Input the labeled natural image dataset into the Vision Transformer network. Cut each natural image in the labeled natural image dataset into w×h image patches, perform convolutional operations through the convolutional layer to become a one-dimensional vector, and obtain N = w×h patch tokens t; then add class tokens to the patch tokens D represents the dimension of each patch token;

[0105] Send all the patch tokens with added class tokens into L cascaded transformer blocks for feature extraction to obtain the feature matrix S of the image c and L attention vectors

[0106] Denote as the input of the l-th transformer block. In the attention operation of the l-th transformer block, the output patch token is calculated by the following formula:

[0107]

[0108] where the parameter matrices and respectively represent the linear transformation layers before the attention operation of the l-th transformer block; the matrix A l is the attention matrix, and its first row is the attention vector of the class token The attention vector records the dependence of the class token on other image patch tokens. When the loss function works, the attention vector Approach to focusing on the object regions useful for the classification task.

[0109] After inputting the feature matrix of the image into the convolutional layer and the global average pooling layer, a multi-label classification result is obtained;

[0110] For L attention vectors calculate the mean and deform according to the positions of the image patches in the natural image to obtain the attention map. The formula is:

[0111]

[0112] A′ * = Γ w×h (A * )

[0113] where Γ w×h (·) is the deformation function;

[0114] Multiply the attention map and the feature matrix of the image element-wise to generate the class activation map TS-CAM, denoted as The element-wise multiplication formula is:

[0115]

[0116] S332. Input the labeled natural image dataset into the candidate region generator, and use the convolutional oriented boundary algorithm and the hierarchical segmentation algorithm to generate COB candidate regions;

[0117] Convolutional Oriented Boundaries (COB) uses a deep convolutional neural network on the basis of MCG to obtain the edge information of the image. This algorithm implements an end-to-end learnable convolutional neural network to detect the edges of the image and the corresponding directions. Only one image-level forward propagation is required to generate multi-scale contour information and estimate the edge directions. Using these multi-scale contour information and edge directions with the MCG hierarchical segmentation algorithm, COB candidate regions can be generated.

[0118] S333. According to the class activation map and the COB candidate regions, obtain the candidate region pseudo-labels generated by the candidate region pseudo-label generator;

[0119] As Figure 2 shown, in the candidate region pseudo-label generator, the way to obtain the local peaks on the class activation map is:

[0120] First, take out a certain class activation map M c according to the multi-label classification result; then in the class activation map M cPerform a maximum pooling operation on it, with the pooling kernel size of m×m. The center of the pooling kernel traverses each position of the class activation map and records a local maximum value and the corresponding position coordinates. When the position coordinates of the local maximum value recorded at a certain pixel on the class activation map are exactly the position coordinates of this pixel, it is recorded as a local peak of the class activation map.

[0121] Utilize each peak and the COB candidate regions to obtain an auxiliary mask The specific operation is as follows:

[0122] For each local peak Find all COB candidate regions that contain this local peak and average these COB candidate regions. By taking a threshold, obtain the auxiliary mask corresponding to this local peak point That is:

[0123]

[0124]

[0125] Where refers to the number of COB candidate regions that contain this local peak point , p∈[0,H] and q∈[0,W] are integers, representing coordinate indices, H and W respectively represent the height and width of the natural image, and the threshold β∈[0,1] is a hyperparameter.

[0126] Sort the auxiliary masks in ascending order according to the size of the local peaks . Calculate the overlap degree n between a certain COB candidate region R and the auxiliary mask in sequence. If the overlap degree exceeds a certain threshold λ, then mark the pseudo-label z n of this COB candidate region as class c, that is, z n =c; where the overlap degree IOU represents the ratio of the overlapping part of the two regions to the union part of the two regions. In this embodiment, the threshold λ = 0.5.

[0127] If the overlap degree of a certain COB candidate region with all the auxiliary masks is lower than the threshold, then mark this COB candidate region as the background class.

[0128] S334. Input the COB candidate regions into the ViT candidate region generator, adopt the SegAlign method and the fully connected layer to generate the feature vectors of the COB candidate regions and pass through the fully connected layer to map them into the classification scores and classes of the COB candidate regions;

[0129] Specifically, as Figure 3 shown, in the ViT candidate region feature generator, the class activation map is divided into n×n image patches, and the input Vision Transformer network obtains the feature vector Token embeddings of each image patch; in this embodiment, the class activation map is divided into 14×14 image patches, and the dimension of the obtained feature vector is 768 dimensions.

[0130] The feature vectors Token embeddings of all image patches are concatenated in order into a 196×768 feature matrix, and then according to the position of each image patch in the corresponding natural image, the concatenated feature matrix, that is, the feature matrix containing n×n (14×14) feature vectors, is reconstructed (reshaped) into a new feature matrix;

[0131] The new feature matrix is input into a 1×1 convolution conv to fuse the features of each channel to obtain the feature layer F (FeatureMaps); the role of the 1×1 convolution is to fuse the features of each channel, only changing the number of channels of the feature map without changing the width and height dimensions of the feature map;

[0132] Use the SegAlign method to obtain and align the features of each COB candidate region on the feature layer F to obtain the aligned features

[0133] Flatten the aligned feature f n into one dimension and input it into a three-layer fully connected layer. After Softmax, the classification score of the COB candidate region is obtained In this embodiment, the number of nodes in the three-layer fully connected layer is 4096, 4096, and C (the number of labels), respectively.

[0134] SegAlign is an improved version of RoiAlign and can be applied to COB candidate regions. As Figure 4 shown, SegAlign outputs the aligned feature f corresponding to the candidate region according to the input feature layer F and the candidate region R corresponding to the graph n , specifically:

[0135] For a candidate region R, its receptive field on the feature map F is R F , for the receptive field R F corresponding to the candidate region R, its circumscribed rectangle is B. If the function is a bilinear transformation from the spatial coordinates (i,j)∈f to (i′,j′)∈B, that is, is a bilinear interpolation function on the feature map F, the SegAlign operator can be expressed as the following formula:

[0136]

[0137] In this formula, for the sake of simplicity and convenience of representation, the channel dimension of the feature layer is not represented in the formula.

[0138] S335. Calculate the loss value and optimize the loss function, and iteratively train until the function converges to obtain a trained weakly supervised instance segmentation model.

[0139] S4. Input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain the instance segmentation result.

[0140] Input a natural image into the COB candidate region generator to generate candidate regions Then input it into the ViT candidate region scoring module to obtain the category and score of each candidate region, and finally obtain the final instance segmentation result through non-maximum suppression (NMS).

[0141] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0142] Based on the same idea as the weakly supervised instance segmentation method based on Vision Transformer in the above embodiments, the present invention also provides a weakly supervised instance segmentation system based on Vision Transformer, which can be used to execute the above weakly supervised instance segmentation method based on Vision Transformer. For the sake of convenience of description, in the structural schematic diagram of the embodiment of the weakly supervised instance segmentation system based on Vision Transformer, only the parts related to the embodiments of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and it may include more or fewer components than those illustrated, or combine certain components, or arrange different components.

[0143] Please refer to Figure 5 , in another embodiment of the present application, a weakly supervised instance segmentation system based on Vision Transformer is provided, and the system includes a data acquisition module, a model construction module, a model training module, and an instance segmentation module;

[0144] The data acquisition module is used to acquire a labeled natural image dataset and a natural image to be segmented;

[0145] The model construction module is used to construct a weakly supervised instance segmentation model, including a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator;

[0146] The model training module is used to initialize the weakly supervised instance segmentation model, construct a loss function and perform iterative training on a labeled natural image dataset, and optimize the loss function to obtain a trained weakly supervised instance segmentation model;

[0147] The instance segmentation module is used to input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain an instance segmentation result.

[0148] It should be noted that the weakly supervised instance segmentation system based on Vision Transformer of the present invention corresponds one-to-one with the weakly supervised instance segmentation method based on Vision Transformer of the present invention. The technical features and their beneficial effects described in the embodiments of the above-mentioned weakly supervised instance segmentation method based on Vision Transformer are applicable to the embodiments of the weakly supervised instance segmentation system based on Vision Transformer. For specific content, please refer to the description in the method embodiments of the present invention, which will not be repeated here. This is hereby declared.

[0149] In addition, in the implementation manner of the weakly supervised instance segmentation system based on Vision Transformer in the above embodiment, the logical division of each program module is only an example. In practical applications, according to needs, for example, considering the configuration requirements of the corresponding hardware or the convenience of software implementation, the above functions can be assigned to different program modules to complete, that is, the internal structure of the weakly supervised instance segmentation system based on Vision Transformer is divided into different program modules to complete all or part of the functions described above.

[0150] Please refer to Figure 6 , in one embodiment, a computer-readable storage medium is provided, storing a program in the memory. When the processor executes the program, the weakly supervised instance segmentation method based on Vision Transformer can be implemented, specifically:

[0151] Obtain a labeled natural image dataset and a natural image to be segmented;

[0152] Construct a weakly supervised instance segmentation model, including a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator;

[0153] Among them, the Vision Transformer network is used to obtain multi-label classification results and generate class activation maps; the candidate region pseudo-label generator generates candidate region pseudo-labels according to the class activation maps; the candidate region generator uses a convolution-guided boundary algorithm and a hierarchical segmentation algorithm to generate COB candidate regions; the ViT candidate region feature generator uses the SegAlign method to generate feature vectors of the COB candidate regions and passes through a fully connected layer to map to the classification scores of the COB candidate regions;

[0154] Initialize the weakly supervised instance segmentation model, construct a loss function and perform iterative training on a labeled natural image dataset, and optimize the loss function to obtain a trained weakly supervised instance segmentation model;

[0155] Input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain the instance segmentation result.

[0156] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0157] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0158] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A weakly supervised instance segmentation method based on the Vision Transformer network, characterized in that, It includes the following steps: Obtain a labeled natural image dataset and a natural image to be segmented; Construct a weakly supervised instance segmentation model; the weakly supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator; The Vision Transformer network is used to obtain a multi-label classification result and generate a class activation map; the candidate region pseudo-label generator generates candidate region pseudo-labels according to the class activation map; The candidate region generator uses a convolution-guided boundary algorithm and a hierarchical segmentation algorithm to generate COB candidate regions; the ViT candidate region feature generator uses the SegAlign method to generate a feature vector of the COB candidate region and passes through a fully connected layer to map it to the classification score of the COB candidate region; Initialize the weakly supervised instance segmentation model, construct a loss function and perform iterative training on the labeled natural image dataset, and optimize the loss function to obtain a trained weakly supervised instance segmentation model; The iterative training on the labeled natural image dataset is specifically as follows: Use the Vision Transformer network to classify the labeled natural image dataset to obtain a multi-label classification result and generate a class activation map; Input the labeled natural image dataset into the candidate region generator, and use a convolution-guided boundary algorithm and a hierarchical segmentation algorithm to generate COB candidate regions; According to the class activation map and the COB candidate regions, use the candidate region pseudo-label generator to obtain candidate region pseudo-labels; Input the COB candidate regions into the ViT candidate region generator, use the SegAlign method and a fully connected layer to generate a feature vector of the COB candidate region and pass through a fully connected layer to map it to the classification score and class of the COB candidate region; Calculate the loss value and optimize the loss function, and perform iterative training until the function converges to obtain a trained weakly supervised instance segmentation model; Input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain an instance segmentation result.

2. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 1, characterized in that, The labeled natural image dataset is expressed as: Among them, X i represents the i-th labeled natural image, and Y i represents the label of the i-th natural image; represents the number of images in the labeled natural image dataset, and C represents the number of labels; Before using the labeled natural image dataset to perform iterative training on the weakly supervised instance segmentation model, randomly crop the natural images in the labeled natural image dataset into images of a set size, perform random horizontal flipping on the images, and then perform normalization processing by channel; The initialization of the weakly supervised instance segmentation model means pre-training the weakly supervised instance segmentation model on a large image dataset and using the model parameters after pre-training as initialization parameters.

3. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 2, characterized in that, The loss function includes a Focal Loss function and a CE Loss function; The Focal Loss function is used to train the ViT multi-label classification module and is expressed as: where y is the true label and p t is the predicted probability, and the definition of p t is as follows: Among them, p is the output value of the Vision Transformer network without any activation function processing, and the acquisition method is: The natural image with an input size of W×H is sliced into w×h image patches, each image patch containing P×P pixels, where w = W / P and h = H / P; the image patches are input into the Vision Transformer network to output a feature matrix, which is then passed through a convolutional layer and a global average pooling layer to map the feature matrix into a C-dimensional prediction score vector, which is the output value p of the Vision Transformer network; The CELoss loss function is used to train the ViT candidate region scoring module, expressed as: Among them, y i,k represents the true label k of the i-th COB candidate region. There are a total of K label values for N COB candidate regions, p i ′ ,k represents the probability that the i-th COB candidate region is predicted as the k-th label value.

4. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 3, characterized in that, The Vision Transformer network includes a convolutional layer, L cascaded transformer blocks, and a global average pooling layer; the transformer blocks include a linear transformation layer, a multi-head self-attention layer, and a multi-layer perceptron block; The obtaining of the multi-label classification result and the generation of the class activation map are specifically as follows: Input the natural image dataset with tags into the Vision Transformer network. Cut each natural image with size W×H in the natural image dataset with tags into w×h image patches, perform convolution operations through the convolutional layer to become one-dimensional vectors, and obtain N block tokens t; add class tokens to the block tokens D represents the dimension of each block token; Send all the block tags with added category tags into L cascaded transformer blocks for feature extraction to obtain the feature matrix S of the image c and L attention vectors The feature matrix S of the image c After being input into the convolutional layer and the global average pooling layer, a multi-label classification result is obtained; For L attention vectors Calculate the mean and deform according to the position of the image patch in the natural image to obtain the attention map. The formula is: A′ * = Γ w×h (A * ) where Γ w×h (·) is a deformation function; Multiply the attention map and the feature matrix of the image element-wise to generate the class activation map TS-CAM, denoted as The element-wise multiplication formula is: 。 5. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 4, wherein, The candidate region pseudo-labels obtained by using the candidate region pseudo-label generator are specifically as follows: In the candidate region pseudo-label generator, obtain local peaks on the class activation map Using each local peak and the positional relationship of the COB candidate regions to obtain an auxiliary mask The auxiliary mask is sorted in ascending order according to the size of the local peak , and the overlap degree between a certain COB candidate region R n and the auxiliary mask is calculated in sequence If the overlap degree exceeds a certain threshold λ, the pseudo-label z of this COB candidate region n is marked as category c, that is, z n = c; the overlap degree IOU represents the ratio of the overlapping part of the two regions to the set part of the two regions If the overlap degree of a certain COB candidate region with all auxiliary masks is lower than the threshold, then this COB candidate region is marked as the background class.

6. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 5, wherein, Obtaining local peaks on the class activation map Specifically: Extract a certain category activation map M according to the multi-label classification result c ; Perform a max pooling operation on the class activation map M c with a pooling kernel size of m×m. The center of the pooling kernel traverses each position of the class activation map and records a local maximum value and the corresponding position coordinates; When the position coordinates of the local maximum recorded at a certain pixel on the class activation map are exactly the position coordinates of that pixel, it is denoted as a local peak The use of each local peak and the position relationship of the COB candidate regions to obtain an auxiliary mask The specific operation is as follows: For each local peak Find all COB candidate regions containing the local peak and average them, and obtain an auxiliary mask corresponding to the local peak point by taking a threshold That is: where refers to the number of COB candidate regions containing the local peak point , p ∈ [0, H] and q ∈ [0, W] are integers representing coordinate indices, and the threshold β ∈ [0, 1] is a hyperparameter.

7. The weakly supervised instance segmentation method based on the Vision Transformer network according to claim 5, wherein, The mapping to the classification score and class of the COB candidate region is specifically as follows: In the ViT candidate region feature generator, the class activation map is divided into n×n image patches, and the feature vectors of each image patch are obtained by inputting them into the Vision Transformer network; The feature vectors of all image patches are concatenated in order to form a feature matrix, and then according to the positions of each image patch in the corresponding natural image, the concatenated feature matrix is reconstructed into a new feature matrix; The new feature matrix is input into a 1×1 convolution to fuse the features of each channel to obtain the feature layer F; Use the SegAlign method to obtain the features of each COB candidate region on the feature layer F and perform alignment to obtain the aligned features Flatten the alignment feature f n into one dimension and input it into a three-layer fully connected layer. After Softmax, the classification scores of the COB candidate regions are obtained 8. The weakly supervised instance segmentation system based on the Vision Transformer network, wherein, Applied to the weakly supervised instance segmentation method based on the Vision Transformer network according to any one of claims 1-7, including a data acquisition module, a model construction module, a model training module, and an instance segmentation module; The data acquisition module is used to acquire a labeled natural image dataset and the natural image to be segmented; The model construction module is used to construct a weakly supervised instance segmentation model; the weakly supervised instance segmentation model includes a ViT multi-label classification module and a ViT candidate region scoring module; the ViT multi-label classification module includes a Vision Transformer network and a candidate region pseudo-label generator; the ViT candidate region scoring module includes a candidate region generator and a ViT candidate region feature generator; The model training module is used to initialize the weakly supervised instance segmentation model, construct a loss function and perform iterative training on the labeled natural image dataset, and optimize the loss function to obtain a trained weakly supervised instance segmentation model; The iterative training on the labeled natural image dataset is specifically as follows: Use the Vision Transformer network to classify the labeled natural image dataset to obtain a multi-label classification result and generate a class activation map; Input the labeled natural image dataset into the candidate region generator, and use the convolutional orientation boundary algorithm and the hierarchical segmentation algorithm to generate COB candidate regions; According to the class activation map and the COB candidate regions, use the candidate region pseudo-label generator to obtain the candidate region pseudo-labels; Input the COB candidate regions into the ViT candidate region generator, and use the SegAlign method and the fully connected layer to generate the feature vectors of the COB candidate regions and pass through the fully connected layer to map to the classification scores and classes of the COB candidate regions; Calculate the loss value and optimize the loss function, and iteratively train until the function converges to obtain the trained weakly supervised instance segmentation model; The instance segmentation module is used to input the natural image to be segmented into the trained weakly supervised instance segmentation model to obtain the instance segmentation result.

9. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the weakly supervised instance segmentation method based on the Vision Transformer network according to any one of claims 1-7.

Citation Information

Patent Citations

  • Weak supervision target positioning method and device based on shallow feature background suppression

    CN114596471A