Self-Supervised Learning-Based Method for Unimodal Image Segmentation and Content Completion

Through the self-supervised learning generative model combined with encoder and generator, the end-to-end training problem of non-modal image segmentation and content completion is solved, efficient feature map matching and parameter optimization are achieved, and segmentation and completion efficiency is improved.

CN115797392BActive Publication Date: 2025-07-22NANJING UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211653377.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-21
Publication Date
2025-07-22
Estimated Expiration
2042-12-21

AI Technical Summary

Technical Problem

The prior art cannot realize end-to-end training of non-modal image segmentation and content completion, intermediate layer features are not effectively utilized, model parameters are large and inference speed is slow.

Method used

The generative model of self-supervised learning is adopted, and the non-modal segmentation and content reconstruction results of the target object are output through the combination of encoder and generator. Self-supervised learning is used to train on the KINS dataset, and the encoder and generator share latent vector and category information, and use a variety of loss functions to constrain the training process.

Benefits of technology

End-to-end training of non-modal segmentation and content completion is realized, with high matching of feature maps, small model parameters, and fast inference speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797392B_ABST
    Figure CN115797392B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for non-modal image segmentation and content completion based on self-supervised learning. This method adopts a model of an encoder plus a generator. In the generator part, it simultaneously outputs the non-modal segmentation of the target object and the result of image reconstruction, and uses the method of self-supervised learning to train on the KINS dataset. The main steps of this method include: using the encoder to connect the input image x and the segmentation eraser_mask of the occluder and map them to the latent space, and output a one-dimensional latent vector z; inputting the vector z and the category c of the target object into the pre-trained large generator BigGAN, and using its prior information to generate the image x<supgt;*< / supgt;, restoring the occluded part of the target object, and at the same time generating the complete segmentation y<supgt;*< / supgt; of the target object through the segmentation branch. Divide the dataset into a training dataset, a validation dataset, and a test dataset. During training, generate training data through the occlusion algorithm and use the self-supervised method to train the entire neural network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of image segmentation and image inpainting, and particularly to an amodal image segmentation and content completion method based on self-supervised learning. Background Art

[0002] Both image segmentation and image inpainting are fundamental and important tasks in the field of computer vision. Image segmentation is to pixel-wisely segment the target object parts in an image, while amodal image segmentation is to segment both the parts of the target object in the image and the occluded parts of the object. Image inpainting is to fill in the damaged parts in an image. The present invention aims to use a complete model to jointly predict both the segmentation and the content, so as to improve the prediction performance of segmentation and the performance of content completion.

[0003] The PCNet method proposed by Xiaohang Zhan et al. This method consists of two independent networks: PCNet-M and PCNet-C, which sequentially implement occlusion relationship reasoning, amodal segmentation, and image inpainting. Huan Ling et al. proposed a variational generation framework for amodal segmentation, called Amodal-VAE, which does not require any amodal labels during training because it can utilize widely available object instance segmentation to generate data for self-supervised learning. Khoi Nguyen et al. proposed the ASBU method, which implicitly learns the object shape prior through uncertainty weighting, uses this uncertainty-weighted loss function to constrain the training, predicts an uncertainty map for the segmentation result, and weights the loss function with the estimated uncertainty to regularize the model training, so as to produce a lower segmentation loss in regions with high uncertainty.

[0004] Although the above methods have made great progress in the results of multi-object tracking, these methods still have the following problems: 1) Implementing amodal segmentation and completion with two separate models cannot achieve end-to-end training; 2) Using two independent models for inference, the repeated features generated in the intermediate layers are not effectively utilized, which easily causes the phenomenon that the segmentation result does not match the image; 3) The number of parameters of multiple models is large, and the step-by-step inference speed is slow. Summary of the Invention

[0005] The purpose of the present invention is to provide an amodal image segmentation and content completion method based on self-supervised learning. This method adopts a generative model, and the model simultaneously outputs the amodal segmentation and content reconstruction results of the target object, and is trained on the KINS dataset using the self-supervised learning method.

[0006] The technical solution to implement the present invention: In the first aspect, the present invention provides an amodal image segmentation and content completion method based on self-supervised learning, including the following steps:

[0007] Step 1: Randomly select an object from the objects in the dataset as the occluder eraser, occlude the image x of the target object, and generate the occluded picture x′;

[0008] Step 2: Connect the occluded picture x′ of the object and the segmentation of the occluder eraser_mask, and encode them into the latent vector z = E([x′, eraser_mask]) through the encoder E;

[0009] Step 3: Input the latent vector z and the category c of the target object into the generator G, and simultaneously output the complete target object and the segmentation result x * , y * = G(z, c);

[0010] Step 4: Use different loss functions Loss img (x * , x), Loss mask (y * , y) to respectively constrain the image and the segmentation, and train the entire neural network model, where x and y are the image and segmentation of the unoccluded target object;

[0011] Step 5: During inference, use all other foreground objects except the target object in the image as the occluder eraser, connect the occluded target object image and the segmentation of the occluder, and input them into the model for inference to output the complete target object and its segmentation.

[0012] In a second aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in the first aspect.

[0013] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in the first aspect.

[0014] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect.

[0015] Compared with the prior art, the present invention has the following remarkable advantages: 1) It can implement non-modal segmentation and content completion of the occluded part through one model, and can achieve end-to-end training; 2) Use one model for inference, and the feature maps generated by the hidden layer are fully reused, and the segmentation result matches the image; 3) The image generation part and the segmentation part of the model share some parameters, with a small number of parameters and a fast inference speed. Description of the Drawings

[0016] Figure 1 This is a schematic diagram of the process of the present invention.

[0017] Figure 2 This is a schematic diagram of dataset processing.

[0018] Figure 3 This is the overall structure diagram of the neural network model. Detailed implementation manners

[0019] A self-supervised learning-based unmodal image segmentation and content completion algorithm includes the following steps:

[0020] Step 1: Randomly select an object from the objects in the dataset as an occluder eraser, and occlude the image x of the target object to generate an occluded picture x';

[0021] Step 2: Connect the occluded picture x' of the object and the segmentation of the occluder eraser_mask, and encode them into a latent vector z = E([x', eraser_mask]) through the encoder E;

[0022] Step 3: Input the latent vector z and the category c of the target object into the generator G, and simultaneously output the complete target object and the segmentation result x * , y * = G(z, c);

[0023] Step 4: Use different loss functions Loss img (x * , x), Loss mask (y * , y) to respectively constrain the image and the segmentation, and train the entire neural network model, where x and y are the image and the segmentation of the unoccluded target object;

[0024] Step 5: During inference, use all other foreground objects except the target object in the image as the occluder eraser, connect the occluded target object image and the segmentation of the occluder, and input them into the model for inference to output the complete target object and its segmentation.

[0025] Preferably, in Step 1, an occlusion algorithm is used to generate the data and labels required for training from the original dataset, and self-supervised methods are used for training without manual annotation.

[0026] Preferably, in Step 3, the same generator model is used to simultaneously generate the complete appearance information and the segmentation result of the target object, reusing the feature information of the intermediate layer, and the segmentation result has a high matching degree with the appearance information.

[0027] Preferably, in Step 3, the segmentation branch uses convolution and cross-layer connections, and the number of model parameters is small.

[0028] Preferably, in step 3, the generator part uses the BigGAN model pre-trained on ImageNet, which has rich prior knowledge and can generate more realistic images.

[0029] Preferably, in step 4, multiple loss functions are used to constrain the training of the generated image part, including the perceptual loss function and the mean square error loss function.

[0030] Preferably, in step 4, multiple loss functions are used to constrain the training of the segmentation branch, including the cross-entropy loss function and the dice loss function.

[0031] The present invention proposes a method for non-modal image segmentation and content completion based on self-supervised learning. This method adopts a generative model that simultaneously outputs the non-modal segmentation and content reconstruction results of the target object, and is trained on the KINS dataset using the self-supervised learning method.

[0032] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0033] Embodiment

[0034] As Figure 1 shown, the main steps of this method include: using an encoder to connect and map the input image and the segmentation of the occluder to the latent space, and output a one-dimensional latent vector; inputting the vector and the category of the target object into the pre-trained large generator BigGAN, and using its prior information to generate an image to restore the occluded part of the target object, and at the same time generating a complete segmentation of the target object through the segmentation branch. The dataset is divided into a training dataset, a validation dataset, and a test dataset. During training, training data is generated through an occlusion algorithm, and the entire neural network model is trained using a self-supervised method. The specific steps are as follows:

[0035] Step 1: Randomly select an object from the objects in the dataset as the occluder eraser, and occlude the image x of the target object to generate the occluded image x':

[0036] x' = x × (1 - eraser_mask) + eraser × eraser_mask

[0037] where the image x before occlusion is as Figure 2 shown in (a), eraser_mask is the segmentation of the occluder eraser, as Figure 2 shown in (b), and the generated occluded image x' is as Figure 2 shown in (c).

[0038] Step 2: The encoder E uses the ResNet50 model. Its input is a tensor of 4×128×128, and its output is a one-dimensional vector of 1×120. The image x′ of the occluded object and the segmentation eraser_mask of the occluder are concatenated in the first dimension using the cat operation in PyTorch and encoded into the latent vector z = E([x′, eraser_mask]) by the encoder E;

[0039] Step 3: The generator G is based on BigGAN and adds a segmentation branch. BigGAN consists of GBlocks, and the segmentation branch is modeled after the generator part and consists of segmentation modules SegBlocks. The SegBlock contains a 1×1 convolutional layer and upsampling. The SegBlock takes the hidden layer feature map output by the same-layer GBlock and the hidden layer feature map output by the previous SegBlock as inputs, and the output feature map dimension of the SegBlock is equal to the output feature map dimension of the GBlock. The latent vector z and the category c of the target object are input into the generator G, and the complete target object and segmentation result x * , y * = G(z, c), where c is an integer, c ∈ [1, 1000], representing the category of the target object. The last layer of the model uses the tanh activation function, and the output images x* and segmentations y* are tensors of 3×128×128 and 1×128×128 respectively, with values in the range (-1, 1);

[0040] Step 4: As Figure 3 shown, different loss functions Loss img (x * , x), Loss mask (y * , y) are used to constrain the image and segmentation respectively to train the entire neural network model, where x and y are the image and segmentation of the unoccluded target object, and where

[0041] Loss img (x * , x) = Loss LPIPS (x * , x) + Loss mse (x * , x)

[0042] Loss mask (y * , y) = Loss bce (y * , y) + Loss Dice (y * , y)

[0043] where x, x* is a 3×128×128 tensor, y, y * is a 1×128×128 tensor. Among them, Loss LPIPS is the perceptual loss. The feature maps of two input images are generated layer by layer using a pre-trained VGG16 network, and the difference between the feature maps is calculated as the loss. Loss mse is the mean squared error loss function, which calculates the error between images pixel by pixel. Loss bce is the binary cross-entropy loss function, which calculates the error between the predicted value and the true value pixel by pixel. Loss Dice is a set similarity metric function used to calculate the similarity between two samples.

[0044] Step 5, during inference, all other foreground objects in the image except the target object are used as the occluder eraser. The segmentation of the occluded target object image and the occluder is connected and input into the model for inference to output the complete target object and its segmentation.

Claims

1. A method for non-modal image segmentation and content completion based on self-supervised learning, characterized in that, It includes the following steps: Step 1: Randomly select an object from the objects in the dataset as an occluder eraser, occlude the image x of the target object, and generate an occluded picture x'; Step 2: Connect the occluded picture x' of the object and the segmentation of the occluder eraser_mask, and encode them into a latent vector z = E([x', eraser_mask]) through the encoder E; Step 3: Input the latent vector z and the category c of the target object into the generator G, and simultaneously output the complete target object and the segmentation result x * , y * = G(z, c); Step 4: Use different loss functions Loss img (x * , x), Loss mask (y * , y) are used to constrain the image and the segmentation respectively to train the entire neural network model, where x and y are the images and segmentations of the unoccluded target objects; Step 5: During inference, use all other foreground objects in the image except the target object as the occluder eraser, connect the occluded target object image and the segmentation of the occluder, and input them into the model for inference to output the complete target object and its segmentation.

2. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 1, an occlusion algorithm is used to generate the data and labels required for training from the original dataset, and self-supervised methods are used for training without manual annotation.

3. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 3, the same generator model is used to generate the complete appearance information and segmentation result of the target object, and the feature information of the intermediate layer is reused.

4. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 3, the segmentation branch uses 1×1 convolution and cross-layer connection.

5. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 3, the generator part uses the BigGAN model pre-trained on ImageNet.

6. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 4, multiple loss functions are used to constrain the training of the generated picture part, including the perceptual loss function and the mean square error loss function.

7. The method for unsupervised image segmentation and content completion based on self-supervised learning according to claim 1, characterized in that in step 4, multiple loss functions are used to constrain the training of the segmentation branch, including the cross-entropy loss function and the sieve loss function.

8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Pedestrian detection method and device based on occlusion perception self-supervised learning

    CN110084146A

  • Self-supervised information extraction method combining depth features and contrast learning

    CN114612685A