A weak annotation method and a video object segmentation method based on text description

By adopting a weak annotation method in video target segmentation, using a combination of dense annotation and bounding box annotation, combining the language-guided cross-frame segmentation module and cross-frame comparison learning module, the problem of high annotation cost in the existing technology is solved, and efficient video target segmentation model training and labeling cost reduction is achieved.

CN116091977BActive Publication Date: 2025-05-30SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310109854.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-13
Publication Date
2025-05-30
Estimated Expiration
2043-02-13

AI Technical Summary

Technical Problem

The existing video target segmentation algorithm based on text description relies on intensive annotation, which leads to high labeling costs and cannot achieve large-scale data set annotation, limiting the full training of deep models.

Method used

A weak labeling method is proposed. By using dense labeling for the first frame of the video to be labeled, other frames are marked with bounding box, combining the language-guided cross-frame segmentation module and cross-frame comparison learning module to train the video target segmentation model.

Benefits of technology

The annotation cost of video target segmentation data set is greatly reduced, making full use of dense annotation, giving full play to the advantages of bounding box annotation, improving the model's ability to distinguish feature prospects and backgrounds, and achieving comparable performance with the full supervision method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116091977B_ABST
    Figure CN116091977B_ABST
Patent Text Reader

Abstract

The present application provides a weak annotation method and a video object segmentation method based on text description. The weak annotation method includes: obtaining a video to be annotated; performing dense annotation on the first frame of the video to be annotated, and performing bounding box annotation on other frames except the first frame to obtain an annotated image for each frame of the video to be annotated; wherein, the first frame is the frame in which the object first appears in the video to be annotated. This annotation method can greatly reduce the annotation cost of the video object segmentation dataset based on text description.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a weak annotation method and a video object segmentation method based on text description. Background Art

[0002] Existing video object segmentation algorithms based on text description are basically based on fully supervised methods and rely on densely annotated datasets. For example, in the YouTube-RVOS dataset, for each video in the dataset, objects are pixel-level annotated frame by frame. However, dense annotation is costly, resulting in the inability to annotate large-scale datasets.

[0003] Therefore, due to the high cost of video annotation, existing datasets are small in scale and cannot fully train deep models well. To solve this problem, existing state-of-the-art methods mainly adopt a strategy of pre-training first and then fine-tuning. For example, first pre-train using some image segmentation datasets based on language description, and then fine-tune using the video dataset YouTube-RVOS. Or directly perform data augmentation on the image dataset to generate "pseudo-videos" and jointly train them with the video dataset. Summary of the Invention

[0004] The purpose of the embodiments of this specification is to provide a weak annotation method and a video object segmentation method based on text description.

[0005] To solve the above technical problems, the embodiments of the present application are implemented as follows:

[0006] In a first aspect, the present application provides a weak annotation method, which includes:

[0007] Obtain a video to be annotated;

[0008] Perform dense annotation on the first frame of the video to be annotated, and perform bounding box annotation on other frames except the first frame to obtain an annotated image for each frame of the video to be annotated; wherein, the first frame is the frame in which the object first appears in the video to be annotated.

[0009] In a second aspect, the present application provides a video object segmentation method based on text description, which includes:

[0010] Obtain a video to be segmented and a language text description;

[0011] Input the video to be segmented and the language text description into a video object segmentation model to perform segmentation prediction on the objects in the video to be segmented; when training the video object segmentation model, use the annotated image for each frame of the video to be annotated obtained by the weak annotation method as described in the first aspect to calculate the loss function and update the video object segmentation model.

[0012] In one embodiment, the video object segmentation model includes:

[0013] A language encoder, configured to receive a language text description and output language features;

[0014] A visual encoder, configured to receive each frame image of the video to be segmented and output corresponding visual features;

[0015] A visual enhanced language module, configured to receive language features and visual features and output enhanced language features;

[0016] A language enhanced visual module, configured to receive language features and visual features and output enhanced visual features;

[0017] A language-guided cross-frame segmentation module, configured to receive enhanced visual features and enhanced language features and output the objects in the video to be segmented for segmentation prediction.

[0018] In one embodiment, the language-guided cross-frame segmentation module includes a dynamic convolution kernel of the current frame and dynamic convolution kernels of other frames;

[0019] The dynamic convolution kernel of the current frame acts on the enhanced visual features of the current frame to generate the segmentation of the current frame;

[0020] The dynamic convolution kernels of other frames act on the enhanced visual features of the current frame to generate predictions for the current frame.

[0021] In one embodiment, the video object segmentation model further includes:

[0022] A cross-frame contrast learning module, configured to enhance the foreground-background discrimination ability in the features;

[0023] The cross-frame contrast learning module adopts a cross-modal contrast learning method or a region consistency contrast learning method.

[0024] In one embodiment, in the cross-modal contrast learning method, the language features and the object region features in the video to be segmented are positive samples, and the language features and the background regions are negative samples.

[0025] In one embodiment, in the region consistency contrast learning method, the mean of the object region features in the video to be segmented and the features of all object regions are positive samples, and the mean of the object region features in the video to be segmented and all background regions are negative samples.

[0026] In one embodiment, the method further includes distinguishing the foreground region and the background region:

[0027] For the first frame image of the video to be segmented with dense annotation, use the dense annotation labels to distinguish the foreground region and the background region;

[0028] For other frame images annotated with bounding boxes in the video to be segmented, the area outside the bounding box is used as the background area, and the segmentation predicted by the video object segmentation model is used as the pseudo-label. The area with a confidence greater than the first threshold is selected as the foreground area, and the area with a confidence less than the second threshold is selected as the background area, where the first threshold is greater than the second threshold.

[0029] In one of the embodiments, after distinguishing the foreground area and the background area, a cross-modal contrast learning method or a region consistency contrast learning method is adopted to add a contrast learning loss function to train the video object segmentation model's ability to distinguish foreground and background features.

[0030] As can be seen from the technical solutions provided in the embodiments of this specification above, the weak annotation method provided in this application can significantly reduce the annotation cost of the video object segmentation dataset based on text descriptions.

[0031] The video object segmentation method based on text descriptions proposed based on the weak annotation method can reduce the annotation cost of the dataset. In this method, a language-guided cross-frame segmentation module is used, so that the annotation of each frame can supervise the learning process of other frames, thereby making full use of valuable dense annotations and also giving full play to the advantages of bounding box annotations; this method uses a cross-frame contrast learning module to improve the model's ability to distinguish foreground and background in features. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 It is a flowchart of the weak annotation method provided in this application;

[0034] Figure 2 It is a comparison schematic diagram of frame-by-frame annotation, all using bounding box annotation, and the weak annotation method of this application;

[0035] Figure 3 It is a flowchart of the video object segmentation method based on text descriptions provided in this application;

[0036] Figure 4 It is an architecture diagram of the video object segmentation model provided in this application;

[0037] Figure 5 It is a working flowchart of the language-guided cross-frame segmentation module provided in this application;

[0038] Figure 6 It is a schematic diagram of the working process of the cross-frame contrast learning module provided for this application. Specific embodiments

[0039] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0040] In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of this application.

[0041] Without departing from the scope or spirit of this application, various improvements and changes can be made to the specific embodiments of this application specification, which are obvious to those skilled in the art. Other embodiments obtained from the specification of this application are obvious to those skilled in the art. The specification and embodiments of this application are only exemplary.

[0042] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, that is, they are intended to include but not limited to.

[0043] Unless otherwise specified, the "parts" in this application are calculated by mass parts.

[0044] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0045] In order to reduce the dependence on densely labeled data when training a deep model, this application proposes a weak labeling method to reduce the labeling cost for subsequent construction of a large-scale video object segmentation dataset based on text descriptions.

[0046] Refer to Figure 1 , which shows a schematic diagram of the process suitable for the weak labeling method provided in the embodiments of this application.

[0047] As Figure 1 shown, a weak labeling method may include:

[0048] S110. Obtain the video to be labeled;

[0049] S120. For the first frame in the video to be annotated, dense annotation is adopted, and for the other frames except the first frame, bounding box annotation is adopted to obtain the annotated image of each frame of the video to be annotated; wherein, the first frame is the frame when the object first appears in the video to be annotated.

[0050] Specifically, the weak annotation method proposed in this application is to adopt dense annotation for the first frame of the video to be annotated, and bounding box annotation for the other frames. As Figure 2 shown in the figure is a comparison schematic diagram of frame-by-frame annotation, all using bounding box annotation and the weak annotation method of this application. Compared with the frame-by-frame annotation method, the time cost of the annotation method in this application is reduced by 7.5 times; compared with the method of all using bounding box supervision, the introduction of dense annotation in the first frame of this application can supervise the model to generate more refined segmentation, and at the same time, using bounding box annotation for the other frames can ensure that the model learns the ability to locate the target object.

[0051] The weak annotation method provided by this application can greatly reduce the annotation cost of the video object segmentation dataset based on text description.

[0052] The weak annotation method provided by this application can be applied to the video object segmentation method based on text description, and can also be applied to other video segmentation tasks, such as video instance segmentation, single-sample video object segmentation.

[0053] For the above-mentioned weak annotation method, this application also proposes a video object segmentation method based on text description, so that the annotation of each frame can supervise the learning process of other frames, thereby making full use of the valuable dense annotation, and at the same time, can also give full play to the advantages of bounding box annotation.

[0054] Refer to Figure 3 , which shows a schematic flowchart of the video object segmentation method based on text description provided by the embodiments of this application.

[0055] As Figure 3 shown, a video object segmentation method based on text description may include:

[0056] S310. Obtain the video to be segmented and the language text description;

[0057] S320. Input the video to be segmented and the language text description into the video object segmentation model to perform segmentation prediction on the target in the video to be segmented; when training the video object segmentation model, use the annotated image of each frame of the video to be annotated obtained by the weak annotation method provided in the above embodiments to calculate the loss function to update the video object segmentation model.

[0058] In one embodiment, as Figure 4 shown, the video object segmentation model includes:

[0059] A language encoder for receiving a language text description and outputting language features;

[0060] A visual encoder for receiving each frame image of the video to be segmented and outputting corresponding visual features;

[0061] A visual enhanced language module for receiving language features and visual features and outputting enhanced language features;

[0062] A language enhanced visual module for receiving language features and visual features and outputting enhanced visual features;

[0063] A language-guided cross-frame segmentation module for receiving the enhanced visual features and the enhanced language features and outputting the objects in the video to be segmented for segmentation prediction.

[0064] Specifically, the language encoder can adopt a popular natural language processing model, such as Bert; while the visual encoder can adopt a popular visual backbone network, such as ResNet. The visual enhanced language module and the language enhanced visual module are implemented by the self-attention mechanism: in the visual enhanced language module, we use the visual features as the Key features and the Value features, and the language features as the Query features and input them into the self-attention mechanism; in the language enhanced visual module, the visual features are used as the Query features, and the language features are used as the Key features and the Value features and input into the self-attention mechanism.

[0065] In one embodiment, the language-guided cross-frame segmentation module includes a dynamic convolution kernel of the current frame and dynamic convolution kernels of other frames;

[0066] The dynamic convolution kernel of the current frame acts on the enhanced visual features of the current frame to generate the segmentation of the current frame;

[0067] The dynamic convolution kernels of other frames act on the enhanced visual features of the current frame to generate predictions for the current frame.

[0068] Specifically, the current frame image passes through the visual encoder and the language enhanced visual module to generate enhanced visual features, and the language text description passes through the language encoder and the visual enhanced language module to obtain enhanced language features. They are input into the dynamic convolution kernel branch, and the generated language-guided dynamic convolution kernel is responsible for acting on the enhanced visual features to generate predictions. Therefore, the language-guided dynamic convolution kernel has a representation related to the target object. Therefore, the dynamic convolution kernels in each frame also encode the representation of the object and can be used for the segmentation of the current frame. As Figure 5 shown, the dynamic convolution kernel f of the current frame θ 1 can act on the enhanced visual features of the current frame Generate the segmentation of the current frame; meanwhile, we can also use the dynamic convolutional kernels of other frames, such as to act on the enhanced visual features of the current frame to generate predictions for the current frame. In this way, we can use the dense annotation of the first frame to supervise all predictions, so as to realize the supervision of the dense annotation of the first frame on the learning of the model in other frames.

[0069] The language-guided cross-frame segmentation module provided in this embodiment can realize that the annotation of the current frame can supervise the learning process of the model in other frames by using the features of other frames for the segmentation of the current frame. That is, by using the language-guided cross-frame segmentation module, the annotation of each frame can supervise the learning process of other frames, so as to make full use of the precious dense annotation and also give full play to the advantages of the bounding box annotation.

[0070] When using bounding box annotations, the ability to distinguish between foreground and background is usually not strong. Therefore, this application proposes a unified contrast learning method to improve the model's ability to distinguish between foreground and background in features, which can be used for both dense annotation frames and bounding box annotation frames.

[0071] In one embodiment, continuing to refer to Figure 4 , the video object segmentation model further includes:

[0072] A cross-frame contrast learning module for enhancing the ability to distinguish between foreground and background in features;

[0073] The cross-frame contrast learning module adopts a cross-modal contrast learning method or a region consistency contrast learning method.

[0074] Among them, in the cross-modal contrast learning method, the language feature and the target region feature in the video to be segmented are positive samples, and the language feature and the background region are negative samples.

[0075] Among them, in the region consistency contrast learning method, the mean of the target region features in the video to be segmented and the features of all target object regions are positive samples, and the mean of the target region features in the video to be segmented and all background regions are negative samples.

[0076] In one embodiment, the method further includes distinguishing the foreground region and the background region:

[0077] For the first frame image of the video to be segmented with dense annotation, use the dense annotation label to distinguish the foreground region and the background region;

[0078] For the other frame images of the video to be segmented with bounding box annotation, take the outside of the bounding box region as the background region, and use the segmentation predicted by the video object segmentation model as a pseudo label. Select the regions with a confidence greater than the first threshold as the foreground region, and the regions with a confidence less than the second threshold as the background region, where the first threshold is greater than the second threshold.

[0079] After distinguishing the foreground region and the background region, a cross-modal contrastive learning method or a regional consistency contrastive learning method is adopted to add a contrastive learning loss function to train the ability of the video object segmentation model to distinguish between foreground and background features.

[0080] Specifically, the first threshold and the second threshold can be set according to actual needs. Exemplarily, the first threshold is 0.9 and the second threshold is 0.1.

[0081] This application proposes two cross-frame contrastive learning schemes, which are applicable to both densely annotated and bounding box annotated frames. One is the cross-modal contrastive learning method, that is, we consider that the language features should be positive samples with the target (i.e., object) region features in the video to be segmented, and the language features and the background region are negative samples. The other is the regional consistency contrastive learning method, that is, we consider that the mean of the features of the target object regions in the video to be segmented should be positive samples with the features of all target object regions, and the mean of the features of the target object regions in the video to be segmented and all background regions are negative samples.

[0082] For the first frame image with dense annotation in the video to be segmented, we can directly use the dense annotation label to distinguish the foreground and background regions; for the frames with only bounding box annotations, we can consider that the areas outside the bounding box are all background, and then use the segmentation predicted by the model as a pseudo-label, select the regions with a confidence greater than 0.9 as the foreground region, and the regions with a confidence less than 0.1 as the background region; the regions between the thresholds are considered as regions that cannot be distinguished and do not participate in subsequent contrastive learning.

[0083] After distinguishing the foreground and background features of the features, we can use the previously defined cross-modal contrast method and regional consistency contrast method to add a contrastive learning loss function to train the model's ability to distinguish between foreground and background features. The schematic diagram is as Figure 6 shown.

[0084] Among them, the formula of the contrastive learning loss function of the cross-modal contrastive learning method is as follows:

[0085]

[0086] Among them, R s is the language feature, and y fore , y back are the features in the foreground region and the background region respectively.

[0087] Among them, the formula of the contrastive learning loss function of the regional consistency contrastive learning method is as follows:

[0088]

[0089] Among them, and They are the mean of the foreground region features and the mean of the background region features respectively.

[0090] During the model training process, the above two loss functions can enable the features to learn the ability to distinguish the foreground region and the background region in the features.

[0091] The video object segmentation method based on text description provided by the embodiments of this application does not rely on densely annotated datasets, can learn from low-cost weakly annotated datasets, and achieves comparable performance to fully supervised methods.

[0092] Table 1 below shows the experimental simulation evaluation score comparison of the basic structure of this application (i.e., the video object segmentation model only includes a language encoder, a visual encoder, a visual enhancement language module, and a language enhancement visual module), the basic structure + language-guided cross-frame segmentation module (i.e., multi-modal cross-frame segmentation), the basic structure + language-guided cross-frame segmentation module + cross-frame contrast learning module (i.e., contrast learning), and the fully supervised method.

[0093] Table 1

[0094]

[0095] Among them, in Table 1, Ref-Youtube-VOS-Train_val is the dataset, J and F are evaluation metrics. Specifically, J is the region similarity score, F is the edge similarity score, and J&F is the mean of the two.

[0096] It should be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitations, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0097] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

Claims

1. A video object segmentation method based on text description, characterized in that, the method includes: obtaining the video to be segmented and the language text description; inputting the video to be segmented and the language text description into a video object segmentation model to perform segmentation prediction on the objects in the video to be segmented; when training the video object segmentation model, the loss function is calculated using the labeled images of each frame of the video to be labeled obtained by the weak labeling method to update the video object segmentation model; the weak labeling method includes: obtaining the video to be labeled; performing dense labeling on the first frame of the video to be labeled, and performing bounding box labeling on other frames except the first frame to obtain the labeled image of each frame of the video to be labeled; wherein, the first frame is the frame in which the object first appears in the video to be labeled; distinguishing the foreground region and the background region, including: for the first frame image of the video to be segmented using dense labeling, using the dense labeling label to distinguish the foreground region and the background region; for other frame images of the video to be segmented using bounding box labeling, taking the area outside the bounding box as the background region, and using the segmentation predicted by the video object segmentation model as a pseudo label, and selecting the region with a confidence greater than the first threshold as the foreground region and the region with a confidence less than the second threshold as the background region, wherein the first threshold is greater than the second threshold.

2. The method according to claim 1, characterized in that, the video object segmentation model includes: a language encoder for receiving the language text description and outputting language features; a visual encoder for receiving each frame image of the video to be segmented and outputting corresponding visual features; a visual enhancement language module for receiving the language features and the visual features and outputting enhanced language features; a language enhancement visual module for receiving the language features and the visual features and outputting enhanced visual features; a language-guided cross-frame segmentation module for receiving the enhanced visual features and the enhanced language features and outputting the objects in the video to be segmented for segmentation prediction.

3. The method according to claim 2, characterized in that, the language-guided cross-frame segmentation module includes a dynamic convolution kernel of the current frame and dynamic convolution kernels of other frames; the dynamic convolution kernel of the current frame acts on the enhanced visual features of the current frame to generate the segmentation of the current frame; the dynamic convolution kernels of other frames act on the enhanced visual features of the current frame to generate predictions for the current frame.

4. The method according to claim 2, characterized in that, the video object segmentation model further includes: a cross-frame contrast learning module for enhancing the foreground-background discrimination ability in the features; the cross-frame contrast learning module adopts a cross-modal contrast learning method or a region consistency contrast learning method.

5. The method according to claim 4, characterized in that, in the cross-modal contrast learning method, the language features and the object region features in the video to be segmented are positive samples, and the language features and the background region are negative samples.

6. The method according to claim 4, characterized in that, In the regional consistency contrast learning method, the mean of the target region features in the video to be segmented and the features of all target object regions are positive samples, and the mean of the target region features in the video to be segmented and all background regions are negative samples.

7. According to the method described in claim 1, it is characterized in that after distinguishing the foreground region and the background region, a cross-modal contrast learning method or a regional consistency contrast learning method is used to add a contrast learning loss function to train the video object segmentation model for the ability to distinguish feature foreground and background.

Citation Information

Patent Citations

  • Video target segmentation method based on weak supervised learning

    CN114743002A

  • Method for enhancing audio-visual association by adopting self-supervised curriculum learning

    US20220165171A1