Zero sample learning-based reference video target segmentation method and system
By adopting zero-sample learning and multi-grained feature fusion methods in the reference video target segmentation task, using pre-trained models to extract multi-grained visual features and perform cross-modal matching, the problem of relying on a large amount of annotated data and difficulty in dealing with the movement of target objects in the video in the prior art is solved, and efficient video target segmentation is achieved.
Patent Information
- Application Number
- CN202510042681.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-10
AI Technical Summary
Existing methods for referring to video target segmentation rely on a large number of densely annotated pixel-level data. The distribution gap between training data and test data leads to performance degradation, and data collection is laborious, especially in the absence of training annotations, it is difficult to effectively perform cross-modal understanding and motion feature extraction of target objects in video.
Using a zero-sample learning method, the transfer ability of the pre-trained model is used to extract multi-grained visual features (target granularity, frame granularity and video granularity), combined with a parameterless cross-modal attention mechanism and a pre-trained action recognition model, the visual features and text features are matched to achieve video target segmentation.
Without training samples, the model's generalization ability in complex video scenes and language descriptions is improved, the motion clues in the video are effectively extracted, visual feature representation is enhanced, and the performance of referential video target segmentation is improved.
Smart Images

Figure CN120125813A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object recognition, and in particular to a reference video object segmentation method and system based on zero-shot learning. Background Art
[0002] With the rapid development of big data technology, the Internet, and social media, multi-modal data such as pictures, texts, and videos are everywhere, and single-modal text or picture information can no longer meet people's needs. Therefore, cross-research between the two deep learning fields of computer vision and natural language processing has become a hot topic. Different from tasks involving only a single modality, multi-modal tasks require more fine-grained information extraction from two or more different modalities of data, and at the same time, analyze and reason about relationship information based on the correlation between different modalities of data. That is, a multi-modal model not only needs to fully understand the content of a single modality but also bridge the semantic gap between different modality information and effectively model and interact with multi-modal information. Among them, one of the representative research tasks is the reference video object segmentation task. Referring video object segmentation (R-VOS) segments the target object described by a given natural language expression in a video sequence and has potential wide applications, such as language-based video editing, intelligent video surveillance, user-friendly human-computer interaction, etc.
[0003] The reference video object segmentation task is very challenging because it involves cross-modal understanding. Most existing reference video object segmentation methods only consider the closed-set reference video object segmentation problem and use a large amount of densely annotated pixel-level data and their corresponding text descriptions for deep training based on convolutional neural network (CNN) or Transformer models. When there is a large distribution gap between the training data and the test data, the performance of the model may drop sharply. In addition, collecting such a large amount of annotated data is very laborious. In extreme cases, the data and labels may not be accessible. To solve these problems, the present invention focuses on the zero-shot reference video object segmentation task without available training annotations and explores multi-grained visual feature expressions by leveraging the significant zero-shot transfer ability of pre-trained models.
[0004] Although CLIP can align visual features and text features in the same semantic space, how to use CLIP to achieve pixel-level prediction is still a major obstacle. When applying CLIP trained on the entire image to the pixel-level prediction task, there is a distribution gap from the whole to the pixel. To narrow this gap, attempts have been made in the semantic segmentation field by relying on a large amount of annotations. However, the high cost of annotations still limits the application and scalability of the method, which obviously goes against the original intention of reducing cumbersome annotations.
[0005] Moreover, when using the image-based method to transfer to the video domain, it will not be able to distinguish target objects with similar appearance features but performing different actions in a video sequence, which may lead to confusion of target objects in consecutive frames. And videos have temporal continuity, and the motion cues in time are often ignored in the reference video object segmentation task. Therefore, how to use temporal information to focus on the motion properties of target objects is a major challenge.
[0006] To further improve the performance of the reference video object segmentation task, a large amount of work focuses on the alignment between the visual modality and the language modality. Previous work has used simple concatenation operations, dynamic convolutions, or cross-modal attention mechanisms to achieve the alignment between modalities. The recent progress is to achieve more effective alignment through a cross-modal Transformer decoder. Although the performance has been improved, these methods require additional training with a large amount of labeled data and may overfit to specific data distributions. And large-scale pre-trained models such as CLIP have powerful representation and transfer capabilities. Therefore, how to make full use of the advantages of CLIP to promote the alignment between the visual modality and the language modality without additional labeled data training is a question worth exploring.
[0007] In traditional reference video object segmentation methods, a bottom-up or top-down process is usually adopted. An intuitive idea is to directly use the image-level reference method to process video frames separately, such as RefVOS. However, the obvious defect of this method is that they cannot utilize the valuable temporal information between frames, and may lead to inconsistent target predictions due to scene or appearance changes. To solve such problems, URVOS proposes to decouple the reference video object segmentation task into a two-stage task of reference image segmentation and mask propagation.
[0008] The structures of the two models, RefVOS and URVOS, are as Figure 1 and Figure 2 shown. RefVOS processes each frame independently, so it is applicable to both images and videos. It uses the optimal visual and language feature extractors, combines them into multi-modal features, and feeds them into the decoder to generate the binary mask of the target object. While URVOS uses ResNet-50 as the encoder, and uses the features of the fourth stage and the fifth stage (Res4 and Res5) respectively to estimate the memory and cross-modal attention. The memory attention feature and the cross-modal attention feature are gradually combined in the decoder.
[0009] In summary, most of the existing referential video object segmentation methods train neural network models based on a large amount of densely annotated pixel-level data and corresponding language expressions. However, collecting such a large amount of annotated data is very time-consuming and labor-intensive. In extreme cases, the data and labels may not be available. In addition, some work has been proposed to decouple the segmentation task into two stages: candidate mask generation and object-level mask retrieval. Such methods usually use candidate masks to cover the original image to obtain the image corresponding to the current candidate mask. However, this often ignores the global semantic information. Worse yet, when applying this method based on object-level visual features to the video domain, this method often fails to distinguish two object targets with similar appearance features but performing different actions in a video clip, which may lead to target confusion in consecutive frames. In addition, in terms of cross-modal alignment, previous work has used simple concatenation, dynamic convolution, or cross-modal attention mechanisms to complete visual-language alignment. However, these methods all require additional data for training. Summary of the Invention
[0010] An object of the present invention is to provide a referential video object segmentation method and system based on zero-shot learning to solve at least one of the technical problems in the above-mentioned background art.
[0011] To achieve the above object, the present invention adopts the following technical solutions:
[0012] In a first aspect, the present invention provides a referential video object segmentation method based on zero-shot learning, including:
[0013] Obtain the referential video data to be segmented;
[0014] Process the obtained referential video data by using a pre-trained referential video object segmentation model to obtain a final object segmentation result; wherein, the referential video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module; the multi-granularity visual feature extraction module is used to extract object-granularity visual features, frame-granularity visual features, and video-granularity visual features, and aggregate the object-granularity visual features, frame-granularity visual features, and video-granularity visual features to obtain multi-granularity temporal visual features; the text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features; the matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
[0015] Further, in the extraction of visual features at the target granularity, the target mask is used to perform a dot product with the image tokens obtained by the CLIP visual encoder to extract the visual features at the target granularity; in the extraction of visual features at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by the text features; in the extraction of visual features at the video granularity, the pre-trained action recognition model X-CLIP of the existing work is loaded, the significant regions of the moving objects are focused on, and the motion features of the target objects are extracted; the visual features at the target granularity, frame granularity, and video granularity extracted are aggregated to obtain the final visual features.
[0016] In a second aspect, the present invention provides a reference video object segmentation system based on zero-shot learning, including:
[0017] An acquisition module, configured to acquire the reference video data to be segmented;
[0018] A segmentation module, configured to process the acquired reference video data by using a pre-trained reference video object segmentation model to obtain a final object segmentation result; wherein, the reference video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module; the multi-granularity visual feature extraction module is configured to extract visual features at the target granularity, frame granularity, and video granularity, and aggregate the visual features at the target granularity, frame granularity, and video granularity to obtain multi-granularity temporal visual features; the text feature extraction module is configured to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features; the matching module is configured to match the multi-granularity visual features and the text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
[0019] Further, in the extraction of visual features at the target granularity, the target mask is used to perform a dot product with the image tokens obtained by the CLIP visual encoder to extract the visual features at the target granularity; in the extraction of visual features at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by the text features; in the extraction of visual features at the video granularity, the pre-trained action recognition model X-CLIP of the existing work is loaded, the significant regions of the moving objects are focused on, and the motion features of the target objects are extracted; the visual features at the target granularity, frame granularity, and video granularity extracted are aggregated to obtain the final visual features.
[0020] In a third aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions, and when the computer instructions are executed by a processor, the reference video object segmentation method based on zero-shot learning as described in the first aspect is implemented.
[0021] Fourthly, the present invention provides a computer device, including a memory and a processor, where the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the zero-shot learning-based referential video object segmentation method as described in the first aspect.
[0022] Fifthly, the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the zero-shot learning-based referential video object segmentation method as described in the first aspect.
[0023] Term Explanation:
[0024] Zero-shot learning: Zero-shot learning aims to enable a pre-trained model to predict class labels of previously unknown data, that is, data samples that do not exist in the training data.
[0025] Referential video object segmentation: The referential video object segmentation task aims to predict the pixel-level mask of the target object referred to by a given language expression on the target frame in a video.
[0026] Instance segmentation: Instance segmentation aims to detect the objects in an image and then assign class labels to each pixel of the objects.
[0027] Multi-granularity features: It refers to combining or integrating features extracted at different levels (such as object level, frame level, and video level) to obtain a more comprehensive and rich feature representation.
[0028] Action recognition: Action recognition aims to classify human action categories from video clips
[0029] Multi-modal feature alignment: It refers to establishing an association or mapping between feature representations of different perceptual modalities, so that features of different modalities are closer or aligned in a shared representation space.
[0030] CLIP: CLIP refers to Contrastive Language-Image Pretraining, which is a pre-trained neural network model trained on a large amount of paired Internet data. This model was initially used to match images and texts and is common in the multi-modal field and can be used for text-image retrieval.
[0031] Transformer: A Transformer is a deep learning model that adopts an encoder-decoder architecture. By introducing the self-attention mechanism, it achieves better sequence modeling capabilities, enabling the model to maintain global attention while processing the input sequence in parallel, thus achieving remarkable success in various tasks.
[0032] Advantages of the present invention: A deep learning network model for referring video object segmentation based on zero-shot learning with multi-granularity feature fusion introduces a multi-granularity visual feature extraction module, including an object-granularity visual feature extraction stage, a frame-granularity visual feature extraction stage, and a video-granularity feature extraction stage. The object-granularity visual feature extraction stage integrates the global semantic information of video frames. The frame-granularity visual feature extraction stage can further align the visual modality and the text modality. The video-granularity visual feature extraction stage can effectively extract the motion cues in the video, enhancing the visual feature representation. The multi-granularity visual features effectively improve the segmentation performance.
[0033] The advantages of the additional aspects of the present invention will be more clearly given in the following description part, or understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0035] Figure 1 It is a framework diagram of the RefVOS model in the prior art.
[0036] Figure 2 It is a framework diagram of the URVOS model in the prior art.
[0037] Figure 3 It is a schematic diagram of the framework structure of the referring video object segmentation model based on zero-shot learning according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The following details the embodiments of the present invention. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention, and should not be construed as a limitation of the present invention.
[0039] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs.
[0040] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0041] Those skilled in the art can understand that, unless specifically stated, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or their groups.
[0042] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. Without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0043] For ease of understanding the present invention, the following further explains the present invention with specific embodiments in conjunction with the accompanying drawings, and the specific embodiments do not constitute a limitation on the embodiments of the present invention.
[0044] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.
[0045] The multi-granularity feature fusion method provided by the present invention for the zero-shot referential video object segmentation task is used to improve the generalization ability of the model in complex video scenarios and language descriptions without training samples.
[0046] The referential video object segmentation task is very challenging because it involves cross-modal understanding. Most existing referential video object segmentation methods consider the closed-set referential video object segmentation problem and train deep models based on convolutional neural networks or Transformers with a large amount of densely annotated pixel-level data and their corresponding language descriptions. When there is a large distribution gap between the training data and the test data, the performance may drop sharply. In addition, collecting such a large amount of annotated data is very laborious. In extreme cases, neither the training data nor the labels may be available. To solve these problems, the present invention focuses on using the zero-shot learning method to solve the referential video object segmentation task without available training annotations, and explores multi-granularity feature representations by leveraging the zero-shot transfer ability of significant pre-trained models.
[0047] In terms of feature representation, most existing work focuses on visual feature representations at the frame granularity and object granularity. Among them, object granularity features can provide object-level information, but object granularity features may lack the context information of the image and cannot effectively extract the fine-grained information of the image. In the video domain, the temporal correlation and action clues in the video are often ignored. Just using object granularity and frame granularity features cannot effectively distinguish two target objects with similar appearance features but performing different actions in the video, which may lead to confusion of the targets in consecutive frames.
[0048] In terms of modality alignment, most existing work adopts concatenation, dynamic filters or cross-modal attention mechanism modules to achieve cross-modal alignment. Although the performance has been improved, these methods still need to use labeled data for additional training and are prone to overfitting to specific data.
[0049] Therefore, the present invention designs a deep learning network model for referential video object segmentation with multi-granularity feature fusion based on zero-shot learning to address the problem that motion clues in the video are often ignored and the problem that modality alignment requires additional training. In the feature extraction stage, a multi-granularity feature extraction strategy is adopted to extract visual features at the object granularity, frame granularity, and video granularity respectively to enhance the visual feature representation. In terms of feature alignment, considering that the visual features and text features extracted by CLIP are already in the same semantic space, it is considered to discard the linear layer in the traditional attention mechanism and use a parameter-free cross-modal attention mechanism for alignment between the visual modality and the text modality.
[0050] The network of the method for reference video object segmentation based on zero - shot learning is built using the deep learning framework PyTorch, and the multi - granularity feature fusion method for the reference video object segmentation task is implemented in three steps: (1) Use the pre - trained instance segmentation model FreeSOLO to segment the input image into several candidate object masks; (2) Use the multi - granularity visual feature extraction method to extract visual features at the object granularity, frame granularity, and video granularity from bottom to top, and aggregate the three types of hierarchical features; (3) In the text feature extraction stage, focus on the sentence features and central word features of the text expression respectively, and fuse the two types of features. (4) Select the mask corresponding to the visual feature with the highest matching degree as the final prediction according to the matching degree between the multi - granularity visual feature and the text feature.
[0051] The network model of the method for reference video object segmentation based on zero - shot learning consists of three parts: a candidate mask generation module, a multi - granularity visual feature and text feature extraction module, and a multi - granularity visual feature and text feature matching module. The candidate mask generation module is used to generate several candidate masks of arbitrary categories. The multi - granularity visual feature and text feature extraction module is used to extract and fuse hierarchical feature representations. The multi - granularity visual feature and text feature matching module is used to match the two - modality features and select the mask corresponding to the visual feature with the highest matching degree as the final prediction according to the matching degree. In the visual feature extraction stage at the object granularity, the dot product is taken between the object mask and the image tokens obtained through the CLIP visual encoder instead of between the mask and the original image to extract the visual features at the object granularity; in the visual feature extraction stage at the frame granularity, a parameter - free cross - modal attention mechanism is used to obtain the frame - level visual features enhanced by the text features; in the visual feature extraction stage at the video granularity, a pre - trained action recognition model is used to focus on the significant regions of moving objects and extract the motion features of the target objects; the visual features at the object granularity, frame granularity, and video granularity extracted are aggregated to further match the object mask.
[0052] Embodiment 1
[0053] In this Embodiment 1, first, a reference video object segmentation system based on zero-shot learning is provided, including: an acquisition module for acquiring the reference video data to be segmented; a segmentation module for processing the acquired reference video data by using a pre-trained reference video object segmentation model to obtain a final object segmentation result. The reference video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module. The multi-granularity visual feature extraction module is used to extract visual features at the object granularity, frame granularity, and video granularity, and aggregate the visual features at the object granularity, frame granularity, and video granularity to obtain multi-granularity temporal visual features. The text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features. The matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result. In the extraction of visual features at the object granularity, the object mask is used to perform a dot product with the image tokens obtained by the CLIP visual encoder to extract the visual features at the object granularity. In the extraction of visual features at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by text features. In the extraction of visual features at the video granularity, the pre-trained action recognition model X-CLIP is loaded to focus on the significant regions of moving objects and extract the motion features of the target objects. The extracted visual features at the object granularity, frame granularity, and video granularity are aggregated to obtain the final visual features.
[0054] In this embodiment, the above system is used to implement a reference video object segmentation method based on zero-shot learning. In this method, a deep learning network model for reference video object segmentation based on multi-granularity feature fusion is proposed. The multi-granularity visual features include object granularity, frame granularity, and video granularity. In the stage of extracting visual features at the object granularity, the object mask is used to perform a dot product with the image tokens obtained by the CLIP visual encoder instead of using the mask to perform a dot product with the original image to extract the visual features at the object granularity. In the stage of extracting visual features at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by text features. In the stage of extracting visual features at the video granularity, a pre-trained action recognition model is used to focus on the significant regions of moving objects and extract the motion features of the target objects. The extracted visual features at the object granularity, frame granularity, and video granularity are aggregated to enrich the visual semantic information and further complete the matching of the target object mask. More importantly, the entire model can be applied to the reference video object segmentation task without any training.
[0055] Such as Figure 3As shown in the figure, the model framework of this embodiment mainly includes a multi-granularity visual feature extraction module, a text feature extraction module, and a cross-modal feature matching module. The multi-granularity visual feature extraction module includes an object-granularity visual feature extraction stage, a frame-granularity visual feature extraction stage, and a video-granularity visual feature extraction stage. In the object-granularity visual feature extraction stage, first, a pre-trained instance segmentation model is used to generate a number of candidate masks. Instead of using the mask to dot product with the original image, the dot product is performed between a single candidate mask and the image tokens obtained through the CLIP visual feature encoder to extract the object-granularity visual features. In the frame-granularity visual feature extraction stage, a parameter-free cross-modal attention mechanism is used. The linear layer in the traditional attention mechanism is discarded, and the visual features and text features obtained through the CLIP visual encoder and text encoder are directly clicked to obtain a cross-modal attention matrix, which weights the visual features to obtain the frame-granularity visual features enhanced by the text features. In the video-granularity visual feature extraction stage, a pre-trained action recognition model is used to focus on the significant regions of moving objects and extract the motion features of the target objects. The object-granularity, frame-granularity, and video-granularity visual features extracted are aggregated to enrich the visual semantic information. In the text feature extraction stage, the sentence features and central word features of the text expression are respectively focused on, and the two features are fused. Finally, according to the matching degree between the multi-granularity visual features and text features, the mask corresponding to the visual feature with the highest matching degree is selected as the final predicted mask.
[0056] The network model mainly consists of three stages: feature extraction, feature fusion, and mask prediction.
[0057] In the feature extraction stage, the image and the corresponding text expression are input into the feature extraction module. For the text expression, it is respectively input into the sentence-level feature extraction branch and the word-level feature extraction branch using the CLIP text encoder to extract hierarchical text features. For the image, it is respectively input into the object-granularity, frame-granularity, and video-granularity visual feature extraction branches. In the object-granularity visual feature extraction stage, first, a pre-trained instance segmentation model is used to generate a number of candidate masks. Instead of using the mask to dot product with the original image, the dot product is performed between a single candidate mask and the image tokens obtained through the CLIP visual feature encoder to extract the object-granularity visual features. In the frame-granularity visual feature extraction stage, a parameter-free cross-modal attention mechanism is used to obtain the frame-granularity visual features enhanced by the text features. In the video-granularity visual feature extraction stage, a pre-trained action recognition model is used to focus on the significant regions of moving objects and extract the motion features of the target objects. In the text feature extraction stage, the sentence features and central word features of the text expression are respectively focused on, and the two features are fused.
[0058] In the feature fusion stage, the visually extracted object-level, frame-level, and video-level features are aggregated to form multi-granularity visual features, enriching the visual semantic information. The sentence-level text features and word-level text features are fused to form hierarchical text features.
[0059] In the mask prediction stage, according to the matching degree between the multi-granularity visual features and text features, the mask corresponding to the visual feature with the highest matching degree is selected as the final predicted mask.
[0060] In the feature extraction stage, the model FreeSOLO, the CLIP visual feature extractor, and the text feature extractor used to generate candidate instance object masks, as well as the action recognition model X-CLIP used to extract video motion features, are all initialized with pre-trained parameters, and the parameter initialization process only needs to be performed once. FreeSOLO is an unsupervised object segmentation model, with the full name Free Self-supervised Object Localization and Segmentation, whose goal is to directly achieve object localization and segmentation in images through unlabeled data. FreeSOLO does not require any annotation or supervision signal and performs excellently in the open-set object detection task. Its core method is to combine self-supervised learning and saliency detection to generate high-quality boundaries and masks of objects. X-CLIP is a multi-modal video understanding model, aiming to tightly combine text and video content to achieve efficient video semantic understanding and reasoning tasks. It is an extension of CLIP, expanding the capabilities of CLIP from image-text matching to video-text matching. Compared with static images, videos contain consecutive frames, so X-CLIP can capture the temporal dynamic information of videos. The present invention leverages the prior knowledge pre-trained by FreeSOLO and X-CLIP and loads the model weights pre-trained in previous work.
[0061] In the Referring Video Object Segmentation (RVOS) task, top-down and bottom-up approaches are two common framework design ideas for segmenting the object referred to by the text description in the video. The bottom-up approach starts from the semantic information of the text instruction and gradually projects it into the feature representation of the video to guide the segmentation of the object. The top-down approach, on the other hand, first performs a basic segmentation on all pixels or target regions in the video and then selects or adjusts the target region through the constraints of the text description. In this embodiment, a top-down two-stage method is used to achieve the zero-shot referring video object segmentation task. First, a pre-trained instance segmentation model, FreeSOLO, is used to generate a series of candidate masks for the target video frames, and the candidate mask with the highest matching degree is selected as the final prediction according to the matching degree between the extracted multi-granularity visual features and text features. Among them, the matching degree between the visual features and text features is calculated by Cosine Similarity. This method measures the directional similarity of the two modal features in the high-dimensional embedding space, thereby evaluating their matching degree. The multi-granularity visual feature extraction proposed in this embodiment includes the target-granularity visual feature extraction stage, the frame-granularity visual feature extraction stage, and the video-granularity visual feature extraction stage.
[0062] Considering that directly using the candidate mask to perform a dot product with the video frame will ignore the background information of the video frame, resulting in a video frame that only contains the candidate mask region, and such training data is not included in the training stage of CLIP. Using the CLIP visual encoder to directly extract features from such pictures with a transparent background will significantly reduce the model performance. Therefore, in order to integrate the global semantic information of the video frame, in the target-granularity visual feature extraction stage, the CLIP visual feature extractor is first used to extract features from the original video frame to obtain a series of feature tokens, and then the candidate mask is used to perform a dot product on the feature tokens. Such an operation can make the extracted visual features focus on the global semantic information.
[0063] Considering that directly applying the method of image segmentation to the video may confuse target objects with similar appearance features but performing different actions, therefore, in the video-granularity feature extraction stage, the pre-trained action recognition model X-CLIP is used to extract the motion cues in the video, enabling the model to find the moving target referred to by the text.
[0064] In terms of cross-modal semantic alignment, this embodiment uses a parameter-free cross-modal attention mechanism. Since the visual features and text features extracted by the large-scale pre-trained model CLIP can already be in the same semantic space, it is possible to consider discarding the linear layer in the traditional attention mechanism, directly taking the dot product of the visual features and text features to obtain a cross-modal attention matrix, and using this matrix to weight the visual features to further align the visual modality and text modality. The specific principle is as follows. First, extract the feature representations of the image and text from the CLIP model, which are F I and F t . For the convenience of subsequent similarity calculation and stability, perform L2 normalization to project the feature vectors onto the unit sphere. Then calculate the cosine similarity between each local image feature and the global text feature to form attention weights. Then use the calculated attention weights to perform weighted aggregation on the local features of the image to generate the global feature representation of the image strengthened by the text features.
[0065] Embodiment 2
[0066] This Embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the above-mentioned reference video object segmentation method based on zero-shot learning is implemented. The method includes:
[0067] Obtain the reference video data to be segmented;
[0068] Use the pre-trained reference video object segmentation model to process the obtained reference video data to obtain the final object segmentation result. Among them, the reference video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module. The multi-granularity visual feature extraction module is used to extract the visual features of the object granularity, frame granularity, and video granularity, and aggregate the visual features of the object granularity, frame granularity, and video granularity to obtain multi-granularity temporal visual features. The text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features. The matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
[0069] Embodiment 3
[0070] Embodiment 3 provides a computer device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute the above-mentioned reference video object segmentation method based on zero-shot learning. The method includes:
[0071] Obtain the reference video data to be segmented;
[0072] Process the obtained reference video data by using a pre-trained reference video object segmentation model to obtain a final object segmentation result. Among them, the reference video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module. The multi-granularity visual feature extraction module is used to extract visual features at the object granularity, frame granularity, and video granularity, and aggregate the visual features at the object granularity, frame granularity, and video granularity to obtain multi-granularity temporal visual features. The text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features. The matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
[0073] Embodiment 4
[0074] Embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program. Among them, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the above-mentioned reference video object segmentation method based on zero-shot learning. The method includes:
[0075] Obtain the reference video data to be segmented;
[0076] Process the obtained reference video data using a pre-trained reference video object segmentation model to obtain the final object segmentation result; wherein, the reference video object segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module, and a matching module; the multi-granularity visual feature extraction module is used to extract visual features at the object granularity, frame granularity, and video granularity, and aggregate the visual features at the object granularity, frame granularity, and video granularity to obtain multi-granularity temporal visual features; the text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features; the matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result.
[0077] In summary, for the reference video object segmentation method according to the embodiments of the present invention, a reference video object segmentation deep learning network model based on multi-granularity feature fusion is trained, wherein the multi-granularity visual features include object granularity, frame granularity, and video granularity. In the visual feature extraction stage at the object granularity, the object mask is used to perform a dot product with the image tokens obtained through the CLIP visual encoder instead of using the mask to perform a dot product with the original image, and the visual features at the object granularity are extracted; in the visual feature extraction stage at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by the text features; in the visual feature extraction stage at the video granularity, a pre-trained action recognition model is used to focus on the significant regions of the moving objects and extract the motion features of the target objects; the visual features at the object granularity, frame granularity, and video granularity extracted are aggregated to enrich the visual semantic information and further complete the matching of the target object mask. More importantly, the entire model can be applied to the reference video object segmentation task without any training.
[0078] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.
[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device to perform a series of operational steps on the computer or other programmable device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in multiple blocks.
[0082] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.
Claims
1. A method for segmenting a referential video object based on zero-shot learning, characterized in that: include: Obtaining reference video data to be segmented; The acquired reference video data is processed using a pre-trained reference video target segmentation model to obtain a final target segmentation result; wherein the reference video target segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module and a matching module; the multi-granularity visual feature extraction module is used to extract visual features of target granularity, visual features of frame granularity and visual features of video granularity, and aggregate the visual features of target granularity, frame granularity and video granularity to obtain multi-granularity visual features; The text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features; the matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
2. The method for segmenting a referential video object based on zero-shot learning according to claim 1, characterized in that: In the visual feature extraction at the target granularity, the target mask is used to perform a dot product with the image token obtained by the CLIP visual encoder to extract the visual features at the target granularity; in the visual feature extraction at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by text features; in the visual feature extraction at the video granularity, the existing pre-trained action recognition model X-CLIP is loaded to focus on the salient areas of the moving object and extract the motion features of the target object; the extracted visual features at the target granularity, frame granularity and video granularity are aggregated to obtain the final visual features.
3. A zero-shot learning-based video object segmentation system, characterized in that: include: An acquisition module, used for acquiring the reference video data to be segmented; A segmentation module is used to process the acquired reference video data using a pre-trained reference video target segmentation model to obtain a final target segmentation result; wherein the reference video target segmentation model includes a multi-granularity visual feature extraction module, a text feature extraction module and a matching module; the multi-granularity visual feature extraction module is used to extract visual features of target granularity, visual features of frame granularity and visual features of video granularity, and aggregate the visual features of target granularity, frame granularity and video granularity to obtain multi-granularity visual features; The text feature extraction module is used to extract the sentence features and central word features of the text expression, and fuse the two features to obtain text features; the matching module is used to match the multi-granularity visual features and text features, and select the mask corresponding to the visual feature with the highest matching degree as the final prediction result according to the matching degree.
4. The method for segmenting a referential video object based on zero-sample learning according to claim 3, characterized in that: In the extraction of visual features at the target granularity, the target mask is used to perform a dot product with the image token obtained by the CLIP visual encoder to extract the visual features at the target granularity; in the extraction of visual features at the frame granularity, a parameter-free cross-modal attention mechanism is used to obtain the frame-level visual features enhanced by text features; in the extraction of visual features at the video granularity, the pre-trained action recognition model is loaded. X -CLIP focuses on the salient areas of moving objects and extracts the motion features of the target objects; it aggregates the extracted visual features of the target granularity, frame granularity, and video granularity to obtain the final visual features.
5. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by the processor, the method for segmenting a reference video object based on zero-sample learning as described in claim 1 or 2 is implemented.
6. A computer device, characterized in that: It includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the zero-sample learning-based reference video target segmentation method as described in claim 1 or 2.
7. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes instructions for implementing the zero-sample learning-based reference video target segmentation method as described in claim 1 or 2.
Citation Information
Cited By
Zero sample anaphora image segmentation method based on text perception adapter
CN121725007A