Motion expression video segmentation method and system based on task decoupling
By decoupling the RVOS task into video instance segmentation and motion expression understanding, and using the decoupled motion expression video segmentation framework to process video and description text, the problem of performance degradation in the existing RVOS method in MEVS tasks is solved, and a more efficient motion expression video segmentation effect is achieved.
Patent Information
- Application Number
- CN202510327481.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-19
AI Technical Summary
Existing RVOS methods have significantly reduced performance when processing motion expression video segmentation (MEVS) tasks, especially on datasets of multi-objective and inter-frame motion-associated features, making it difficult to effectively capture the features of motion across time.
Using a task-decoupled motion expression video segmentation method, by decoupling the task into video instance segmentation and motion expression understanding, video and description text are input to the decoupled motion expression video segmentation framework for processing. The frozen video instance segmentation and decoupled motion expression video segmentation module, including a motion expression encoder and a motion query decoder, generate motion expression features and identify the object specified by the description text.
At lower training costs, better performance is achieved and the motion expression video segmentation task can be effectively handled, especially on datasets of multi-objective and inter-frame motion-associated features.
Smart Images

Figure CN120182894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video segmentation, and particularly relates to a motion expression video segmentation method and system based on task decoupling. Background Art
[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.
[0003] RVOS refers to video object segmentation, which is a multi-modal video task aiming to segment a specified target object throughout a video according to a given language expression. Current RVOS methods can be divided into two types: multi-stage methods and single-stage methods. The multi-stage methods use image-level segmentation models to independently process each frame of a video clip. Representative works such as URVOS (Unified Referring video object segmentation) first perform initial mask prediction through an image-level segmentation model and then use a semi-supervised VOS method for mask propagation refinement. Inspired by DETR (Detection Transformer), single-stage methods based on the Transformer architecture have been widely proposed. The prior art models the RVOS task as a sequence prediction problem and processes the video and text simultaneously through a single multi-modal Transformer model. For example, the proposed MTTR (Multimodal Tracking Transformer) is trained end-to-end, without text-related inductive bias components and without the need for additional mask optimization post-processing steps; ReferFormer is proposed, which introduces a small set of object queries related to the language input condition, and all queries are forced to only search for the referred object, and object tracking is naturally achieved by connecting the corresponding queries across frames.
[0004] Existing RVOS datasets mainly focus on single prominent objects and static attributes, and these target objects have clear and invariant features. Therefore, the above RVOS methods can achieve good results on these datasets. To overcome the limitation of only focusing on single prominent objects and static attributes, the prior art proposes motion expression video segmentation (MEVS), which emphasizes the importance of spatio-temporal motion features in videos; the dataset proposed along with the paper is called MeViS, which contains a large number of motion expressions for referring to target objects in complex environments, and one expression may refer to multiple target objects. Current RVOS methods are designed specifically for traditional RVOS datasets and their performance drops significantly when applied to the MEVS task.
[0005] Due to the multi-objective and inter-frame motion correlation characteristics of the dataset, researchers have tried to inject text information into the video instance segmentation (VIS) model to handle this new task. The VIS task aims to detect, segment, and track all object instances in a video simultaneously. One of the main challenges in MEVS is to accurately locate and align the capture of cross-time motion, so offline VIS methods are much more popular. LMPM (Language-guided Motion Perception and Matching) replaces the random initial query with a language-conditioned query and selects the matching object trajectories through thresholding, simply converting the offline VIS model VITA into a MEVS model. The DsHmp framework introduces an expression decoupling method, uses Mask2Former to segment possible targets based on static clues, and utilizes motion clues to segment the referential targets; and proposes a hierarchical motion perception module to effectively capture temporal information at different time scales, and also uses contrastive learning methods to distinguish the motion between visually similar objects.
[0006] DsHmp uses static clues to segment as many candidate objects as possible, but does not have the true labels of all objects specified by the static clues. For example, the text description is "a bird flying away". DsHmp trains Mask2Former to first segment all the birds, but in fact, only the flying away bird is supervised, which is inconsistent with the goal of decoupling video segmentation into static perception and motion perception. In addition, existing methods train the entire framework from scratch, which is difficult to optimize and has a high training cost. Summary of the Invention
[0007] To overcome the deficiencies of the above-mentioned prior art, the present invention provides a motion expression video segmentation method and system based on task decoupling, constructs a decoupled motion expression video segmentation framework based on the existing VIS model, emphasizes decoupling the task into video instance segmentation and motion expression understanding, and achieves better performance with a lower training cost compared to previous methods.
[0008] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:
[0009] In a first aspect, the present invention provides a motion expression video segmentation method based on task decoupling, including:
[0010] Obtain a video and a description text;
[0011] Input the video and the description text into a decoupled motion expression video segmentation framework for processing, and identify the objects specified by the description text in the video;
[0012] The decoupled motion expression video segmentation framework includes a frozen video instance segmenter and a decoupled motion expression video segmentation module. The decoupled motion expression video segmentation module includes a motion expression encoder and a motion query decoder. The frozen video instance segmenter receives the video, tracks and segments all candidate objects in the video, generates frame queries and video queries, and inputs them into the motion expression encoder. Motion word features and static word features are extracted from the description text and input into the motion expression encoder. The motion expression encoder interacts the frame queries, video queries with the motion word features and static word features to generate motion expression features. At the same time, the motion query is initialized using the video query. The initialized motion query and the motion expression features are input into the motion query decoder for decoding to obtain the decoded motion query, and the object specified in the description text is recognized based on the decoded motion query.
[0013] In a further technical solution, the decoupled motion expression video segmentation framework further includes a classification head and a mask generation head. The decoded motion queries are respectively input into the classification head and the mask generation head to obtain classification scores and mask embeddings.
[0014] In a further technical solution, the target decoder in the frozen video instance segmenter generates mask features, and the mask features are multiplied by the mask embeddings to obtain predicted masks.
[0015] In a further technical solution, the description text is input into a text encoder to extract word features, and the word features are input into a text decomposer to respectively extract motion word features and static word features.
[0016] In a further technical solution, the motion expression encoder interacts the frame queries, video queries with the motion word features and static word features to generate motion expression features specifically as follows:
[0017] The static word features are injected into the frame queries using cross-attention to obtain enhanced frame queries, and the motion word features are injected into the video queries using cross-attention to obtain enhanced video queries.
[0018] The word features are pooled to obtain sentence features, and the sentence features and the enhanced frame queries are interacted using serial cross-attention to obtain sentence features integrating frame-level static target information.
[0019] The sentence features integrating frame-level static target information and the enhanced video queries are interacted using serial cross-attention to obtain motion expression features.
[0020] In a further technical solution, initializing the motion query using the video query specifically means: the frozen video instance segmenter outputs the classification scores corresponding to each video query, and a set number of queries with the highest classification scores are selected from the video queries to form the motion query.
[0021] In a further technical solution, the motion query decoder decodes the motion query, the word features output by the text encoder, and the motion expression features to obtain a decoded motion query.
[0022] In a second aspect, the present invention provides a motion expression video segmentation system based on task decoupling, including:
[0023] A data acquisition module configured to: acquire a video and a description text;
[0024] A segmentation prediction module configured to: respectively input the video and the description text into a decoupled motion expression video segmentation framework for processing, and identify the object specified by the description text in the video;
[0025] The decoupled motion expression video segmentation framework includes a frozen video instance segmenter and a decoupled motion expression video segmentation module. The decoupled motion expression video segmentation module includes a motion expression encoder and a motion query decoder. The frozen video instance segmenter receives the video, tracks and segments all candidate objects in the video, generates frame queries and video queries, and inputs them into the motion expression encoder. Motion word features and static word features are extracted from the description text and input into the motion expression encoder. The motion expression encoder interacts the frame queries, video queries with the motion word features and static word features to generate motion expression features. At the same time, the motion query is initialized with the video query, and the initialized motion query and the motion expression features are input into the motion query decoder for decoding to obtain a decoded motion query, and the object specified by the description text is identified based on the decoded motion query.
[0026] In a third aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps in a motion expression video segmentation method based on task decoupling as described in the first aspect are implemented.
[0027] In a fourth aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in a motion expression video segmentation method based on task decoupling as described in the first aspect are implemented.
[0028] The above one or more technical solutions have the following beneficial effects:
[0029] The present invention constructs a decoupled motion representation video segmentation framework based on an off-the-shelf query-based VIS model, emphasizing decoupling the task into video instance segmentation and motion representation understanding. Specifically, first, a frozen video instance segmenter is used to track and segment all candidate objects, and then the specified objects are recognized according to the motion representation. Only a few parameters are required to train the decoupled motion representation video segmentation module; in order to further encode visually enhanced text information, a motion representation encoder is proposed, which interacts the frame queries with static cues, focusing on potential candidate objects of a specific category, and interacts the video queries with motion cues, focusing on the targets of a specific motion in a category-independent manner; in order to reduce the optimization difficulty, a method for initializing the motion queries based on video queries is proposed, and the video queries are selected according to the classification score Top K to initialize the motion queries to eliminate the interference of background queries; finally, the mask and classification results are generated through the motion query decoder.
[0030] The present invention proposes a simple and effective decoupled motion representation video segmentation framework DMVS, which is built on an off-the-shelf query-based VIS model and emphasizes decoupling motion representation video segmentation into video instance segmentation and motion representation understanding.
[0031] The present invention proposes a motion representation encoder that interacts the frame queries and video queries with static and motion cues respectively to further encode motion representations with enhanced visual information.
[0032] The present invention proposes a novel query initialization strategy, that is, using video queries guided by classification priors to initialize motion queries, which greatly reduces the optimization difficulty and training cost. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0034] Figure 1 is the architecture diagram of the decoupled motion representation video segmentation framework according to the embodiment of the present invention;
[0035] Figure 2 is the qualitative comparison diagram of the motion representation video segmentation method of the embodiment of the present invention and DsHmp on the MeViS validation set. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0037] It should be noted that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly dictates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0038] In the case of no conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0039] Term explanation: Motion expression video segmentation aims to segment the referred video objects according to the input motion description. Task decoupling refers to decoupling the motion expression video segmentation task into video instance segmentation and motion expression understanding.
[0040] Motion expression video segmentation (MEVS) aims to segment the objects in a video according to the input motion description. Compared with the traditional referring video object segmentation (RVOS) that focuses on a single salient object and static expression, it pays more attention to motion expression and multi-object segmentation, so it is more challenging. The classic RVOS model performs poorly on this task, and existing work completes this task by simply injecting text information into the video instance segmentation (VIS) model. However, this requires retraining the entire model, which is difficult to optimize and has poor performance.
[0041] Embodiment 1
[0042] As Figure 1 shown, this embodiment discloses a motion expression video segmentation method based on task decoupling, and the method includes the following steps:
[0043] S1: Obtain a video and a description text;
[0044] S2: Input the video and the description text into a decoupled motion representation video segmentation framework for processing, and identify the objects specified in the description text in the video; the decoupled motion representation video segmentation framework includes a frozen video instance segmentation model and a decoupled motion representation video segmentation module, and the decoupled motion representation video segmentation module includes a motion representation encoder and a motion query decoder; the frozen video instance segmentation model receives the video, tracks and segments all candidate objects in the video, generates frame queries and video queries and inputs them into the motion representation encoder; extract motion word features and static word features from the description text and input them into the motion representation encoder; the motion representation encoder interacts the frame queries, video queries with the motion word features and static word features to generate motion representation features; at the same time, initialize the motion query with the video query, input the initialized motion query and the motion representation features into the motion query decoder for decoding to obtain the decoded motion query, and identify the objects specified in the description text based on the decoded motion query.
[0045] The decoupled motion representation video segmentation framework DMVS adopts a motion representation video segmentation method, which is divided into three stages. In the first stage, a frozen video instance segmentation model (VITA) is used as the video instance segmentation model to identify, segment and track all objects. In the second stage, based on the frame query Q f and video query Q v generated by VITA, the motion representation encoder is used to interact the frame query Q f , video query Q v with the motion cue F m , static cue F s to generate motion representation features F me with enhanced visual information. In the third stage, the motion query Q m is initialized with the video query guided by the classification prior, and then the motion query decoder is used to decode the motion query layer by layer for classification and mask prediction. Finally, the motion query Q m generates mask embeddings through the mask generation head, multiplies the mask embeddings with the mask feature F mask to obtain the predicted mask, and selects those with scores higher than the threshold as the output.
[0046] (I) Video Instance Segmentation Model (VITA)
[0047] In this embodiment, VITA is adopted as the video instance segmenter, which is an offline VIS method based on the image instance segmentation model. VITA achieves video-level understanding by associating frame-level object queries without using spatio-temporal backbone features. Specifically, given a video input with a resolution of H×W consisting of T frames, VITA first processes each frame in a completely frame-independent manner using Mask2Former, generating two parts: 1) frame queries containing object information for each frame where N f is the number of object queries and C is the number of channels. 2) Mask features (image features) are generated from the object decoder where S is the stride of the feature map. Then, the object encoder establishes temporal communication between the frame queries by using the self-attention mechanism along the time axis. Finally, the object decoder aggregates the information from the frame queries Q f to the video query which is finally used to predict the class and mask of the objects in the video in one go.
[0048] The video instance segmenter VITA consists of three parts: a frame-level detector, an object encoder, and an object decoder. First, VITA uses the frame-level detector Mask2Former to process each frame in a completely frame-independent manner and outputs frame queries. Then the object encoder uses the window self-attention mechanism to perform inter-frame interaction between the frame queries. Finally, the object decoder aggregates the frame query information into the video query to output the final video query.
[0049] (2) Decoupled Motion Representation Video Segmentation Module
[0050] (1) Motion Representation Encoder
[0051] The interaction between text and visual features before query decoding can enhance the information of each modality. The present invention believes that frame queries and video queries can provide sufficient object-specific visual information without interacting with dense backbone features. Frame queries independently represent all objects in each frame and require static information to determine the object category. Video queries, on the other hand, integrate global context information and are more suitable for interacting with motion information in the time domain. Therefore, the present invention decomposes the given expression into static and action words, which serve as clues for static-aware frame queries and motion-aware video queries.
[0052] As Figure 1 shown, for the descriptive text "a cat turning around to play with a toy", first use the text encoder BERT to extract word features where K wRepresents the maximum number of words in a sentence in the dataset. Use a text decomposer to identify nouns, adjectives, and prepositions in the sentence from word features to obtain static clues (static word features), such as "cat" and "toy". At the same time, extract verbs and adverbs to obtain motion clues (motion word features), such as "turn around" and "play". Therefore, extract static word features and motion word features where K s / K m represents the length of static / motion words. The word feature F w is the static word feature F s plus the motion word feature F m , and the sentence feature (sentence-level feature) F se ∈R C .
[0053] First, use static word features and motion word features respectively to enhance the target query. Specifically, use the cross-attention mechanism to inject static clues into the frame query:
[0054]
[0055] where is the frame query enhanced by static clues.
[0056] Similarly, inject motion clues into the video query through cross-attention:
[0057]
[0058] where is the video query enhanced by motion clues.
[0059] Dshmp only allows static clues and dynamic clues to interact with the target query at the beginning of initialization. The randomly initialized query does not contain any object clues, so effective text information cannot be extracted. In fact, it relies on the combination of all text features in the decoding stage to establish modal associations. What this invention uses are the frame query and video query that already contain all object information, combines them with static and motion clues respectively, and fully utilizes the important role of separated text features.
[0060] Then, this invention uses the target query containing specific hint clues to enhance the sentence feature. In particular, use serial cross-attention to encode the existing sentence feature:
[0061]
[0062] where F′ se ∈R C represents the sentence feature integrating frame-level static target information, Fme ∈R C represents the sentence features integrating video-level moving object information, and it is also the final encoded motion expression used in the subsequent decoding process.
[0063] (2) Motion Query Decoder
[0064] The premise of the motion query decoder is to initialize the motion query. Due to the lack of additional information, the object queries in the VIS model are mainly randomly initialized. The RVOS method usually treats the language as a query, for example, initializing the object query by copying the sentence embedding. However, no matter which method is adopted, due to the significant gap between the output query after final decoding and the object, there are still certain difficulties in optimization. Ideally, RVOS mainly selects the VIS output based on text information. Therefore, the present invention proposes to initialize the motion query through video queries. In VITA, the default number of queries is relatively large. The present invention selects the object query based on the Top K of the classification score to eliminate the interference of background queries:
[0065] Q m = TopK(Q v , S v , N m )
[0066] where is the motion query, S v is the classification score corresponding to each video query output by the VIS model, and Top K represents selecting the highest N v classification scores from Q m queries. Through this initialization method, the motion query almost contains the queries corresponding to all instances, and the decoding process changes from identifying and segmenting objects from scratch to matching the most suitable object from all instances and refining it. This enables the optimization process to focus on the reference subtask, thereby reducing the optimization difficulty and training cost.
[0067] Next, the motion query decoder is used to extract information from the action expression features instead of frame queries or video queries. Specifically, the action query uses the cross-attention mechanism to interact with the word features, extracts the most basic lexical-level information, and determines the importance of each word. Then, it interacts with the visually enhanced sentence embedding through the cross-attention mechanism to extract the overall action expression representation:
[0068] Q′ m = Decoder(Q m , F w , F me )
[0069] where, Q′ mDecoded motion queries for final classification and mask prediction.
[0070] A decoder consists of two cross-attention layers, one self-attention layer, and one feed-forward neural network layer. The motion query decoder can effectively capture the video context and aggregate the motion expression information into motion queries. Therefore, the motion query decoder exhibits a fast convergence speed, achieves high accuracy, and significantly reduces the training memory compared to previous RVOS methods.
[0071] (3) Model Training and Inference
[0072] Finally, the motion query Q' output from the motion query decoder m is passed to the classification head H c and the mask generation head H m :
[0073] S = H c (Q' m )
[0074] where is the binary classification score, and H c is a single-layer linear layer.
[0075] M = F mask · H m (Q' m )
[0076] where is the predicted mask, and H m is three MLP layers to generate mask embeddings, and · represents the dot product operation.
[0077] Finally, the masks with classification scores greater than the threshold in all predicted masks are output to identify the object masks specified by the description text. In this embodiment, the threshold is selected as 0.7.
[0078] The proposed DMVS module is attached to the video instance segmenter in the present invention, and the entire model is trained end-to-end. It should be noted that the frame-level output of Mask2Former and the video-level output of VITA are not used for loss calculation, and only the video-level output loss of the DMVS module is considered. The total learning loss of the model is as follows:
[0079] L total = L mask + λ cls L cls
[0080] where L mask is the mask loss, which consists of binary cross-entropy loss and Dice loss; L cls is the classification loss.
[0081] The present invention uses the entire video as the input for the inference process. DMVS learning represents the motion queries that refer to instance in the whole video. In traditional object video segmentation, when dealing with a dataset with single-object expressions, the mask with the highest prediction score is selected as the final prediction result. For the MeViS dataset with multi-object expressions, the masks with prediction confidence scores greater than the confidence threshold σ are selected as the final prediction results.
[0082] The present invention designs a decoupled motion representation video segmentation (DMVS) method, which is constructed based on the existing VIS model and emphasizes decoupling the task into video instance segmentation and motion representation understanding. It first uses a frozen video instance segmenter to track and segment all candidate objects, and then identifies the specified objects according to the motion representation. Only the DMVS module with fewer parameters needs to be trained. The DMVS module proposes a motion representation encoder, which interacts with frame queries and video queries using static and motion cues respectively to further encode the motion representation enhanced by visual information; a novel query initialization strategy is proposed, which uses video queries guided by classification priors to initialize the motion queries, greatly reducing the optimization difficulty. Compared with the best existing technologies, the present invention achieves better performance with lower training costs, as shown in the following table:
[0083] Table 1
[0084] Method Number of learnable parameters GPU memory occupied during training Training time J&F metric DsHmp 102.3M 22G 17 hours 46.4 DMVS 12.7M 6G 7 hours 48.6
[0085] Extensive experiments on three benchmark datasets show that the DMVS of the present invention is superior to the state-of-the-art method DsHmp in both qualitative and quantitative metrics. Figure 2 It is a qualitative comparison between the method of the present invention and the main comparison method DsHmp on the MeViS validation set.
[0086] Example 2
[0087] The present embodiment discloses a motion representation video segmentation system based on task decoupling, including:
[0088] A data acquisition module, which is configured to: acquire a video and a description text;
[0089] A segmentation prediction module, which is configured to: respectively input the video and the description text into a decoupled motion representation video segmentation framework for processing, and identify the objects specified in the description text in the video;
[0090] The decoupled motion representation video segmentation framework includes a frozen video instance segmenter and a decoupled motion representation video segmentation module. The decoupled motion representation video segmentation module includes a motion representation encoder and a motion query decoder. The frozen video instance segmenter receives the video, tracks and segments all candidate objects in the video, generates frame queries and video queries, and inputs them into the motion representation encoder. Motion word features and static word features are extracted from the description text and input into the motion representation encoder. The motion representation encoder interacts the frame queries, video queries with the motion word features and static word features to generate motion representation features. At the same time, the motion query is initialized with the video query, and the initialized motion query and the motion representation features are input into the motion query decoder for decoding to obtain the decoded motion query, and the object specified in the description text is recognized based on the decoded motion query.
[0091] Embodiment III
[0092] The purpose of this embodiment is to provide a computing device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method in Embodiment I are implemented.
[0093] Embodiment IV
[0094] The purpose of this embodiment is to provide a computer-readable storage medium. A computer-readable storage medium stores a computer program, and when the program is executed by a processor, the steps of the method in Embodiment I are executed.
[0095] The steps involved in the devices in the above Embodiments III and IV correspond to those in Method Embodiment I. For specific implementation manners, reference may be made to the relevant description part of Embodiment I. The term "computer-readable storage medium" should be understood to include a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0096] Those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computer device. Optionally, they can be implemented by program codes executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. The present invention is not limited to any specific combination of hardware and software.
[0097] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0098] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A motion expression video segmentation method based on task decoupling, characterized in that: include: Get the video and description text; Inputting the video and the description text into the decoupled motion expression video segmentation framework for processing, respectively, to identify the object specified by the description text in the video; The decoupled motion expression video segmentation framework includes a frozen video instance segmenter and a decoupled motion expression video segmentation module, and the decoupled motion expression video segmentation module includes a motion expression encoder and a motion query decoder; the frozen video instance segmenter receives the video and tracks and segments all candidate objects in the video, generates frame queries and video queries and inputs them into the motion expression encoder; extracts motion word features and static word features from the description text and inputs them into the motion expression encoder; the motion expression encoder interacts the frame query, video query with the motion word features and static word features to generate motion expression features; at the same time, the video query is used to initialize the motion query, the initialized motion query and the motion expression features are input into the motion query decoder for decoding, and the decoded motion query is obtained, and the object specified by the description text is identified based on the decoded motion query.
2. The motion expression video segmentation method based on task decoupling as claimed in claim 1, characterized in that: The decoupled motion expression video segmentation framework also includes a classification head and a mask generation head. The decoded motion query is input into the classification head and the mask generation head respectively to obtain a classification score and a mask embedding.
3. The motion expression video segmentation method based on task decoupling as claimed in claim 2, characterized in that: The target decoder in the frozen video instance segmenter generates mask features, and the mask features are multiplied with the mask embedding to obtain a predicted mask.
4. The motion expression video segmentation method based on task decoupling as claimed in claim 1, characterized in that: The description text is input into a text encoder to extract word features, and the word features are input into a text decomposer to respectively extract motion word features and static word features.
5. The motion expression video segmentation method based on task decoupling as claimed in claim 1, characterized in that: The motion expression encoder interacts the frame query, video query, motion word features, and static word features to generate motion expression features: Using cross attention to inject static word features into frame queries to obtain enhanced frame queries, and using cross attention to inject motion word features into video queries to obtain enhanced video queries; Pooling word features to obtain sentence features, using serial cross attention to interact the sentence features with enhanced frame queries to obtain sentence features that integrate frame-level static target information; Serial cross attention is used to interact sentence features that integrate frame-level static object information and enhanced video queries to obtain motion expression features.
6. The motion expression video segmentation method based on task decoupling as claimed in claim 1, characterized in that: The specific method of using video query to initialize motion query is as follows: the frozen video instance segmenter outputs a classification score corresponding to each video query, and a set number of queries with the highest classification scores are selected from the video queries to form a motion query.
7. The motion expression video segmentation method based on task decoupling as claimed in claim 6, characterized in that: The motion query decoder decodes the motion query, word features and motion expression features output by the text encoder to obtain a decoded motion query.
8. A motion expression video segmentation system based on task decoupling, characterized in that: include: A data acquisition module is configured to: acquire video and description text; A segmentation prediction module is configured to: input the video and the description text into a decoupled motion expression video segmentation framework for processing, and identify the object specified by the description text in the video; The decoupled motion expression video segmentation framework includes a frozen video instance segmenter and a decoupled motion expression video segmentation module, and the decoupled motion expression video segmentation module includes a motion expression encoder and a motion query decoder; the frozen video instance segmenter receives the video and tracks and segments all candidate objects in the video, generates frame queries and video queries and inputs them into the motion expression encoder; extracts motion word features and static word features from the description text and inputs them into the motion expression encoder; the motion expression encoder interacts the frame query, video query with the motion word features and static word features to generate motion expression features; at the same time, the video query is used to initialize the motion query, the initialized motion query and the motion expression features are input into the motion query decoder for decoding, and the decoded motion query is obtained, and the object specified by the description text is identified based on the decoded motion query.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a motion expression video segmentation method based on task decoupling as described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the motion expression video segmentation method based on task decoupling as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Semantic decoupling-based combined action recognition method for self-attention model
CN115953832A
Subject Tracking Systems for a Movable Imaging System
US20180025498A1
System and method for self-supervised video transformer
US20240169692A1
Cited By
Abnormality monitoring processing method and system based on artificial intelligence
CN121259750A