Long video target inference segmentation method based on context mark prompt

Through the CiCVS method, contextual labeling prompts and multimodal feature fusion modules are used to solve the consistency and accuracy problems of target segmentation in long videos, and high-quality target tracking and segmentation in long-term video sequences are achieved.

CN120707859AActive Publication Date: 2025-09-26DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510857459.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-26
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

Existing technologies find it difficult to accurately track and segment targets in long video sequences, especially when faced with challenges such as occlusion, intense motion, and long-term temporal dependencies, making it difficult to maintain consistency and high precision.

Method used

The Context-infused Consistent Video Segmentor (CiCVS) method is adopted to generate accurate and coherent long-term target trajectories through a contextual labeling prompt strategy, combined with a pre-trained image encoder, a multi-layer perceptron mapping module, a multimodal feature fusion module and a mask propagator.

Benefits of technology

It effectively solves the problem of consistent tracking and segmentation in long videos, improves the accuracy and robustness of target segmentation in long-term videos, and especially achieves more accurate target tracking in complex scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707859A_ABST
    Figure CN120707859A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of image segmentation, and discloses a context mark prompt-based long video target reasoning segmentation method, which comprises a pre-training image encoder, a multi-layer perceptron mapping module, a multi-modal feature fusion module, a large language model and a mask spreading device. The method comprises the following steps: sampling support frames from equally divided video clips, and processing the support frames and key frames together through a pre-trained image encoder and a multi-layer perceptron mapping module to obtain corresponding visual features; the multi-modal feature fusion module injects the visual features of the reference expression and the support frame into potential queries through a plurality of fusion modules to generate enriched potential queries; the enriched potential queries guide the large language model to generate key frames and full video level lt; sEGgt, SEGgt; and finally, the SAM2-based mask spreading device is used for accurately decoding and continuously and consistently spreading the SAM2-based mask spreading device in all frames. According to the method, the problems of long-distance dependence modeling and consistency tracking are solved through context mark prompt and a multi-modal feature fusion module.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image segmentation and relates to a long video target reasoning and segmentation method based on contextual labeling prompts. Background Art

[0002] Reasoning Video Object Segmentation leverages multimodal large language models to accurately segment and track objects in videos based on semantically rich and often implicit expressions of human intent. Unlike traditional referential video object segmentation methods, the core challenge of reasoning video object segmentation lies in how to interpret expressions that encode subtle human intent or complex spatiotemporal relationships, and based on this, achieve pixel-level, consistent, and accurate segmentation of implicit objects throughout the entire video sequence. Several recently proposed benchmark datasets have further advanced the development of this emerging field. Such technological advances not only improve the robustness and sophistication of video understanding, but also open up broader prospects for downstream applications such as autonomous driving and embodied intelligence.

[0003] Early approaches in this field include VISA, which combines a multimodal large language model with SAM for keyframe selection and segmentation in “Visa: Reasoning video objects segmentation via large language models”, and then achieves mask propagation through object tracking; Bai et al. introduced a sparse-dense sampling strategy in “One token to segthem all: Language instructed reasoning segmentation in videos” to capture richer spatiotemporal features and use a single special marker to complete multi-frame segmentation; Zheng et al. built on this by incorporating visual cues into text embedding in “Villa: Video reasoning segmentation with large language model” and strengthening temporal modeling with the help of a video frame decoder; Wei et al. achieved the unification of pixel-level perception in images and videos by mixing entity recognizers and fine-grained visual perceptrons in “Hyperseg: Towards universal visual segmentation with large language model”, thereby improving object localization accuracy; At the same time, Gong et al. used temporal tokens in “The devil is in temporal token: High quality video reasoning” to achieve the goal of segmenting objects in videos. Segmentation" emphasizes the importance of temporal tokenization and focuses on efficient global context extraction and optimized key frame selection to obtain better segmentation performance.

[0004] Long-term video object segmentation aims to achieve accurate tracking and segmentation of objects in extended video sequences, addressing challenges such as occlusion, intense motion, and long-term temporal dependencies. This task was first proposed by Hong et al. in "Lvos: A benchmark for long-term video objects segmentation." They also proposed the DDMemory method, which consists of three complementary memory banks designed to effectively capture temporal information. They also introduced the LVOS benchmark, which covers longer video scenes, to further advance related research. In "Learning spatial-semantic features for robust video objects segmentation," Li et al. proposed a robust video object segmentation framework that effectively addresses object ambiguity in long videos by learning spatial-semantic features and discriminative object queries. In "Livos: Light video objects segmentation with gated linear matching," Liu et al. introduced a lightweight memory network to address the high memory consumption associated with long-term mask propagation. In “Sam2: Segment anything in images and videos,” Ravi et al. proposed a flexible and hintable video segmentation framework that can perform segmentation at arbitrary granularity, enabling more accurate object tracking in complex scenes. Building on this, Ding et al. further extended SAM2 by introducing a memory tree in “Sam2long: Enhancing sam2 for long video segmentation with a training-free memory tree,” improving segmentation performance without the need for additional training. Furthermore, in “Panovos: Bridging non-panoramic and panoramic views with transformer for video segmentation,” Yan et al. proposed a specialized scheme for panoramic video object segmentation, extending the application of long-term VOS to specialized fields. While most previous methods have focused on single-modality video segmentation, our method uniquely combines long-term video sequences with implicit language instructions, establishing a new paradigm for long-term reasoning in video object segmentation. Summary of the Invention

[0005] To overcome these limitations, we propose the Context-infused Consistent Video Segmentor (CiCVS). CiCVS explicitly leverages full video-level reasoning context to guide a multimodal large language model (MLLM) in generating accurate and coherent long-term target trajectories.

[0006] The technical solution of the present invention:

[0007] A long video object reasoning and segmentation method based on contextual labeling prompts is proposed. A contextual labeling prompt strategy is proposed, which includes a pre-trained image encoder, a multi-layer perceptron mapping module, a multimodal feature fusion module, a large language model and a mask propagator. The contextual labeling prompt strategy first samples support frames from equally divided video clips, and together with the key frames that need to be explicitly segmented, encodes the support frames and key frames into corresponding visual features through the pre-trained image encoder and the multi-layer perceptron mapping module. Subsequently, the multimodal feature fusion module injects the reference expression and the visual features of the support frames into the potential query through multiple fusion modules, thereby generating enriched potential queries. These enriched potential queries guide the large language model to generate key frames and full video level <seg>The markers are finally accurately decoded by the SAM2-based mask propagator and propagated consistently across all frames; specifically:

[0008] (a) Supports frame- and keyframe-driven video clip and visual feature extraction

[0009] Input video with text prompts Divide the video into multiple segments, randomly extract one frame from each segment as the support frame, and form a support set ; In the support set Four frames are uniformly sampled from each video as key frames of the current video to form a key frame subset , keep the input video Diverse temporal dynamics captured in;

[0010] The support frames contain rich contextual information of the video. Key frames are randomly extracted from the support frames. The support frames and key frames are used to guide the large language model (llava-phi3-v) to generate key frames and full video level. <seg>markers, thus enabling more precise spatial awareness in equally divided video segments.

[0011] First, use the pre-trained image encoder (CLIP, contrastive language-image pre-training model) and multi-layer perceptron mapping module (MLP) from the support set and a subset of keyframes Extract visual features from the support frame features and record them as , the key frame features are recorded as :

[0012]

[0013] (b) Multimodal fusion and latent query injection with text priors

[0014] In order to distill temporal information from support frames and inject target-specific semantics, a multimodal feature fusion module is used here; the multimodal feature fusion module introduces input videos with text prompts The text prompt in the reference expression , and integrate the reference expression With support for frame features Injected into learnable latent queries, it provides more detailed prior information for large language models, thereby providing compact and expressive representations for subsequent reasoning.

[0015] Specifically, for each support frame, 64 potential queries are randomly initialized and generated, and all potential queries are collectively referred to as Next, the reference expression Text embedding is formed by encoding the embedding layer of a large language model , the text embedding is copied and compared with the corresponding support frame features Splicing to form a multimodal sequence . Then, the multimodal sequence Integrate into In [1], the process is completed by a set of fusion modules, each of which consists of a multi-head cross attention layer and a feedforward network layer; The layer process is expressed as:

[0016]

[0017]

[0018] in, and denote the randomly initialized multi-head cross attention layer and feedforward network layer, 、 and Respectively represent The input, intermediate, and output potential queries of the layers; this design enables the potential queries to gradually absorb information from the semantic space and context, thereby generating more expressive and goal-oriented representations.

[0019] For the network structure of the contextual tag prompt strategy, the output of the multimodal feature fusion module is expressed as:

[0020]

[0021] in, represents the potential query after enrichment, and MIC stands for multimodal feature fusion module.

[0022] (c) Segmentation token embedding aggregation of large language models and mask propagation of SAM2

[0023] Then, the enriched potential queries , keyframe features and dialogue templates (i.e., “segment the target in each target frame and throughout the video”) is fed into a large language model, which autoregressively generates responses , which embeds keyframes and full video level <seg>Tag; select function It is based on keyframe and full video level <seg>The size of the index value of the tag is extracted from the large language model at the key frame and full video level. <seg>Label and form frame-level embedding and video-level embedding :

[0024] In order to input video In order to form a unified understanding of the target, the strategy based on cosine similarity will <seg>Mark and aggregate to obtain an information-enriched <seg>Mark; Finally, a mask propagator based on SAM2 is used to pass the input video Image features interact to decode information-rich <seg>Mark, generate segmentation mask, and then input video to generate a dense mask sequence, and then The mask is propagated to produce dense mask tracks.

[0025] Training loss:

[0026] In the training phase, a hybrid loss function of DICE loss and binary cross entropy loss is used to supervise the predicted mask sequence, and the text output is optimized by standard cross entropy loss. In the inference phase, key frames are identified based on the occlusion score of SAM2, and its memory mechanism is used to propagate the segmentation mask in the video. The mask loss is combined with binary cross entropy loss and DICE loss:

[0027]

[0028]

[0029] in, represent the real text and the real mask respectively, is the corresponding predicted value, and the weight coefficient and Set as and 0.5; , and They represent the text binary cross entropy loss function, image binary cross entropy loss function and DICE loss function respectively.

[0030] The present invention provides a contextually labeled object segmentation method that addresses the inherent challenges of long-term video reasoning and segmentation. This method effectively addresses the problems of long-range dependency modeling and consistency tracking through contextual labeling and a multimodal feature fusion module. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 Schematic diagram of the network structure of the long video object reasoning and segmentation method based on contextual labeling prompts.

[0032] Figure 2 Schematic diagram of the multimodal feature fusion module. DETAILED DESCRIPTION

[0033] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0034] like Figure 1 As shown in the paper, a long video object reasoning segmentation method based on contextual labeling hints combines two key components: contextual labeling hint strategy and multimodal feature fusion module.

[0035] Step 1: Structured Sampling Strategy

[0036] Input video with text prompts After uniformly dividing the video into sub-sets, a frame is randomly sampled from each sub-set as the support frame of the sub-set to form a support set ; In the support set Four frames are uniformly sampled from each video as key frames of the current video to form a key frame subset .

[0037] Step 2: Image encoding

[0038] The obtained key frame and support frame are input into the pre-trained image encoder And in the multi-layer perceptron mapping module, the corresponding key frame features are obtained and support frame features , and then support frame features Input into the multimodal feature fusion module.

[0039] Step 3: Multimodal Fusion and Latent Query Injection with Text Priors

[0040] like Figure 2 As shown, in this module, we first support frame for each frame All are randomly initialized to generate 64 potential queries Next, the reference expression Encoded as text embedding through the embedding layer of the multimodal large model , the text embedding is copied and compared with the corresponding support frame features Splicing to form a multimodal sequence .

[0041] Then, Integrate into In [1], the process is completed by a set of fusion blocks, each of which consists of a multi-head cross attention layer and a feedforward network. The process of layer l can be expressed as:

[0042]

[0043]

[0044] in, and Represent the multi-head cross attention layer and feedforward network layer respectively, 、 and Respectively represent Potential queries for the input, intermediate, and output of a layer.

[0045] For the network structure of the contextual tag prompt strategy, the output of the multimodal feature fusion module can be expressed as:

[0046]

[0047] represents the potential query after enrichment, and MIC stands for multimodal feature fusion module.

[0048] Step 4: Embedding Generation and <seg>Marker extraction

[0049] We will enrich the query Keyframe features and customized conversation prompts (i.e. “segment the target in each target frame and throughout the video”) is fed into a large language model, which autoregressively generates responses , which embeds keyframe-level and video-level <seg>Mark. We select the function Extract keyframe and full video level information from large language models <seg>Mark, get frame-level embedding and video-level embedding .

[0050] Step 5: Loss Function Design

[0051] In order to maintain the consistency of segmentation between video frames, we adopt a robust SAM2-based segmentation module to generate <seg>The labels are decoded into pixel-level segmentation masks covering the entire video sequence. The loss function of the entire network combines binary cross entropy loss (BCE) and DICE loss:

[0052]

[0053]

[0054] in, represent the real text and the real mask respectively, is the corresponding predicted value, and the weight coefficient and Set as and 0.5. , and They represent the text binary cross entropy loss function, image binary cross entropy loss function and DICE loss function respectively.

[0055] In summary, this method designs a long video object reasoning and segmentation method based on contextual labeling cues, which can perform high-quality long video object reasoning and segmentation.

[0056] Datasets and evaluation criteria

[0057] LLaVA-Phi-3V, at 3.8B, was used as the large language model and fine-tuned using LoRA. During training, only the mask decoder and multimodal feature fusion module of SAM2 were optimized; all other parameters remained frozen. Images were treated as single-frame videos, and VideoQA input was directly passed to the large language model after encoding, omitting the multimodal feature fusion stage.

[0058] The number of potential queries was set to 64, and the number of layers in the fusion module was set to 3. The optimizer used was AdamW, with a learning rate of 0.0003, combined with the Warmup DecayLR learning rate scheduler, and a 100-iteration warmup phase. Training was performed on four NVIDIA A100 GPUs, with a total of 6,000 iterations and a batch size of 128, implemented with a batch size of 1 per device and 32 steps of gradient accumulation.

[0059] Experimental verification results

[0060] Table 1: Comparative analysis on the ReVOS dataset: A detailed comparison of the performance of the method of the present invention and the existing methods on the ReVOS dataset

[0061]

[0062] In this table, the data in bold represent the best results, and the data in horizontal lines represent the second-best results. The same applies to the following tables. Notably, CiCVS-3.8B outperforms VRS-HQ-13B by 1.4% on both the J&F metric and the reasoning metric. These performance improvements highlight the effectiveness of our method in leveraging temporal context to capture fine-grained object dynamics and complex spatial relationships, and also validate the efficiency of the proposed multimodal frame compression strategy.

[0063] Table 2: Evaluation on RVOS dataset: Detailed comparison of the performance of the proposed method and existing methods on three RVOS datasets. The indicators are all comprehensive performance indicators.

[0064]

[0065] The table above compares CiCVS with the previous state-of-the-art RVOS method. On the Ref-Youtube-VOS and Ref-DAVIS17 datasets, CiCVS outperforms VRS-HQ-13B by 1.3% and 0.8% in the J and F metrics, respectively, highlighting its strong generalization capabilities in typical scenarios involving explicit spatial relationships and diverse object interactions. Furthermore, on the motion-aware MeViS dataset, CiCVS improves J by 2.3%, F by 2.8%, and J&F by 2.5%, demonstrating its ability to capture complex object motion while maintaining temporal consistency.< / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg>

Claims

1. A method for object segmentation in long videos based on contextual labeling, characterized in that: This method for object segmentation inference in long videos proposes a contextual labeling hinting strategy, which includes a pre-trained image encoder, a multi-layer perceptron mapping module, a multimodal feature fusion module, a large language model, and a mask propagator. The contextual labeling hinting strategy first samples support frames from equally divided video segments. The support frames and key frames are then encoded into visual features of the support frames and key frames through a pre-trained image encoder and a multi-layer perceptron mapping module. Subsequently, the multimodal feature fusion module injects the reference expression and the visual features of the support frame into the potential query through multiple fusion modules to generate an enriched potential query; the enriched potential query guides the large language model to generate key frames and full video level <seg> The markers are finally accurately decoded by the SAM2-based mask propagator and propagated consistently across all frames.< / seg> 2. The long video target reasoning and segmentation method according to claim 1, characterized in that: Here are the steps: (a) Supports frame- and keyframe-driven video clip and visual feature extraction; Input video with text prompts Divide the video into multiple segments, randomly extract one frame from each segment as the support frame, and form a support set ; In the support set The key frames of the current video are uniformly sampled from each video to form a key frame subset ; Then use the pre-trained image encoder and multilayer perceptron mapping modules from the support set and a subset of keyframes The visual features are extracted from the support frame and the visual features are recorded as , the visual features of the key frames are recorded as : ,(b) Multimodal fusion with text prior and latent query injection; Multimodal feature fusion module introduces input video The text prompt in the reference expression , and the reference expression With support for frame features After integration, it is injected into the learnable potential query; For each frame support frame All potential queries are randomly initialized and all potential queries are collectively referred to as Next, the reference expression Text embedding is formed by encoding the embedding layer of a large language model , the text is embedded in The visual features of the corresponding supporting frames are copied Splicing to form a multimodal sequence ; Then, the multimodal sequence Integrate into In the process, a fusion module is used to complete the process. Each fusion module consists of a multi-head cross attention layer and a feed-forward network layer. The layer process is expressed as: , , in, and denote the randomly initialized multi-head cross attention layer and feedforward network layer, 、 and Respectively represent Potential queries at the layer's input, intermediate, and output; For the network structure of the contextual tag prompt strategy, the output of the multimodal feature fusion module is expressed as: , in, represents the potential query after enrichment, and MIC stands for multimodal feature fusion module; (c) Segmentation token embedding aggregation of large language models and mask propagation of SAM2; Then, the enriched potential queries , keyframe features and dialogue templates Input into the large language model, which generates responses autoregressively , which embeds keyframes and full video level <seg>Mark; according to <seg>The size of the index value of the token extracted from the large language model <seg>Mark, through the function Forming frame-level embedding and video-level embedding : ,< / seg> < / seg> < / seg> In order to input video In order to form a unified understanding of the target, the strategy based on cosine similarity will <seg>Mark and aggregate to obtain an information-enriched <seg>Mark; Finally, a mask propagator based on SAM2 is used to pass the input video Image features interact to decode information-rich <seg>Mark, generate segmentation mask, and then input video The network is propagated to produce a dense mask sequence.< / seg> < / seg> < / seg> 3. The long video target reasoning and segmentation method according to claim 2, characterized in that: In the training phase, a hybrid loss function of DICE loss and binary cross entropy loss is used to supervise the predicted mask sequence, and the text output is optimized by standard cross entropy loss. In the inference phase, key frames are identified based on the occlusion score of SAM2, and its memory mechanism is used to propagate the segmentation mask in the video. The mask loss is combined with binary cross entropy loss and DICE loss: , ,in, represent the real text and the real mask respectively, is the corresponding predicted value, and the weight coefficient and Set as and 0.5; , and They represent the text binary cross entropy loss function, image binary cross entropy loss function and DICE loss function respectively.

Citation Information

Patent Citations

  • Language-guided video target anaphora segmentation method

    CN118658091A

  • Long video understanding method based on iterative hierarchical key frame selection

    CN119785258A

  • High-performance video reasoning segmentation method based on time sequence marking

    CN120107854A

  • Unified referring video object segmentation network

    US20210383171A1

  • Systems and methods for video and language pre-training

    US20230154146A1

Cited By

  • SAM2-based multi-small-target tracker and tracking method

    CN121074767A

  • A multi-small-target tracker and tracking method based on SAM2

    CN121074767B

  • Image screening method and system based on visual large model

    CN121366340A

  • Behavior labeling method based on end-to-end humanoid robot large model and related equipment

    CN121705817A