High-performance video reasoning segmentation method based on time sequence marking

By using a multimodal large language model to generate timing marks in video inference segmentation, and performing timing dynamic aggregation and keyframe selection, the problems of insufficient timing context information capture and inaccurate keyframe positioning in the prior art are solved, and higher video inference segmentation performance is achieved.

CN120107854APending Publication Date: 2025-06-06DALIAN UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510174580.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Existing video inference segmentation methods are difficult to effectively capture the rich timing context information in videos, and the keyframe positioning is not accurate enough, which limits the application of the model in downstream tasks.

Method used

A high-performance video inference segmentation method based on timing marking is proposed, using a multimodal large language model to generate timing markings, and through timing dynamic aggregation and keyframe selection strategies based on timing marking, the model's perception of spatial features and temporal dynamics is improved.

Benefits of technology

Through the encoding and dynamic aggregation of timing marks, the model's perception of video timing information is significantly improved, and the accuracy of keyframe positioning is improved, thus performing better than the previous method in multiple video and image inference segmentation tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107854A_ABST
    Figure CN120107854A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of video processing methods, and discloses a high-performance video reasoning segmentation method based on time sequence marking. The method comprises the following specific steps: carrying out time sequence mark coding by using a multi-mode large language model, carrying out time sequence dynamic aggregation on a coded special mark, and carrying out mask decoding and propagation by using an SAM2 based on a key frame selection strategy of the time sequence mark. According to the method, multiple module designs are combined, and the time sequence marks generated by the multi-modal large model are utilized, so that the perception capability of the model on spatial features and time dynamics is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of video processing methods and relates to a method for performing high-performance video reasoning and segmentation using temporal tags generated by a multimodal large language model. Background Art

[0002] In recent years, the rapid development of artificial intelligence and multimodal large language models has brought new opportunities to the field of video segmentation. Traditional referential video segmentation tasks usually rely on users to provide clear language descriptions, and then segment the referred objects in the video based on them. However, with the diversification of user needs and the increase in scene complexity, segmentation technology that relies solely on explicit reference is difficult to meet the higher requirements for model reasoning capabilities in practical applications.

[0003] Recently, Lai et al. first proposed the reasoning segmentation task in LISA: Reasoning Segmentation via Large Language Model, which aims to use the implicit world knowledge in large language models to reason about text and images containing human intent and output the target referred to by the text in the form of pixel-level masks. This work pioneered a new paradigm, first introducing a new segmentation marker [SEG] to expand the vocabulary of the language model, then fine-tuning the language model through carefully designed prompt words and images to force it to include the segmentation marker in the output, and finally extracting the marker and inputting it and the image into the segmentation everything model (SAM) to obtain the segmentation mask. More work has emerged since then to further unlock the segmentation capabilities of large multimodal models. Yang et al. proposed in LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model to enhance the instruction fine-tuning data by combining key segmentation datasets, thereby addressing the limitations of LISA in distinguishing individual instances and supporting both semantic and instance segmentation tasks. In PixelLM:Pixel Reasoning with Large Multimodal Model, Ren et al. combined a novel pixel decoder and a segmentation code base with learnable annotations to generate high-quality segmentation masks. In GLaMM:Pixel Grounding Large Multimodal Model, Rasheed et al. proposed a grounded dialogue task and set a new benchmark, allowing users to interact with models at various granularities in text and vision domains.

[0004] Although much progress has been made in the research of segmentation by reasoning, most of the work still focuses on performing segmentation on a single image. Yan et al. proposed a video segmentation reasoning task in VISA: Reasoning Video Object Segmentation via Large Language Models. They used an external multimodal large model to select key frames based on the content of the video and prompt text, used another internal multimodal large model to perform reasoning and generate segmentation tags [SEG], input the tags into the segmentation model to obtain the segmentation mask of the key frame, and finally used the pre-trained video tracker to propagate the segmentation mask. However, this method only uses a single segmentation tag [SEG] to represent the segmentation mask, which makes it difficult to extract dynamic changes between frames and complex spatiotemporal representations. In addition, the use of an external multimodal large model for key frame selection may lead to inaccurate results in some complex scenarios and make the model unable to perform end-to-end reasoning, which greatly limits its application in downstream tasks. Summary of the invention

[0005] In order to fully capture the rich temporal context information in the video and improve the accuracy of key frame positioning, we proposed a high-performance video reasoning segmentation method based on temporal labeling. This method combines multiple module designs and uses the temporal labels generated by a multimodal large model to effectively improve the model's perception of spatial features and temporal dynamics.

[0006] The technical solution of the present invention:

[0007] A high-performance video inference segmentation method based on temporal tagging uses a multimodal large language model to encode temporal tags, dynamically aggregates the encoded special tags, selects key frames based on temporal tags, and uses Segmentation Some Model 2 (SAM2) for mask decoding and propagation. The specific steps are as follows:

[0008] Step 1: Use a multimodal large language model for temporal tag encoding

[0009] (1) Generate hierarchical tags

[0010] In order to make the multimodal large language model have the ability to perceive contextual information at the frame level and the time level in the video segmentation task, a method is proposed to encode the spatial information within the frame and the temporal relationship between frames into hierarchical tags. First, the frame-level special tag [SEG] and the time-level special tag [TAK] are used to expand the vocabulary of the multimodal large language model. Then a structured prompt word template is designed: "Please find {target} in the video and segment {target} at the frame level and time level respectively", where {target} represents the reference description of the object. Then, a given input video X is manually V ∈R T×3×H×WSampling is performed to obtain the sampled video frame X V′ , and convert the video frame X V′ and the prompt word template X after tokenization by the tokenizer txt The multimodal language model outputs a response y containing a frame-level special tag [SEG] and a time-level special tag [TAK] through autoregressive iteration. txt ; Where T is the total length of the input video, H and W represent the height and width of each frame respectively;

[0011] (2) Extracting and mapping hierarchical tags

[0012] After the multimodal large language model integrates the frame-level and time-level video context information into the frame-level special tag [SEG] and the time-level special tag [TAK] respectively, the special embedding corresponding to the frame-level special tag [SEG] and the time-level special tag [TAK] is extracted from the last hidden layer of the multimodal large language model. and Where T ′ is the total number of video frames after sampling, d ′ is the hidden layer dimension of the multimodal large language model, and the extracted special embedding is mapped through the learnable multi-layer perceptron MLP to project it into the feature space of SAM2, as shown in formula (1):

[0013]

[0014] Among them, h seg ,h tak They represent the embeddings corresponding to the mapped frame-level special tags [SEG] and time-level special tags [TAK] respectively; finally, the mapped embeddings are fed to the time-series dynamic aggregation module for spatiotemporal feature extraction;

[0015] Step 2: Timing Dynamic Fusion

[0016] (1) Select key frames by similarity of different tags

[0017] Since the frame-level special mark [SEG] reflects the position change of the target in each video frame, and the temporal-level special mark [TAK] reflects the semantic content of the target in the video, by calculating the cosine similarity of the special embedding corresponding to the temporal-level special mark [TAK] and each frame-level special mark [SEG], the video frame corresponding to the frame-level special mark [SEG] with the maximum cosine similarity is selected as the key frame of the overall model in the training stage;

[0018] (2) Weighted fusion based on cosine similarity

[0019] After calculating the cosine similarity of the special embeddings corresponding to the temporal level special tag [TAK] and each frame level special tag [SEG], the cosine similarity is normalized, and then the embedding h corresponding to the mapped frame level special tag [SEG] is converted into the weighted aggregation strategy shown in formula (2) with the normalized cosine similarity as the proportional coefficient. seg Aggregated to the embedded h corresponding to the time-level special tag [TAK] after mapping tak middle:

[0020]

[0021] Among them, α represents the fusion coefficient, which is set to 0.1; λ i represents the proportional coefficient, which is the normalized cosine similarity; h′ tak represents the fused temporal tag embedding, which is used in the key frame selection and segmentation mask generation process; i represents the index value of the video frame;

[0022] Step 3: Keyframe selection strategy based on temporal markers

[0023] In the inference stage of the overall model, a key frame selection strategy based on time series markup is proposed to decode the time series information embedded in the fused time series markup to select more accurate key frames.

[0024] In the inference stage, the CLIP model is first used to calculate the video frame with the highest similarity to the target description, and this video frame is used as the central frame for the given input video X. V Perform global sampling to obtain the sampled video frame

[0025] in, represents the number of video frames sampled in the inference stage, and f is the identifier of the inference stage; the sampled video frames The structured prompt word template is sent to the multimodal large language model to perform steps 1 and 2. Each sampled video frame is embedded with the fused temporal tag h′ tak The target occlusion score is generated by SAM2. The target occlusion score indicates the probability of the segmented target existing in the sampled video frame. The calculation method is shown in formula (3):

[0026]

[0027] Where E and MD represent the image encoder and mask decoder in SAM2, respectively. Indicates the target occlusion score corresponding to each sampled video frame;

[0028] Finally, the target occlusion score normalized by Softmax and the h corresponding to each video frame calculated in step 2 are used.seg and h tak The cosine similarity S t The sum is used as the criterion for determining the key frame, and the video frame with the highest sum value is X k Selected as key frames in the inference phase;

[0029] Step 4: Mask decoding and propagation using SAM2

[0030] The key frame is input into SAM2. The image encoder of SAM2 is responsible for extracting the multi-scale visual features of the key frame. The mask decoder of SAM2 embeds the multi-scale visual features of the key frame and the fused temporal mark into h′. tak Interact to decode the segmentation mask of the key frame As shown in formula (4):

[0031]

[0032] In the training phase, the two video frames adjacent to the key frame are extracted as unconditional frames and input into SAM2, and propagated using its mask memory mechanism to obtain the segmentation mask sequence. In the inference stage, all remaining video frames are input into SAM2 as unconditional frames for mask propagation to obtain temporal level segmentation masks;

[0033] Step 5: Loss function of the overall model during the training phase

[0034] Trained in an end-to-end manner, its loss function includes the text generation loss L of the multimodal large language model txt and the mask loss L of the segmentation model mask , the overall loss function L total As shown in formula (5):

[0035]

[0036] Mask loss L mask By binary cross entropy loss L bce and dice loss L dice Composition, as shown in formula (6):

[0037]

[0038] Among them, y txt and denote the text and segmentation mask sequences predicted by the model, respectively. and denote the true value text and mask respectively, λ txt , mask , bce and λ diceThey represent the weighting coefficients of different losses, which are set to 1, 1, 2 and 0.5 in the experiment.

[0039] Beneficial effects of the present invention: After comparative experimental verification, the effects of the present invention in multiple video and image reasoning segmentation and reference segmentation benchmarks significantly surpass previous methods, which fully proves the effectiveness of the proposed method. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is the overall structural diagram of the model of the present application method.

[0041] Figure 2 It is a time series dynamic aggregation graph.

[0042] Figure 3 It is a marker-based keyframe selection graph.

[0043] Figure 4 is the mask decoding and propagation graph.

[0044] Figure 5 It is a comparison chart of qualitative results between the present invention and previous video reasoning segmentation methods. DETAILED DESCRIPTION

[0045] The specific implementation of the present invention is further described below in conjunction with the accompanying drawings and technical solutions.

[0046] Example

[0047] 1. Introduction to the data sets used for training and testing:

[0048] 1) Image semantic segmentation datasets: ADE20K, COCO-Stuff, PACO-LVIS, and PASCALPart.

[0049] 2) Image referent segmentation datasets: refCLEF, refCOCO, refCOCO+, refCOCOg.

[0050] 3) Image reasoning segmentation dataset: ReasonSeg.

[0051] 4) Video referent segmentation datasets: Ref-YouTube-VOS, Ref-DAVIS17, MeViS.

[0052] 5) Video Inference Segmentation Dataset: ReVOS

[0053] 6) Video and image question answering datasets: LLaVA-Instruct-150k and Video-ChatGPT.

[0054] 2. Introduction to evaluation indicators:

[0055] The evaluation indicators for video segmentation include region similarity J, contour accuracy F, and their average J&F. The evaluation indicators for image segmentation include generalized intersection over union (gIoU) and complete intersection over union (cIoU). In addition, this scheme uses the robustness score Evaluate the model's hallucination capabilities.

[0056] 3. The steps of the training phase are as follows:

[0057] (1) Encoding time series tags using a multimodal large model

[0058] First, given an input video X V ∈R T×3×H×W (where T is the total length of the input video, and H and W represent the height and width of each frame respectively) uniformly sample the sampled video frame X V′ and compare it with the tokenized prompt X txt The multimodal language model outputs the response y containing the special tags [SEG] and [TAK] through autoregressive iteration. txt Then, the embeddings corresponding to the two special tags are extracted from the last hidden layer of the large language model. (where T ′ is the total number of video frames after sampling, d ′ is the hidden layer dimension of the multimodal large language model) and The extracted special embedding is mapped through a learnable multi-layer perceptron MLP and projected into the feature space of the perception model SAM2, as shown in Formula 1:

[0059]

[0060] Among them, h seg ,h tak They represent the embedding after mapping respectively.

[0061] (2) Embed the frame level into h seg Aggregate by cosine similarity to temporal embedding h tak middle

[0062] First, the cosine similarity of the embedding corresponding to the video level [TAK] and each frame [SEG] is calculated, and the frame corresponding to the maximum similarity is selected as the key frame of the model in the training stage. Then, the embedding corresponding to the frame level [SEG] is aggregated into the temporal level embedding [TAK] with the similarity as the proportional coefficient through the weighted aggregation strategy shown in Formula 2:

[0063]

[0064] Among them, α represents the fusion coefficient, λ irepresents the normalized cosine similarity between different tags, h′ tak Represents the fused time series mark.

[0065] (3) Mask decoding and propagation using timing markers

[0066] After obtaining the fused temporal embedding h′ tak And select the appropriate key frame X k After that, it is input into the segmentation model 2. The image encoder of the segmentation model 2 is responsible for extracting the multi-scale visual features of the key frame, while the mask decoder interacts with the visual features and the temporal embedding to decode the segmentation mask of the key frame. As shown in Formula 3:

[0067]

[0068] In the training phase, two video frames adjacent to the key frame are extracted as unconditional frames and input into the segmentation model 2, and propagated using its mask memory mechanism to obtain the segmentation mask sequence

[0069] (4) Loss function in the model training phase

[0070] This method is trained in an end-to-end manner, and its loss function includes the text generation loss L of the multimodal large language model. txt and the mask loss L of the segmentation model mask , the overall loss function L total As shown in Formula 4:

[0071]

[0072] Mask loss L mask By binary cross entropy loss L bce and dice loss L dice Composition, as shown in Formula 5:

[0073]

[0074] in and denote the true value text and mask respectively, λ txt , mask , bce and λ dice They represent the weighting coefficients of different losses, which are set to 1, 1, 2 and 0.5 in the experiment.

[0075] 4. The steps of the reasoning phase are as follows:

[0076] In the inference stage, the CLIP model is first used to calculate the video frame with the highest similarity to the object description, and then the video frame is used as the central frame to perform global sampling on the entire video to obtain the sampled video frame. in Indicates the number of sampled video frames. The video frame and text description are sent to the multimodal large model for temporal tag encoding and temporal dynamic aggregation (the same as the training stage). Then, each sampled video frame and the fused [TAK] embedding are segmented to generate an object occlusion score by splitting everything 2. The score indicates the probability that a segmented target exists in the frame. The calculation method is shown in Formula 6:

[0077]

[0078] Among them, E and MD represent the image encoder and mask decoder in segmentation 2, respectively. Indicates the occlusion score corresponding to each sampling frame.

[0079] Finally, the occlusion score normalized by Softmax and the cosine similarity S between the markers calculated in step 2 are used t The sum is used as the criterion for determining the key frame, and the frame with the highest sum X k is selected as the key frame in the inference stage. Finally, the fused temporal marker h′ is also used tak Segment the keyframes and propagate the masks across all remaining video frames,

[0080] 5. Use this example solution for experimental verification:

[0081] (1) Quantitative results

[0082] Table 1: Comparison of quantitative results of this method and previous methods in the ReVOS dataset

[0083]

[0084]

[0085] Table 1 shows the performance comparison between our method and previous methods on the ReVOS video reasoning segmentation benchmark. Our method significantly outperforms the previous state-of-the-art method VISA on all ten metrics. In particular, on the J&F metric, our method using the 13B language model improves by 5.9% on the reference subset and 12.5% ​​on the reasoning subset over VISA-13B. These improvements highlight the role of our temporal aggregation strategy and the [TAK] tag in keyframe selection and object segmentation, which enhances the reasoning ability of the model.

[0086] Table 2: Comparison of the results of this method and previous methods on the general reference video segmentation dataset

[0087]

[0088] Table 2 shows the comparison results of our method with the state-of-the-art RVOS method. On the Ref-YouTube-VOS and Ref-DAVIS17 datasets, our method with the 13B language model improves the J&F metrics of VideoLISA and VISA-13B by 7.3% and 5.6%, respectively. In addition, on the motion-intensive MeViS dataset, VRS-HQ-13B achieves significant improvements of 6.2%, 6.1%, and 6.4% in J, F, and J&F metrics, respectively. This demonstrates its outstanding performance in capturing motion and maintaining temporal consistency. Compared with the single-label representation of VideoLISA, these improvements are attributed to the enhanced inter-frame perception ability of the temporal aggregation strategy and the optimized keyframe localization ability of the label-based keyframe selection strategy.

Claims

1. A high-performance video inference segmentation method based on temporal labeling, characterized in that: Here are the steps: Step 1: Use a multimodal large language model for temporal tag encoding (1) Generate hierarchical tags In order to enable the multimodal large language model to have the ability to perceive contextual information at the frame level and time level in the video segmentation task, a method is proposed to encode the spatial information within the frame and the temporal relationship between frames into hierarchical tags; First, we use frame-level special tags [SEG] and time-level special tags [TAK] to expand the vocabulary of the multimodal large language model; Then design a structured prompt word template: "Please find {target} in the video and segment {target} at the frame level and time level respectively", where {target} represents the reference description of the object; then manually V ∈R T×3×H×W Sampling is performed to obtain the sampled video frame X V′ , and convert the video frame X V′ and the prompt word template X after tokenization by the tokenizer txt The multimodal language model is fed into the large multimodal language model; the multimodal language model outputs frame-level special tags [SEG] and time-series-level special tags through autoregressive iteration. [TAK]'s reply txt ; Where T is the total length of the input video, H and W represent the height and width of each frame respectively; (2) Extracting and mapping hierarchical tags After the multimodal large language model integrates the frame-level and time-level video context information into the frame-level special tag [SEG] and the time-level special tag [TAK] respectively, the special embedding corresponding to the frame-level special tag [SEG] and the time-level special tag [TAK] is extracted from the last hidden layer of the multimodal large language model. and Where T ′ is the total number of video frames after sampling, d ′ is the hidden layer dimension of the multimodal large language model, and the extracted special embedding is mapped through the learnable multi-layer perceptron MLP to project it into the feature space of SAM2, as shown in formula (1): Among them, h seg ,h tak They represent the embeddings corresponding to the mapped frame-level special tags [SEG] and time-level special tags [TAK] respectively; finally, the mapped embeddings are fed to the time-series dynamic aggregation module for spatiotemporal feature extraction; Step 2: Timing Dynamic Fusion (1) Select key frames by similarity of different tags Since the frame-level special mark [SEG] reflects the position change of the target in each video frame, and the temporal-level special mark [TAK] reflects the semantic content of the target in the video, by calculating the cosine similarity of the special embedding corresponding to the temporal-level special mark [TAK] and each frame-level special mark [SEG], the video frame corresponding to the frame-level special mark [SEG] with the maximum cosine similarity is selected as the key frame of the overall model in the training stage; (2) Weighted fusion based on cosine similarity After calculating the cosine similarity of the special embeddings corresponding to the temporal level special tag [TAK] and each frame level special tag [SEG], the cosine similarity is normalized, and then the embedding h corresponding to the mapped frame level special tag [SEG] is converted into the weighted aggregation strategy shown in formula (2) with the normalized cosine similarity as the proportional coefficient. seg Aggregated to the embedded h corresponding to the time-level special tag [TAK] after mapping tak middle: Among them, α represents the fusion coefficient, which is set to 0.1; λ i represents the proportional coefficient, which is the normalized cosine similarity; h′ tak represents the fused temporal tag embedding, which is used in the key frame selection and segmentation mask generation process; i represents the index value of the video frame; Step 3: Keyframe selection strategy based on temporal markers In the inference stage of the overall model, a key frame selection strategy based on time series markup is proposed to decode the time series information embedded in the fused time series markup to select more accurate key frames. In the inference stage, the CLIP model is first used to calculate the video frame with the highest similarity to the target description, and this video frame is used as the central frame for the given input video X. V Perform global sampling to obtain the sampled video frame in, represents the number of video frames sampled in the inference stage, and f is the identifier of the inference stage; the sampled video frames The structured prompt word template is sent to the multimodal large language model to perform steps 1 and 2. Each sampled video frame is embedded with the fused temporal tag h′ tak The target occlusion score is generated by SAM2. The target occlusion score indicates the probability of the segmented target existing in the sampled video frame. The calculation method is shown in formula (3): Where E and MD represent the image encoder and mask decoder in SAM2, respectively. Indicates the target occlusion score corresponding to each sampled video frame; Finally, the target occlusion score normalized by Softmax and the h corresponding to each video frame calculated in step 2 are used. seg and h tak The cosine similarity S t The sum is used as the criterion for determining the key frame, and the video frame with the highest sum value is X k Selected as key frames in the inference phase; Step 4: Mask decoding and propagation using SAM2 The key frame is input into SAM2. The image encoder of SAM2 is responsible for extracting the multi-scale visual features of the key frame. The mask decoder of SAM2 embeds the multi-scale visual features of the key frame and the fused temporal mark into h′. tak Interact to decode the segmentation mask of the key frame As shown in formula (4): In the training phase, the two video frames adjacent to the key frame are extracted as unconditional frames and input into SAM2, and propagated using its mask memory mechanism to obtain the segmentation mask sequence. In the inference stage, all remaining video frames are input into SAM2 as unconditional frames for mask propagation to obtain temporal level segmentation masks; Step 5: Loss function of the overall model during the training phase Trained in an end-to-end manner, its loss function includes the text generation loss L of the multimodal large language model txt and the mask loss L of the segmentation model mask , the overall loss function L total As shown in formula (5): Mask loss L mask By binary cross entropy loss L bce and dice loss L dice Composition, as shown in formula (6): Among them, y txt and denote the text and segmentation mask sequences predicted by the model, respectively. and denote the true value text and mask respectively, λ txt , mask , bce and λ dice They represent the weighting coefficients of different losses respectively.

Citation Information

Cited By

  • Multi-modal target tracking method and system based on time sequence modeling and prompt fine tuning

    CN120510480A

  • Long video target inference segmentation method based on context mark prompt

    CN120707859A

  • Video space-time understanding method and device based on multi-modal large model, and medium

    CN120877194A

  • A video spatio-temporal understanding method and device based on a multi-modal large model and a medium

    CN120877194B

  • Video reasoning method and device, electronic equipment, storage medium and program product

    CN121330580A