A video spatio-temporal understanding method and device based on a multi-modal large model and a medium
By using a multimodal large-scale video spatiotemporal understanding method, joint training of temporal understanding and spatial awareness is achieved, solving the problem that existing technologies cannot perform multi-task joint training, improving the ability of video large-scale language models to interpret fine-grained spatiotemporal events, and achieving more accurate spatiotemporal event localization.
Patent Information
- Application Number
- CN202511365934.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Existing technologies are insufficient to fully explain complex spatiotemporal events in real-world scenarios, especially in terms of the inability to perform explicit temporal localization during spatial segmentation and the inability to conduct end-to-end joint training of multiple tasks, which limits the ability of video large language models to achieve fine-grained spatiotemporal understanding.
By using a multimodal large model, target videos are sampled at different frame rates. Visual and textual tags are integrated using a multimodal encoder and a large language model to achieve joint training of temporal understanding and spatial awareness. A single-stage training fine-tuning method is adopted, combining timestamp and spatial mask segmentation, and using a multilayer perceptron and memory attention mechanism for video spatiotemporal understanding.
Joint training for fine-grained video spatiotemporal understanding was achieved, enabling simultaneous time-of-flight retrieval and reference video segmentation, thus improving the accuracy of temporal understanding and spatial awareness. A unified video spatiotemporal understanding model was constructed, which can more accurately locate the spatiotemporal position of events.
Smart Images

Figure CN120877194B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision, and in particular to a method and apparatus for video spatiotemporal understanding based on a multimodal large model. Background Technology
[0002] The continuous evolution of Multimodal Large Language Models (MLLMs) has elevated visual content understanding to new heights. Early visual large language models, such as LLaVA, achieved deep fusion of visual and linguistic information through training on large-scale image and text data, enabling the model to fully understand and reason about images using the knowledge of large language models. Building on this, researchers have further extended the capabilities of MLLMs from images to videos, achieving holistic understanding of video content. VideoChat developed a video-centric dialogue system that combines spatial awareness and temporal reasoning, demonstrating strong generalization capabilities across a wide range of multimodal benchmarks. LLaMA-VID goes a step further, compressing video representations to retain key visual information to support long video understanding, breaking through the limits of existing frameworks. However, these works mainly focus on holistic video understanding and fall short in capturing fine-grained spatial and temporal details.
[0003] Recently, an increasing number of researchers have focused on equipping video language models with fine-grained understanding capabilities, such as time-lapse retrieval, reference video segmentation, and object tracking. These tasks go beyond video description or general video question answering, requiring models to locate specific actions or objects within precise temporal and spatial boundaries. However, existing research still focuses on the temporal or spatial dimensions in isolation, limiting its ability to comprehensively explain complex spatiotemporal events in real-world scenes.
[0004] Chinese patent application CN202411636948.6 discloses a video analysis and processing system and method based on a multimodal large model; Chinese patent application CN202510174580.4 discloses a high-performance video reasoning and segmentation method based on temporal tagging. The above-mentioned existing solutions have the problems of not being able to perform end-to-end joint training of multiple tasks and not being able to perform explicit temporal localization while performing spatial segmentation. On the one hand, the existing patent solutions lack sufficient exploration of spatiotemporal joint training; on the other hand, the inability to perform end-to-end training of multiple tasks hinders the mutual promotion between different tasks. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies in fully explaining complex spatiotemporal events in real-world scenarios, and to provide a video spatiotemporal understanding method and apparatus based on a multimodal large model.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] As a first aspect of the present invention, a video spatiotemporal understanding method based on a multimodal large model is provided, comprising the following steps:
[0008] Acquire the target video and the query text for temporal understanding and spatial awareness;
[0009] For the temporal understanding task and the spatial awareness task, the target video is sampled at different frame rates, and the video is encoded into visual tags using a multimodal encoder; for the temporal understanding task, a timestamp text description is added after the visual tag of each frame.
[0010] Use a text tag generator to encode the query input text into text tags;
[0011] After integrating visual and textual tags, the data is input into a large language model. For time understanding tasks, the large language model outputs the timestamp text to locate the time position of the event query based on the query text and timestamp text descriptions. For spatial awareness tasks, the large language model directly determines the task requirements based on the query text and outputs the data. <seg>The mark, the <seg>The corresponding hidden layer features are labeled to indicate the object space mask of the sampled frame generated by the mask segmentation decoder.
[0012] For time understanding tasks, extract the timestamp of the query from the text response; for spatial awareness tasks, locate the location from the text response. <seg>The sampled frames are labeled and feature embeddings are obtained; the visual encoder of the mask segmentation model is used to encode the sampled frames to obtain visual features; the feature embeddings are used as cues and input together with the visual features into the mask decoder of the mask segmentation model to generate a mask for the sampled frames and propagate the sampled frame mask across the entire video.
[0013] As a preferred technical method, the visual marker encoding is as follows:
[0014] Different frame rates of sampled video were used for temporal understanding tasks and spatial awareness tasks;
[0015] The sampled video frames are input into a multimodal encoder and encoded into visual markers through image block embedding;
[0016] For time-based tasks, add a timestamp text description after the visual markers of each frame; for spatial tasks, no text description is added.
[0017] As a preferred technical method, the text tag encoding is as follows: the query text is segmented into words, and the word sequence is encoded into text tags through language embedding.
[0018] As a preferred technical method, the process of decoding the text response using the large language model is as follows:
[0019] Visual tags are projected onto the text space through a multilayer perceptron, and then directly linked with text tags and input together into a large language model for decoding and generation.
[0020] The decoded tags are then processed by a multilayer perceptron to reconstruct the corresponding text response.
[0021] As a preferred technical method, the visual feature encoding is as follows:
[0022] For time understanding tasks and spatial awareness tasks, different frame numbers of sampled video are used, and the resolution of the sampled frames is higher than that of the sampled frames input to the multimodal encoder.
[0023] Visual features are obtained by encoding sampled frames using a visual encoder.
[0024] As a preferred technical method, the steps for generating the sampling frame mask are as follows:
[0025] For the output of large language models <seg>Tagging, extracting the last Transformer layer of the large language model <seg>The output features of the marked locations are embedded and projected onto the language cue embedding space using a multilayer perceptron to obtain the cue embedding.
[0026] The visual features of the sampled frames extracted by the visual encoder of the cue embedding and segmentation model are input into the mask decoder to generate a mask for the sampled frames.
[0027] As a preferred technique, the mask propagation process of the sampled frame is as follows:
[0028] A memory encoder is used to generate a memory based on the mask of the sampled frame;
[0029] The mask for the unsampled frames in the entire video is generated using a memory attention mechanism.
[0030] As a preferred technique, the method employs a single-stage training fine-tuning approach, utilizing various image and video datasets covering temporal and spatial understanding tasks for end-to-end training. The training loss function is set as follows:
[0031]
[0032] For general question-and-answer sessions and question-and-answer sessions containing timestamps, the standard text generation loss is used. That is, the autoregressive cross-entropy loss for text generation; for mask segmentation, it combines pixel-level binary cross-entropy loss. and DICE losses .
[0033] As a second aspect of the present invention, a video spatiotemporal understanding device based on a multimodal large model is provided, including a memory, a processor, and a program stored in the memory, wherein the processor executes the program to implement the video spatiotemporal understanding method based on a multimodal large model as described above.
[0034] As a third aspect of the present invention, a storage medium is provided having a program stored thereon, which, when executed, implements the video spatiotemporal understanding method based on a multimodal large model as described above.
[0035] Compared with the prior art, the present invention has the following beneficial effects:
[0036] This invention achieves joint training for fine-grained video spatiotemporal understanding, constructing a unified model capable of simultaneously performing temporal understanding (represented by time-of-day retrieval) and spatial awareness (represented by reference video segmentation). Besides providing an overall understanding of the video, it enables more precise event spatiotemporal localization. Furthermore, the excellent performance achieved by this invention verifies that joint training of temporal understanding and spatial awareness can simultaneously improve both capabilities. This invention realizes fine-grained joint spatiotemporal understanding of videos by constructing a model that can simultaneously process fine-grained temporal and spatial information from videos, integrating spatiotemporally labeled data, and learning consistent spatiotemporal correspondences through joint training. Attached Figure Description
[0037] Figure 1 This is a flowchart of a video spatiotemporal understanding method based on a multimodal large model according to the present invention.
[0038] Figure 2 This is a diagram of the video joint spatiotemporal understanding network architecture constructed in this invention. Detailed Implementation
[0039] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0040] Example 1
[0041] This invention proposes a video spatiotemporal understanding method based on a multimodal large-scale model. It connects a multimodal large-scale language model and a mask segmentation model. First, a multimodal encoder encodes video features, employing different numbers of sampling frames and representation formats for temporal and spatial tasks. After aligning the video features to the text space, they are input into the large-scale language model along with text tags, and decoded to obtain the corresponding text response. Timestamp information is directly aligned to the text space; that is, for time-based queries, the timestamp can be directly extracted from the text response. Spatial information is obtained through... <seg>The token encoding, whose embedding serves as a cue input to the mask decoder, enables mask generation for sampled frames and mask propagation throughout the video. For example... Figure 1 As shown, the method specifically includes the following steps:
[0042] S1. Encode the video into visual markers using a multimodal encoder;
[0043] S2. Encode the text into text tags using a text tag generator;
[0044] S3. The visual encoder of the mask segmentation model encodes the sampled frames in step S1. Compared with the visual encoding in step S1, a higher resolution frame is used here to provide finer-grained visual features for mask segmentation encoding.
[0045] S4. Integrate the visual tags obtained in step S1 and the text tags obtained in step S2, input them into the large language model, and decode to obtain the text response;
[0046] S5. For time-based tasks, extract the query timestamp directly from the text response; for spatial tasks, locate the location from the text response. <seg>Label the data and obtain its feature embeddings;
[0047] S6. Embed the features obtained in step S5 as prompts, and input them together with the visual features obtained in step S3 into the mask decoder of the segmentation model to generate the masks for these sampled frames.
[0048] S7. Using the mask segmentation model, the sample frame mask obtained in step S6 is propagated across the entire video.
[0049] Specifically, in step S1, the visual features are encoded using an InternVL-2.5-4B multimodal encoder. The input frame resolution is 448. For the temporal task, 32 frames are uniformly sampled from the video, and a timestamp text description is added after the visual marker in each frame. For the spatial task, 5 frames are uniformly sampled from the video, without any additional text description processing. The specific steps are as follows:
[0050] S1.1 Uniformly sample video frames. For time-based tasks, sample 32 frames evenly from the video; for spatial tasks, sample 5 frames evenly from the video. The frame resolution is 448. More frames are used for time-based tasks to provide longer-range, finer-grained temporal location information.
[0051] S1.2 Input the sampled video frames into the InternVL-2.5-4B multimodal encoder and encode them into visual tags through PatchEmbedding.
[0052] S1.3 For time-based tasks, add a timestamp text description after the visual marker of each frame, such as "This frame was sampled at the 10th second". For spatial tasks, no additional text description is required.
[0053] In step S2, the query text is segmented, and then the word sequence is encoded into text tags through Language Embedding.
[0054] In step S3, the SAM-2-large visual encoder encodes the sampled frame using the same video frame as in step S1, but at a high resolution of 1024.
[0055] In step S4, the visual tags obtained in S1 and the text tags obtained in S2 are integrated and input into the Qwen-2.5-3B large language model to decode and obtain the text answer:
[0056] S4.1 Project the visual tags obtained in S1 onto the text space using an MLP, and directly connect them with the text tags obtained in S2. Then, input them together into the Qwen-2.5-3B large language model for decoding and generation.
[0057] S4.2 The decoded tokens are then used by the MLP to restore the corresponding text response.
[0058] For time-related tasks, the large language model uses the query text and the timestamp text added in step S1 to output a timestamp text to locate the time position of the event query. For spatial tasks, the large language model directly determines the task requirements based on the query text and outputs... <seg>Special text markers are used to determine the position of the output text sequence corresponding to the marker, and the last Transformer layer of the large language model is extracted. <seg>The output feature embedding at the marked location is used to prompt the mask segmentation decoder to generate the object space mask for the sampled frame.
[0059] In step S5, for time-related tasks, the timestamp of the query is directly extracted from the text response; for spatial tasks, the location is determined from the text response. <seg>The marker is used to extract the feature embedding of the large language model's decoding output.
[0060] In step S6, the feature embedding obtained in S5 is used as a cue and, together with the visual features obtained in S3, is input into the SAM-2-large mask decoder to generate masks for these sampled frames:
[0061] S6.1 Project the feature embeddings obtained in S5 into the language prompt embedding space using an MLP;
[0062] S6.2. The cues are embedded together with the visual features obtained in S3 and input into the SAM-2-large mask decoder to generate masks for these sampled frames.
[0063] Specifically, for the encoding and text response generation of multimodal large models, let The basic model of language is expressed as follows:
[0064]
[0065] in, , indicates the input text query. and These represent the number and dimension of the text tags, respectively. , indicating the input video, , and These represent the number, height, and width of the video frames, respectively. In addition to video, the model also receives images. As input, integrating visual and textual tags, the large language model outputs a textual response. For time-based tasks, the text response will include the required timestamp information; for space-based tasks, the text response will include special markers. <seg>This marker will be used to guide the mask decoder in generating the mask sequence for the sampled frames. .
[0066] For the output <seg>The present invention extracts the output feature embedding of the last Transformer layer of the large language model at the corresponding position and applies MLP projection to generate... As a cue embedding in the mask decoder, it also serves as the visual encoder of the segmentation model. It also extracts visual features from the sampled frames. It will be embedded with the prompt. They are fed together into the mask decoder to generate a mask sequence. .make For the mask decoder, its expression is as follows:
[0067]
[0068] In step S7, the SAM-2-large mask segmentation model is used to propagate the sampled frame mask obtained in S6 across the entire video:
[0069] S7.1. Use the SAM-2-large memory encoder to generate memory based on the mask of the sampled frame.
[0070] S7.2 Utilize the memory attention mechanism to propagate the generated mask across other non-sampled frames of the entire video.
[0071] Furthermore, this invention also collects open-source time-annotated and spatially annotated video data, and automatically annotates some videos with fine-grained joint spatiotemporal annotations from ActivityNet. The data sources include, but are not limited to, open-source datasets Charades, MeVIS, and ReVOS. At the same time, it automatically annotates some videos with fine-grained joint spatiotemporal annotations from ActivityNet as a supplement to the open-source data.
[0072] The collected training data is input into a video joint spatiotemporal understanding network based on a multimodal large language model for training, resulting in the trained model weights. All hyperparameters and loss functions are set according to the Sa2VA model settings. All visual encoders do not participate in weight updates during training, and the large language model undergoes LoRA fine-tuning.
[0073] The model of this invention adopts a single-stage training and fine-tuning approach, and performs end-to-end training using various image and video datasets covering temporal and spatial understanding tasks. The training process uses three loss functions, the expressions of which are as follows:
[0074]
[0075] For general question-and-answer sessions and question-and-answer sessions containing timestamps, the standard text generation loss is used. This refers to the autoregressive cross-entropy loss for text generation. For mask segmentation, a pixel-level binary cross-entropy (BCE) loss is used. and DICE losses .
[0076] Finally, the trained model weights are loaded into the network, and the corresponding test data is input into the network for testing, thereby achieving performance evaluation of real-time retrieval and reference video segmentation.
[0077] Compared to existing video language models that focus only on temporal understanding or spatial awareness, this invention achieves joint training of fine-grained video spatiotemporal understanding. It constructs a unified model capable of simultaneously performing temporal understanding (represented by time-of-day retrieval) and spatial awareness (represented by reference video segmentation). In addition to overall video understanding, it enables more accurate event spatiotemporal localization. Furthermore, the superior performance achieved by this invention verifies that joint training of temporal understanding and spatial awareness can simultaneously enhance both capabilities, laying the foundation for further development of more general video language models.
[0078] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0079] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.< / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg> < / seg>
Claims
1. A video spatiotemporal understanding method based on a multimodal large model, characterized by the following steps: include: Acquire the target video and the query text for temporal understanding and spatial awareness; For the temporal understanding task and the spatial awareness task, the target video is sampled at different frame rates, and the video is encoded into visual tags using a multimodal encoder; for the temporal understanding task, a timestamp text description is added after the visual tag of each frame. Use a text tag generator to encode the query input text into text tags; After integrating visual and textual tags, the data is input into a large language model. For time understanding tasks, the large language model outputs the timestamp text to locate the time position of the event query based on the query text and timestamp text descriptions. For spatial awareness tasks, the large language model directly determines the task requirements based on the query text and outputs the data. <seg>The mark, the <seg> The corresponding hidden layer features are labeled to indicate the object space mask of the sampled frame generated by the mask segmentation decoder.< / seg> < / seg> For time understanding tasks, extract the timestamp of the query from the text response; For spatial perception tasks, locating from text responses <seg> Label and obtain feature embeddings; use the visual encoder of the mask segmentation model to encode sampled frames to obtain visual features;< / seg> Feature embeddings are used as cues and, together with visual features, are input into the mask decoder of the mask segmentation model to generate a mask for the sampled frame and propagate the sampled frame mask across the entire video.
2. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The visual marker encoding is as follows: Different frame rates of sampled video were used for temporal understanding tasks and spatial awareness tasks; The sampled video frames are input into a multimodal encoder and encoded into visual markers through image block embedding; For time-based tasks, add a timestamp text description after the visual markers for each frame; For space missions, no textual descriptions are provided.
3. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The text tag encoding is as follows: the query text is segmented, and the word sequence is encoded into text tags through language embedding.
4. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The process of decoding the text response using the large language model is as follows: Visual tags are projected onto the text space through a multilayer perceptron, and then directly linked with text tags and input together into a large language model for decoding and generation. The decoded tags are then processed by a multilayer perceptron to reconstruct the corresponding text response.
5. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The steps of the visual encoder in the mask segmentation model to encode sampled frames and obtain visual features are as follows: For the time understanding task and the spatial perception task, the target video is sampled at different frame rates, and the resolution of the sampled frames is higher than that of the sampled frames input to the multimodal encoder; the visual features are obtained by encoding the sampled frames through the visual encoder.
6. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The mask decoder of the mask segmentation model generates the sample frame mask in the following steps: For the output of large language models <seg>Tagging, extracting the last Transformer layer of the large language model <seg> The output features of the marked locations are embedded and projected onto the language cue embedding space using a multilayer perceptron to obtain the cue embedding.< / seg> < / seg> The visual features of the sampled frames extracted by the visual encoder of the cue embedding and segmentation model are input into the mask decoder to generate a mask for the sampled frames.
7. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The mask propagation process of the sampled frame is as follows: A memory encoder is used to generate a memory based on the mask of the sampled frame; The mask for the unsampled frames in the entire video is generated using a memory attention mechanism.
8. The video spatiotemporal understanding method based on a multimodal large model according to claim 1, characterized in that, The method employs a single-stage training and fine-tuning approach, utilizing various image and video datasets covering temporal and spatial understanding tasks for end-to-end training. The training loss function is set as follows: , For question-and-answer processing, a standard text generation loss is used. That is, the autoregressive cross-entropy loss for text generation; for mask segmentation, it combines pixel-level binary cross-entropy loss. and DICE losses .
9. A video spatiotemporal understanding device based on a multimodal large model, comprising a memory, a processor, and a program stored in the memory, characterized in that, When the processor executes the program, it implements the video spatiotemporal understanding method based on a multimodal large model as described in any one of claims 1-8.
10. A storage medium having a program stored thereon, characterized in that, When the program is executed, it implements the video spatiotemporal understanding method based on a multimodal large model as described in any one of claims 1-8.
Citation Information
Patent Citations
Video analysis processing system and method based on multi-modal large model
CN119540831A
Video time sequence sentence positioning method based on quadruple constraint and partial supervision
CN116881502A
High-performance video reasoning segmentation method based on time sequence marking
CN120107854A