Sampling method and system for long video understanding
The token sampling method is used to select long video segments and keyframes. The multimodal capability of the video large language model is used to solve the problems of semantic discontinuity of video clips and excessive consumption of computing resources in long video understanding, and efficient video information processing is achieved.
Patent Information
- Application Number
- CN202510873208.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing long video understanding methods perform poorly in video timing segmentation and keyframe selection, and excessive computing resources consumption, resulting in the problems of video clip semantic discontinuity and excessive memory overhead.
The token sampling method is used to segment the long video into semantic video clips through the video large language model, a text summary is generated, and the token number is dynamically allocated based on the relative weight of the frame, keyframes are selected, and matching degree selection is performed based on visual and language information.
It realizes that while ensuring the integrity of video information, it reduces video memory overhead and computing resource consumption, improves the efficiency and performance of long video understanding, and solves the inaccuracy of video timing segmentation and keyframe selection.
Smart Images

Figure CN120388323A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular to an efficient sampling method and system for long video understanding. Background Art
[0002] In recent years, with the rapid development of information technology and the continuous improvement of network infrastructure, global video data has seen explosive growth. These videos cover a wide range of scenarios, from daily life records to professional applications. Faced with this massive and diverse video data, how to effectively understand and extract key information has become a research focus in the field of video understanding. Intelligent video understanding technology is a cutting-edge technology that has rapidly developed in this context and has attracted widespread attention from academia and industry. By leveraging the latest research results in deep learning, natural language processing, and computer vision, this technology is dedicated to automatically analyzing complex video content, extracting key information, and performing structured annotation and summarization of video content, greatly improving the understandability and usability of video data.
[0003] The Large Video Language Model (LLM) is one of the latest breakthroughs in video understanding. By introducing a large-scale pre-trained model and combining video understanding technology with a large language model, the LLM is able to generate text based on video content, achieving both semantic understanding of video content and natural language generation. This technology not only improves the intelligence of video understanding but also demonstrates excellent performance in tasks such as video question-answering, video summary generation, and multimodal interaction. However, despite its widespread application and significant success in video understanding, the LLM still faces the following challenges when it comes to understanding long videos: 1) insufficient video temporal segmentation capabilities; 2) inaccurate keyframe selection; and 3) high computational resource consumption.
[0004] Video temporal segmentation is a key step in understanding long videos. Leveraging the powerful semantic analysis capabilities of large video language models, temporal segmentation can divide long videos into multiple semantically independent segments based on content themes, laying the foundation for subsequent understanding. Video temporal segmentation aims to divide videos into semantically coherent and independent segments. Compared to the original long video, these segments contain fewer frames and are more suitable for further processing. Existing methods typically combine scene detection tools with similarity calculations or employ specialized time series labeling techniques for temporal localization. Most existing models use fixed time intervals or scene-based segmentation for temporal segmentation. While these methods are effective for short videos or videos with relatively simple content structures, in long videos, fixed time intervals or scene-based segmentation methods struggle to accurately identify semantic boundaries between segments, as they often arise from topic transitions or content changes. This results in segmented segments lacking coherence and logical logic. Furthermore, semantic transitions in videos may not always sync with changes in the visual scene. Relying solely on visual features for temporal segmentation often results in discontinuous semantic information within the segmented segments.
[0005] After completing the temporal segmentation of the video, each video clip may still contain a large amount of redundant information, so it is necessary to further screen out the most representative key frames to condense the information of the clip. Key frame selection is the core link in generating high-quality video understanding, which directly determines the simplicity of video content presentation and its information carrying capacity. Many studies have been devoted to reducing the number of input frames required for downstream tasks such as action recognition, video temporal localization, and video question answering. These methods include compressing frames by calculating inter-frame similarity, utilizing temporal attention mechanisms, and retrieving key frames through large language model outputs. Existing methods usually adopt a fixed-step frame sampling strategy or select key frames based on visual similarity, but this method ignores the differences in information density between frames, which may result in missing key frames in clips with concentrated key information, or collecting a large number of redundant frames in clips with sparse information.
[0006] In addition to the aforementioned limitations, existing models consume significant computational resources when completing long video understanding tasks. A notable characteristic of long videos is their length and the sheer volume of information they contain. Therefore, models must process a large number of frames, leading to a sharp increase in video memory consumption and computational complexity. For example, at one frame per second, a 5-minute video will generate 300 frames. If each frame generates approximately 700 visual tokens, the total number of tokens will exceed 200,000, exceeding the processing capacity of most models. Summary of the Invention
[0007] The purpose of the present invention is to overcome the problems of the above-mentioned existing long video understanding methods, such as poor performance in video timing segmentation and key frame selection tasks and consumption of a large amount of computing resources, and to provide a sampling method for long video understanding.
[0008] The purpose of the present invention can be achieved by the following technical solutions: As a first aspect of the present invention, a sampling method for long video understanding is provided, comprising the following steps: Take a long video and its subtitles as input, and use the video model to split the long video into multiple semantically consistent video segments; For each video clip, the visual encoder is used to map the input sequence, which is then fed into the video language model to generate a text summary of the video clip. Perform token sampling on the video clip, calculate the relative weights between the frames of the video clip, and assign different numbers of tokens to each frame according to the weights; The video frames and text summaries after token sampling are input into the video language model, and the key frames of the video clips are obtained based on the matching degree between the video frame tokens and the text summary.
[0009] As a preferred technical solution, the process of obtaining the video clip is as follows: The text sequence of the input video subtitles is processed by the word segmenter and split into basic language units; The text sequence passes through the embedding layer, mapping each language unit to a corresponding vector representation; The vector representation undergoes dimension transformation through the projection layer and is adjusted to the input dimension of the video large language model to obtain the first input sequence; Input the first input sequence into the video language model, perform context modeling and semantic understanding based on the vector representation, and output segmentation information of the video; Split a long video into multiple video clips based on segmentation information.
[0010] As a preferred technical solution, the process of generating the text summary is as follows: The video clip is converted into a visual feature representation with semantic information through a visual encoder; The visual features are mapped from the original visual space to a space consistent with the input dimension of the large language model through the connection layer to obtain the second input sequence; The second input sequence is fed into the video language model, which fuses the visual information with the language information in a unified representation space to obtain a text summary of the video clip.
[0011] As a preferred technical solution, the process of obtaining the key frames of the video clip is as follows: Extract the video clip to obtain a frame set; Calculate the weights of each frame in the frame set, and perform token sampling according to the weights to obtain a set of video frame tokens; Input the set of video frame tokens and the text summary into the video large language model. According to the matching degree between the video frame tokens and the text summary, find the several frames in the set of video frame tokens that best match the text summary.
[0012] As a preferred technical solution, the token sampling consists of two steps: weight calculation and sampling, which are specifically as follows: Use a visual information encoder to obtain the tokens corresponding to the original video frames , for the tokens corresponding to the original video frames perform compression to obtain compressed visual tokens ; Calculate the attention weights between the compressed visual tokens of all frames, and use the sum of the attention weights of a certain frame relative to all other frames as the final weight of this frame; For frames with high weights, sample the tokens at the corresponding positions from the tokens corresponding to the original video frames ; for frames with low weights, sample the tokens at the corresponding positions from the compressed visual tokens .
[0013] As the second aspect of the present invention, a sampling system for long video understanding is provided. The system executes the sampling method for long video understanding as described above, including: Video segmentation module: Obtain a long video and video subtitles as inputs, and use a video large model to divide the long video into multiple semantically consistent video segments; Text summary generation module: For each video segment, map the input sequence through a visual encoder, and input the input sequence into the video large language model to generate a text summary of the video segment; Token sampling module: Perform token sampling on the video segment, calculate the relative weights between the frames of the video segment, and allocate different numbers of tokens to each frame according to the weights; Key frame acquisition module: Input the video frames after token sampling and the text summary into the video large language model, and obtain the key frames of the video segment based on the matching degree between the video frame tokens and the text summary.
[0014] As a preferred technical solution, the video segmentation module specifically includes: Tokenizer, which splits the text sequence of the input video subtitles into basic language units; Embedding layer, which maps each language unit of the text sequence into a corresponding vector representation; Projection layer, which performs dimensional transformation on the vector representation and adjusts it to the input dimension of the video large language model to obtain the first input sequence; For the first input sequence, the video large language model performs context modeling and semantic understanding based on the vector representation, and outputs the segmentation information of the video. The video processing tool divides the long video into multiple video segments according to the segmentation information.
[0015] As a preferred technical solution, the text summary generation module specifically includes: The visual encoder converts the video segment into a visual feature representation with semantic information. The connection layer maps the visual features from the original visual space to a space consistent with the input dimension of the large language model to obtain the second input sequence. For the input of the second input sequence, the video large language model fuses the visual information and the language information in a unified representation space to obtain the text summary of the video segment.
[0016] As a preferred technical solution, the token sampling module specifically includes: The token acquisition unit uses the visual information encoder to obtain the tokens corresponding to the original video frames , for the tokens corresponding to the original video frames to obtain compressed visual tokens through compression ; The weight calculation unit calculates the attention weights among the compressed visual tokens of all frames, and takes the sum of the attention weights of a certain frame relative to all other frames as the final weight of this frame. The token allocation unit samples the tokens of the frames with high weights from the positions corresponding to the tokens of the original video frames ; for the frames with low weights, its tokens are sampled from the positions corresponding to the compressed video visual tokens .
[0017] As a preferred technical solution, the key frame acquisition module specifically includes: The frame extraction unit extracts the video segment to obtain a frame set. The sampling unit inputs the video frame token set and the text summary obtained by the token sampling module to the video large language, and finds the several frames in the video frame token set that best match the text summary according to the matching degree between the video frame tokens and the text summary.
[0018] Compared with the prior art, the present invention has the following beneficial effects: 1) The long video understanding sampling method of the present invention is based on the token sampling method and the multimodal large model. The method of the present invention first segments the long video and ensures the semantic coherence of each segment. Secondly, it generates a text summary of the video segment for different long video segments. Finally, based on the token sampling method, it calculates the importance of the frames of all video segments using the text summary, and dynamically allocates the number of tokens according to the importance of the video frames, improving the frame processing efficiency while ensuring the integrity of video information. The method of the present invention can reduce the video memory overhead caused by large-scale video frame processing and achieve efficient long video understanding of the multimodal large model.
[0019] 2) The present invention uses a video large language model as the backbone and utilizes the semantic understanding ability of the large language model, so that the long video is divided into multiple video segments with consistent semantics, avoiding the problem of discontinuous semantics of video segments caused by video segmentation according to scene changes and other factors in the prior art solutions; and the input sequence is sent into the large language model component in the video large language model to obtain a text summary of the video segment. The present invention can utilize its multimodal capabilities to complete understanding tasks in text modality, picture modality, and video modality.
[0020] 3) The present invention proposes a novel token sampling method, calculates the weight of each frame in the video segment frame set, and then performs token sampling according to the weight; according to the calculated weights of the compressed video frames, sampling is performed, and the tokens corresponding to the original video frames are assigned to the frames with high weights, and the tokens corresponding to the compressed video frames are assigned to the frames with low weights. This makes up for the problem of excessive computational resource consumption of the video large language model in the long video understanding task. By using the token sampling method to sample the tokens of video frames, the present invention can retain complete video information on the basis of greatly reducing the video memory overhead, and improve the performance and efficiency of the model in the long video understanding task. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic flow chart of the sampling method for long video understanding proposed by the present invention; Figure 2 It is a schematic diagram of the token sampling of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives the detailed implementation manner and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0023] Embodiment 1 The present invention proposes a sampling method for long video understanding, based on token sampling and multi-modal large models. First, the long video is segmented to ensure the semantic coherence of each segment. Secondly, text summaries of different long video segments are generated. Finally, based on the token sampling method, the importance of frames of all video segments is calculated using the text summaries, and the number of tokens is dynamically allocated according to the importance of the video frames, improving the frame processing efficiency while ensuring the integrity of video information. The method of the present invention can reduce the video memory overhead caused by large-scale video frame processing and achieve efficient long video understanding of multi-modal large models. As Figure 1 described, the method of the present invention includes the following steps: Step S1: Using the subtitle file of the long video as input and leveraging the semantic understanding ability of the multi-modal large model, the long video is divided into multiple video segments with consistent semantics.
[0024] In this embodiment, the input video subtitle is encoded through a tokenizer and an embedding vector layer to obtain an input sequence. This input sequence is fed into the large language model component in the video large language model to obtain the corresponding output. Specifically, the input text sequence is first processed by the tokenizer and split into basic language units (i.e., tokens). The tokenized text sequence is then fed into the embedding layer, which maps each token to a corresponding vector representation, i.e., token embedding. Next, these vector representations undergo a dimensionality transformation through a projection layer to adapt to the input dimensionality requirements of the large language model, and the large language model performs subsequent context modeling and semantic understanding based on these vectors. The present invention utilizes the semantic understanding ability of the large language model to divide the long video into multiple video segments with consistent semantics, avoiding the problem of discontinuous semantics in video segments caused by video segmentation according to scene changes and other methods in the prior art solutions. The output of the large language model contains the segmentation information of the video by the video large language model, and the long video is divided into multiple video segments using video processing tools according to the segmentation information.
[0025] Step S2: Process each video segment to generate a text summary of the video segment.
[0026] In this embodiment, the input video segment is mapped through a visual encoder and a connection layer to obtain an input sequence, which is fed into the large language model component in the video large language model to obtain the corresponding output. Specifically, the visual encoder processes the video segment and converts it into a visual feature representation with semantic information. Subsequently, the connection layer maps the visual features from the original visual space to a space consistent with the input dimensionality of the large language model to ensure that visual information can be fused with language information in a unified representation space. The output contains the text summary of the video segment by the video large language model.
[0027] Step S3: Use the token sampling method to input the sampled video frames into the multimodal large model to improve its long video understanding efficiency.
[0028] In this embodiment, the input is a text summary corresponding to the video clip. and the video clip is uniformly sampled at one frame per second. The token sampling method adopted by the present invention first calculates the frame set Every frame The weight of , and then the token sampling is performed according to this weight, aiming to Identify frame sets The key frame containing the most information is selected to achieve a balance between video memory overhead and video information integrity. Then, the present invention calculates the weight of the compressed video frame. , perform sampling, assign the token corresponding to the original video frame to the frame with high weight, and assign the token corresponding to the compressed video frame to the frame with low weight. Finally, the set of video frame tokens (in or ) and text summary The large language model component that is input to the video large language model enables the large language model to and Make a choice based on the degree of matching and find the set Zhongyu The most appropriate number of frames. Specifically, the large language model selects key frames based on visual and linguistic information. It first uses the relationships between video frames to calculate frames with high information density, retains the number of tokens in these frames, compresses the number of tokens in the remaining frames, and then selects key frames from all processed frames.
[0029] Furthermore, the token sampling method consists of two steps: weight calculation and sampling, aiming to balance the model memory overhead and video information integrity.
[0030] In this embodiment, the input of the token sampling method is a frame set , which is in the form of the token corresponding to the original video frame (in is the number of video frames, is the number of tokens contained in a single original video frame, is the token hidden state size, the same below), the original video frame token is the output of the visual information encoder (mapped by the connection layer). First, for Compress to get compressed visual token (in is the number of tokens contained in a single compressed video frame) In this embodiment, the original video frame token There are 729 tokens per frame, and the 729 tokens per frame are compressed. The specific compression method is to divide the 729 tokens into n parts, and then take the average of the parts, each part gets 1 token, and finally the 729 tokens are compressed into n tokens, that is, compressed visual tokens. The present invention first calculates The relative weight between frames , and then filter out the frames with high relative weights. The process of calculating relative weights is:
[0031] in, and for The result after mapping is Respectively represent The compressed visual tokens of all frames calculate attention weights with each other, and the sum of the attention weights of a frame relative to all other frames is the final relative weight of the frame. Figure 2 For convenience, we assume there are five frames. When calculating the weights, we calculate the attention weight of each frame relative to all frames. Then, we average the five attention weights corresponding to each frame to obtain the final attention weight for that frame, which is also known as the relative weight of that frame. In the figure, dark blue represents a high similarity score between two frames, while light blue represents a low similarity score. The average of the scores in the same column represents the relative weight of the frames corresponding to that column. The relative weight of each frame is then calculated based on the calculations for each column, and frames with high relative weights are selected.
[0032] For frames with relatively high weights, their tokens are converted from the tokens corresponding to the original video frames. For frames with low weights, their tokens are sampled from the compressed visual tokens The token sampling method adopted by the present invention can retain the number of tokens of video frames with high information content and reduce the number of tokens of frames with low information content, thereby reducing video memory overhead while retaining the integrity of video information.
[0033] Example 2 As another embodiment of the present invention, the present invention further provides a sampling system for long video understanding, which executes the sampling method for long video understanding described in Example 1 above, including: Video segmentation module: takes a long video and video subtitles as input, and uses a large video model to split the long video into multiple semantically consistent video segments; Text summary generation module: For each video segment, an input sequence is obtained through mapping by a visual encoder, and the input sequence is input into the video large language model to generate a text summary of the video segment; Token sampling module: Perform token sampling on the video segment, calculate the relative weights between frames of the video segment, and allocate different numbers of tokens to each frame according to the weights; Key frame acquisition module: Input the video frames after token sampling and the text summary into the video large language model, and based on the matching degree between the video frame tokens and the text summary, obtain the key frames of the video segment.
[0034] Furthermore, the video segmentation module specifically includes: a tokenizer that splits the text sequence of the input video caption into basic language units; an embedding layer that maps each language unit of the text sequence into a corresponding vector representation; a projection layer that performs dimensionality transformation on the vector representation and adjusts it to the input dimension of the video large language model to obtain a first input sequence; a video large language model that performs context modeling and semantic understanding based on the vector representation for the first input sequence and outputs the segmentation information of the video; a video processing tool that divides the long video into multiple video segments according to the segmentation information.
[0035] Furthermore, the text summary generation module specifically includes: a visual encoder that converts the video segment into a visual feature representation with semantic information; a connection layer that maps the visual features from the original visual space to a space consistent with the input dimension of the large language model to obtain a second input sequence; a video large language model that fuses the visual information and the language information in a unified representation space for the input of the second input sequence to obtain a text summary of the video segment.
[0036] Furthermore, the token sampling module specifically includes: Token acquisition unit, which uses a visual information encoder to obtain the tokens corresponding to the original video frames , for the tokens corresponding to the original video frames to be compressed to obtain compressed visual tokens ; Weight calculation unit, which calculates the attention weights between the compressed visual tokens of all frames, and the sum of the attention weights of a certain frame relative to all other frames is the final weight of this frame; Token allocation unit, for frames with high weights, their tokens are sampled from the corresponding positions in the tokens corresponding to the original video frames ; for frames with low weights, their tokens are sampled from the corresponding positions in the compressed visual tokens .
[0037] Further, the key frame acquisition module specifically includes: a frame extraction unit that extracts a video segment to obtain a frame set; a sampling unit that inputs the video frame token set obtained by token sampling and the text summary to the video large language, and finds the several frames in the video frame token set that best match the text summary according to the matching degree between the video frame tokens and the text summary.
[0038] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art should be within the protection scope determined by the claims.
Claims
1. A sampling method for long video understanding, characterized in that the steps Including: Obtain a long video and its video subtitle as input, and use a video large model to divide the long video into multiple video segments with consistent semantics; For each video segment, obtain an input sequence by mapping through a visual encoder, and input the input sequence into a video large language model to generate a text summary of the video segment; Perform token sampling on the video segment, calculate the relative weights between frames of the video segment, and allocate different numbers of tokens to each frame according to the weights; Input the video frames after token sampling and the text summary into the video large language model, and obtain the key frames of the video segment based on the matching degree between the video frame tokens and the text summary.
2. The sampling method for long video understanding according to claim 1, characterized in that The process of obtaining the video segment is as follows: The text sequence of the input video subtitle is processed by a tokenizer and split into basic language units; The text sequence passes through an embedding layer to map each language unit into a corresponding vector representation; The vector representation undergoes dimensional transformation through a projection layer and is adjusted to the input dimension of the video large language model to obtain a first input sequence; Input the first input sequence into the video large language model, perform context modeling and semantic understanding based on the vector representation, and output the segmentation information of the video; Divide the long video into multiple video segments according to the segmentation information.
3. The sampling method for long video understanding according to claim 1, wherein The process of generating the text summary is as follows: The video segment is converted into a visual feature representation with semantic information through a visual encoder; The visual features pass through a connection layer and are mapped from the original visual space to a space consistent with the input dimension of the large language model to obtain a second input sequence; Input the second input sequence into the video large language model, and fuse the visual information and the language information in a unified representation space to obtain the text summary of the video segment.
4. A sampling method for long video understanding according to claim 1, characterized in that The process of obtaining the key frames of the video segment is as follows: Extract the video segment to obtain a set of frames; Calculate the weight of each frame in the set of frames, perform token sampling according to the weights, and obtain a set of video frame tokens; Input the set of video frame tokens and the text summary into the video large language, and find the several frames in the set of video frame tokens that best match the text summary according to the matching degree between the video frame tokens and the text summary.
5. The sampling method for long video understanding according to claim 4, wherein The token sampling consists of two steps: weight calculation and sampling, which are as follows: Use a visual information encoder to obtain tokens corresponding to the original video frames For the tokens corresponding to the original video frames perform compression to obtain compressed visual tokens ; Calculate the attention weights between the compressed visual tokens of all frames, and take the sum of the attention weights of a certain frame relative to all other frames as the final weight of this frame; For frames with high weights, their tokens are sampled from the corresponding positions in the tokens corresponding to the original video frames ; For frames with low weights, their tokens are sampled from the corresponding positions in the compressed visual tokens wherein.
6. A sampling system for long video understanding, characterized in that, The system executes the sampling method for long video understanding as described in any one of claims 1-5, including: Video segmentation module: Obtain a long video and its video subtitle as input, and use a video large model to divide the long video into multiple video segments with consistent semantics; Text summary generation module: For each video segment, obtain an input sequence by mapping through a visual encoder, and input the input sequence into a video large language model to generate a text summary of the video segment; Token sampling module: Perform token sampling on the video segment, calculate the relative weights between frames of the video segment, and allocate different numbers of tokens to each frame according to the weights; Key frame acquisition module: Input the video frames after token sampling and the text summary into the video large language model, and obtain the key frames of the video segment based on the matching degree between the video frame tokens and the text summary.
7. The sampling system for long video understanding according to claim 6, characterized in that, The video segmentation module specifically includes: A tokenizer that splits the text sequence of the input video subtitles into basic language units; An embedding layer that maps each language unit of the text sequence to a corresponding vector representation; A projection layer that performs a dimensionality transformation on the vector representation and adjusts it to the input dimension of the video large language model to obtain a first input sequence; A video large language model that, for the first input sequence, performs context modeling and semantic understanding based on the vector representation and outputs the segmentation information of the video; A video processing tool that divides the long video into multiple video segments according to the segmentation information.
8. The sampling system for long video understanding according to claim 6, wherein, The text summary generation module specifically includes: A visual encoder that converts the video segment into a visual feature representation with semantic information; A connection layer that maps the visual features from the original visual space to a space consistent with the input dimension of the large language model to obtain a second input sequence; A video large language model that, for the input of the second input sequence, fuses the visual information and the language information in a unified representation space to obtain the text summary of the video segment.
9. The sampling system for long video understanding according to claim 6, wherein The token sampling module specifically includes: A token acquisition unit that uses a visual information encoder to acquire tokens corresponding to the original video frames For the tokens corresponding to the original video frames perform compression to obtain compressed visual tokens ; A weight calculation unit that calculates the attention weights among the compressed visual tokens of all frames, and takes the sum of the attention weights of a certain frame relative to all other frames as the final weight of that frame; Token allocation unit. For frames with high weights, their tokens are sampled from corresponding positions in the tokens corresponding to the original video frames. For frames with low weights, their tokens are sampled from corresponding positions in the compressed video perception tokens. 10. A sampling system for long video understanding according to claim 9, characterized in that, The key frame acquisition module specifically includes: A frame extraction unit that extracts the video segment to obtain a frame set; A sampling unit that inputs the video frame token set and the text summary obtained by the token sampling module to the video large language, and finds the several frames in the video frame token set that best match the text summary according to the matching degree between the video frame tokens and the text summary.
Citation Information
Patent Citations
Video processing method and device, computer equipment and storage medium
CN117336525A
Video automatic generation system based on Internet
CN119299802A
Long video understanding method and device, equipment and storage medium
CN119380240A
Multi-agent-based automatic generation method, system and terminal for micro-drama
CN119383413A
Text conditioned video resampler for video understanding
US20250166379A1
Cited By
Video semantic token compression method, video identification method and electronic equipment
CN120881297A