Multi-modal large-model adaptive video frame compression method and system

By decoupling user text instructions into three-dimensional instructions to generate a dynamic semantic weight matrix, and combining cross-modal attention and spatial importance gating technology, the redundancy and efficiency problems of large multimodal models in long video understanding are solved, and efficient and accurate video frame compression and real-time processing are achieved.

CN120751130APending Publication Date: 2025-10-03XIAMEN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511016440.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The Transformer-based multimodal large model has problems in long video understanding, such as spatiotemporal redundancy accumulation, semantic fragmentation, sensitivity to long-tail distribution, and computational efficiency bottlenecks, which leads to waste of computing resources and difficulty in ensuring real-time performance.

Method used

By decoupling user text instructions into three-dimensional instructions of time, space and context, a dynamic semantic weight matrix is ​​generated. Combined with the cross-modal attention mechanism and spatial importance gating technology, the number of visual features and spatial resolution are dynamically adjusted to perform adaptive video frame compression.

Benefits of technology

It significantly reduces computational complexity, improves the accuracy and real-time performance of video understanding, reduces visual-text semantic alignment errors, improves the detection recall rate of key frames, and reduces GPU memory usage, supporting efficient processing of hours of long videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751130A_ABST
    Figure CN120751130A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large model adaptive video frame compression method and system, and relates to the field of multi-modal video analysis, and the method comprises the steps: S1, obtaining a user text instruction and a sampling video frame of an original video; s2, converting the user text instruction into a space-time semantic instruction through hierarchical thinking chain reasoning; s3, extracting visual features of the sampled video frames, and performing importance scoring on the visual features through a space-time semantic instruction to obtain a semantic weight matrix; and S4, based on the semantic weight matrix, dynamically adjusting the number of visual features and the spatial resolution of each frame, and based on the new spatial resolution, adjusting adaptive pooling parameters and performing adaptive weighted pooling to obtain compressed and refined features. According to the method, a user text instruction is decoupled into a time, space and context three-dimensional instruction, a dynamic semantic weight matrix is generated, and a vision-text semantic alignment error is reduced; the token density is adaptively adjusted based on the weight matrix, the redundant region is compressed and merged, and the calculation complexity is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal video analysis, and in particular to a multimodal large model adaptive video frame compression method and system. Background Art

[0002] Currently, large multimodal models based on Transformer (such as Video-LLaMA, InternVL, QwenVL, etc.) have been fine-tuned on a large amount of video data and have performed well in video understanding. Recently, understanding models for long videos have expanded the context window to process longer video sequences, thereby achieving stable performance in long-time video understanding tasks. As the length of videos increases, more and more studies have also paid attention to the information redundancy problem in long video understanding, and thus proposed visual feature pruning schemes to make the model focus on highly responsive visual information, thereby reducing the impact of information redundancy on model performance.

[0003] Although the Transformer-based multimodal large model performs well in video understanding, it still has the following bottlenecks: ① Accumulation of spatiotemporal redundancy: The visual features of consecutive frames in long videos are highly repetitive. As the video sampling frame rate increases, the visual redundant content will increase exponentially. The visual features of consecutive frames in long videos are highly repetitive, resulting in a token utilization rate of less than 30% with the traditional uniform sampling strategy, and a large amount of computing resources are wasted on redundant information processing; ② Semantic fragmentation: Existing methods rely on attention weights to dynamically discard visual features, ignoring the spatiotemporal correlation between user instructions and video content, and key frames are easily compressed or omitted incorrectly; ③ Long-tail distribution sensitivity: Key frames are sparsely and unevenly distributed in long videos, and traditional sampling strategies cannot effectively cover them, resulting in a missed detection rate of more than 25%; ④ Computational efficiency bottleneck: High-density frame sampling causes a surge in video memory usage, especially when processing hour-level videos, real-time performance is difficult to guarantee. Summary of the Invention

[0004] To address the above problems, the present invention proposes a multimodal large-model adaptive video frame compression method and system. By decoupling user text instructions into three-dimensional instructions of time, space and context, the cross-modal attention mechanism is driven to generate a dynamic semantic weight matrix, accurately quantifying the correlation strength between video frames and user text instructions, and reducing the visual-text semantic alignment error; based on the weight matrix, the token density is adaptively adjusted, and combined with the spatial importance gating technology, the redundant areas are mathematically driven to compress and merge, significantly reducing the computational complexity.

[0005] On the one hand, a multimodal large model adaptive video frame compression method, the specific steps are as follows:

[0006] S1, a user text instruction and a sampled video frame acquisition step, acquiring a user text instruction and a sampled video frame of an original video;

[0007] S2, the multi-granularity semantic parsing step, extracts the time constraints, spatial entities and relationships, and the background knowledge that needs to be associated in the user's text instructions through hierarchical thinking chain reasoning to obtain spatiotemporal semantic instructions;

[0008] S3, the cross-modal feature distillation step, extracts the visual features of the sampled video frames, scores the importance of the visual features based on spatiotemporal semantic instructions, and obtains a semantic weight matrix representing the importance of each frame in relation to the user's text instructions;

[0009] S4, the spatiotemporal gated compression step, dynamically adjusts the number of visual features and spatial resolution of each frame based on the semantic weight matrix, adjusts the adaptive pooling parameters based on the new spatial resolution, and uses the adaptive pooling parameters to perform adaptive weighted pooling to obtain compressed and refined features.

[0010] Preferably, the multi-granularity semantic parsing step is specifically as follows:

[0011] The user text instructions are subjected to three-step hierarchical chain reasoning, and the output of each step of the hierarchical chain reasoning is the decoupled user instruction; the first step of the hierarchical chain reasoning uses the timestamp extraction algorithm to parse the time constraints in the user instruction and generate the time mask matrix M t The second step of the hierarchical thinking chain reasoning uses a cross-modal pre-training model to extract spatial entities and relationships and construct an entity relationship graph G s =(V s ,E s ), where node V s is a physical entity, edge E s Describe spatial relationships; the third step of the hierarchical thinking chain reasoning uses the text encoder to infer the background knowledge to be associated and generate the knowledge matrix K c ;

[0012] Use a cross-modal pre-trained model to integrate the decoupled user instructions and encode them into a unified multimodal representation to obtain spatiotemporal semantic instructions. .

[0013] Preferably, the visual features of the sampled video frames are extracted, and the importance of the visual features is scored using the spatiotemporal semantic instructions to obtain a semantic weight matrix representing the importance of each frame associated with the user text instruction, as follows:

[0014] Use CLIP-ViT-L / 14 visual feature extractor to extract visual features from the sampled video frames to obtain visual features;

[0015] The spatiotemporal semantic instructions are encoded into semantic vectors, and the semantic vectors, time masks and knowledge masks are fused to generate enhanced text features, which are expressed as:

[0016]

[0017] in, represents enhanced text features; E t Represents the semantic vector; M t Indicates the time mask; K c represents the knowledge mask; It means element-by-element addition;

[0018] Based on visual features and enhanced text features, a semantic weight matrix is ​​generated through cross-modal attention, which is expressed as:

[0019]

[0020] Among them, W i represents the semantic weight matrix; Softmax represents normalization along the spatial dimension to ensure that the sum of the weights is 1; (E i ) T Represents the visual feature E i The transpose of ; D represents the text feature channel.

[0021] Preferably, the text encoder is T5-XL; the cross-modal pre-trained model is BLIP-2, and ViT-L / 14 is used as a visual encoder to extract spatial entities and relationships; the decoupled user instructions are integrated using the cross-modal pre-trained model and encoded into a unified multimodal representation to obtain spatiotemporal semantic instructions, which are expressed as:

[0022]

[0023] in, Represents spatiotemporal semantic instructions, Indicates timing instructions, Represents a spatial instruction, Indicates contextual instructions; P H-CoT represents decoupled user instructions; Q represents video query; VLM represents cross-modal pre-training model.

[0024] Preferably, the parameters, gradients and optimizer states of the cross-modal pre-trained model are sharded to several GPUs to reduce CPU memory usage; the visual encoding operations performed by the visual encoder, the text encoding operations performed by the text encoder, and the feature fusion operations for integrating decoupled user instructions using the cross-modal pre-trained model are executed in a staged pipeline to reduce latency.

[0025] Preferably, the spatiotemporal gating compression step is specifically:

[0026] Perform dynamic feature allocation and calculate the number of features per frame based on the semantic weight matrix, expressed as:

[0027]

[0028] Among them, N i represents the number of features per frame; α represents the number of bases, and β represents the scaling factor; represents the L2 norm square; W i represents the semantic weight matrix;

[0029] Calculate the target resolution, expressed as:

[0030]

[0031] in, represents the target size, Indicates the height of the target size; Indicates the width of the target size; C indicates the number of visual feature channels;

[0032] Determine the adaptive pooling parameters and calculate the pooling kernel size (k h ,k w ) and step size (s h ,s w ), expressed as:

[0033]

[0034] Where H×W represents the original space size, H represents the height of the original space size, and W represents the width of the original space size; Indicates rounding down;

[0035] Adaptive pooling parameters are used to perform adaptive weighted pooling on the weight area corresponding to each pooling window, which is expressed as

[0036]

[0037] Where ∈ = 1e-8 to avoid division by zero; B v Represents the adaptive pooling visual feature block corresponding to the pooling window P. The set of all adaptive pooling visual feature blocks is {B v};W P (m,n) represents the weight of the pooling window P at position (m,n),

[0038] E i is the visual feature; E i (m,n) represents the visual feature at position (m,n);

[0039] {B v All feature blocks in} are concatenated into compressed and refined features.

[0040] Preferably, after S4, the step further includes: S5, a long-tail distribution optimization step, which is specifically as follows:

[0041] Dynamically resample the original video for the areas corresponding to the sampled video frames whose weights in the semantic weight matrix are less than a preset weight threshold, and select new key frames; the resampling parameters are specifically:

[0042]

[0043] Where Δt represents the time interval of resampling; T total Indicates the total length of the video; N retry Indicates the number of resampling times; Indicates rounding down;

[0044] After each resampling, the new keyframe is dynamically extended to perform visual feature assignment, which is expressed as:

[0045]

[0046] in, N represents the number of features per frame after dynamic expansion of visual feature allocation; i represents the number of features per frame; δ represents the increment coefficient to ensure gradual refinement.

[0047] On the other hand, a multimodal large model adaptive video frame compression system includes the following:

[0048] A user text instruction and sampled video frame acquisition module, used to acquire user text instructions and sampled video frames of original videos;

[0049] The multi-granularity semantic parsing module is used to extract the time constraints, spatial entities and relationships, and the background knowledge required to be associated in the user's text instructions through hierarchical thought chain reasoning to obtain spatiotemporal semantic instructions;

[0050] The cross-modal feature distillation module extracts visual features from sampled video frames, scores the importance of these visual features based on spatiotemporal semantic instructions, and generates a semantic weight matrix representing the importance of each frame in relation to the user's textual instructions.

[0051] The spatiotemporal gated compression module is used to dynamically adjust the number of visual features and spatial resolution of each frame based on the semantic weight matrix, adjust the adaptive pooling parameters based on the new spatial resolution, and use the adaptive pooling parameters to perform adaptive weighted pooling to obtain compressed and refined features.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] (1) The multi-granularity semantic parsing step of the present invention is responsible for converting user natural language instructions into structured spatiotemporal semantic instructions. Through the Hierarchical Chain-of-Thoughts (H-CoT) technology, complex queries are decomposed into executable instructions in the three dimensions of time, space, and context, providing precise guidance for subsequent processing.

[0054] (2) The cross-modal feature distillation step of the present invention is responsible for fusing visual and text features, constructing a dynamic semantic weight matrix, and quantifying the association strength between each region in the video frame and the user's natural language instructions; through the cross-modal attention mechanism, fine-grained visual-text alignment is achieved, providing a quantitative basis for token allocation. The introduced cross-modal attention weight matrix reduces the visual-text semantic alignment error by 22% (from 0.35 to 0.27), improves the accuracy by 15.3% (68.2% to 71.1%) in the MLVU long video benchmark test, and the detection recall rate of key entities (such as smoke and specific people) exceeds 90%;

[0055] (3) The spatiotemporal gating compression step of the present invention dynamically adjusts the number of visual features and spatial resolution of each frame based on the semantic weight matrix to achieve a balance between redundancy suppression and key information retention. Low-weight regions are adaptively compressed and refined through spatial importance gating (SIG) technology.

[0056] (4) The long-tail distribution optimization step of the present invention improves the coverage of low-weight frames and reduces the risk of missed detection by using a resampling strategy and a progressive enhancement mechanism, targeting the sparse distribution of key frames in long videos;

[0057] (5) The present invention reduces GPU memory usage and improves inference speed by sharding model parameters, gradients, and optimizer states onto multiple GPUs. In particular, when processing videos longer than one hour, the performance degradation rate is much better than that of the traditional uniform sampling scheme. Through distributed computing and memory optimization technology, the present invention supports efficient processing of hour-long videos and solves the resource bottleneck problem of a single card. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The present invention will be described in further detail below with reference to the accompanying drawings;

[0059] Figure 1 This is a flowchart of a multimodal large model adaptive video frame compression method according to an embodiment of the present invention;

[0060] Figure 2 Schematic diagram of a flow chart of a multimodal large model adaptive video frame compression method according to an embodiment of the present invention;

[0061] Figure 3This is a structural block diagram of a multimodal large model adaptive video frame compression system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0062] The present invention is further described below through specific embodiments.

[0063] like Figure 1 and Figure 2 As shown in FIG, a multimodal large model adaptive video frame compression method, the specific steps are as follows:

[0064] S1, user text instruction and sample video frame acquisition step.

[0065] Get user text instructions and sampled video frames of the original video.

[0066] S2, multi-granularity semantic parsing step.

[0067] In this stage, the text instructions input by the user are decomposed into spatiotemporal semantic instructions through hierarchical chain of thought reasoning (H-CoT) This helps the model understand user instructions and provide higher quality guidance for subsequent adaptive adjustment of frame resolution. It specifically includes the following three main parts: (1) Timing instructions Use timestamp extraction algorithms (such as regular expression matching and context inference) to parse the time constraints in user instructions (such as "first 10 minutes", "last scene", etc.) to generate the time mask matrix M t , marking key time periods; (2) spatial instructions The entity detection capability based on the vision-language model (BLIP-2) extracts spatial entities and relationships (such as "table on the left" and "person holding an object") and constructs the entity relationship graph G s =(V s ,E s ), where node V s is a physical entity (such as "people, objects"), and edge E s Describing spatial relationships (e.g., “on the left,” “holding an object”); (3) Contextual instructions Infer the background knowledge that needs to be associated (such as "holiday decoration tradition") and generate the knowledge matrix K c Based on the above three-step hierarchical thinking chain reasoning, we can get the decoupled user instruction P H-CoT , and integrate to get semantic instructions Expressed as:

[0068]

[0069] Among them, VLM uses BLIP-2 (cross-modal pre-training model), the text encoder is T5-XL, and the visual encoder is ViT-L / 14.

[0070] S3, cross-modal feature distillation step.

[0071] In this stage, the semantic instructions obtained in the previous stage The importance of the sampled video frames is scored to obtain the importance of each frame in relation to the user's instructions. This importance is then used as a guide for adaptive compression and refinement of visual features. Specifically, this involves the following three steps:

[0072] Visual feature extraction: for video frame F i Extract features (CLIP-ViT-L / 14), where the video frame spatial resolution H = W = 14 and the number of visual feature channels C = 768;

[0073] Text feature encoding: semantic instructions Encoded as a vector (text feature channel D = 1024), fusion time mask M t With the knowledge mask K c , generate enhanced text features

[0074] Dynamic frame importance weight calculation: Generating frame-level semantic weight matrix via cross-modal attention Expressed as:

[0075]

[0076] Among them, Softmax is normalized along the spatial dimension to ensure that the sum of weights is 1.

[0077] S4, spatiotemporal gated compression step.

[0078] The spatial importance gating (SIG) technology proposed in this embodiment is based on the frame-level semantic weight W obtained in the previous step. i Dynamic compression refines the number of visual features of each video frame. The specific steps are as follows:

[0079] Dynamic feature allocation: First, calculate the number of features per frame based on the weights Wherein, α is the base number, the scaling factor β is 0.5, and the global semantic strength is quantified by the L2 norm square; in this embodiment, α=64, β=0.5.

[0080] Target resolution calculation: The new spatial resolution is expressed as

[0081] Adaptive pooling parameter determination: based on the original space size H×W and the target size Calculate the pooling kernel size (k h ,k w ) and step size (s h ,s w ),in, k w With s w Similarly,

[0082] Adaptive weighted pooling: For each pooling window According to its corresponding weight area Calculate the weighted average, expressed as:

[0083]

[0084] Where ∈ = 1e-8 to avoid division by zero, {B v} represents the adaptive pooling visual feature block;

[0085] Merge compressed features: Combine all block features {B v}Spliced ​​into compressed and refined features

[0086] S5, long-tail distribution optimization step.

[0087] Due to the refinement of visual features, the model can accept more potential key visual information. Therefore, in this step, the original video frames are dynamically sampled through the long-tail distribution optimization step. The specific steps are as follows:

[0088] Keyframe resampling: For low-weight frames (W i <τ, τ=0.1), then resample according to the time interval Δt, expressed as:

[0089]

[0090] Among them, T total is the total length of the video (seconds), N retry is the number of resampling times (for example, for a video of 1 hour duration, T total =3600, N retry =5, Δt=720 seconds);

[0091] Progressive enhancement: After each resampling, if a new keyframe is detected, the visual feature allocation is dynamically expanded, expressed as:

[0092]

[0093] Among them, δ=32 is the incremental coefficient to ensure gradual refinement.

[0094] Distributed Computing: To improve VLM inference speed, the Deepspeed ZeRO-3 strategy is used to shard model parameters, gradients, and optimizer states across multiple GPUs, reducing video memory usage by 40%. Visual encoding, text encoding, and feature fusion are executed in a phased pipeline, reducing latency by 35%.

[0095] In summary, this embodiment breaks through the efficiency and accuracy bottlenecks in long video understanding through a multimodal collaborative dynamic semantic alignment and resource optimization mechanism. Compared to existing technologies that rely on pre-trained models to expand upper and lower windows or visual feature compression algorithms based on attention weight analysis, this solution, based on hierarchical chain reasoning (H-CoT) and cross-modal feature distillation technology, achieves full-process innovation from semantic parsing to efficient reasoning. At the technical implementation level, the multi-granularity semantic parsing step decouples user queries into three-dimensional instructions of time, space, and context, driving the cross-modal attention mechanism to generate a dynamic semantic weight matrix, accurately quantifying the correlation strength between video frames and queries. On this basis, the spatiotemporal gated compression step adaptively adjusts the token density based on the weight matrix, and combines the spatial importance gating (SIG) technology to perform mathematically driven compression and merging of redundant regions, significantly reducing computational complexity. In view of the sparse distribution of key frames in long videos, a long-tail distribution optimization step is introduced. Through resampling strategies and progressive token expansion, the key frame recall rate is increased to 92% and the missed detection rate is compressed to below 8%. Furthermore, the distributed computing engine leverages ZeRO-3 memory sharding and pipeline parallelization technology to support real-time processing of two hours of video on a single graphics card, overcoming the resource limitations of traditional solutions. This solution, based entirely on open-source models and a modular design, adapts to mainstream multimodal large models (such as QwenVL and InternVL) without fine-tuning, ensuring flexibility and scalability while reducing deployment costs.

[0096] By adopting the above technical solutions, this embodiment shows significant advantages in terms of efficiency, accuracy and practicality. In terms of processing efficiency, the coordinated optimization of spatiotemporal gated compression and distributed computing reduces GPU memory usage by 45% (from 48GB to 28GB) and increases the inference speed by 2.5 times (45FPS to 112FPS). Especially when processing videos of more than 1 hour, the performance degradation rate is less than 5%, which is much better than the traditional uniform sampling scheme. In terms of understanding accuracy, the introduction of the cross-modal attention weight matrix reduces the visual-text semantic alignment error by 22% (from 0.35 to 0.27), and improves the accuracy by 15.3% (68.2% to 71.1%) in the MLVU long video benchmark test. The detection recall rate of key entities (such as smoke and specific people) exceeds 90%. This embodiment shows strong robustness in complex scenarios and can flexibly adapt to multiple types of video content such as news, education, and security, with a generalization error of less than 3%. In actual deployment, through an open-source toolchain and modular design, it supports full-stack deployment from cloud clusters to edge devices (such as NVIDIA Jetson AGX), reducing operation and maintenance costs by 60% compared to commercial API solutions, providing an efficient and economical solution for large-scale video analysis applications. Experiments have shown that the solution of this embodiment not only overcomes the problems of redundant computation and semantic fragmentation in long video understanding, but also opens up a new technical path for the practical application of large multimodal models through mathematically driven compression and distributed optimization.

[0097] like Figure 3 As shown, the present invention also discloses a multimodal large model adaptive video frame compression system, comprising:

[0098] A user text instruction and sample video frame acquisition module 301 is used to acquire user text instructions and sample video frames of the original video;

[0099] The multi-granularity semantic parsing module 302 is used to extract the time constraints, spatial entities and relationships, and the background knowledge to be associated in the user text instructions through hierarchical thought chain reasoning to obtain spatiotemporal semantic instructions;

[0100] Cross-modal feature distillation module 303 is used to extract visual features of sampled video frames, score the importance of the visual features based on spatiotemporal semantic instructions, and obtain a semantic weight matrix representing the importance of each frame in relation to the user text instruction;

[0101] The spatiotemporal gated compression module 304 is used to dynamically adjust the number of visual features and spatial resolution of each frame based on the semantic weight matrix, adjust the adaptive pooling parameters based on the new spatial resolution, and use the adaptive pooling parameters to perform adaptive weighted pooling to obtain compressed and refined features.

[0102] The long-tail optimization module 305 is used to dynamically resample the original video for the areas corresponding to the sampled video frames whose weights in the semantic weight matrix are less than a preset weight threshold, and screen new key frames; after each resampling, dynamically expand the visual feature allocation for the new key frames.

[0103] The specific implementation of a multimodal large model adaptive video frame compression system and the multimodal large model adaptive video frame compression method are not repeated in this embodiment.

[0104] The above is only a specific implementation of the present invention, but the design concept of the present invention is not limited to this. Any non-substantial changes to the present invention using this concept shall be deemed as an infringement of the protection scope of the present invention.

Claims

1. A multimodal large model adaptive video frame compression method, characterized in that: The steps include: S1, a user text instruction and a sampled video frame acquisition step, acquiring a user text instruction and a sampled video frame of an original video; S2, the multi-granularity semantic parsing step, extracts the time constraints, spatial entities and relationships, and the background knowledge that needs to be associated in the user's text instructions through hierarchical thinking chain reasoning to obtain spatiotemporal semantic instructions; S3, the cross-modal feature distillation step, extracts the visual features of the sampled video frames, scores the importance of the visual features based on spatiotemporal semantic instructions, and obtains a semantic weight matrix representing the importance of each frame in relation to the user's text instructions; S4, the spatiotemporal gated compression step, dynamically adjusts the number of visual features and spatial resolution of each frame based on the semantic weight matrix, adjusts the adaptive pooling parameters based on the new spatial resolution, and uses the adaptive pooling parameters to perform adaptive weighted pooling to obtain compressed and refined features.

2. The multimodal large model adaptive video frame compression method according to claim 1, characterized in that: The multi-granularity semantic parsing steps are as follows: The user text instructions are subjected to three-step hierarchical chain reasoning, and the output of each step of the hierarchical chain reasoning is the decoupled user instruction; the first step of the hierarchical chain reasoning uses the timestamp extraction algorithm to parse the time constraints in the user instruction and generate the time mask matrix M t The second step of the hierarchical thinking chain reasoning uses a cross-modal pre-training model to extract spatial entities and relationships and construct an entity relationship graph G s =(V s ,E s ), where node V s is a physical entity, edge E s Describe spatial relationships; the third step of the hierarchical thinking chain reasoning uses the text encoder to infer the background knowledge to be associated and generate the knowledge matrix K c ; Use a cross-modal pre-trained model to integrate the decoupled user instructions and encode them into a unified multimodal representation to obtain spatiotemporal semantic instructions.

3. The multimodal large model adaptive video frame compression method according to claim 2, characterized in that: The visual features of the sampled video frames are extracted, and the importance of the visual features is scored using the spatiotemporal semantic instructions to obtain a semantic weight matrix representing the importance of each frame associated with the user text instruction, as follows: Use CLIP-ViT-L / 14 visual feature extractor to extract visual features from the sampled video frames to obtain visual features; The spatiotemporal semantic instructions are encoded into semantic vectors, and the semantic vectors, time masks and knowledge masks are fused to generate enhanced text features, which are expressed as: in, represents enhanced text features; E t Represents the semantic vector; M t Indicates the time mask; K c represents the knowledge mask; It means element-by-element addition; Based on visual features and enhanced text features, a semantic weight matrix is ​​generated through cross-modal attention, which is expressed as: Among them, W i represents the semantic weight matrix; Softmax represents normalization along the spatial dimension to ensure that the sum of the weights is 1; (E i ) T Represents the visual feature E i The transpose of ; D represents the text feature channel.

4. The multimodal large model adaptive video frame compression method according to claim 2, characterized in that: The text encoder is T5-XL; the cross-modal pre-trained model is BLIP-2, and ViT-L / 14 is used as the visual encoder to extract spatial entities and relationships; the decoupled user instructions are integrated using the cross-modal pre-trained model and encoded into a unified multimodal representation to obtain spatiotemporal semantic instructions, which are expressed as: in, Represents spatiotemporal semantic instructions, Indicates timing instructions, Represents a spatial instruction, Indicates contextual instructions; P H-CoT represents decoupled user instructions; Q represents video query; VLM represents cross-modal pre-training model.

5. The multimodal large model adaptive video frame compression method according to claim 3, characterized in that: The parameters, gradients, and optimizer states of the cross-modal pre-trained model are sharded onto several GPUs to reduce CPU memory usage. The visual encoding operations performed by the visual encoder, the text encoding operations performed by the text encoder, and the feature fusion operations for integrating decoupled user instructions using the cross-modal pre-trained model are executed in a staged pipeline to reduce latency.

6. The multimodal large model adaptive video frame compression method according to claim 1, characterized in that: The spatiotemporal gating compression step is specifically as follows: Perform dynamic feature allocation and calculate the number of features per frame based on the semantic weight matrix, expressed as: Among them, N i represents the number of features per frame; α represents the number of bases, and β represents the scaling factor; represents the L2 norm square; W i represents the semantic weight matrix; Calculate the target resolution, expressed as: in, represents the target size, Indicates the height of the target size; Indicates the width of the target size; C indicates the number of visual feature channels; Determine the adaptive pooling parameters and calculate the pooling kernel size (k h ,k w ) and step size (s h ,s w ), expressed as: Where H×W represents the original space size, H represents the height of the original space size, and W represents the width of the original space size; Indicates rounding down; Adaptive pooling parameters are used to perform adaptive weighted pooling on the weight area corresponding to each pooling window, which is expressed as Where ∈ = 1e-8 to avoid division by zero; B v Represents the adaptive pooling visual feature block corresponding to the pooling window P. The set of all adaptive pooling visual feature blocks is {B v };W P (m,n) represents the weight of the pooling window P at position (m,n), E i is the visual feature; E i (m,n) represents the visual feature at position (m,n); {B v All feature blocks in} are concatenated into compressed and refined features.

7. The multimodal large model adaptive video frame compression method according to claim 1, characterized in that: After S4, the following further steps are included: S5, a long-tail distribution optimization step, specifically as follows: Dynamically resample the original video for the areas corresponding to the sampled video frames whose weights in the semantic weight matrix are less than a preset weight threshold, and select new key frames; the resampling parameters are specifically: Where Δt represents the time interval of resampling; T total Indicates the total length of the video; N retry Indicates the number of resampling times; Indicates rounding down; After each resampling, the new keyframe is dynamically extended to perform visual feature assignment, which is expressed as: in, N represents the number of features per frame after dynamic expansion of visual feature allocation; i represents the number of features per frame; δ represents the increment coefficient to ensure gradual refinement.

8. A multimodal large model adaptive video frame compression system, comprising: A user text instruction and sampled video frame acquisition module, used to acquire user text instructions and sampled video frames of original videos; The multi-granularity semantic parsing module is used to extract the time constraints, spatial entities and relationships, and the background knowledge required to be associated in the user's text instructions through hierarchical thought chain reasoning to obtain spatiotemporal semantic instructions; The cross-modal feature distillation module extracts visual features from sampled video frames, scores the importance of these visual features based on spatiotemporal semantic instructions, and generates a semantic weight matrix representing the importance of each frame in relation to the user's textual instructions. The spatiotemporal gated compression module is used to dynamically adjust the number of visual features and spatial resolution of each frame based on the semantic weight matrix, adjust the adaptive pooling parameters based on the new spatial resolution, and use the adaptive pooling parameters to perform adaptive weighted pooling to obtain compressed and refined features.

Citation Information

Cited By

  • Multi-mode large model video content understanding reasoning acceleration method and system

    CN121305451A

  • Dynamic redundancy compression and multi-modal optimization method and system based on AI intelligent technology

    CN121486582A