Dynamic redundancy compression and multi-modal optimization method and system based on AI intelligent technology

By employing a dynamic redundancy compression and multimodal optimization method based on AI intelligent technology, the problem of poor quality in traditional video compression for complex video content processing is solved, achieving efficient video data compression and reconstruction, and improving coding efficiency and visual effects.

CN121486582APending Publication Date: 2026-02-06TIANJIN JUXIN GUANGHE TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511761104.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Traditional video compression methods often fail to maintain good reconstruction quality when processing complex video content, especially in low bitrate environments. The compressed image may lose important details and exhibit blurring and compression artifacts.

Method used

We employ a dynamic redundancy compression and multimodal optimization method based on AI technology. By using an AI model to analyze the spatiotemporal redundancy and knowledge redundancy of video in real time, we utilize feature sets of image, audio and text modalities, combined with spatiotemporal attention maps and hierarchical compression strategies, to optimize the encoding and decoding process of video frames.

Benefits of technology

It significantly improves video processing capabilities, reduces the proportion of redundant information, maintains high-quality video visual effects, and improves coding efficiency and reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486582A_ABST
    Figure CN121486582A_ABST
Patent Text Reader

Abstract

The invention relates to a dynamic redundancy compression and multi-modal optimization method and system based on an AI intelligent technology. Comprising the following steps: acquiring video data, carrying out feature extraction on the video data, and respectively acquiring a feature set of each modal in a video and a spatio-temporal feature set between continuous frames of the video; according to the feature set of each modal, determining the feature priority of each modal based on a modal feature importance evaluation rule; performing optimization processing through an encoder according to the spatio-temporal feature set, and reconstructing video continuous frames through decoding; calculating a space-time attention map according to the space-time features, wherein the value of each position in the space-time attention map represents the importance degree of the feature of each position; processing the spatio-temporal features by using a spatio-temporal attention map to obtain a coded representation of each position; and according to the feature priority of each mode in the video frame, compressing the video frame by using a hierarchical compression strategy to obtain a compressed video frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video compression processing technology, and in particular to a dynamic redundancy compression and multimodal optimization method and system based on AI intelligent technology. Background Technology

[0002] Video compression technology plays a crucial role in multimedia processing, primarily reducing the storage and transmission requirements of video data. Uncompressed video files are typically large in size, placing significant strain on storage devices and network bandwidth. The goal of video compression technology is to simplify data representation by removing redundant information, while maintaining high-quality video visuals as much as possible in the process. Traditional video compression methods typically rely on standard coding techniques such as H.264 and HEVC, which mainly improve compression efficiency through inter- and intra-frame redundancy removal, transform coding, and entropy coding.

[0003] While these traditional methods perform well in some applications, they often fail to maintain good reconstruction quality when processing complex video content, especially in low bitrate environments. As video content becomes more complex and bitrate requirements decrease, the limitations of traditional video coding techniques become increasingly apparent; compressed images may lose important details or even exhibit noticeable blurring and compression artifacts.

[0004] To address this, the present invention proposes a dynamic redundancy compression and multimodal optimization method and system based on AI intelligent technology. Based on dynamic redundancy identification technology, the method analyzes the spatiotemporal redundancy and knowledge redundancy of video in real time through AI models, significantly improving the long video processing capability, adapting to multimodal large models and optimizing KV Cache, and reducing the proportion of redundant information. Summary of the Invention

[0005] This invention addresses the technical problems existing in the prior art by providing a method and system for dynamic redundancy compression and multimodal optimization based on AI intelligent technology.

[0006] The technical solution of this invention to solve the above-mentioned technical problems is as follows: a method and system for dynamic redundancy compression and multimodal optimization based on AI intelligent technology; the method includes: S1: Acquire video data and extract its features, obtaining the feature sets of each modality in the video and the spatiotemporal feature sets between consecutive frames of the video; S2: Based on the feature sets of each mode, determine the feature priority of each mode according to the modality feature importance evaluation rules; S3: Based on the spatiotemporal feature set, the encoder is used to optimize the set, and the video frames are reconstructed by decoding. S4: Calculate a spatiotemporal attention map based on the spatiotemporal features, where the value at each position in the spatiotemporal attention map represents the importance of the feature at each position; process the spatiotemporal features using the spatiotemporal attention map to obtain the encoded representation of each position; S5: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame to obtain the compressed video frame.

[0007] Furthermore, the dynamic redundancy compression and multimodal optimization method based on AI intelligent technology includes image modality, audio modality, and text modality. The feature sets of each modality are as follows: the image modality includes the pixel area ratio of the main object and the number of mouse clicks on the image; the audio modality includes the effective speech duration ratio and the number of audio segment tags; and the text modality includes the ratio of the number of keywords to the total number of words and the number of text copies.

[0008] Furthermore, in the aforementioned dynamic redundancy compression and multimodal optimization method based on AI intelligent technology, the specific analysis process of the feature importance of the image modality is as follows: dynamically calculate the difference between the pixel area ratio of the main object and the number of mouse clicks on the image and preset values ​​respectively. Normalization of each dynamic difference yields the content contribution and user attention of the image modality; The feature importance of the image modality is generated by summing the weighted calculations of content contribution and user attention with their respective percentage weights.

[0009] Furthermore, in the aforementioned dynamic redundancy compression and multimodal optimization method based on AI intelligent technology, the encoder includes a feature extraction module, which extracts temporal features from consecutive video frames. The feature extraction module includes several spatiotemporal convolutional units and several temporal downsampling units, with each temporal downsampling unit positioned after one of the spatiotemporal convolutional units. The spatiotemporal convolutional units are used to perform convolution processing on the data input to the spatiotemporal convolutional units. The temporal downsampling unit is used to perform temporal downsampling on the feature data input to the temporal downsampling unit in order to reduce the temporal resolution of the feature data.

[0010] Furthermore, the dynamic redundancy compression and multimodal optimization method based on AI intelligent technology optimizes the spatiotemporal features through the spatiotemporal attention module in the encoder to obtain the encoded representation of the spatiotemporal block to be processed, including: calculating a spatiotemporal attention map based on the spatiotemporal features through the spatiotemporal attention module, wherein the value of each position in the spatiotemporal attention map represents the importance of the feature at each position; The spatiotemporal features are processed using the spatiotemporal attention map to obtain the encoded representation of the spatiotemporal block to be processed.

[0011] Furthermore, in the AI-based dynamic redundancy compression and multimodal optimization method, the encoder is constructed using a 3D convolutional neural network, and the decoder uses a temporal context modeling mechanism to reconstruct each frame based on the context information of the preceding and following frames in the spatiotemporal block to obtain the decoding result of the spatiotemporal block to be processed.

[0012] A dynamic redundancy compression and multimodal optimization system based on AI technology, applied to any of the described dynamic redundancy compression and multimodal optimization methods based on AI technology, wherein the system comprises: Data Acquisition and Feature Extraction Module: This module is used to acquire video data and extract its features. It can acquire feature sets for each modality in the video and spatiotemporal feature sets between consecutive video frames. The feature sets for each modality are determined according to the characteristics of different modalities. For image modalities, the feature sets include the pixel area ratio of the main object and the number of mouse clicks on the image. For audio modalities, the feature sets include the effective speech duration ratio and the number of audio segment tags. For text modalities, the feature sets include the ratio of the number of keywords to the total number of words and the number of text copies. Feature Priority Determination Module: Based on the feature sets of each modality, the feature priority of each modality is determined according to the modality feature importance evaluation rules; different feature importance analysis processes are used for different modalities. Encoding Optimization Module: The encoder in this module optimizes the spatiotemporal feature set. The encoder includes a feature extraction module, which consists of several spatiotemporal convolutional units and several temporal downsampling units, each placed after a spatiotemporal convolutional unit. The spatiotemporal convolutional units perform convolution processing on the input data, and the temporal downsampling units perform temporal downsampling on the input feature data to reduce the temporal resolution of the feature data. Simultaneously, the encoder also calculates a spatiotemporal attention map based on the spatiotemporal features using a spatiotemporal attention module. This attention map is then used to process the spatiotemporal features to obtain the encoded representation of the spatiotemporal block to be processed. The encoder can be constructed using a 3D convolutional neural network. Decoding and Reconstruction Module: The decoder uses a temporal context modeling mechanism to reconstruct each frame based on the context information of the preceding and following frames in the spatiotemporal block to obtain the decoding result of the spatiotemporal block to be processed, thereby reconstructing the continuous frames of the video.

[0013] Furthermore, the dynamic redundancy compression and multimodal optimization system based on AI intelligent technology also includes: Spatiotemporal attention processing module: Calculates a spatiotemporal attention map based on spatiotemporal features. The value of each position in the spatiotemporal attention map represents the importance of the feature at each position. The spatiotemporal features are processed using this spatiotemporal attention map to obtain the encoded representation of each position. Video frame compression module: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame, and finally the compressed video frame is obtained.

[0014] The beneficial effects of this invention are: The system can acquire feature sets from three different modalities in videos: image, audio, and text. It can also extract spatiotemporal feature sets between consecutive video frames. This comprehensive feature coverage allows the system to analyze and process video data from multiple dimensions, providing a rich and detailed information foundation for subsequent optimization and compression.

[0015] Different feature importance analysis processes are used for different modalities. For example, for the image modality, the pixel area ratio of the main object and the number of mouse clicks on the image are dynamically differencing preset values. After normalization, the content contribution and user attention are obtained, and finally, a weighted sum is used to generate the feature importance. This modal evaluation method considers the unique attributes of different modalities, making the determination of feature priorities more scientific and reasonable, and highlighting the truly important features in each modality.

[0016] The feature extraction module in the encoder consists of a spatiotemporal convolutional unit and a temporal downsampling unit. The spatiotemporal convolutional unit performs convolution processing to extract features, while the temporal downsampling unit reduces the temporal resolution of the feature data, effectively reducing the amount of data. Simultaneously, the spatiotemporal attention module calculates and processes spatiotemporal attention maps based on spatiotemporal features, highlighting important spatiotemporal features and improving encoding efficiency. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating a dynamic redundancy compression and multimodal optimization method based on AI intelligent technology. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0020] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0021] In one embodiment, the dynamic redundancy compression and multimodal optimization method based on AI intelligent technology includes: S1: Acquire video data and extract its features, obtaining the feature sets of each modality in the video and the spatiotemporal feature sets between consecutive frames of the video; S2: Based on the feature sets of each mode, determine the feature priority of each mode according to the modality feature importance evaluation rules; S3: Based on the spatiotemporal feature set, the encoder is used to optimize the set, and the video frames are reconstructed by decoding. S4: Calculate a spatiotemporal attention map based on the spatiotemporal features, where the value at each position in the spatiotemporal attention map represents the importance of the feature at each position; process the spatiotemporal features using the spatiotemporal attention map to obtain the encoded representation of each position; S5: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame to obtain the compressed video frame.

[0022] Specifically, the modalities include image modality, audio modality, and text modality, and the feature sets of each modality are as follows: the image modality includes the pixel area ratio of the main object and the number of mouse clicks on the image; the audio modality includes the effective speech duration ratio and the number of audio segment tags; and the text modality includes the ratio of the number of keywords to the total number of words and the number of text copies.

[0023] It should be noted that the method for obtaining the pixel area ratio of the main object is as follows: 1) Preprocessing: Convert the video frame to RGB format and adjust the resolution to meet the algorithm input requirements; 2) Object detection: Input the object detection model and obtain the bounding boxes and category labels of all detected objects; 3) Subject filtering: Filter irrelevant objects according to the subject categories preset in the video business scenario (such as "face" and "license plate" for security videos, and "podium" and "teacher" for educational videos); 4) Area calculation: Calculate the sum of the pixel areas of the filtered objects and divide it by the total pixel area of ​​the video frame to obtain the proportion of the main object.

[0024] It should be noted that the effective speech duration percentage is obtained in the following ways: 1) Audio framing: the audio signal is divided into frames of fixed duration; 2) VAD detection: the energy value and zero-crossing rate of each frame are calculated and input into the VAD classifier to determine whether it is a speech frame; 3) Speech segment merging: consecutive speech frames are merged into speech segments and noise segments shorter than a set threshold are filtered out; 4) Activity calculation: the total duration of effective speech segments is divided by the total duration of audio to obtain the speech activity.

[0025] It should be noted that the ratio of the number of keywords to the total number of words is obtained as follows: 1) Text preprocessing: the text (subtitles, OCR recognition results) is segmented, stop words are removed, and part-of-speech tagging is performed; 2) Keyword extraction: entity words are extracted through the text named entity recognition model, and professional terms are selected by combining the video's pre-set business dictionary (such as "symptoms" and "drugs" in medical videos, and "knowledge points" and "concepts" in educational videos); 3) Density calculation: the number of keywords is divided by the total number of words in the text to obtain the ratio of the number of keywords to the total number of words.

[0026] Specifically, the specific analysis process for the feature importance of the image modality is as follows: the pixel area ratio of the main object and the number of mouse clicks on the image are dynamically calculated with preset values; It should be noted that the dynamic difference calculation specifically refers to: the difference between the pixel area ratio of the main object and the preset pixel area ratio of the main object, and the difference between the number of mouse clicks on the image and the preset number of mouse clicks on the image.

[0027] Normalization of each dynamic difference yields the content contribution and user attention of the image modality; It should be noted that the specific process for obtaining the content contribution and user attention of the image modality is as follows: the ratio of the difference between the pixel area ratio of the main object and the preset pixel area ratio of the main object to the preset pixel area ratio of the main object is used as the content contribution, and the ratio of the difference between the number of mouse clicks on the image and the preset number of mouse clicks on the image to the preset number of mouse clicks on the image is used as the user attention.

[0028] The feature importance of the image modality is generated by summing the weighted calculations of content contribution and user attention with their respective percentage weights.

[0029] Specifically, the encoder includes a feature extraction module, which extracts time features from consecutive frames of the video. The feature extraction module includes several spatiotemporal convolution units and several temporal downsampling units. Each temporal downsampling unit is set after a spatiotemporal convolution unit. The spatiotemporal convolution unit is used to perform convolution processing on the data input to the spatiotemporal convolution unit. The temporal downsampling unit is used to perform temporal downsampling on the feature data input to the temporal downsampling unit in order to reduce the temporal resolution of the feature data.

[0030] Specifically, the spatiotemporal features are optimized by the spatiotemporal attention module in the encoder to obtain the encoded representation of the spatiotemporal block to be processed, including: calculating a spatiotemporal attention map based on the spatiotemporal features by the spatiotemporal attention module, wherein the value at each position in the spatiotemporal attention map represents the importance of the feature at each position; The spatiotemporal features are processed using the spatiotemporal attention map to obtain the encoded representation of the spatiotemporal block to be processed.

[0031] After extracting the spatiotemporal features of the spatiotemporal block to be processed using the feature extraction module in the 3D convolutional neural network, the spatiotemporal features can be optimized using the spatiotemporal attention module in the 3D convolutional neural network to obtain the encoded representation of the spatiotemporal block to be processed. The encoded representation can be used to decode and reconstruct the spatiotemporal block to be processed.

[0032] Optionally, a spatiotemporal attention mechanism can be introduced to adaptively weight features at different spatiotemporal locations, highlighting key spatiotemporal information and suppressing redundant information. Specifically, a spatiotemporal attention module first calculates a spatiotemporal attention map based on the spatiotemporal features; then, the spatiotemporal features are processed using this attention map to obtain optimized spatiotemporal features. In this process, the spatiotemporal attention mechanism learns a set of weight matrices to adaptively weight features at different spatiotemporal locations, emphasizing important spatiotemporal regions and suppressing redundant or irrelevant information. Specifically, the mechanism calculates a spatiotemporal attention map based on the content of the feature map, where the value at each position represents the importance of the feature at that position; then, the original feature map is multiplied element-wise with the spatiotemporal attention map to obtain a weighted feature map. In this way, the spatiotemporal attention mechanism can automatically focus on key frames and key regions in a video sequence, improving the representational power and robustness of multi-frame joint coding. Spatiotemporal attention mechanisms can be introduced at different stages of the encoding process, such as after spatiotemporal convolution or temporal downsampling, to achieve adaptive weighting of spatiotemporal features at different levels. This mechanism not only enhances the network's sensitivity to important features, but also greatly improves the accuracy and resolution of feature representation.

[0033] Specifically, the encoder is constructed using a 3D convolutional neural network. Based on the context information of the preceding and following frames of each frame in the spatiotemporal block to be processed, the encoder reconstructs each frame using the temporal context modeling mechanism in the decoder to obtain the decoding result of the spatiotemporal block to be processed.

[0034] In one embodiment, during the convolution of the spatiotemporal block to be processed using a spatiotemporal convolutional unit, local spatiotemporal features at different locations within the spatiotemporal block can be continuously extracted through a spatiotemporal sliding window operation. Specifically, the spatiotemporal convolutional kernel slides across the spatiotemporal block with a certain stride, covering a local spatiotemporal neighborhood with each slide. Within each local spatiotemporal neighborhood, the convolutional kernel performs element-wise multiplication and summation with the pixel values ​​in the neighborhood to generate a new feature value. By sliding along the temporal and spatial dimensions, the convolutional kernel can progressively extract local spatiotemporal features at different locations and time steps. The stride of the sliding window controls the extraction density and overlap of spatiotemporal features; a smaller stride generates denser, finer-grained feature maps, while a larger stride improves the efficiency of feature extraction and reduces redundancy. Simultaneously, to handle the edge regions of the spatiotemporal block, an appropriate padding strategy (such as zero-padding) can be employed to maintain the spatial size of the feature map.

[0035] Through spatiotemporal sliding window operations, convolutional kernels extract a series of local features within different local spatiotemporal neighborhoods. To obtain higher-level spatiotemporal feature representations, these local features can be aggregated. Optional aggregation methods include element-wise addition, max pooling, and average pooling. Element-wise addition sums the feature values ​​extracted from different locations element-wise to generate a summarized feature map; max pooling selects the maximum value of feature values ​​at different locations, highlighting significant spatiotemporal patterns; average pooling averages the feature values ​​at different locations to obtain a smooth feature representation. Through spatiotemporal feature aggregation, spatiotemporal convolutional units can synthesize local features and extract more robust and abstract spatiotemporal representations. Simultaneously, aggregation operations can reduce the spatial resolution of the feature map, progressively compressing the feature size, enabling the network to capture spatiotemporal patterns over larger scales and longer time spans.

[0036] Spatiotemporal convolutional units (SCUs) employ a 3D convolutional kernel that slides across a spatiotemporal block, weighting and summing pixel values ​​within a local spatiotemporal neighborhood to generate new feature values. In this way, SCUs can simultaneously capture temporal dependencies between multiple frames and spatial structures within a single frame. The temporal dimension of the 3D convolutional kernel (e.g., the first dimension of a 3×3×3 kernel) determines the number of frames processed in each convolution operation, embodying the core idea of ​​multi-frame joint encoding. For example, a 3×3×3 kernel covers three consecutive frames temporally and a 3×3 region spatially. By performing convolution operations simultaneously in both temporal and spatial dimensions, 3D convolutional neural networks can capture dynamic changes across frames and detailed intra-frame structures in a single convolution operation. Multiple such convolutional kernels can work in parallel, extracting features from different perspectives, thereby improving the network's ability to perceive complex spatiotemporal patterns.

[0037] Unlike 2D convolutional kernels, 3D convolutional kernels are three-dimensional tensors, typically represented by the shape (T, H, W), where T represents the length of the temporal dimension, and H and W represent the height and width of the spatial dimensions, respectively. The parameters of the 3D convolutional kernel are learned by the network during training to capture key spatiotemporal patterns in video sequences. The temporal dimension T of the convolutional kernel determines the number of frames processed in each convolution operation; a larger T means the kernel can cover a longer time span and capture longer-distance temporal dependencies. Simultaneously, the spatial dimension (H, W) controls the size of the receptive field within a single frame; a larger spatial dimension allows the kernel to extract spatial features at a larger scale. By adjusting the shape and size of the spatiotemporal convolutional kernel, the feature extraction capabilities of both temporal and spatial dimensions can be flexibly balanced, enabling joint encoding of information from multiple frames. In the specific compression process, spatiotemporal convolution operations also involve setting other parameters, including the kernel's stride and padding method. The stride determines the distance the convolutional kernel slides; a larger stride reduces the resolution of the feature map, thereby reducing computational and storage requirements. The padding method (such as zero padding) determines how edge pixels are processed; appropriate padding can preserve the spatial size of the feature map. In implementation, the reasonable selection of these parameters is crucial for improving the network's feature extraction capability and computational efficiency. Preferably, the 3D convolutional neural network is pre-trained, thus determining the aforementioned parameters of the relatively optimal spatiotemporal convolutional units through training.

[0038] On the other hand, a dynamic redundancy compression and multimodal optimization system based on AI intelligent technology is provided, applied to any of the AI ​​intelligent technology-based dynamic redundancy compression and multimodal optimization methods described above, the system comprising: Data Acquisition and Feature Extraction Module: This module is used to acquire video data and extract its features. It can acquire feature sets for each modality in the video and spatiotemporal feature sets between consecutive video frames. The feature sets for each modality are determined according to the characteristics of different modalities. For image modalities, the feature sets include the pixel area ratio of the main object and the number of mouse clicks on the image. For audio modalities, the feature sets include the effective speech duration ratio and the number of audio segment tags. For text modalities, the feature sets include the ratio of the number of keywords to the total number of words and the number of text copies. Feature Priority Determination Module: Based on the feature sets of each modality, the feature priority of each modality is determined according to the modality feature importance evaluation rules; different feature importance analysis processes are used for different modalities. Encoding Optimization Module: The encoder in this module optimizes the spatiotemporal feature set. The encoder includes a feature extraction module, which consists of several spatiotemporal convolutional units and several temporal downsampling units, each placed after a spatiotemporal convolutional unit. The spatiotemporal convolutional units perform convolution processing on the input data, and the temporal downsampling units perform temporal downsampling on the input feature data to reduce the temporal resolution of the feature data. Simultaneously, the encoder also calculates a spatiotemporal attention map based on the spatiotemporal features using a spatiotemporal attention module. This attention map is then used to process the spatiotemporal features to obtain the encoded representation of the spatiotemporal block to be processed. The encoder can be constructed using a 3D convolutional neural network. Decoding and Reconstruction Module: The decoder uses a temporal context modeling mechanism to reconstruct each frame based on the context information of the preceding and following frames in the spatiotemporal block to obtain the decoding result of the spatiotemporal block to be processed, thereby reconstructing the continuous frames of the video.

[0039] Specifically, the system also includes: Spatiotemporal attention processing module: Calculates a spatiotemporal attention map based on spatiotemporal features. The value of each position in the spatiotemporal attention map represents the importance of the feature at each position. The spatiotemporal features are processed using this spatiotemporal attention map to obtain the encoded representation of each position. Video frame compression module: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame, and finally the compressed video frame is obtained.

[0040] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0041] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A dynamic redundancy compression and multimodal optimization method based on AI intelligent technology, characterized in that, Includes the following steps: S1: Acquire video data and extract its features, obtaining the feature sets of each modality in the video and the spatiotemporal feature sets between consecutive frames of the video; S2: Based on the feature sets of each mode, determine the feature priority of each mode according to the modality feature importance evaluation rules; S3: Based on the spatiotemporal feature set, the encoder is used to optimize the set, and the video frames are reconstructed by decoding. S4: Calculate a spatiotemporal attention map based on the spatiotemporal features, where the value at each position in the spatiotemporal attention map represents the importance of the feature at each position; process the spatiotemporal features using the spatiotemporal attention map to obtain the encoded representation of each position; S5: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame to obtain the compressed video frame.

2. The dynamic redundancy compression and multimodal optimization method based on AI intelligent technology according to claim 1, characterized in that, The modalities include image modality, audio modality, and text modality. The feature sets of each modality are as follows: the image modality includes the pixel area ratio of the main object and the number of times the image is clicked by the mouse; the audio modality includes the effective speech duration ratio and the number of times audio segments are marked; and the text modality includes the ratio of the number of keywords to the total number of words and the number of times the text is copied.

3. The dynamic redundancy compression and multimodal optimization method based on AI intelligent technology according to claim 1, characterized in that, The specific analysis process of the feature importance of the image modality is as follows: dynamically calculate the difference between the pixel area ratio of the main object and the number of mouse clicks on the image and preset values ​​respectively. Normalization of each dynamic difference yields the content contribution and user attention of the image modality; The feature importance of the image modality is generated by summing the weighted calculations of content contribution and user attention with their respective percentage weights.

4. The dynamic redundancy compression and multimodal optimization method based on AI intelligent technology according to claim 1, characterized in that, The encoder includes a feature extraction module, which extracts time features from consecutive video frames. The feature extraction module includes several spatiotemporal convolution units and several temporal downsampling units. Each temporal downsampling unit is set after a spatiotemporal convolution unit. The spatiotemporal convolution unit is used to perform convolution processing on the data input to the spatiotemporal convolution unit. The temporal downsampling unit is used to perform temporal downsampling on the feature data input to the temporal downsampling unit in order to reduce the temporal resolution of the feature data.

5. The dynamic redundancy compression and multimodal optimization method based on AI intelligent technology according to claim 1, characterized in that, The spatiotemporal features are optimized by the spatiotemporal attention module in the encoder to obtain the encoded representation of the spatiotemporal block to be processed, including: calculating a spatiotemporal attention map based on the spatiotemporal features by the spatiotemporal attention module, wherein the value of each position in the spatiotemporal attention map represents the importance of the feature at each position; The spatiotemporal features are processed using the spatiotemporal attention map to obtain the encoded representation of the spatiotemporal block to be processed.

6. The dynamic redundancy compression and multimodal optimization method based on AI intelligent technology according to claim 1, characterized in that, The encoder is constructed using a 3D convolutional neural network. Based on the context information of the preceding and following frames of each frame in the spatiotemporal block to be processed, the encoder reconstructs each frame using the temporal context modeling mechanism in the decoder to obtain the decoding result of the spatiotemporal block to be processed.

7. A dynamic redundancy compression and multimodal optimization system based on AI intelligent technology, characterized in that, The system, applied to the dynamic redundancy compression and multimodal optimization method based on AI intelligent technology as described in any one of claims 1-6, comprises: Data Acquisition and Feature Extraction Module: This module is used to acquire video data and extract its features. It can acquire feature sets for each modality in the video and spatiotemporal feature sets between consecutive video frames. The feature sets for each modality are determined according to the characteristics of different modalities. For image modalities, the feature sets include the pixel area ratio of the main object and the number of mouse clicks on the image. For audio modalities, the feature sets include the effective speech duration ratio and the number of audio segment tags. For text modalities, the feature sets include the ratio of the number of keywords to the total number of words and the number of text copies. Feature Priority Determination Module: Based on the feature sets of each modality, the feature priority of each modality is determined according to the modality feature importance evaluation rules; different feature importance analysis processes are used for different modalities. Encoding Optimization Module: The encoder in this module is used to optimize the spatiotemporal feature set. The encoder includes a feature extraction module, which consists of several spatiotemporal convolutional units and several temporal downsampling units, with each temporal downsampling unit following one spatiotemporal convolutional unit. The spatiotemporal convolutional units perform convolution processing on the input data, and the temporal downsampling units perform temporal downsampling on the input feature data to reduce the temporal resolution of the feature data. Simultaneously, the encoder also uses a spatiotemporal attention module to calculate a spatiotemporal attention map based on the spatiotemporal features, and uses this attention map to process the spatiotemporal features to obtain the encoded representation of the spatiotemporal block to be processed. The encoder can be constructed using a 3D convolutional neural network. Decoding and Reconstruction Module: The decoder uses a temporal context modeling mechanism to reconstruct each frame based on the context information of the preceding and following frames in the spatiotemporal block to obtain the decoding result of the spatiotemporal block to be processed, thereby reconstructing the continuous frames of the video.

8. The dynamic redundancy compression and multimodal optimization system based on AI intelligent technology according to claim 7, characterized in that, The system also includes: Spatiotemporal attention processing module: Calculates a spatiotemporal attention map based on spatiotemporal features. The value of each position in the spatiotemporal attention map represents the importance of the feature at each position. The spatiotemporal features are processed using this spatiotemporal attention map to obtain the encoded representation of each position. Video frame compression module: Based on the feature priority of each modality in the video frame, a layered compression strategy is used to compress the video frame, and finally the compressed video frame is obtained.

Citation Information

Patent Citations

  • Video encoding and decoding method and system

    CN113099228A

  • Multi-mode video data compression and transmission method

    CN120091137A

  • Video compression method, video decompression method and related devices

    CN120128719A

  • Multi-modal large-model adaptive video frame compression method and system

    CN120751130A

  • System and method for homomorphic multi-modal data processing using variational autoencoders

    US20250192981A1