Efficient video compression method based on multivariate space-time entropy network
By combining adaptive channel pruning and multivariate spatiotemporal entropy models with the Swin Transformer module, the video compression model is dynamically optimized, solving the problems of insufficient modeling of complex spatiotemporal characteristics and multi-scene adaptation capabilities of existing technologies, and achieving efficient video compression and reconstruction quality improvement.
Patent Information
- Application Number
- CN202510794068.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-05
AI Technical Summary
Existing video compression technology has shortcomings in modeling complex spatiotemporal characteristics, improving compression efficiency and multi-scene adaptation capabilities. It is difficult to fully exploit the spatiotemporal redundancy characteristics of video data, resulting in reduced compression efficiency and poor visual effects.
Adaptive channel pruning technology is combined with a multivariate spatiotemporal information entropy model and the Swin Transformer module. The bit rate is adjusted through adaptive mask generation, a multi-scale window mechanism, and random sampling, and the model structure is dynamically optimized to adapt to different video content and scene requirements.
It significantly improves video compression performance, reduces spatiotemporal redundancy, improves reconstruction quality, supports flexible bit rate adjustment, adapts to diverse application scenarios, and demonstrates excellent performance on different video datasets.
Smart Images

Figure CN120602668A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an efficient video compression method based on a multivariate spatiotemporal entropy network, belonging to the technical field of video processing and compression. Background Art
[0002] With the rapid development of internet technology, video data has become one of the largest components of network traffic. In recent years, the rapid adoption of applications such as short video platforms, online video conferencing, distance education, and live streaming has further fueled explosive growth in demand for video transmission and storage. However, limited network bandwidth and storage device costs have made achieving efficient compression while maintaining video quality a core challenge in video technology.
[0003] After years of development, traditional video compression methods (such as H.264 / AVC, H.265 / HEVC, and H.266 / VVC) have achieved remarkable results in block-matching motion estimation and entropy coding. These techniques rely on manually designed rules and algorithms to minimize video data redundancy through methods such as block-based motion compensation and residual compression. However, the compression performance of these methods is gradually reaching saturation. With the continuous increase in video resolution and the increasing complexity of dynamic scenes, traditional methods are facing numerous shortcomings in meeting the requirements of new video applications. Their compression performance is approaching theoretical bottlenecks, making further improvement difficult, and they fail to fully utilize the complex characteristics of video data. Specifically, traditional methods primarily rely on block-matching motion estimation for inter-frame prediction and compression. However, this process typically uses a fixed strategy and cannot dynamically adapt to changes in video content, resulting in performance limitations in scenes with high dynamics and complex backgrounds. Furthermore, these methods have limited ability to model the spatiotemporal characteristics of video, only able to handle basic temporal and spatial redundancy and failing to fully exploit the potential multidimensional data characteristics. This limitation not only restricts compression efficiency but also causes some details to be lost in video reconstruction, affecting the final visual quality.
[0004] In recent years, deep learning-based video compression methods have garnered widespread attention. By incorporating neural network technology, these methods can automatically learn the underlying characteristics of video data, gradually replacing the manually designed modules used in traditional compression methods. Deep learning methods demonstrate significant advantages in terms of reconstruction quality and compression efficiency. For example, end-to-end video compression models can effectively reduce spatiotemporal redundancy in video data by jointly optimizing modules such as motion estimation, entropy coding, and inter-frame prediction. However, these deep learning-based technologies are not without their challenges. First, existing deep learning methods typically employ a fixed number of latent representation channels and fail to adaptively adjust the model structure to the dynamic changes in video content. This fixed design can easily lead to information redundancy and reduced compression efficiency. Second, these methods are inadequate in modeling the spatial characteristics of complex scenes and are unable to fully eliminate fine-grained spatial redundancy in videos. Furthermore, in real-world applications requiring multiple scenes and multiple bitrates, existing deep learning models lack flexibility and are unable to effectively balance compression efficiency and reconstruction quality across different scenarios. This lack of adaptability makes current technologies difficult to meet diverse practical needs.
[0005] In summary, despite significant progress in video compression technology, driven by both traditional and deep learning methods, significant deficiencies remain in modeling complex spatiotemporal features, improving compression efficiency, and enhancing multi-scenario adaptability. Key challenges in the field of video compression include further exploring the spatiotemporal redundancy of video data, dynamically optimizing model structures, improving compression performance, and adapting to diverse practical application scenarios. Summary of the Invention
[0006] Purpose of the Invention: To address the shortcomings of traditional video compression methods and existing deep learning approaches in terms of spatiotemporal feature modeling, compression efficiency, and multi-scenario adaptability, this paper proposes a highly efficient video compression method based on a multivariate spatiotemporal entropy network. Through adaptive channel pruning, multivariate spatiotemporal information entropy modeling, and a context enhancement module, this method significantly improves video compression performance and reconstruction quality, while enabling flexible bitrate adjustment to accommodate diverse application needs.
[0007] Technical solution: A high-efficiency video compression method based on a multivariate spatiotemporal entropy network, specifically including the following steps: Step 1: Input the video frame to be compressed. Using adaptive channel pruning technology, an adaptive mask is generated using multivariate spatiotemporal priors to dynamically remove redundant channel information from the video frame. This significantly reduces data redundancy while maintaining high-quality image feature reconstruction. The multivariate spatiotemporal priors include latent priors, temporal context priors, and super priors.
[0008] Step 2: Based on the multivariate spatiotemporal information entropy model, the optical flow information and the feature priors of the current frame and the previous frame are integrated to optimize the initial encoding content of the video and significantly reduce spatiotemporal redundancy.
[0009] Step 3: A spatial context enhancement network including the Swin Transformer module is introduced to capture global and local information through a multi-scale window mechanism, effectively enhancing the spatial information prediction capability and further reducing spatial redundancy.
[0010] In step 4, the optimized spatiotemporal features are integrated, the range of the hyperparameter λ is expanded based on a unified training strategy, and random sampling is used to achieve smooth bitrate adjustment to generate efficient and compressed video feature encoding results to adapt to different bitrate requirements and complex video scenarios.
[0011] The step 1 generates an adaptive mask by multivariate spatiotemporal prior to process redundant information of the channel, including the following steps: Step 1-1: Use the optical flow estimation network to obtain the current frame x t Reconstructed frame with previous frame Sports information , then combine the motion information by using the DCVC_TCM spatiotemporal context module and the feature information F of the previous frame t Obtaining spatiotemporal context information , and use the super prior information extraction model to obtain super prior information z t . The spatiotemporal context information , super prior information z t and the latent representation of the previous frame Perform feature fusion to obtain multivariate spatiotemporal prior M prior .
[0012] Step 1-2: Multivariate spatiotemporal prior M prior Input channel mask generation module, which generates channel adaptive soft mask through convolution layer and activation layer The purpose of this soft mask is to normalize the channel importance value to a limited range so that subsequent quantization and pruning operations can be processed more efficiently.
[0013] Steps 1-3: Soft Mask Quantize and generate the final adaptive mask Mask c , soft mask the continuous values Converted to discrete integer values. The purpose of quantization is to further reduce data complexity by simplifying the representation while ensuring efficient application of masks in subsequent pruning operations. Discrete value masks can effectively reduce redundant computations and improve pruning and encoding efficiency.
[0014] Step 1-4: Using the generated adaptive mask c Dynamically prune the latent feature channels. During the pruning process, channels with contributions lower than the set value are removed to significantly reduce data redundancy. The specific operation involves multiplying the channel feature tensor pixel by pixel by the adaptive mask value to suppress redundant information and enhance the important channel information that contributes to video reconstruction, and finally obtain the optimized latent representation. This step reduces computational and storage costs while ensuring a balance between compression efficiency and image quality.
[0015] Step 2: Obtain the priority encoding content as the spatial context of the content to be encoded through a video efficient compression method based on a multivariate spatiotemporal entropy model.
[0016] Step 2-1: First, due to the potential representation Adaptive Mask c The impact of obtaining the priority encoding data mask Mask y When multivariate spatiotemporal prior M is required prior Based on the intermediate state M of the fusion adaptive mask i,j,k Information. The mask is converted into a discrete binary value. This binary mask acts as a filter, optimizing the data retained during compression.
[0017] Step 2-2: Use Mask y The potential representation after removing the channel of the current frame Select the first encoding content y1.
[0018] Step 2-3: In order to obtain the probability distribution of the first encoded content y1, it is necessary to transform the multivariate spatiotemporal prior M prior and the intermediate state M of the channel adaptive mask i,j,k Input parameter probability estimation network E1 and estimate its probability distribution parameters and , to achieve efficient entropy coding and compression.
[0019] In step 3, spatial context information is predicted using the Swin Transformer module, which specifically includes the following steps: Step 3-1: By using the priority encoding data mask Mask y Obtain the content y2 to be encoded.
[0020] Step 3-2: Input the quantized y1 into the spatial context enhancement network introduced into the Swin Transformer module. Through the multi-scale window mechanism, the Swin Transformer module further captures global and local information and extracts the spatial context prior S prior.
[0021] Step 3-3: Incorporate spatial context prior S prior , multivariate spatiotemporal prior M prior and the intermediate state M of the adaptive mask i,j,k Obtain the probability distribution of the content to be encoded y2. Context priors can more accurately predict the remaining information in the latent features, thereby optimizing data distribution, further reducing spatial redundancy, and improving compression efficiency.
[0022] Step 4: Achieve single-model multi-bitrate functionality by combining the multi-granularity quantization mechanism with the mapping of the global quantization step size.
[0023] Step 4-1: Integrate the optimized spatiotemporal features (i.e., spatial context prior S prior and the multivariate spatiotemporal prior M prior ), using a neural network-based video compression training strategy and giving the extended hyperparameters By global quantization step size global sq and hyperparameters , perform function mapping, and then in the hyperparameters Random sampling within the range to obtain the corresponding global quantization step size global sq , to adapt to various scenarios and different compression requirements, ensuring that the model has higher flexibility and adaptability.
[0024] Step 4-2: In the process of training the network, use the loss function Achieve a balance between compression rate and reconstruction quality.
[0025] Step 4-3: To reduce the impact of error accumulation and propagation between frames in the video sequence on the training results, a cascade training loss mechanism is introduced in the later stages of training. By assigning a weight w to each frame t , to achieve a hierarchical structure of inter-frame quality, which is used to enhance the quality contribution of specific frames and ensure a reasonable distribution of compression quality between key frames and non-key frames.
[0026] Step 4-4: During the training process, in order to achieve smooth bit rate adjustment, a random sampling strategy is adopted and the The value range is (64, 2048), and Apply the corresponding quantization step global sq , where global sq The value range of is (0.4, 2.4). By uniformly sampling within this range, the model can be trained evenly at different bitrates, effectively improving its compression adaptability in low and high bitrate scenarios.
[0027] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the efficient video compression method of the multivariate spatiotemporal entropy network as described above are implemented.
[0028] A computer-readable storage medium stores a computer program for executing the above-mentioned high-efficiency video compression method using a multivariate spatiotemporal entropy network.
[0029] In summary, the present invention proposes an efficient video compression method based on a multivariate spatiotemporal entropy network, which improves on the limitations of traditional video compression methods and existing neural network compression methods. By introducing adaptive channel pruning technology and a multivariate spatiotemporal entropy network, this method effectively reduces spatiotemporal and spatial redundancy, significantly improving rate-distortion performance. Experimental results show that compared with existing state-of-the-art video compression methods, this method exhibits significant advantages on multiple datasets, achieving bitrate savings of up to 14.86% compared to VTM, and is particularly outstanding in low-resolution and low-motion scenes (such as HEVC Classes C, D, and E). It also has significant advantages in PSNR and MS-SSIM indicators. This method also supports smooth bitrate adjustment, demonstrating its strong adaptability and robustness in various video compression tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of the overall execution flow in an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall architecture of the multivariate spatiotemporal entropy network; Figure 3 It is a schematic diagram of the multivariate spatiotemporal entropy model in video compression; Figure 4 It is the detailed structure of the channel removal module and the Swin Transformer module; Figure 5 This is an example of the quality comparison of the video image before and after channel removal; Figure 6 This is the rate-distortion performance curve of different methods based on PSNR; Figure 7 These are the rate-distortion performance curves of different methods based on MS-SSIM. DETAILED DESCRIPTION
[0031] The present invention is further illustrated below with reference to specific examples. It should be understood that these examples are only used to illustrate the present invention and are not used to limit the scope of the present invention. After reading the present invention, modifications of various equivalent forms of the present invention made by those skilled in the art all fall within the scope defined by the claims attached to this application.
[0032] The detailed process of this case is as follows Figure 2 The paper proposes an efficient video compression method based on a multivariate spatiotemporal entropy network, including: Stage 1: Dynamically optimize the potential representation of the current frame through the channel pruning module.
[0033] First, input the video frame to be compressed. Figure 4 As shown in the module on the left, adaptive channel pruning technology is used to generate adaptive masks using multivariate spatiotemporal priors to dynamically remove redundant channel information in video frames while maintaining high-quality image feature reconstruction. Multivariate spatiotemporal priors include potential priors, temporal context priors, and super priors. Through the feature fusion network F , the spatiotemporal context information , super prior information z t Latent representation of the previous frame Fusion is performed to generate multivariate spatiotemporal prior M prior ,Right now: .
[0034] Next, the input channel mask generation module is passed through the convolution layer and the activation layer to generate the intermediate state of the channel adaptive mask. This step is achieved by calculating the average activation value A of each channel. k , and quantify the contribution of different channels to video reconstruction. The formula is as follows: , Among them, H and W represent the height and width of the video frame respectively, and M i,j,k Represents the activation value of channel k at position (i, j).
[0035] Then, through nonlinear transformation, the Sigmoid activation function is used to generate the soft mask, and the formula is as follows: .
[0036] The soft mask normalizes the channel importance value to a limited range (0 to 1), allowing for more efficient processing in subsequent quantization and pruning operations. The soft mask is quantized to generate the final adaptive mask, which is expressed as: , Among them, Q(x) is a quantization function that converts continuous values into discrete binary values.
[0037] After the adaptive mask is generated, the latent feature channels are dynamically pruned using the mask. During the pruning process, channels with low importance are removed, while channels with high importance are retained, thereby significantly reducing data redundancy. The specific operation is to multiply the channel feature tensor pixel by pixel by the adaptive mask value to obtain the optimized latent representation. , the formula is as follows: , Among them, y t represents the potential features of the current frame, Represents element-by-element multiplication operation, Mask c is an adaptive mask.
[0038] Phase 2: Probabilistic modeling and optimization based on multivariate spatiotemporal information.
[0039] Next, Figure 3 As shown in Figure 1, the spatial representation y1 of the priority coding content is obtained through the multivariate spatiotemporal entropy model as the spatial feature input of the content to be coded. First, the mask data of the priority coding content is generated by combining the multivariate spatiotemporal prior information and the intermediate state of the channel adaptive mask. The mask data is represented as: , , Where σ represents the sigmoid activation function, Q is the quantization function, and the quantized data is converted into discrete values. This process selects the optimal encoding content by applying a two-step screening. In this process, the input multivariate spatiotemporal prior parameters M prior Estimate the probability distribution of the priority coding content through the network and calculate its mean and standard deviation , in order to achieve efficient encoding and compression. The formula is as follows: , Among them, E1 is the probability estimation network Stage 3: Spatial context enhancement with Swin Transformer.
[0040] In order to make full use of spatial context information, the Swin Transformer module is introduced. The structure of Swin Transformer is as follows Figure 4 The module on the right is shown. This module captures global and local contextual information through a multi-scale window mechanism, enhancing the ability to remove spatial redundancy of latent features. The content to be encoded is y2, and its formula is: .
[0041] The fluxed y1 is input into the spatial context enhancement network that introduces the existing Swin Transformer module. Through the multi-scale window mechanism, the module further captures global and local information and extracts the spatial context information S of the current frame prior , the formula is as follows: , , Among them, Q is the quantization function and ST represents the Swin Transformer module, which is used to retain important information in the latent features.
[0042] Through the sliding window multi-head self-attention (SW-MSA) module and the window multi-head self-attention (W-MSA) module, the global context information across windows is further captured, and the multivariate information context in the time dimension is combined to generate the final compressed latent features. In order to obtain the probability distribution of y2, we transform the spatial information S prior Integrate it into the probability estimation network E2 to be encoded, the formula is as follows: .
[0043] Phase 4: Multi-bitrate adaptation.
[0044] After the optimized potential features are integrated, the global quantization step size is global sq and hyperparameters , perform function mapping, and then in the hyperparameters Random sampling within the range to obtain the corresponding global quantization step size global sq , to adapt to various scenarios and different compression requirements, ensuring that the model has higher flexibility and adaptability. A unified loss function is used in the training process to balance the compression rate and reconstruction quality. The formula is as follows:
[0045] in, is the original frame x t and reconstructed frames The distortion between t ) and R(y t ) represent the bit rates of motion features and latent features respectively, is the Lagrange multiplier that controls the trade-off between compression rate and distortion. By randomly sampling and expanding the bit rate range to (64, 2048), Apply function mapping to obtain the corresponding quantization step size to ensure that the model is evenly trained at different bit rate points, thereby improving its adaptability in low bit rate and high bit rate scenarios. sq The value range is (0.4, 2.4), where the higher the value, the higher the compression rate.
[0046] Through the above four stages, the method proposed in the present invention effectively reduces spatiotemporal redundancy, significantly improves video compression performance, and demonstrates good reconstruction quality and adaptability in different scenarios.
[0047] from Figure 6 and Figure 7 To validate the effectiveness of our proposed video compression method, we conducted comparative experiments on three datasets: HEVC, UVG, and MCL-JCV. These datasets cover videos of varying resolutions, enabling a comprehensive evaluation of the algorithm's performance. We used PSNR and SSIM as evaluation metrics: PSNR reflects video quality, while SSIM assesses video structural similarity. The proposed method was compared with existing learning-based video compression schemes and Versatile Video Coding (VTM). The results demonstrate that our method outperforms existing methods across multiple test sets, particularly on low-resolution videos (e.g., HEVC Classes C and D), achieving bitrate savings of 24.16% and 21.48%, respectively. The performance is even more significant on videos with less motion (e.g., HEVC Class E). For high-resolution datasets (e.g., UVG and MCL-JCV), the proposed method's advantage is relatively modest due to the greater complexity and complexity of the compression task, owing to the greater amount of video detail. However, it still outperforms VTM in some high-resolution scenarios, demonstrating its strong scalability. Overall, the proposed method performs well in low-resolution and less-motion videos, can effectively reduce the bitrate and maintain high video quality, and has broad application potential.
[0048] Obviously, those skilled in the art should understand that the various steps of the video compression method based on the multivariate spatiotemporal entropy network proposed in the above-mentioned embodiment of the present invention can be implemented by a general-purpose computing device. These steps need to be run on at least two computing devices, and can also be distributed in a network composed of multiple computing devices. Optionally, these steps can also be implemented by a program code executable by a computing device, so that the program code is stored in a storage device for the computing device to call and execute. In some cases, the above steps can be performed in an order different from that described in this specification, or each step can be independently made into a separate integrated circuit module, or multiple steps can be integrated into a single hardware module for implementation. Therefore, the embodiments of the present invention are not limited to any specific hardware or software implementation form.
[0049] The above description clearly illustrates the basic principles, main features and technical advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the specific contents described in the above embodiments, which are only used to explain the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention can have various forms of modifications and improvements, and these changes and improvements also fall within the scope of protection claimed in the present invention. The scope of protection of the present invention is defined by the appended claims and their equivalents.
Claims
1. An efficient video compression method based on a multivariate spatiotemporal entropy network, characterized in that: The following steps are involved: Step 1: Input the video frame to be compressed, and use the adaptive channel pruning technology to generate an adaptive mask through multivariate spatiotemporal priors to dynamically remove redundant channel information in the video frame; Step 2: Based on the multivariate spatiotemporal information entropy model, the initial encoding content of the video is optimized by fusing the optical flow information and the feature priors of the current and previous frames. Step 3: Introduce a spatial context enhancement network containing the Swin Transformer module to capture global and local information through a multi-scale window mechanism; In step 4, the optimized spatiotemporal features are integrated, the range of the hyperparameter λ is expanded based on a unified training strategy, and random sampling is used to achieve smooth bitrate adjustment to generate efficient and compressed video feature encoding results to adapt to different bitrate requirements and complex video scenarios.
2. The efficient video compression method of the multivariate spatiotemporal entropy network according to claim 1, characterized in that: The multivariate spatiotemporal prior includes latent prior, temporal context prior and super prior information.
3. The efficient video compression method of the multivariate spatiotemporal entropy network according to claim 1, characterized in that: The step 1 generates an adaptive mask by multivariate spatiotemporal prior to process redundant information of the channel, including the following steps: Step 1-1: Use the optical flow estimation network to obtain the current frame x t Reconstructed frame with previous frame Sports information , then combine the motion information by using the DCVC_TCM spatiotemporal context module and the feature information F of the previous frame t Obtaining spatiotemporal context information , and use the super prior information extraction model to obtain super prior information z t ; Spatiotemporal context information , super prior information z t and the latent representation of the previous frame Perform feature fusion to obtain multivariate spatiotemporal prior M prior ; Step 1-2: Multivariate spatiotemporal prior M prior Input channel mask generation module, which generates channel adaptive soft mask through convolution layer and activation layer ; Steps 1-3: Soft Mask Quantize and generate the final adaptive mask Mask c , soft mask the continuous values Convert to discrete integer values; Step 1-4: Using the generated adaptive mask c Dynamically prune the potential feature channels and finally obtain the optimized potential representation .
4. The efficient video compression method of the multivariate spatiotemporal entropy network according to claim 1, characterized in that: Step 2: Obtain the priority encoding content as the spatial context of the content to be encoded through a video efficient compression method based on a multivariate spatiotemporal entropy model; Step 2-1: First, due to the potential representation Adaptive Mask c The impact of obtaining the priority encoding data mask Mask y When multivariate spatiotemporal prior M is required prior Based on the intermediate state M of the fusion adaptive mask i,j,k information; converting the mask into discrete binary values; Step 2-3: Use Mask y The potential representation after removing the channel of the current frame Select the first encoding content y1; Step 2-4: In order to obtain the probability distribution of the first encoded content y1, it is necessary to transform the multivariate spatiotemporal prior M prior and the intermediate state M of the channel adaptive mask i,j,k Input parameter probability estimation network E1 and estimate its probability distribution parameters and , to achieve efficient entropy coding and compression.
5. The efficient video compression method of the multivariate spatiotemporal entropy network according to claim 1, characterized in that: In step 3, spatial context information is predicted using the Swin Transformer module, which specifically includes the following steps: Step 3-1: By using the priority encoding data mask Mask y Get the content y2 to be encoded; Step 3-2: Input the quantized y1 into the spatial context enhancement network introduced into the Swin Transformer module; through the multi-scale window mechanism, the Swin Transformer module further captures global and local information and extracts the spatial context prior S prior ; Step 3-3: Incorporate spatial context prior S prior , multivariate spatiotemporal prior M prior and the intermediate state M of the adaptive mask i,j,k Get the probability distribution of the content y2 to be encoded.
6. The high-efficiency video compression method of the multivariate spatiotemporal entropy network according to claim 1, characterized in that: Step 4: Achieve single-model multi-bitrate functionality by combining the multi-granularity quantization mechanism with the global quantization step size mapping, including: Step 4-1: Integrate the optimized spatiotemporal features, use the neural network-based video compression training strategy, and give the expansion hyperparameters range; by global quantization step size global sq and hyperparameters , perform function mapping, and then in the hyperparameters Random sampling within the range to obtain the corresponding global quantization step size global sq , to adapt to various scenarios and different compression requirements; Step 4-2: In the process of training the network, use the loss function Achieve a balance between compression rate and reconstruction quality; Step 4-3: In order to reduce the impact of error accumulation and propagation between frames in the video sequence on the training results, a cascade training loss mechanism is introduced in the later stages of training; by assigning a weight w to each frame t , achieving a hierarchical structure of inter-frame quality; Step 4-4: During training, adopt a random sampling strategy and expand The value range is (64, 2048), and Apply the corresponding quantization step global sq .
7. A computer device, characterized in that: The computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the efficient video compression method of the multivariate spatiotemporal entropy network as described in any one of claims 1 to 6 are implemented.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program for executing the efficient video compression method of the multivariate spatiotemporal entropy network according to any one of claims 1 to 6.
Citation Information
Patent Citations
Flexible deep learning network model compression method based on channel gradient pruning
CN112396179A
Model compression method and device
CN119494381A
High-precision large-proportion structured network compression method based on minimum entropy
CN119862918A
Cited By
Double-path parallel motion data optimization method fusing space-time prior and attention
CN121148022A