Implicit Video Representation Compression for Lower Bitrate Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video compression technologies, whether traditional hybrid coding or deep learning based, fail to achieve ideal video compression effect and bitrate efficiency due to limitations in block or frame processing methods.
Innovation Solution
A method involving dividing the original video sequence into image blocks, encoding these blocks using a neural network model to obtain implicit representation parameters, adjusting these parameters based on a loss function until preset requirements are met, and reconstructing the video sequence to optimize compression quality and bitrate.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If block to block or frame to frame processing methods are used to remove time redundancy, then the processing can be simplified and implemented, but the video compression effect and bitrate performance are limited
Solution Approach 1:
The video sequence is divided into multiple image blocks, which are then processed through the encoding network model to obtain implicit representation video parameters. This segmentation allows the system to process video data in manageable units while capturing temporal redundancy across the entire sequence through the neural network's implicit representation capability.
Solution Approach 2:
The patent transforms video data into implicit representation video parameters through the encoding network model, changing the representation parameters from traditional pixel-based data to compressed latent space parameters. This parameter transformation enables more efficient compression by capturing temporal redundancy in a lower-dimensional space.
2Ease of manufacture
If traditional hybrid coding technologies such as H.264/AVC, H.265/HEVC and H.266/VVC are used, then the coding standards are well-established and easy to implement, but the video compression effect and bitrate are not ideal
Solution Approach 1:
The patent replaces traditional mechanical coding systems (H.264, H.265, H.266) with a neural network-based implicit representation system. The encoding network model and reconstruction network model work together to transform video data into compressed implicit parameters, substituting conventional algorithmic approaches with deep learning-based methods that achieve superior compression performance.
Solution Approach 2:
The system changes the fundamental parameters of video representation from traditional coding formats to implicit representation parameters generated by the encoding network. This parameter transformation enables the system to achieve better compression ratios and bitrate efficiency compared to conventional coding standards.
3Productivity
If deep learning based video coding methods are used, then the compression performance can be improved, but the complexity of the system increases and implementation becomes more difficult
Solution Approach 1:
The deep learning system is segmented into distinct functional components: the encoding network model that processes image blocks to generate implicit parameters, the reconstruction network model that decodes these parameters back to video sequences, and the loss function optimization module. This segmentation makes the complex system more manageable and easier to implement while maintaining high compression performance.
4Productivity
If the implicit representation video parameters are adjusted using loss function optimization, then the compression efficiency can be improved, but the processing time and computational cost increase
Solution Approach 1:
The system uses a loss function that compares the reconstructed video sequence with the original video sequence to provide feedback on compression quality. This feedback mechanism guides the optimization of implicit representation parameters, allowing the system to achieve high compression efficiency by iteratively adjusting parameters based on reconstruction accuracy while controlling processing time through effective optimization.
Data Source
AI summary
A video processing method includes: dividing an original video sequence into a plurality of image blocks; inputting the plurality of image blocks into an encoding network model to obtain implicit representation video parameters corresponding to the original video sequence; inputting the implicit representation video parameters into a reconstruction network model to obtain a reconstructed video sequence corresponding to the original video sequence; calculating a loss function value based on the original video sequence and the reconstructed video sequence; based on the loss function value, adjusting the implicit representation video parameters until they meet preset requirements, and obtaining target implicit representation video parameters. The method compresses the video sequence, and adjusts the parameters of the compressed implicit representation video by reconstructing the video sequence to fully consider the redundant information of the original video sequence and reduce the bitrate of the implicit representation video parameters.


