AI Video Encoding Using Deep Neural Networks for Bitrate Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current image encoding and decoding methods face challenges in efficiently processing high-resolution images, particularly in reducing bitrates while maintaining image quality, especially with the increasing demand for high-quality image reproduction and storage.
Innovation Solution
The use of deep neural networks (DNNs) for both down-scaling high-resolution images to low resolution and up-scaling them back, with AI-based encoding and decoding processes that include DNN information for target up-scaling, allowing for efficient bitrate reduction and quality preservation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high-resolution images are encoded using traditional codecs, then image quality is maintained, but bitrate increases significantly
Solution Approach 1:
The patent replaces traditional mechanical/image-processing codec systems with an AI-based neural network system. The encoder uses a neural network to learn efficient representations of high-resolution images, while the decoder uses another neural network to reconstruct images from compressed representations. This substitution of AI/ML systems for traditional signal processing methods enables superior compression efficiency while maintaining or improving image quality, directly resolving the contradiction between bitrate reduction and quality preservation.
Solution Approach 2:
The patent transforms the encoding problem from traditional pixel-manipulation parameters to learned feature representations through neural networks. By changing the fundamental parameters from spatial domain pixel values to transformed domain latent representations, the system achieves more efficient compression. The neural networks learn optimal parameter transformations that preserve perceptually important information while discarding redundant data, thereby reducing bitrate without sacrificing image quality.
2Productivity
If traditional encoding methods are used, then processing is simple, but efficiency in handling high-resolution images deteriorates
Solution Approach 1:
The patent replaces traditional mechanical/image-processing codec systems with an AI-based neural network system. The encoder uses a neural network to learn efficient representations of high-resolution images, while the decoder uses another neural network to reconstruct images from compressed representations. This substitution of AI/ML systems for traditional signal processing methods enables superior compression efficiency while maintaining or improving image quality, directly resolving the contradiction between bitrate reduction and quality preservation.
Solution Approach 2:
The patent performs preliminary learning and feature extraction during the encoding phase, where the neural network pre-processes the high-resolution image to extract essential features and patterns. This preliminary action enables the decoder to efficiently reconstruct the image with fewer computational resources during playback, improving overall encoding efficiency while managing processing complexity through distributed computation between encode and decode stages.
Data Source
AI summary
Provided is a computer-recordable recording medium having stored thereon a video file including artificial intelligence (AI) encoding data, wherein the AI encoding data includes: image data including encoding information of a low resolution image generated by AI down-scaling a high resolution image; and AI data about AI up-scaling of the low resolution image reconstructed according to the image data, wherein the AI data includes: AI target data indicating whether AI up-scaling is to be applied to at least one frame; and AI supplementary data about up-scaling deep neural network (DNN) information used for AI up-scaling of the at least one frame from among a plurality of pieces of pre-set default DNN configuration information, when AI up-scaling is applied to the at least one frame.


