Neural Network Image Encoder Mode Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image and video encoding technologies face inefficiencies in selecting optimal encoding modes to optimize rate-distortion properties, often requiring computationally expensive hand-crafted algorithms and leading to inflexible and non-standard-compliant solutions.
Innovation Solution
A computer-implemented method using an artificial neural network to process image data and select optimal encoding modes from a plurality of modes supported by an external encoder, trained with differentiable functions to emulate the encoding process and optimize rate and quality scores.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If hand-crafted algorithms are used to select encoding modes, then rate-distortion optimization can be achieved, but computational complexity and processing time increase significantly
Solution Approach 1:
The patent replaces hand-crafted mechanical algorithms with a neural network system that learns encoding mode selection automatically. The neural network is trained to predict optimal encoding modes directly from input data, substituting the traditional step-by-step hand-crafted algorithmic approach with a learned model that can be evaluated more efficiently during encoding.
Solution Approach 2:
The neural network is trained in advance on a large dataset of encoding scenarios to learn the optimal encoding mode selections. This preliminary training phase allows the network to capture complex patterns and relationships, so that during actual encoding, mode selection can be performed quickly by simply feeding the current input through the trained network without executing complex optimization algorithms.
2Productivity
If encoding modes are modified to improve efficiency, then rate-distortion properties are optimized, but compliance with standard bitstream formats is lost
Solution Approach 1:
The neural network learns to select from the existing set of standard-compliant encoding modes defined by the video coding standard. By adjusting the network's learned parameters and selection strategy rather than modifying the encoding modes themselves, the system achieves improved efficiency while maintaining compatibility with standard bitstream formats and existing decoders.
3Measurement precision
If bespoke encoding algorithms are implemented, then rate-distortion optimization is achieved, but flexibility and adaptability to different standards are reduced
Solution Approach 1:
The neural network is designed with a universal architecture that can be trained and applied to different video coding standards and scenarios. The same network framework can learn to select encoding modes for various standards by training on appropriate datasets, providing flexibility and adaptability without requiring bespoke algorithms for each standard.
Data Source
AI summary
A method of processing, prior to encoding using an external encoder, image data using an artificial neural network is provided. The external encoder is operable in a plurality of encoding modes. At the neural network, image data representing one or more images is received. The image data is processed using the neural network to generate output data indicative of an encoding mode selected from the plurality of encoding modes of the external encoder. The neural network trained to select using image data an encoding mode of the plurality of encoding modes of the external encoder using one or more differentiable functions configured to emulate an encoding process. The generated output data is outputted from the neural network to the external encoder to enable the external encoder to encode the image data using the selected encoding mode.


