Real-Time Video Quality Metrics for Human-Perception Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding methods struggle to efficiently generate visual quality metrics that accurately reflect human perception while maintaining real-time processing capabilities, leading to inefficiencies and increased latency due to high computational overhead.
Innovation Solution
Implementing a dedicated hardware-based solution within a graphics processing unit to compute visual impairment metrics on-the-fly during encoding, using a machine learning model to aggregate and analyze per-pixel data, thereby generating enhanced real-time visual quality metrics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If complex visual quality metrics are computed using software-based methods, then measurement precision of human perception is improved, but productivity and processing speed deteriorate due to high computational overhead
Solution Approach 1:
The patent replaces software-based computational mechanisms with hardware-based mechanisms (FPGA, ASIC, or GPU circuits) to compute visual quality metrics. This substitution of the computational system enables real-time metric generation during video encoding without the processing delays inherent in software implementations, thereby resolving the contradiction between measurement precision and productivity.
Solution Approach 2:
The patent performs visual quality metric computations in parallel with the video encoding process itself, rather than as a subsequent post-processing step. By computing metrics concurrently with encoding operations using dedicated hardware circuits, the system obtains quality measurements without delaying the encoding throughput, thus maintaining both high precision and high productivity.
2Measurement precision
If comprehensive per-pixel visual quality metrics are generated, then measurement precision is improved, but loss of time increases due to post-processing requirements
Solution Approach 1:
The patent computes per-pixel visual quality metrics concurrently with the video encoding process using hardware circuits, rather than performing comprehensive analysis after encoding completes. This preliminary computation approach eliminates post-processing latency while maintaining high measurement precision, as the metrics are generated during the encoding operation itself.
Solution Approach 2:
The patent maintains continuous computation of visual quality metrics throughout the encoding process without interruption or separate processing stages. The hardware circuits continuously generate per-pixel quality data in real-time alongside the encoding operations, ensuring no time loss occurs between encoding completion and metric generation.
3Measurement precision
If multiple visual quality metrics are computed and aggregated, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The patent merges multiple visual quality metric computations (PSNR, SSIM, MS-SSIM, VMAF) into a single integrated hardware processing pipeline. By combining these metric calculations and their aggregation into unified circuitry within the encoding device, the system achieves high measurement precision through comprehensive metric analysis while avoiding the complexity increase that would result from separate software-based processing systems.
Solution Approach 2:
The patent designs the hardware processing circuitry to perform multiple functions: computing various visual quality metrics, aggregating per-pixel data, generating quality scores, and providing feedback to the encoder. This multi-functional hardware architecture achieves comprehensive quality assessment precision without requiring separate dedicated systems for each function, thereby controlling overall device complexity.
Data Source
AI summary
This disclosure describes systems, methods, and devices related to generating visual quality metrics for encoded video frames. A method may include generating respective first visual quality metrics for pixels of an encoded video frame; generating respective second visual quality metrics for the pixels, the respective first visual quality metrics and the respective second visual quality metrics indicative of estimated human perceptions of the encoded video frame; generating a pixel block-based weight for the respective first visual quality metrics; generating a frame-based weight for the respective second visual quality metrics; and generating, based on the respective first visual quality metrics, the pixel block-based weight, the respective second visual quality metrics, and the frame-based weight, a human visual score indicative of a visual quality of the encoded video frame.


