Machine-Learned Perceptual Metrics for Cross-Device Video Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods struggle to accurately evaluate the perceptual quality of video content across different devices and network conditions, affecting user experience.
Innovation Solution
A computer-implemented method using machine-learned models to generate features from video frames, which are then processed to determine a perceptual quality score, correlating to consumer experience, and used to improve video encoding and decoding processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If video content is processed for compact delivery to improve transmission efficiency, then network bandwidth consumption is reduced, but video quality deteriorates
Solution Approach 1:
The patent transforms the video quality assessment from traditional objective parameters (PSNR, SSIM) to perceptual quality parameters that reflect human visual system characteristics. This involves changing the assessment parameters to include perceptual features such as naturalness, realism, and aesthetic quality, allowing for better quality evaluation at lower bitrates.
Solution Approach 2:
The patent replaces traditional mathematical models of video quality assessment with machine learning-based perceptual models. These models use deep neural networks to learn perceptual features directly from video content and human feedback, substituting conventional signal-processing approaches with data-driven perceptual assessment.
2Ease of operation
If traditional objective quality metrics are used to evaluate video content, then measurement simplicity is maintained, but measurement precision deteriorates
Solution Approach 1:
The patent introduces machine learning models as intermediaries between the video content and quality assessment. These models act as perceptual mediators that translate video signals into meaningful quality scores based on learned perceptual patterns, bridging the gap between objective measurement and subjective perception.
Solution Approach 2:
The patent fundamentally changes the parameter space from traditional signal-based metrics to perceptual features extracted by machine learning models. This includes using features related to human visual perception, cognitive processing, and aesthetic judgment, thereby improving measurement precision while maintaining automated assessment capability.
3Manufacturing precision
If video is encoded with high quality parameters to maintain perceptual quality, then video quality is preserved, but data transmission volume increases
Solution Approach 1:
The patent employs feedback mechanisms where perceptual quality scores are used to iteratively optimize encoding parameters. The system assesses perceptual quality at different compression levels and uses this feedback to determine the optimal encoding settings that achieve target quality at minimal bitrate, enabling adaptive quality-preserving compression.
4Measurement precision
If machine learning models are trained to improve perceptual quality assessment, then measurement precision is enhanced, but device complexity increases
Solution Approach 1:
The patent segments the video quality assessment task into multiple specialized machine learning models, each handling specific aspects such as spatial quality, temporal quality, and perceptual naturalness. This modular approach improves measurement precision for different quality dimensions while managing overall system complexity through division of labor.
Data Source
AI summary
An example computer-implemented method for determining a perceptual quality of a subject video content item is provided. The example method can include inputting a subject frame set from the subject video content item into a first machine-learned model. The example method can also include generating, using the first machine-learned model, a feature based at least in part on the subject frame set. The example method can also include outputting, using a second machine-learned model, a score indicating the perceptual quality of the subject video content item based at least in part on the feature.


