Perceptual Video Quality Prediction via Genre-Specific Model Partitioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for predicting the perceived quality of decoded video content are inadequate, as they either rely on unreliable fidelity metrics or subjective user ratings that are not robust enough to handle diverse video content, making it impractical to accurately assess visual quality across various genres.
Innovation Solution
A computer-implemented method that generates partitions for metric values based on video genres, optimizes hyperparameters through cross-validation, and trains a model to compute a perceptual video quality metric, which mitigates overfitting and accurately predicts quality across a wide range of video content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full-reference quality metrics like PSNR are used to compare source video content to encoded video content, then signal fidelity is accurately reflected, but human perception of video quality is not reliably predicted
Solution Approach 1:
The patent segments video content into different genre categories (e.g., nature documentaries, action movies, dramas) and creates separate predictive models for each genre. This segmentation allows the system to capture genre-specific characteristics that affect perceived quality, resolving the contradiction by making measurements more relevant to human perception within each genre context while maintaining overall measurement precision across diverse content types.
Solution Approach 2:
The patent transforms traditional signal fidelity parameters into perceptual quality parameters by incorporating genre-specific weighting factors and adjustment parameters. The system changes the measurement parameters dynamically based on detected video genre, allowing the same encoded video to be evaluated with different parameters depending on its genre, thereby improving alignment with human perception while preserving measurement precision.
2Reliability
If perceptive quality metrics are generated based on subjective user ratings of simple cartoons, then the metric works for that specific content type, but it fails to accurately predict visual quality of complex action movies
Solution Approach 1:
The patent creates a universal quality assessment system that functions across multiple video genres by training separate models on genre-specific data and then selecting the appropriate model based on the input video's genre classification. This multi-functional approach allows the system to achieve high reliability for each specific content type while maintaining versatility across the entire spectrum of video content through model selection and ensemble techniques.
Solution Approach 2:
The patent adds a genre dimension to the quality assessment process by classifying videos into different genres and applying genre-specific predictive models. This dimensional expansion transforms a single-dimension approach (general quality metrics) into a multi-dimensional approach (genre-specific metrics), enabling the system to achieve both specificity for individual genres and versatility across all genres simultaneously.
3Productivity
If automated video quality assessment is implemented using traditional metrics, then encoding efficiency can be monitored, but the assessment does not accurately reflect actual viewing experience across different video genres
Solution Approach 1:
The patent implements a dynamic quality assessment system that adapts its evaluation criteria based on the detected genre of the encoded video content. Rather than using static traditional metrics, the system dynamically selects and applies genre-specific predictive models that have been trained on appropriate content types. This dynamic approach maintains encoding efficiency monitoring capabilities while significantly improving the reliability of viewing experience predictions across diverse video genres.
Data Source
AI summary
In various embodiments, a quality trainer trains a model that computes a value for a perceptual video quality metric for encoded video content. During a pre-training phase, the quality trainer partitions baseline values for metrics that describe baseline encoded video content into partitions based on genre. The quality trainer then performs cross-validation operations on the partitions to optimize hyperparameters associated with the model. Subsequently, during a training phase, the quality trainer performs training operations on the model that includes the optimized hyperparameters based on the baseline values for the metrics to generate a trained model. The trained model accurately tracks the video quality for the baseline encoded video content. Further, because the cross-validation operations minimize any potential overfitting, the trained model accurately and consistently predicts perceived video quality for non-baseline encoded video content across a wide range of genres.


