SPQR Video Quality Metric Face Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video quality metrics are not suitable for real-time monitoring due to high computational and transmission overheads, and they fail to accurately reflect user opinion across various video decoders and network conditions.
Innovation Solution
A novel video quality metric, Simplified Perceptual Quality Region (SPQR), which focuses on detecting the location of a speaker's face in video frames and comparing these locations between sent and received frames to identify discrepancies, providing a lightweight and accurate method for detecting video quality degradation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If full-reference techniques are used to measure video quality, then measurement precision is improved, but computational overhead and transmission overhead increase significantly
Solution Approach 1:
The patent extracts only the essential feature (face location) from the complex video stream, ignoring all other visual information. By taking out only the critical component needed for quality assessment, the system achieves accurate measurement with minimal computational and transmission overhead, resolving the contradiction between precision and complexity
Solution Approach 2:
The patent segments the video quality assessment into two independent parts: face detection (performed locally at the receiver) and location comparison (performed with transmitted reference features). This segmentation allows the system to process only the necessary portion of the video data, reducing overall computational burden while maintaining measurement accuracy
2Device complexity
If reduced-reference techniques are used to reduce computational overhead, then computational overhead is reduced, but transmission overhead increases due to need to send extracted features
Solution Approach 1:
The patent extracts only the face location information from the video stream and transmits only this extracted feature to the reference endpoint. By taking out only the essential feature rather than transmitting complete video frames or multiple extracted features, the system minimizes transmission overhead while maintaining sufficient information for accurate quality assessment
Solution Approach 2:
The patent uses partial action by transmitting only the necessary face location data rather than complete video features. This partial transmission approach provides just enough information for quality measurement without the excessive overhead associated with transmitting full extracted feature sets, resolving the contradiction between computational reduction and transmission cost
3Measurement precision
If pixel-based techniques are used to detect distortions, then measurement precision is improved, but adaptability decreases because they cannot handle unanticipated distortions
Solution Approach 1:
The patent extracts the face location feature which is robust to various distortion types. By focusing on the location of the face rather than pixel-level details, the system can detect quality degradation caused by different network conditions and distortions without requiring specific pixel-based distortion models, thereby improving adaptability while maintaining precision
Solution Approach 2:
The patent creates a universal quality metric that works across different video codecs, decoders, and network conditions by using face location comparison. This universal approach can handle anticipated and unanticipated distortions alike, providing both precision and adaptability simultaneously
Data Source
AI summary
System and method to detect video quality degradation in a video stream received by a telecommunications endpoint, the method including: locating reference features characteristic of content in the received video stream; calculating reduced reference features from the located reference features; receiving reduced reference features of a transmitted video stream, the transmitted video stream corresponding to the received video stream; calculating a distance between the reduced reference features in the received video stream and the reduced reference features of the transmitted video stream; and detecting video quality degradation when the calculated distance exceeds a predetermined threshold.


