Video Packet Loss Visibility Assessment Using Spatio-Temporal Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for assessing packet loss impairment in video sequences, such as full-reference, picture buffer no-reference, bitstream no-reference, and hybrid no-reference techniques, have limited accuracy in determining the perceptual impact of packet loss on video quality, particularly in IP-based broadcast systems where retransmission is not possible, as they do not effectively consider factors like error location, size, duration, and masking properties.
Innovation Solution
A method that combines dynamic temporal and spatial modeling to determine error visibility in video sequences by identifying affected blocks, calculating expected and actual temporal and spatial measures, and comparing these to estimate the perceptual impact of packet loss, using a receiver with modules for video decoding, temporal analysis, spatial analysis, and PLI visibility classification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If full-reference measures like MSE and PSNR are used to assess packet loss impairment, then the assessment is convenient and tractable, but the accuracy as an indicator of perceived video quality is limited
Solution Approach 1:
The patent transforms the assessment from using simple pixel-difference parameters (MSE, PSNR) to using spatio-temporal modeling parameters that capture error propagation characteristics, location, size, and duration. This changes the nature of the measurement parameters to better reflect perceptual quality while maintaining computational feasibility through systematic modeling approaches.
2Adaptability or versatility
If NR-P techniques like slice boundary mismatch are used to measure PLI factors, then the methods can operate without reference to source video, but they suffer from misclassification of natural image variation as error leading to inaccuracy
Solution Approach 1:
The patent performs preliminary classification of macroblocks as errored or unerrored based on bitstream information before applying spatial-temporal analysis. This preliminary action allows the system to focus analysis only on regions with actual errors, preventing misclassification of natural image variations while maintaining the no-reference advantage.
Solution Approach 2:
The patent introduces dynamic temporal modeling that tracks error propagation across multiple frames, allowing the assessment to adapt to changing error patterns over time. This dynamic approach distinguishes between natural image variations and actual error propagation by analyzing temporal consistency, thereby improving measurement precision while maintaining operational versatility.
3Measurement precision
If FR structural similarity SSIM based techniques are used to compare reference and distorted signals, then structural information changes are captured, but the method does not directly consider PLI factors such as error location, size and duration
Solution Approach 1:
The patent merges SSIM-based structural similarity assessment with spatio-temporal error modeling by combining the structural comparison framework with additional layers that track error location, size, and duration. This integration allows the system to maintain the strengths of SSIM while incorporating PLI-specific factors through a unified assessment model that operates on both structural and error-characteristic dimensions.
4Measurement precision
If NR-H models use errored bitstream to measure error extent with macroblock type and motion information, then the accuracy is enhanced, but the device complexity increases
Solution Approach 1:
The patent extracts only the essential error-related information from the bitstream (macroblock type and motion information) needed for accurate assessment, rather than processing all bitstream data. This extraction approach maintains measurement precision by focusing on the most relevant parameters while reducing device complexity by eliminating unnecessary processing of non-essential data elements.
Data Source
AI summary
The invention presents a new NR-H method for assessment of packet loss visibility measure for a video sequence, where the measure is indicative of the effect on the perceptual quality of the video. Packet loss can occur as a result of the video being transmitted over an imperfect network. The invention combines dynamic modelling of temporal and spatial properties of the decoded pictures with bitstream information revealing location, extent and propagation of any errors. Analysis is performed on blocks of pixels, and preferably the macroblocks defined in the particular video encoding scheme. Knowledge of the error extent from the bitstream information is used to target spatial analysis around the specific error locations. Perceptual impact is estimated by utilizing spatio-temporal modelling to predict the properties of a missing block, and comparing those predictions with the actual properties of the missing block.


