Video Encoding Using Co-sited Gradient and Variance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video encoding methods fail to efficiently differentiate between regions of varying complexity in video frames, leading to suboptimal compression quality as they rely on gradient and variance metrics that are not sensitive enough to directional patterns and do not distinguish between directional and non-directional activities.
Innovation Solution
The proposed solution utilizes a combination of variance and gradient metrics to classify frame portions into different categories, adjusting encoding qualities (such as quantization parameters) based on these classifications to optimize encoding for regions with lines, text, and object borders, implemented in both single and multi-pass encoding processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional gradient and variance metrics are used for encoding, then the encoding process is simple, but the ability to differentiate between regions of varying complexity is insufficient
Solution Approach 1:
The patent segments the frame into multiple regions and applies different encoding strategies to each region based on its complexity characteristics. By dividing the frame into regions with different complexity levels (simple, medium, complex), the encoding process can target specific areas requiring higher precision without uniformly increasing complexity across the entire frame.
Solution Approach 2:
The patent implements local quality enhancement by applying different quantization parameters and encoding presets to different regions of the frame. Complex regions (containing text, lines, or objects) receive higher quality encoding with lower quantization, while simple regions use higher quantization for compression, thereby improving overall measurement precision for critical regions without uniformly increasing system complexity.
2Manufacturing precision
If uniform encoding quality is applied to all regions, then the encoding process is simple, but the visual quality and compression efficiency are suboptimal
Solution Approach 1:
The patent applies local quality optimization by assigning different quantization parameters to different regions based on their complexity. Critical regions (text, lines, objects) use lower quantization values for higher visual quality, while simple regions use higher quantization for better compression, thereby simultaneously improving visual quality and compression efficiency.
Solution Approach 2:
The patent dynamically changes encoding parameters (quantization parameter, preset level) based on region complexity classification. By adjusting these parameters locally rather than uniformly, the system achieves both high visual quality in important regions and high compression efficiency in less critical regions, resolving the contradiction between quality and productivity.
3Measurement precision
If gradient-based line detection is used, then directional patterns can be detected, but the method cannot distinguish between directional and non-directional activities
Solution Approach 1:
The patent combines multiple detection metrics (gradient, variance, and their ratio) to create a composite classification system. By using the ratio of gradient to variance, the system can distinguish between directional activities (lines, text with high gradient relative to variance) and non-directional activities (complex textures with comparable gradient and variance), thereby enhancing both measurement precision and adaptability.
Solution Approach 2:
The patent introduces an intermediary classification step that uses the gradient-variance ratio to categorize regions before applying encoding. This intermediary classification mechanism enables the system to differentiate between various content types (text, lines, objects, complex regions) based on their mathematical characteristics, improving both directional detection precision and overall content differentiation versatility.
Data Source
AI summary
Methods and devices are provided for encoding video. By using co-sited gradient and variance values to detect text and line in frames of the video. A processor is configured to receive a plurality of frames of video, determine, for a portion of a frame, a variance of the portion of the frame and a gradient of the portion of the frame and encode, using one of a plurality of different encoding qualities, the portion of the frame based on the gradient and the variance of the portion of the frame. Encoding is performed at both the sub-frame level and frame level. The portion of the frame is classified into one of a plurality of categories based on the gradient and variance and encoded based on the category.


