Per-title encoding spatial temporal downscaling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video streaming technologies use a single or limited fixed bitrate ladders for all video content, leading to suboptimal bitrate allocation, resulting in bandwidth waste and lower Quality of Experience (QoE), as they do not consider the unique characteristics of different video contents and the impact of spatial and temporal resolutions on perceived video quality.
Innovation Solution
The method involves per-title encoding using spatial and temporal resolution downscaling, where video segments are downscaled to lower resolutions and frame rates, and then upscaled back to optimize the bitrate ladder over both spatial and temporal resolutions, selecting bitrate-resolution-framerate triples that form a convex hull for optimal quality, using machine learning to predict perceived video quality and reduce encoding costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a fixed bitrate ladder is used for all video content, then device complexity is reduced and ease of operation is improved, but bitrate allocation efficiency deteriorates leading to bandwidth waste and lower QoE
Solution Approach 1:
The patent implements dynamic bitrate ladder optimization by analyzing video content characteristics (spatial complexity, temporal complexity, motion content) and adapting the bitrate allocation strategy accordingly. Different video segments receive customized bitrate ladders based on their specific characteristics, transforming the static fixed bitrate approach into a dynamic adaptive system that optimizes bandwidth utilization while maintaining ease of operation through automated analysis and selection.
2Manufacturing precision
If higher spatial resolution is used, then video quality is improved, but encoding cost and storage requirements increase
Solution Approach 1:
The patent applies local quality optimization by analyzing spatial complexity variations within different regions and segments of video content. Instead of uniformly applying high spatial resolution encoding across the entire video, the system identifies regions with lower spatial complexity and allocates lower bitrates to those areas, while concentrating higher bitrates on regions with high spatial complexity. This localized approach maintains overall video quality while significantly reducing encoding costs and storage requirements.
3Manufacturing precision
If higher temporal resolution is used, then video quality is improved, but encoding cost and bandwidth consumption increase
Solution Approach 1:
The patent implements temporal quality optimization by analyzing temporal complexity variations across different video segments. The system identifies segments with low temporal complexity (e.g., static scenes) and reduces the framerate or skips frames in those segments, while maintaining higher framerates in segments with high temporal complexity (e.g., fast motion scenes). This localized temporal downscaling maintains perceived video quality while significantly reducing bandwidth consumption and encoding costs.
4Productivity
If per-title encoding with multiple bitrate-resolution pairs is implemented, then bitrate allocation efficiency is improved, but device complexity and encoding time increase
Solution Approach 1:
The patent applies preliminary action by performing comprehensive video content analysis and pre-computing optimized bitrate ladders during the encoding phase. The system analyzes video characteristics (spatial complexity, temporal complexity, motion content) beforehand and generates customized bitrate allocation strategies for each video title or segment. This preliminary optimization reduces runtime complexity during playback, as the adaptive streaming system can simply select from pre-computed options rather than performing complex real-time analysis, thus balancing bitrate allocation efficiency with device complexity.
Data Source
AI summary
Techniques relating to per-title encoding using spatial and temporal resolution downscaling is disclosed. A method for per-title encoding includes receiving a video input comprised of video segments, spatially downscaling the video input, temporally downscaling the video input, encoding the video input to generate an encoded video, then temporally and spatially upscaling the encoded video. Spatially downscaling may include reducing a resolution of the video input, and temporally downscaling may include reducing a framerate of the video input. Objective metrics for the upscaled encoded video show improved quality over conventional methods.


