Video Decoder SEI Message for Spatial-Temporal Resolution Trade-offs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing technologies fail to effectively communicate the intended balance between frame rate and resolution to decoders, leading to suboptimal user experience due to bit rate constraints and content type variability.
Innovation Solution
Incorporating a supplemental enhancement information (SEI) message to transmit the spatial-to-temporal resolution ratio within H.264 or H.265 video streams, allowing decoders to adjust jitter buffers and handle missing packets based on content type, such as live camera or synthetic data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If high spatial resolution is used for synthetic data, then image quality is improved, but frame rate decreases
Solution Approach 1:
The patent applies local quality by enabling different processing strategies for different content types within the same video stream. The decoder identifies whether content is synthetic data or live camera input and applies appropriate processing: high spatial resolution for synthetic data and high frame rate for live input, thus optimizing each content type locally rather than using a uniform approach
2Speed
If high frame rate is used for live camera input, then temporal resolution is improved, but spatial resolution decreases
Solution Approach 1:
The system applies different quality priorities to different content types: live camera input receives high temporal resolution processing while synthetic data receives high spatial resolution processing. This local quality differentiation resolves the contradiction by matching processing characteristics to content requirements
3Productivity
If video encoder tunes to specific content type, then bandwidth efficiency is improved, but decoder cannot adapt processing to content type
Solution Approach 1:
The patent implements feedback by having the encoder send content type information (synthetic data or live camera input) to the decoder. This feedback loop enables the decoder to adapt its processing strategy to match the encoded content type, resolving the contradiction between bandwidth efficiency and decoder adaptability
Solution Approach 2:
The content type information acts as an intermediary that carries encoding intentions from the encoder to the decoder. This intermediary message enables the decoder to understand how the video was encoded and adjust its processing accordingly, maintaining both bandwidth efficiency and decoder adaptability
4Device complexity
If uniform processing is applied to all video content, then decoder complexity is reduced, but user experience deteriorates
Solution Approach 1:
The patent applies local quality by implementing content-type-specific processing in the decoder. Rather than uniform processing, the decoder adapts its behavior based on the content type (synthetic data vs. live input), improving user experience through optimized processing while maintaining manageable complexity through rule-based adaptation
Data Source
AI summary
In one embodiment, a decoder or transcoder of a video conference network receives the first video stream and an indication of the ratio of the spatial-to-temporal resolution of the tuning of the encoding. The behavior of the decoder or transcoder is set based on the indication of the ratio. The behavior is for use of the first video stream.


