Video Stream Segmentation for Scalable Resolution Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for transmitting and storing video content with high spatial and temporal resolutions are resource-intensive and inefficient, lacking a solution that balances resource usage with perceived quality and scalability.
Innovation Solution
A method that processes a sequence of images by dividing it into subsequences, applying temporal and spatial subsampling based on image frequency, reducing spatial resolution while maintaining temporal resolution, and inserting the processed subsequences into an output container for efficient encoding and storage, allowing for scalable reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If variable bitrate compression is used to adapt bit rate to content, then video quality is improved within reduced flow rate range, but the technique is only applicable in a reduced flow rate range
Solution Approach 1:
The video stream is segmented into multiple layers (base layer and enhancement layers), each with different spatial and temporal resolutions. This segmentation allows the system to adapt to different flow rates by transmitting only the necessary layers, thus expanding the applicable flow rate range while maintaining video quality.
Solution Approach 2:
The patent introduces a new dimension of scalability by decomposing video content into multiple spatial and temporal layers. This dimensional decomposition allows flexible adaptation to different network conditions and device capabilities, resolving the contradiction between quality and flow rate range applicability.
2Adaptability or versatility
If adaptive bit rate stores several encoded streams with different resolutions, then scalability to different device capabilities is improved, but resource cost increases due to coding, storing and transmitting several versions
Solution Approach 1:
Instead of storing and transmitting completely separate encoded streams for different resolutions, the patent merges multiple resolutions into a single scalable stream structure. The base layer and enhancement layers are combined in one bitstream, reducing the total quantity of data that needs to be stored and transmitted while maintaining adaptability to different device capabilities.
Solution Approach 2:
The patent implements dynamic scalability where the receiver can dynamically select which layers to decode based on available device resources. This dynamic approach allows a single encoded stream to serve multiple device capabilities without requiring pre-storage of multiple complete versions, thus reducing resource consumption.
3Manufacturing precision
If scalable compression produces base layer and enhancement layers, then resolution/quality can be increased with sufficient resources, but implementation complexity increases with large memory size and increased latency
Solution Approach 1:
The patent applies local quality enhancement by providing different quality levels for different parts of the video stream through layered encoding. The base layer provides essential quality information, while enhancement layers provide additional detail only where needed. This allows receivers to allocate memory and processing resources locally and efficiently, reducing overall system complexity while maintaining high resolution capability when resources are available.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a method for forming an image sequence, called the output sequence (Ss), from an input image sequence (SE), said input image sequence having an input spatial resolution (RSE) and an input temporal resolution (RTE), said output sequence having an output temporal resolution (RTs) equal to the input temporal resolution (RTE) and an output spatial resolution (RSs) equal to a preset fraction (1/N) of the input spatial resolution (RSE), set by an integer number (N) higher than or equal to 2, characterised in that the method comprises the following steps, which are implemented for a sub-sequence (SSE) of the input sequence, called the current sub-sequence, comprising a preset number (N) of images: obtaining (El) a temporal frequency, called the image frequency (FI), associated with said sub-sequence; processing (E2) the current input sub-sequence to obtain an output sub-sequence (SSs), inserting (E4) the output sub-sequence (SSs) and the associated image frequency (FI) into an output container (Cs).