Multi-layered Video Streaming Composition via Inter-layer Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video streaming technologies face inefficiencies in composing multiple videos for devices with single hardware decoders, such as low-cost TVs and mobile devices, due to signal quality degradation and computational complexity in full transcoding, and experience bitrate peaks during layout changes in video conferencing applications.
Innovation Solution
A video streaming apparatus that forms a multi-layered data stream by copying and synthesizing video content in the compressed domain, using a copy former and synthesizer to create a single scalable bitstream with inter-layer prediction, allowing for flexible composition and reduced bitrate consumption during layout changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If full transcoding is used to generate a single video bitstream from multiple videos, then a single bitstream suitable for devices with single decoders is achieved, but signal quality degrades due to repeated encoding and computational complexity increases
Solution Approach 1:
The patent segments the video processing into multiple layers (base layer and enhancement layers) where each layer processes different components of the video content. The base layer contains essential information while enhancement layers add additional quality, allowing progressive transmission without full re-encoding of the entire content.
Solution Approach 2:
The patent implements a nested structure where enhancement layers are embedded within the scalable video coding framework. Each enhancement layer is nested within the previous layer, with lower layers providing foundation for higher layers, enabling efficient composition without breaking the compression chain.
2Adaptability or versatility
If full transcoding is used to compose multiple videos into a single bitstream, then composition flexibility is achieved, but computational complexity increases significantly
Solution Approach 1:
The patent divides the composition task into separate layer processing operations. Instead of fully re-encoding the entire composited video, it processes different layers independently with appropriate complexity levels, reducing overall computational burden while maintaining composition flexibility.
Solution Approach 2:
The patent changes the processing parameters by working in the compressed domain rather than pixel domain. By operating on encoded bitstreams and using scalable video coding parameters, it achieves composition with significantly reduced computational complexity compared to traditional full transcoding approaches.
3Device complexity
If traditional video stitching in compressed domain is used for scalable video coding, then computational complexity is reduced, but bitrate peaks occur during layout changes
Solution Approach 1:
The patent performs preliminary actions by pre-planning the layout changes and preparing the necessary layer structures in advance. When layout changes are anticipated, the scalable video coding structure is already in place to accommodate them, preventing sudden bitrate peaks during actual transitions.
Solution Approach 2:
The patent implements dynamic adaptation of the video composition by allowing flexible switching between different layer configurations. The scalable video coding structure enables dynamic adjustment of which layers are transmitted and how they are composed, optimizing bitrate usage during layout changes rather than requiring full re-encoding.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Video streaming concepts are presented. According to a first aspect, the video stream is formed as a multi-layered data stream with forming a set of one or more layers of the multi-layered data stream by copying from the coded version of the video content, while a composition of the at least one video is synthesized in at least a portion of pictures of a predetermined layer of the multi-layer data stream by means of inter-layer prediction. According to a second aspect, inter-layer prediction is used to either substitute otherwise missing referenced pictures of a newly encompassed video by inserting replacement pictures, or portions of the newly encompassed video referencing, by motion-compensated prediction, pictures which are missing are replaced by inter-layer prediction. According to a third aspect, output pictures inserted into the composed video stream so as to synthesize the composition of the video content by copying from a no-output portion of the composed data stream by temporal prediction, are inserted into the composed data stream so that output pictures are arranged in the data stream in the presentation time order rather than the coded picture order.