Video Bitstream Tranche Coding for Low-Latency Parallel Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing technologies face challenges in parallelizing encoding and decoding processes efficiently, particularly with the HEVC standard, due to increased processing requirements and higher video resolutions, leading to delays in transmission and decoding times, especially in time-sensitive applications like video conferencing and gaming.
Innovation Solution
The approach involves segmenting data into smaller tranches within WPP substreams or tiles, allowing for continued context-adaptive binary arithmetic coding probability adaptation across tranche boundaries, enabling earlier transmission and decoding of these tranches, which are smaller than original slices or substreams, thus reducing delay without compromising coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If video data is processed using traditional slice-based encoding, then coding efficiency is maintained, but parallel processing capability is limited and decoding delays increase
Solution Approach 1:
The patent divides video data into multiple tiles, each tile being independently encodable and decodable. This segmentation enables parallel processing across multiple cores while maintaining spatial coherence within each tile, thus improving productivity without significantly increasing decoding delay
Solution Approach 2:
The patent performs preliminary organization of video data into tile structures during encoding, preparing the data in advance for parallel decoding. This preliminary action allows the decoder to immediately utilize multiple cores without additional processing overhead, reducing loss of time
2Speed
If wavefront processing is used to enable parallel decoding, then processing speed improves, but spatial dependencies between adjacent LCUs must be interrupted
Solution Approach 1:
The patent segments the video picture into tiles with clear boundaries, where each tile is self-contained with all necessary spatial dependencies satisfied within the tile. This segmentation allows parallel processing while maintaining spatial dependency integrity within each segment
Solution Approach 2:
The patent ensures that each tile has local completeness by including all necessary reference data within the tile boundaries. This local quality approach maintains spatial dependency integrity within each tile while enabling parallel processing across different tiles
3Loss of time
If data is transmitted in larger units, then transmission efficiency is higher, but transmission delay increases and parallel decoding cannot start earlier
Solution Approach 1:
The patent segments video data into smaller tile units that can be transmitted independently and in parallel. This segmentation reduces transmission delay by allowing earlier start of parallel decoding while maintaining reasonable transmission efficiency through optimized tile sizing
Solution Approach 2:
The patent dynamically adjusts the granularity of data units based on processing requirements and transmission constraints. By making data units adaptable in size and structure, the system optimizes both transmission efficiency and parallel processing capability
Data Source
AI summary
A raw byte sequence payload describing a picture in slices, WPP substreams or tiles and coded using context-adaptive binary arithmetic coding is subdivided into tranches with continuing the context-adaptive binary arithmetic coding probability adaptation across tranche boundaries. Thereby, tranche boundaries additionally introduced within slices, WPP substreams or tiles do not lead to a reduction in the entropy coding efficiency of these entities. However, the tranches are smaller than the original slices, WPP substreams or tiles and accordingly they may be transmitted with a lower delay, than the un-chopped original entities. According to another aspect combinable with the first aspect, substream marker NAL units are used within a sequence of NAL units of a video bitstream to enable a transport demultiplexer to assign data of slices within NAL units to the corresponding substreams or tiles to be able to, in parallel, serve a multi-threaded decoder with the corresponding substreams or tiles.


