Tree-coded Video Compression with Coupled Parallel Pipelines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video coding standards face increased complexity due to dependencies between neighboring coding units, leading to inefficient processing and hardware utilization in exploring all combinations of coding unit sizes for optimal video quality.
Innovation Solution
The implementation of parallel pipelines that operate in a coupled, tile-interleaved fashion to generate and select coefficients for different coding unit sizes within a coding tree unit, maintaining near 100% hardware utilization and allowing for efficient exploration of all size options in tree-coded video compression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple processing devices operate in parallel to explore mode decisions for coding units of different sizes, then video coding quality is improved, but device complexity and hardware area increase
Solution Approach 1:
The patent divides the processing of different coding unit sizes into separate parallel pipelines. Each pipeline is dedicated to a specific coding unit size (e.g., 64×64, 32×32, 16×16, 8×8), allowing simultaneous exploration of mode decisions for all sizes without requiring a single complex device to handle all sizes sequentially. This segmentation enables quality improvement through parallel processing while managing hardware complexity by creating specialized, simpler pipeline units.
Solution Approach 2:
The patent introduces a new dimension of processing by organizing pipelines in a hierarchical structure corresponding to the quad-tree partitioning levels. Instead of using multiple independent devices at the same level, the solution creates a multi-level pipeline architecture where pipelines at different hierarchy levels process coding units of different sizes simultaneously. This dimensional organization allows efficient resource utilization and reduces overall hardware area requirements.
2Manufacturing precision
If multiple processing devices operate in parallel to explore all coding unit size combinations, then video coding quality is improved, but processing time and synchronization overhead increase
Solution Approach 1:
The patent implements periodic action through its pipelined architecture where each pipeline processes coding units in regular, rhythmic intervals. The pipelines are designed to complete processing of one coding unit size and immediately transition to the next, creating a periodic flow of processed data. This periodic operation enables efficient time management and reduces synchronization overhead by establishing predictable processing cycles across all pipelines.
Solution Approach 2:
The patent ensures continuity of useful action by designing the pipelines to operate continuously without idle periods. Each pipeline maintains a steady stream of processing operations, and the hierarchical structure allows higher-level pipelines to continuously receive and process results from lower-level pipelines. This continuous operation maximizes throughput and minimizes processing time while maintaining quality exploration across all coding unit sizes.
3Device complexity
If conventional sequential processing is used for coding units, then hardware area is minimized, but productivity and processing efficiency decrease
Solution Approach 1:
The patent merges multiple processing functions into a unified hierarchical pipeline architecture. Instead of using separate hardware for each coding unit size, the solution combines all size-specific processing into an integrated system where pipelines at different hierarchy levels work together. This merging approach improves productivity through parallel processing while controlling hardware area by sharing common resources and structures across the unified architecture.
Solution Approach 2:
The patent implements universality by designing the hierarchical pipeline architecture to handle multiple coding unit sizes within a single system. Each pipeline is designed with multi-functional capabilities to process different types of data and perform various operations depending on the coding unit size being processed. This universal design improves processing efficiency without requiring separate specialized hardware for each function, thereby controlling overall hardware area.
Data Source
AI summary
An apparatus includes a circuit and a processor. The circuit may be configured to (i) generate a plurality of sets of coefficients by compressing a tile in a picture in a video signal at each of a plurality of different sizes of a plurality of coding units in a coding tree unit and (ii) reconstruct the tile based on a particular one of the sets of coefficients. The sets of coefficients may be generated at two or more of the different sizes of the coding units in parallel. Each of the sets of coefficients may be generated in a corresponding one of a plurality of pipelines that operate in parallel. Each of the sets of coefficients may have a same number of the coefficients. The processor may be configured to select the particular set of coefficients in response to the compression of the tile.


