Parallel Image Decoding with Predictive Memory Bandwidth Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image decoding devices experience a significant increase in memory bandwidth requirements due to the momentary transfer of large data amounts, which can hinder the processing of other devices sharing the memory and lead to delays, especially when handling super high-definition images like 4K2K.
Innovation Solution
An image decoding device that decodes coded image data on a block-by-block basis, using a storage unit to store reference images, a pre-decoding unit to decode reference information, an amount-of-transferred-data prediction unit to calculate predictive data amounts, and a block determination unit to determine blocks for parallel decoding, thereby reducing data variation and memory bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If parallel decoding is performed using multiple decoding devices to achieve high operation capability for 4K2K images, then processing speed is improved, but memory bandwidth requirements increase significantly due to momentary large data transfers
Solution Approach 1:
The patent applies preliminary action by performing pre-decoding of reference information (macroblock type, motion information, block partition type) before actual decoding. This allows the system to predict the data amount of reference images in advance, enabling proactive memory bandwidth management and preventing momentary large data transfers that would occur with conventional parallel decoding approaches.
Solution Approach 2:
The patent introduces an intermediary mechanism through the amount-of-transferred-data prediction unit and block determination unit. These units act as mediators between the storage unit and block decoding units, predicting data amounts and determining optimal blocks for parallel decoding to smooth memory bandwidth usage and eliminate the need for high-capacity memory bandwidth infrastructure.
2Productivity
If conventional parallel decoding is implemented without data prediction, then processing efficiency is improved, but data transfer bottlenecks occur due to unmanaged memory bandwidth usage
Solution Approach 1:
The patent implements feedback through the pre-decoding unit that decodes reference information and feeds this data to the amount-of-transferred-data prediction unit. This feedback loop enables the system to adjust block decoding decisions based on predicted data amounts, creating a self-regulating mechanism that maintains processing efficiency while automatically managing memory bandwidth without external intervention.
Solution Approach 2:
The system applies self-service by enabling block decoding units to autonomously determine which blocks to decode in parallel based on predictions from the amount-of-transferred-data prediction unit. Each decoding unit independently manages its own operation, selecting blocks that optimize processing efficiency while collectively maintaining smooth memory bandwidth usage without requiring centralized control.
Data Source
AI summary
An image decoding device capable of performing parallel decoding of coded image data with a small memory bandwidth while suppressing momentary increase in the amount of data transferred for the decoding. An image decoding device (100) includes: an external memory (110) which stores data of reference images; a stream parser unit (120) which decodes reference information indicating the number of reference images to be referred to on a block-by-block basis; an amount-of-transferred-data prediction unit (131) which calculates, on a block-by-block basis using the reference information, a predictive data amount of a reference image to be read out from the external memory (110); a block determination unit (132) which determines, using the predictive data amount, multiple blocks to be decoded in parallel, so as to reduce variation in amounts of data read out from the external memory (110); and macroblock decoding units (140 to 160) which decode the determined blocks in parallel.


