Convolver Unit Latch Architecture for Low Power Image Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current processing devices face inefficiencies in performing convolution operations, particularly in reducing power consumption while handling larger images, as they require frequent memory access and lack optimized methods for minimizing external memory operations.
Innovation Solution
The design of a processing device with a convolver unit that includes interconnected convolution circuits, utilizing latches to store data for reuse across neighboring pixels and minimizing external memory access by grouping convolution operations based on pooling sample sizes, thereby optimizing memory access and reducing power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional processing devices perform convolution operations on larger images, then image processing capability is improved, but power consumption increases due to frequent memory access
Solution Approach 1:
The patent applies preliminary action by pre-loading input image data into on-chip memory buffers before convolution processing begins. This allows the convolution engine to access data from fast on-chip memory rather than repeatedly accessing external memory, significantly reducing power consumption while maintaining the ability to process larger images.
Solution Approach 2:
The patent implements nesting by integrating multiple convolution engines and memory buffers within a single image processing pipeline. The convolution engines are nested within the processing device, and data is nested between external memory, on-chip memory buffers, and the convolution engines, creating a hierarchical memory structure that reduces external memory access frequency.
2Adaptability or versatility
If convolution operations are performed with frequent memory access, then processing flexibility is maintained, but processing efficiency deteriorates due to increased power consumption
Solution Approach 1:
The patent applies segmentation by dividing the image processing task into distinct stages: data loading into on-chip buffers, convolution processing by multiple engines, and result output. This segmentation allows the system to optimize each stage independently, maintaining flexibility while improving overall processing efficiency by reducing external memory access frequency.
Solution Approach 2:
The patent implements multi-functionality through convolution engines that can process multiple types of convolution operations (standard convolution, depthwise convolution, pointwise convolution) and support various image formats and resolutions. This universal design maintains processing flexibility while the on-chip memory architecture improves efficiency by keeping data locally available.
3Use of energy by moving object
If external memory access is minimized by grouping convolution operations, then power consumption is reduced, but memory access optimization complexity increases
Solution Approach 1:
The patent applies dynamics by implementing a pooling sample generator that dynamically determines grouping strategies based on the specific convolution operation parameters, image size, and memory availability. This dynamic approach allows the system to optimize memory access patterns for different scenarios without requiring complex fixed configurations, reducing power consumption while managing optimization complexity through adaptive control.
Data Source
AI summary
Systems and methods for performing convolution operations. An example processing system comprises: a processing core; and a convolver unit to apply a convolution filter to a plurality of input data elements represented by a two-dimensional array, the convolver unit comprising a plurality of multipliers coupled to two or more sets of latches, wherein each set of latches is to store a plurality of data elements of a respective one-dimensional section of the two-dimensional array.


