Convolver Unit Latch Architecture for Low Power Image Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current processing devices face inefficiencies in performing convolution operations, particularly in reducing power consumption while handling larger images, as they require frequent memory access and lack optimized methods for minimizing external memory operations.

Innovation Solution

The design of a processing device with a convolver unit that includes interconnected convolution circuits, utilizing latches to store data for reuse across neighboring pixels and minimizing external memory access by grouping convolution operations based on pooling sample sizes, thereby optimizing memory access and reducing power consumption.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional processing devices perform convolution operations on larger images, then image processing capability is improved, but power consumption increases due to frequent memory access

Engineering Contradiction:
Improveimage processing capabilityVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies preliminary action by pre-loading input image data into on-chip memory buffers before convolution processing begins. This allows the convolution engine to access data from fast on-chip memory rather than repeatedly accessing external memory, significantly reducing power consumption while maintaining the ability to process larger images.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements nesting by integrating multiple convolution engines and memory buffers within a single image processing pipeline. The convolution engines are nested within the processing device, and data is nested between external memory, on-chip memory buffers, and the convolution engines, creating a hierarchical memory structure that reduces external memory access frequency.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Adaptability or versatility

If convolution operations are performed with frequent memory access, then processing flexibility is maintained, but processing efficiency deteriorates due to increased power consumption

Engineering Contradiction:
Improveprocessing flexibilityVSAvoidprocessing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent applies segmentation by dividing the image processing task into distinct stages: data loading into on-chip buffers, convolution processing by multiple engines, and result output. This segmentation allows the system to optimize each stage independently, maintaining flexibility while improving overall processing efficiency by reducing external memory access frequency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements multi-functionality through convolution engines that can process multiple types of convolution operations (standard convolution, depthwise convolution, pointwise convolution) and support various image formats and resolutions. This universal design maintains processing flexibility while the on-chip memory architecture improves efficiency by keeping data locally available.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Use of energy by moving object

If external memory access is minimized by grouping convolution operations, then power consumption is reduced, but memory access optimization complexity increases

Engineering Contradiction:
Improvepower consumptionVSAvoidmemory access optimization complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent applies dynamics by implementing a pooling sample generator that dynamically determines grouping strategies based on the specific convolution operation parameters, image size, and memory availability. This dynamic approach allows the system to optimize memory access patterns for different scenarios without requiring complex fixed configurations, reducing power consumption while managing optimization complexity through adaptive control.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10402468B2Processing device for performing convolution operations
Publication Date: 2019.09.03 INTEL CORP
  • US10402468B2 patent drawing
  • US10402468B2 patent drawing
  • US10402468B2 patent drawing

AI summary

Systems and methods for performing convolution operations. An example processing system comprises: a processing core; and a convolver unit to apply a convolution filter to a plurality of input data elements represented by a two-dimensional array, the convolver unit comprising a plurality of multipliers coupled to two or more sets of latches, wherein each set of latches is to store a plurality of data elements of a respective one-dimensional section of the two-dimensional array.