In-Memory Compute Array Routing for Faster CNN MAC Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional CNN hardware accelerators rely on complex digital circuits, which pose challenges in optimizing die size, power consumption, and cost, and are inefficient in performing MAC operations due to the need for large numbers of clock cycles.

Innovation Solution

The implementation of an in-memory compute circuit with a memory circuit that generates products by combining input values with weight values stored in rows, using a control circuit to route input values to different rows over multiple clock cycles, and digital-to-analog converters to generate voltage levels for efficient MAC operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional digital circuits are used for CNN hardware acceleration, then high-speed throughput can be achieved, but die size, power consumption, and cost increase significantly

Engineering Contradiction:
ImproveCNN computation speedVSAvoidcircuit complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional digital circuit-based MAC operations with an in-memory compute array that performs computations directly within the memory structure. This substitution eliminates the need for complex digital logic circuits by utilizing the memory array's inherent parallelism and analog computing capabilities, thereby reducing circuit complexity while maintaining computational functionality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The in-memory compute array serves multiple functions: it acts as both the storage medium for weight values and the processing unit for MAC operations. This multi-functionality eliminates the need for separate digital circuitry dedicated to computation, reducing overall device complexity while achieving high-speed throughput through parallel analog computing

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If traditional digital circuits are used for CNN hardware acceleration, then MAC operations can be performed, but the number of clock cycles required is undesirable

Engineering Contradiction:
ImproveMAC operation throughputVSAvoidclock cycles per MAC operation
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces sequential digital circuit-based MAC operations with parallel analog computing in the memory array. The analog nature of the in-memory compute array allows multiple MAC operations to be performed simultaneously across the array, dramatically reducing the number of clock cycles required compared to traditional digital implementations that must process operations sequentially or with limited parallelism

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

Weight values are pre-loaded into the memory array before computation begins. This preliminary action allows the compute array to immediately perform MAC operations using the pre-positioned weights, eliminating the need for repeated weight fetching and enabling high-speed parallel computation with minimal clock cycles

Inventive Principle:
Principle #10Preliminary action

3Productivity

If more rows are opened in the IMC array for parallel processing, then computation speed increases, but power consumption increases

Engineering Contradiction:
Improvecomputation parallelismVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

The patent implements dynamic row activation where the number of active rows in the IMC array is adjusted based on the specific computation requirements. Rather than keeping all rows continuously active, the system dynamically opens only the necessary number of rows for each MAC operation, enabling parallelism when needed while conserving power during operations that require fewer computational resources

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the operational parameters of the IMC array by adjusting the number of active rows based on computational demands. This parameter change allows the system to optimize the trade-off between computation speed and power consumption, achieving high parallelism when fast computation is required while reducing power consumption during lower-demand operations

Inventive Principle:
Principle #35Parameter changes

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables faster and more power-efficient MAC operations by maximizing the number of open rows in the IMC array, allowing for high parallelism and efficient data routing, thereby reducing the time and power required for CNN computations.

Implementation Method 1

digital-to-analog converters to generate voltage levels for efficient MAC operations

Methodology Applied
Scientific EffectDigital-to-analog conversion:

Data Source

PatentUS11694733B2Acceleration of in-memory-compute arrays
Publication Date: 2023.07.04 APPLE INC
  • US11694733B2 patent drawing
  • US11694733B2 patent drawing
  • US11694733B2 patent drawing

AI summary

An apparatus includes an in-memory compute circuit that includes a memory circuit configured to generate a set of products by combining received input values with respective weight values stored in rows of the memory circuit, and to combine the set of products to generate an accumulated output value. The in-memory compute circuit may further include a control circuit and a plurality of routing circuits, including a first routing circuit coupled to a first set of rows of the memory circuit. The control circuit may be configured to cause the first routing circuit to route groups of input values to different ones of the first set of rows over a plurality of clock cycles, and the memory circuit to generate, on a clock cycle following the plurality of clock cycles, a particular accumulated output value that is computed based on the routed groups of input values.