Compute-in-Memory ML Accelerator for RNN and Transformer Workloads

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional compute-in-memory (CIM) architectures are unable to efficiently process advanced machine learning model architectures such as recurrent neural networks, attention models, and BERT models, which are crucial for applications like natural language processing and speech recognition, due to limitations in performing a wide range of processing operations.

Innovation Solution

A machine learning task accelerator is developed, comprising mixed signal processing units with compute-in-memory circuits, analog-to-digital converters, nonlinear operation circuits, and direct memory access controllers, capable of performing matrix multiplication, element-wise operations, and nonlinear processing, thereby supporting advanced machine learning architectures like convolutional neural networks, recurrent neural networks, and transformers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional CIM architectures are used, then processing speed and energy efficiency are improved, but the ability to process advanced machine learning model architectures (RNNs, attention models, BERT) deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidability to process advanced ML architectures
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal processing unit that can handle multiple types of operations including matrix multiplication, element-wise operations, and reduction operations. The processing unit is designed with configurable functional units that can be dynamically programmed to support different ML architectures (CNNs, RNNs, attention models, BERT) through a unified instruction set architecture, eliminating the need for architecture-specific hardware designs.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs dynamic reconfiguration capabilities where the processing units can change their operational mode based on the specific ML model being executed. The functional units are designed to be dynamically programmable, allowing the same hardware to adapt its behavior for different computational patterns required by various ML architectures through control signals and configuration registers.

Inventive Principle:
Principle #15Dynamics

2Productivity

If dedicated hardware accelerators are used, then machine learning processing capacity is improved, but space and power consumption increase

Engineering Contradiction:
ImproveML processing capacityVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent merges computation and memory functions into a single integrated structure where processing units are directly coupled with memory arrays. This compute-in-memory architecture eliminates data movement between separate processing and memory components, significantly reducing power consumption while maintaining high ML processing capacity. The integration allows weights and activations to be stored locally and processed in-place.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The processing units are designed to perform computations using data already present in the memory arrays without requiring external data transfer. The system serves its own computational needs by leveraging the stored weights and activations directly in the memory fabric, eliminating the energy cost of data movement and external interface operations.

Inventive Principle:
Principle #25Self-service

3Loss of energy

If conventional CIM processes are used, then energy efficiency is improved, but the range of supported processing operations deteriorates

Engineering Contradiction:
Improveenergy efficiencyVSAvoidrange of supported operations
Core Design Contradiction:
Loss of energyVSAdaptability or versatility

Solution Approach 1:

The patent segments the processing functionality into multiple specialized functional units within each processing element, including matrix multiplication units, element-wise operation units, and reduction units. Each functional unit is optimized for specific operations but collectively they provide comprehensive support for diverse ML workloads. This segmentation allows energy-efficient specialized processing for each operation type while maintaining overall versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The processing units incorporate universal functional blocks that can be configured to perform different operations through programming. The same physical hardware can execute matrix multiplication, element-wise additions, activations, and reduction operations by changing control parameters, providing both broad operational support and energy efficiency through dedicated hardware acceleration for each operation type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20220414443A1Compute in memory-based machine learning accelerator architecture
Publication Date: 2022.12.29 QUALCOMM INC
  • US20220414443A1 patent drawing
  • US20220414443A1 patent drawing
  • US20220414443A1 patent drawing

AI summary

Certain aspects of the present disclosure provide techniques for processing machine learning model data with a machine learning task accelerator, including: configuring one or more signal processing units (SPUs) of the machine learning task accelerator to process a machine learning model; providing model input data to the one or more configured SPUs; processing the model input data with the machine learning model using the one or more configured SPUs; and receiving output data from the one or more configured SPUs.