Compute-in-Memory ML Accelerator for RNN and Transformer Workloads
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional compute-in-memory (CIM) architectures are unable to efficiently process advanced machine learning model architectures such as recurrent neural networks, attention models, and BERT models, which are crucial for applications like natural language processing and speech recognition, due to limitations in performing a wide range of processing operations.
Innovation Solution
A machine learning task accelerator is developed, comprising mixed signal processing units with compute-in-memory circuits, analog-to-digital converters, nonlinear operation circuits, and direct memory access controllers, capable of performing matrix multiplication, element-wise operations, and nonlinear processing, thereby supporting advanced machine learning architectures like convolutional neural networks, recurrent neural networks, and transformers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional CIM architectures are used, then processing speed and energy efficiency are improved, but the ability to process advanced machine learning model architectures (RNNs, attention models, BERT) deteriorates
Solution Approach 1:
The patent implements a universal processing unit that can handle multiple types of operations including matrix multiplication, element-wise operations, and reduction operations. The processing unit is designed with configurable functional units that can be dynamically programmed to support different ML architectures (CNNs, RNNs, attention models, BERT) through a unified instruction set architecture, eliminating the need for architecture-specific hardware designs.
Solution Approach 2:
The patent employs dynamic reconfiguration capabilities where the processing units can change their operational mode based on the specific ML model being executed. The functional units are designed to be dynamically programmable, allowing the same hardware to adapt its behavior for different computational patterns required by various ML architectures through control signals and configuration registers.
2Productivity
If dedicated hardware accelerators are used, then machine learning processing capacity is improved, but space and power consumption increase
Solution Approach 1:
The patent merges computation and memory functions into a single integrated structure where processing units are directly coupled with memory arrays. This compute-in-memory architecture eliminates data movement between separate processing and memory components, significantly reducing power consumption while maintaining high ML processing capacity. The integration allows weights and activations to be stored locally and processed in-place.
Solution Approach 2:
The processing units are designed to perform computations using data already present in the memory arrays without requiring external data transfer. The system serves its own computational needs by leveraging the stored weights and activations directly in the memory fabric, eliminating the energy cost of data movement and external interface operations.
3Loss of energy
If conventional CIM processes are used, then energy efficiency is improved, but the range of supported processing operations deteriorates
Solution Approach 1:
The patent segments the processing functionality into multiple specialized functional units within each processing element, including matrix multiplication units, element-wise operation units, and reduction units. Each functional unit is optimized for specific operations but collectively they provide comprehensive support for diverse ML workloads. This segmentation allows energy-efficient specialized processing for each operation type while maintaining overall versatility.
Solution Approach 2:
The processing units incorporate universal functional blocks that can be configured to perform different operations through programming. The same physical hardware can execute matrix multiplication, element-wise additions, activations, and reduction operations by changing control parameters, providing both broad operational support and energy efficiency through dedicated hardware acceleration for each operation type.
Data Source
AI summary
Certain aspects of the present disclosure provide techniques for processing machine learning model data with a machine learning task accelerator, including: configuring one or more signal processing units (SPUs) of the machine learning task accelerator to process a machine learning model; providing model input data to the one or more configured SPUs; processing the model input data with the machine learning model using the one or more configured SPUs; and receiving output data from the one or more configured SPUs.


