Configurable Compute-in-Memory Accelerator for Advanced ML Operations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional compute-in-memory (CIM) architectures struggle to support the full range of machine learning operations for advanced architectures like deep neural networks, leading to external data movement and increased power and latency due to reliance on external processors.
Innovation Solution
A dynamically configurable and flexible CIM-based accelerator that integrates a wide range of processing capabilities, including GEMM, GEMV, multiplication, addition, and nonlinear operations, supporting advanced machine learning architectures such as convolutional neural networks, recurrent neural networks, and transformers, by performing computations directly in memory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional CIM architectures are used, then processing capability is improved, but support for advanced machine learning operations deteriorates
Solution Approach 1:
The CIM accelerator is designed with multiple processing modes and configurable architectures that enable it to perform various machine learning operations including matrix-vector multiplication, matrix-matrix multiplication, and support for different network architectures (CNN, RNN, Transformer), making a single system capable of handling diverse ML workloads
Solution Approach 2:
The system employs dynamically reconfigurable computing elements that can change their operation mode and connectivity patterns based on the specific ML task being executed, allowing the architecture to adapt its structure and functionality to match the requirements of different advanced machine learning models
2Adaptability or versatility
If external processors are used for advanced ML operations, then processing versatility is improved, but power consumption and latency increase
Solution Approach 1:
The patent combines memory storage and processing functions into a single integrated CIM accelerator, eliminating the need for separate external processors and reducing data movement between memory and processing units, thereby lowering power consumption while maintaining processing versatility
3Adaptability or versatility
If external processors are used for advanced ML operations, then processing versatility is improved, but latency increases
Solution Approach 1:
By merging memory and processing functions into the CIM accelerator, the system eliminates external data movement and communication overhead, enabling faster execution of advanced ML operations while maintaining support for multiple network architectures and operations
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
Certain aspects of the present disclosure provide techniques for processing machine learning model data with a machine learning task accelerator, including: configuring one or more signal processing units (SPUs) of the machine learning task accelerator to process a machine learning model; providing model input data to the one or more configured SPUs; processing the model input data with the machine learning model using the one or more configured SPUs; and receiving output data from the one or more configured SPUs.