Configurable Compute-in-Memory Accelerator for Advanced ML Operations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional compute-in-memory (CIM) architectures struggle to support the full range of machine learning operations for advanced architectures like deep neural networks, leading to external data movement and increased power and latency due to reliance on external processors.

Innovation Solution

A dynamically configurable and flexible CIM-based accelerator that integrates a wide range of processing capabilities, including GEMM, GEMV, multiplication, addition, and nonlinear operations, supporting advanced machine learning architectures such as convolutional neural networks, recurrent neural networks, and transformers, by performing computations directly in memory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional CIM architectures are used, then processing capability is improved, but support for advanced machine learning operations deteriorates

Engineering Contradiction:
Improveprocessing capabilityVSAvoidsupport for advanced machine learning operations
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The CIM accelerator is designed with multiple processing modes and configurable architectures that enable it to perform various machine learning operations including matrix-vector multiplication, matrix-matrix multiplication, and support for different network architectures (CNN, RNN, Transformer), making a single system capable of handling diverse ML workloads

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system employs dynamically reconfigurable computing elements that can change their operation mode and connectivity patterns based on the specific ML task being executed, allowing the architecture to adapt its structure and functionality to match the requirements of different advanced machine learning models

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If external processors are used for advanced ML operations, then processing versatility is improved, but power consumption and latency increase

Engineering Contradiction:
Improveprocessing versatilityVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by stationary object

Solution Approach 1:

The patent combines memory storage and processing functions into a single integrated CIM accelerator, eliminating the need for separate external processors and reducing data movement between memory and processing units, thereby lowering power consumption while maintaining processing versatility

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If external processors are used for advanced ML operations, then processing versatility is improved, but latency increases

Engineering Contradiction:
Improveprocessing versatilityVSAvoidlatency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

By merging memory and processing functions into the CIM accelerator, the system eliminates external data movement and communication overhead, enabling faster execution of advanced ML operations while maintaining support for multiple network architectures and operations

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4360002B1Compute in memory-based machine learning accelerator architecture
Publication Date: 2026.02.04 QUALCOMM INC
  • EP4360002B1 patent drawingFigure 1
  • EP4360002B1 patent drawingFigure 2A~2B
  • EP4360002B1 patent drawingFigure 3

AI summary

Certain aspects of the present disclosure provide techniques for processing machine learning model data with a machine learning task accelerator, including: configuring one or more signal processing units (SPUs) of the machine learning task accelerator to process a machine learning model; providing model input data to the one or more configured SPUs; processing the model input data with the machine learning model using the one or more configured SPUs; and receiving output data from the one or more configured SPUs.