Unified Compute-in-Memory Processor Architecture for CPU and DNN Modes
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional compute-in-memory (CIM) architectures suffer from processor stall and underutilization issues due to significant pre/post-processing and data movement, with deep neural network computing being bottlenecked by CPU processing and data transfer, and current CIM developments do not support general-purpose CPU operations effectively.
Innovation Solution
A unified compute-in-memory processor architecture that operates in both central processing unit (CPU) and deep neural network (DNN) modes, featuring a data activation memory and data cache output memory configuration that optimizes data flow and instruction set for seamless data sharing between CPU and DNN operations, reducing inter-core data transfer overhead and enhancing energy efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a conventional architecture with separate CPU core and CIM accelerator is used, then specialized DNN computing can be achieved, but processor stall and underutilization occur due to significant pre/post-processing and data movement
Solution Approach 1:
The patent merges the CPU core and CIM accelerator into a unified processor architecture where the CPU core includes integrated CIM units. This integration allows seamless cooperation between general-purpose computing and specialized DNN operations, eliminating the need for separate processing units and reducing data transfer overhead between them.
Solution Approach 2:
The CPU core is designed with multi-functionality to handle both general-purpose computing tasks and specialized DNN computations. The integrated CIM units within the CPU core enable the same processing unit to perform diverse functions, reducing processor underutilization and improving overall system efficiency.
2Adaptability or versatility
If separate CPU and DMA engine are used for data transfer, then flexible data movement is achieved, but deep neural network computing is bottlenecked by CPU processing and data transfer
Solution Approach 1:
The patent integrates data transfer capabilities directly into the CPU core architecture, merging what were previously separate CPU and DMA engine functions. This integration enables efficient data movement for DNN computations without the bottlenecks of traditional separate DMA engines, while maintaining flexibility for various data transfer needs.
3Productivity
If current CIM developments are used, then compute-in-memory functionality is achieved, but general-purpose CPU operations are not supported effectively
Solution Approach 1:
The patent creates a universal processor architecture where the CPU core can effectively perform both general-purpose operations and compute-in-memory functions. The integrated design ensures that CIM functionality is seamlessly incorporated into the CPU core, enabling it to handle diverse computational workloads including both CPU-specific and DNN-specific tasks.
4Adaptability or versatility
If conventional architecture with multiple separate components is used, then functional specialization is achieved, but end-to-end latency is increased due to significant pre/post-processing and data movement
Solution Approach 1:
The patent combines multiple separate functional components (CPU core, CIM accelerator, data transfer units) into a unified processor architecture. This integration reduces the number of interfaces and data transfer steps between components, thereby reducing end-to-end latency while preserving functional specialization through the integrated design.
Data Source
AI summary
In certain aspects, a compute-in-memory processor includes central computing units configured to operate in a central processing unit mode and a deep neural network mode. A data activation memory and a data cache output memory are in communication with the compute-in-memory processor. In the deep neural network mode, the data activation memory is configured as input memory and the data cache output memory is configured as output memory. In the central processing unit mode, the data activation memory is configured as a first data cache and the data cache output memory is configured as a register file and a second data cache.


