Unified Compute-in-Memory Processor Architecture for CPU and DNN Modes

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional compute-in-memory (CIM) architectures suffer from processor stall and underutilization issues due to significant pre/post-processing and data movement, with deep neural network computing being bottlenecked by CPU processing and data transfer, and current CIM developments do not support general-purpose CPU operations effectively.

Innovation Solution

A unified compute-in-memory processor architecture that operates in both central processing unit (CPU) and deep neural network (DNN) modes, featuring a data activation memory and data cache output memory configuration that optimizes data flow and instruction set for seamless data sharing between CPU and DNN operations, reducing inter-core data transfer overhead and enhancing energy efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a conventional architecture with separate CPU core and CIM accelerator is used, then specialized DNN computing can be achieved, but processor stall and underutilization occur due to significant pre/post-processing and data movement

Engineering Contradiction:
ImproveDNN computing efficiencyVSAvoidprocessor stall time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent merges the CPU core and CIM accelerator into a unified processor architecture where the CPU core includes integrated CIM units. This integration allows seamless cooperation between general-purpose computing and specialized DNN operations, eliminating the need for separate processing units and reducing data transfer overhead between them.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The CPU core is designed with multi-functionality to handle both general-purpose computing tasks and specialized DNN computations. The integrated CIM units within the CPU core enable the same processing unit to perform diverse functions, reducing processor underutilization and improving overall system efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If separate CPU and DMA engine are used for data transfer, then flexible data movement is achieved, but deep neural network computing is bottlenecked by CPU processing and data transfer

Engineering Contradiction:
Improvedata movement flexibilityVSAvoidDNN computing throughput
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent integrates data transfer capabilities directly into the CPU core architecture, merging what were previously separate CPU and DMA engine functions. This integration enables efficient data movement for DNN computations without the bottlenecks of traditional separate DMA engines, while maintaining flexibility for various data transfer needs.

Inventive Principle:
Principle #5Merging (Combining)

3Productivity

If current CIM developments are used, then compute-in-memory functionality is achieved, but general-purpose CPU operations are not supported effectively

Engineering Contradiction:
Improvecompute-in-memory efficiencyVSAvoidCPU operation support
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal processor architecture where the CPU core can effectively perform both general-purpose operations and compute-in-memory functions. The integrated design ensures that CIM functionality is seamlessly incorporated into the CPU core, enabling it to handle diverse computational workloads including both CPU-specific and DNN-specific tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If conventional architecture with multiple separate components is used, then functional specialization is achieved, but end-to-end latency is increased due to significant pre/post-processing and data movement

Engineering Contradiction:
Improvefunctional specializationVSAvoidend-to-end latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent combines multiple separate functional components (CPU core, CIM accelerator, data transfer units) into a unified processor architecture. This integration reduces the number of interfaces and data transfer steps between components, thereby reducing end-to-end latency while preserving functional specialization through the integrated design.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240412781A1Compute-in-memory processor supporting both general-purpose CPU and deep learning
Publication Date: 2024.12.12 NORTHWESTERN UNIV
  • US20240412781A1 patent drawing
  • US20240412781A1 patent drawing
  • US20240412781A1 patent drawing

AI summary

In certain aspects, a compute-in-memory processor includes central computing units configured to operate in a central processing unit mode and a deep neural network mode. A data activation memory and a data cache output memory are in communication with the compute-in-memory processor. In the deep neural network mode, the data activation memory is configured as input memory and the data cache output memory is configured as output memory. In the central processing unit mode, the data activation memory is configured as a first data cache and the data cache output memory is configured as a register file and a second data cache.