AI Processor NVM Cores Power-Gating Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI processors face challenges in reducing power consumption, especially when applied to deep-learning applications, as they often require significant energy to access external memory, limiting their efficiency in AIoT applications.

Innovation Solution

The proposed AI processor incorporates multiple Non-Volatile Memory (NVM) AI cores and an AI core that perform basic unit operations for deep-learning tasks within the processor, using SRAM for results and storing only nonzero weights in NVM, with power-gating to minimize power usage based on bit width requirements for each layer.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If AI processors use external memory for data storage, then data capacity is improved, but power consumption increases due to frequent memory accesses

Engineering Contradiction:
Improvedata capacityVSAvoidpower consumption
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent combines memory and processing functions into a single integrated structure where NVM AI cores perform deep-learning operations directly on data stored in NVM, eliminating the need for separate external memory accesses and reducing power consumption while maintaining data capacity

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces SRAM as an intermediary buffer between NVM and processing units, allowing frequently accessed data to be cached locally, thereby reducing the frequency of high-power NVM accesses while maintaining data availability

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If AI processors operate all components for deep-learning operations, then processing capability is improved, but power consumption increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidpower consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent implements dynamic power-gating control that selectively activates or deactivates NVM AI cores based on the specific deep-learning operation being performed and the bit width requirements, allowing the system to maintain processing capability while minimizing power consumption by operating only essential components

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent divides the processing system into multiple independent NVM AI cores that can be selectively powered on or off, allowing granular control over which processing units are active based on computational needs, thereby optimizing the balance between processing capability and power consumption

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If AI processors store all weights in memory, then model accuracy is improved, but power consumption increases due to storing and reading zero weights

Engineering Contradiction:
Improvemodel accuracyVSAvoidpower consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts and stores only nonzero weights in NVM, eliminating the storage and access of zero weights entirely. This sparsity exploitation technique maintains model accuracy by preserving all meaningful weight values while reducing memory usage and power consumption associated with storing and reading redundant zero values

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240062809A1Artificial intelligence processor and method of processing deep-learning operation using the same
Publication Date: 2024.02.22 ELECTRONICS & TELECOMM RES INST
  • US20240062809A1 patent drawing
  • US20240062809A1 patent drawing
  • US20240062809A1 patent drawing

AI summary

Disclosed herein is an Artificial Intelligence (AI) processor. The AI processor includes multiple NVM AI cores for respectively performing basic unit operations required for a deep-learning operation based on data stored in NVM; SRAM for storing at least some of the results of the basic unit operations; and an AI core for performing an accumulation operation on the results of the basic unit operation.