CNN Memory Subsystem Using VCMA-MRAM for Real-Time AI Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing CNN-based ICs for AI are inefficient due to slow computational speed and high costs, particularly when processing large amounts of imagery data, as they often rely on software solutions or hardware designed for general computation, which are impractical for real-time image processing.

Innovation Solution

A CNN-based digital IC with multiple processing units, each coupled to a memory subsystem containing magnetic random access memory (MRAM) cells with voltage-controlled magnetic anisotropy (VCMA) based magnetic tunnel junction (MTJ) elements for storing filter coefficients and imagery data, optimizing memory-in-processor architecture for low power consumption and high read/write speed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If software solutions or general-purpose hardware are used for CNN processing, then versatility is maintained, but computational speed becomes too slow and cost becomes too high for real-time image processing

Engineering Contradiction:
Improvecomputational speedVSAvoidsoftware solution flexibility
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent replaces general-purpose software-based CNN processing with specialized hardware architecture. Specifically, it substitutes software execution on general-purpose processors with dedicated CNN processing units that have fixed functionality optimized for convolution operations, achieving real-time processing speeds while maintaining acceptable versatility through configurable filter banks and processing modes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the operational parameters of the processing system by transitioning from software execution to hardware implementation. This involves changing the computational paradigm from sequential software instructions to parallel hardware circuits, and from general-purpose computing to specialized AI processing, thereby achieving the required speed improvement for real-time image processing.

Inventive Principle:
Principle #35Parameter changes

2Speed

If data is stored far from CNN processing logic, then memory capacity is sufficient, but read/write speed becomes too slow for efficient processing

Engineering Contradiction:
Improveread/write speedVSAvoidmemory capacity
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent implements a hierarchical memory architecture where small, fast SRAM buffers are nested within each CNN processing unit for immediate data access, while larger capacity memories are nested at higher levels of the hierarchy. This nested structure allows frequently accessed filter coefficients and input data to be stored in close proximity to processing logic, achieving high read/write speeds without sacrificing overall memory capacity.

Inventive Principle:
Principle #7Nested doll (Nesting)

Solution Approach 2:

The patent resolves the speed-capacity tradeoff by adding a spatial dimension to the memory architecture. Instead of relying solely on distance from processing logic, it creates a multi-dimensional memory hierarchy with different levels of cache and storage, allowing data to be accessed from multiple locations and paths, thereby achieving both high speed and sufficient capacity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If filter coefficients are stored in the same memory as imagery data, then memory structure is simplified, but memory access efficiency decreases due to different access patterns

Engineering Contradiction:
Improvememory access efficiencyVSAvoidmemory subsystem structure
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the memory subsystem into functionally distinct components: filter coefficient memory, input data memory, and output buffer memory. Each segment is optimized for its specific access pattern - filter coefficients are stored in associative memory for parallel lookup, while input data is stored in sequential memory structures for row-by-row processing. This segmentation improves memory access efficiency despite increased structural complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different memory technologies and access mechanisms to different data types based on their specific requirements. Filter coefficients, which require parallel random access, are stored in associative memory or ROM, while imagery data requiring sequential access is stored in SRAM buffers. This local optimization of memory quality for each data type enhances overall memory access efficiency.

Inventive Principle:
Principle #3Local quality

4Use of energy by moving object

If conventional memory is used in CNN processing units, then manufacturing is easier, but power consumption increases and read/write speed decreases

Engineering Contradiction:
Improvepower consumptionVSAvoidmemory integration difficulty
Core Design Contradiction:
Use of energy by moving objectVSEase of manufacture

Solution Approach 1:

The patent employs a composite memory architecture that integrates multiple memory technologies - SRAM for fast temporary storage, DRAM for intermediate buffering, and non-volatile memory for filter coefficient storage. Each memory type is selected for its optimal power and speed characteristics, creating a composite system that achieves low power consumption while remaining manufacturable through standard semiconductor fabrication processes.

Inventive Principle:
Principle #40Composite materials

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution enables efficient processing of large input signals by providing dedicated memory for filter coefficients and imagery data, enhancing computational speed and reducing power consumption, making it suitable for real-time AI applications like image processing.

Implementation Method 1

Each MRAM cell contains a voltage-controlled magnetic anisotropy (VCMA) based magnetic tunnel junction (MTJ) element. Magnetization direction in VCMA based MTJ element can be in-plane or out-of-plane.

Methodology Applied
Scientific EffectVoltage-controlled magnetic anisotropy (VCMA):

Data Source

PatentUS10534996B2Memory subsystem in CNN based digital IC for artificial intelligence
Publication Date: 2020.01.14 GYRFALCON TECHNOLOGY INC
  • US10534996B2 patent drawing
  • US10534996B2 patent drawing
  • US10534996B2 patent drawing

AI summary

CNN (Cellular Neural Networks or Cellular Nonlinear Networks) based digital Integrated Circuit for artificial intelligence contains multiple CNN processing units. Each CNN processing unit contains CNN logic circuits operatively coupling to a memory subsystem having first and second memories. The first memory contains magnetic random access memory (MRAM) cells for storing weights (e.g., filter coefficients) while the second memory is for storing input signals (e.g., imagery data). The first memory may store one-time-programming weights or filter coefficients. The memory subsystem may contain a third memory that contains MRAM cells for storing one-time-programming data for security purpose. The second memory contains MRAM cells or static random access memory cells. Each MRAM cell contains a voltage-controlled magnetic anisotropy (VCMA) based magnetic tunnel junction (MTJ) element. Magnetization direction in VCMA based MTJ element can be in-plane or out-of-plane.