Analog FLASH Computing Array for In-Memory Neural Network Acceleration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional computing architectures for deep neural networks face limitations in energy efficiency and hardware consumption due to data transmission bottlenecks between the central processing unit (CPU) and memory, restricting computing speed and increasing energy consumption.

Innovation Solution

A deep neural network based on an analog FLASH computing array that utilizes FLASH cells to perform matrix-vector multiplication operations, reducing the need for analog-to-digital or digital-to-analog conversion circuits and optimizing energy and hardware resource utilization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If conventional computing architecture is used with CPU and memory separation, then data transmission between processing and storage is enabled, but computing speed is restricted and energy consumption increases due to data transmission bottleneck

Engineering Contradiction:
Improvecomputing speedVSAvoidenergy consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent merges the memory storage function and computing processing function into a single integrated structure. FLASH cells that store data are directly connected to computing circuits, eliminating the separate memory-CPU architecture. This allows data to be processed in-place without transmission, simultaneously improving computing speed and reducing energy consumption associated with data movement between separate components.

Inventive Principle:
Principle #5Merging (Combining)

2Adaptability or versatility

If conventional computing architecture with separate CPU and memory is used, then data storage and processing are enabled, but hardware requirements and hardware consumption become very huge

Engineering Contradiction:
Improvedeep neural network operation capabilityVSAvoidhardware overhead
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The FLASH computing array serves multiple functions simultaneously: it acts as memory storage, performs analog-to-digital conversion, executes matrix-vector multiplication operations, and supports various deep neural network layer types (convolutional, pooling, fully connected). This multi-functional integration reduces hardware overhead while maintaining adaptability for different deep neural network operations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces traditional digital computing mechanics with analog computing using FLASH cell threshold voltages. Instead of moving data between separate memory and processing units, the system uses analog voltage representations and FLASH cell electrical characteristics to perform computations directly in the memory array, significantly reducing hardware complexity for deep neural network operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The solution significantly improves energy efficiency and reduces hardware overhead by enabling efficient matrix-vector multiplication operations within the neural network, enhancing the hardware realization of artificial intelligence systems.

Implementation Method 1

When a reasonable gate bias is applied to the FLASH cell and the drain-source voltage Vds is less than a specific value, the drain current Id of the FLASH cell and the drain-source voltage Vds may illustrate an approximately linear growth relationship

Methodology Applied
Scientific EffectOhm's Law: Ohm's Law

Data Source

PatentUS12619863B2Deep neural network based on flash analog flash computing array
Publication Date: 2026.05.05 PEKING UNIV
  • US12619863B2 patent drawing
  • US12619863B2 patent drawing
  • US12619863B2 patent drawing

AI summary

A deep neural network based on analog FLASH computing array, includes a number of computing arrays, a number of subtractors, a number of activation circuit units and a number of integral-recognition circuit units. The computing array includes a number of computing units, a number of word lines, a number of bit lines and a number of source lines. Each of the computing units includes a FLASH cell. The gate electrodes of the FLASH cells in the same column are connected to the same word line. The source electrodes of the FLASH cells in the same column are connected to the same source line, and the drain electrodes of the FLASH cells in the same row are connected to the same bit line. Each of the subtractors includes a positive terminal, a negative terminal and an output terminal.