NOR Flash Array Convolution via Voltage Modulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for performing convolution operations, such as using general-purpose processors or GPUs, face inefficiencies in computational complexity and hardware resource utilization, which hinder real-time data processing, especially in applications like image recognition and artificial intelligence.

Innovation Solution

A method and device utilizing an NOR flash array to perform convolution operations by storing convolution kernel elements in NOR flash cells, applying voltages to gate and source terminals to generate output currents, and aggregating these currents through operational amplifiers to determine the convolution result, enabling efficient multiplication and parallel execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional general-purpose processors (CPUs) are used for convolution calculation, then the system is simple and flexible, but computational efficiency is low and processing speed is slow

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidhardware complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces conventional electronic computing systems (CPUs, GPUs) with a magnetic computing system using TMR effect-based crossbar arrays. Magnetic memory cells and spin-torque mechanisms substitute traditional electronic transistors and capacitors, enabling parallel multiplication-accumulation operations through magnetoresistive switching and spin-current modulation, thereby achieving high computational efficiency for convolution operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent implements a two-dimensional crossbar array architecture where weight parameters are stored in magnetic memory cells at the intersections of word lines and bit lines. This spatial arrangement enables simultaneous parallel computation across multiple dimensions, with input signals propagating through the crossbar network to perform matrix multiplication in a single operation cycle, dramatically improving computational throughput.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If GPU is used to implement convolution, then computational speed improves, but hardware overhead resources increase significantly

Engineering Contradiction:
Improveprocessing speedVSAvoidhardware overhead
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The magnetic crossbar array is designed as a universal computing platform that can perform various neural network operations (convolution, fully-connected layers, transpose operations) by reconfiguring input patterns and weight distributions. The same hardware structure handles different computational tasks without requiring dedicated specialized circuits for each operation type, reducing overall hardware overhead while maintaining high processing speed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The computational task is segmented into parallel operations across multiple independent memory cells in the crossbar array. Each cell performs an individual multiplication operation, and the aggregate result emerges from the parallel summation of currents from all active cells, achieving high throughput without requiring complex sequential control logic or large buffer memory structures.

Inventive Principle:
Principle #1Segmentation

3Speed

If real-time data processing is required, then processing speed must be high, but computational complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoidcomputational complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

Weight parameters are pre-loaded into the magnetic memory cells of the crossbar array before computation begins. This preliminary configuration allows the system to perform real-time inference operations at high speed without requiring complex dynamic weight updates during processing, simplifying the computational complexity while maintaining high processing speed for real-time applications.

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach achieves high calculation efficiency and meets real-time data processing requirements by leveraging the NOR flash array's ability to perform multiplication and execute convolution operations efficiently, reducing hardware overhead.

Implementation Method 1

if the element of convolution kernel matrix is 1 or −1, the corresponding NOR flash cell is programmed to low threshold voltage; and if the element of the convolution kernel matrix is 0, the corresponding NOR flash cell is programmed to high threshold voltage

Methodology Applied
Scientific EffectThreshold voltage modulation:

Implementation Method 2

applying voltages corresponding to elements of an input matrix to gate terminals of the NOR flash cells while applying a driving voltage to source terminals of the NOR flash cells, such that currents are output, at drain terminals of the NOR flash cells

Methodology Applied
Scientific EffectField effect transistor current modulation:

Data Source

PatentUS11309026B2Convolution operation method based on NOR flash array
Publication Date: 2022.04.19 PEKING UNIV
  • US11309026B2 patent drawing
  • US11309026B2 patent drawing
  • US11309026B2 patent drawing

AI summary

The present disclosure relates to the field of semiconductor integrated circuits and manufacturing technologies thereof, and discloses a method and device for realizing a convolution operation based on an NOR flash storage structure. The NOR flash array has a structure of an array composed of a plurality of NOR flash cells. The convolution operation method includes: storing elements of a convolution kernel matrix into the NOR flash cells; converting elements of an input matrix into voltages and applying the voltages to gate terminals of the NOR flash cells; applying a driving voltage to source terminals of the NOR flash cells; and collecting, via drain terminals of the NOR flash cells, current values of each column to obtain a convolution operation result.