Flash VMM Arrays Using Control Gates for Accurate In-Memory Computing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital hardware struggles to achieve efficient performance in implementing Machine Learning or Deep Neural Networks due to high computational complexity, especially in Vector-by-Matrix Multiplication (VMM) operations, as they require large systems and numerous operations to process inputs effectively.

Innovation Solution

The implementation of Vector-by-Matrix Multiplication using analog circuitry with non-volatile memory devices, specifically multi-gate flash transistors, where the control gate or a combination of control and word line is used instead of the word line alone for improved performance, and the re-routing of gates responsible for erasing or sourcing across different rows for faster programming.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If digital hardware (CPUs/GPUs) is used to implement Machine Learning and Deep Neural Networks, then the system can perform computations, but the computational complexity becomes unmanageably high and performance is insufficient

Engineering Contradiction:
Improvecomputational performanceVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces digital computational systems (CPUs/GPUs) with an analog neuromorphic system using nonvolatile memory devices. The analog circuitry directly performs Vector-by-Matrix Multiplication operations that mimic neural network computations, substituting the mechanical/digital computation process with an analog physical process that naturally performs the required mathematical operations through current flow and conductance modulation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the operational parameters from digital discrete states to analog continuous states. By utilizing the analog conductance states of nonvolatile memory devices, the system can represent and process neural network weights and activations as continuous physical quantities, enabling direct hardware acceleration of neural network operations without the overhead of digital computation.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If analog circuitry with nonvolatile memory devices is used for in-memory computation, then Vector-by-Matrix Multiplication accuracy improves, but device programming complexity increases

Engineering Contradiction:
ImproveVMM accuracyVSAvoidprogramming complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the programming process into distinct phases: initialization phase (setting default conductance states), training phase (programming weight values during network training), and inference phase (using programmed weights for computation). This segmentation allows complex programming tasks to be broken down into manageable steps, with each phase having specific programming requirements and optimization strategies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements preliminary initialization of nonvolatile memory devices to known conductance states before programming. This preliminary action ensures that devices start from a consistent baseline state, simplifying subsequent programming operations and improving the reliability of weight programming during neural network training. The initialization step prepares the memory array for efficient weight programming without requiring complex in-situ calibration.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If erase gates and source lines are re-routed in a zigzag form across different rows, then array programming density and speed improve, but circuit routing complexity increases

Engineering Contradiction:
Improveprogramming speedVSAvoidrouting complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces a zigzag routing pattern that utilizes both horizontal and vertical dimensions for signal distribution. Instead of simple row-by-row or column-by-column routing, the erase gates and source lines are routed in a zigzag pattern that efficiently covers multiple rows and columns, improving programming throughput by enabling parallel operations across different memory regions while distributing the routing complexity systematically.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Measurement precision

If control gate and word line are used together instead of word line alone for computation, then VMM performance and accuracy improve, but device structure complexity increases

Engineering Contradiction:
ImproveVMM accuracyVSAvoidtransistor structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the nonvolatile memory devices multi-functional by utilizing both the control gate and word line for computational purposes. The control gate is used for weight programming and the word line for both programming and computation operations. This multi-functionality allows the same hardware structure to perform multiple operations (programming, erasing, and analog computation) without requiring separate dedicated circuits for each function, thereby improving VMM accuracy while minimizing additional device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enhances the accuracy and speed of VMM operations, allowing for more efficient and accurate computation in larger neural networks by leveraging the properties of non-volatile memory devices, particularly through the use of multi-gate flash transistors and innovative programming protocols.

Implementation Method 1

a second gate comprising a floating gate disposed over a second portion of the channel, the floating gate controlling a second conductivity of the second portion in response to an amount of charge (electrons or holes) stored on the floating gate

Methodology Applied
Scientific EffectCharge storage on floating gate: Capacitance

Implementation Method 2

a first gate disposed over a first portion of the channel and insulated from the first portion of the channel, the first gate controlling a first conductivity of the first portion in response to a first voltage applied to the first gate

Methodology Applied
Scientific EffectField effect transistor conduction control: Electric Field

Implementation Method 3

a third gate comprising a gate coupled to the floating gate so as to control the amount of charge transferred to the floating gate during programming

Methodology Applied
Scientific EffectCharge transfer control: Electric Field

Implementation Method 4

The FETs are operated in a subthreshold regime

Methodology Applied
Scientific EffectSubthreshold conduction: Conduction (electrical)

Data Source

PatentUS12106211B2Mixed signal neuromorphic computing with nonvolatile memory devices
Publication Date: 2024.10.01 RGT UNIV OF CALIFORNIA
  • US12106211B2 patent drawing
  • US12106211B2 patent drawing
  • US12106211B2 patent drawing

AI summary

Building blocks for implementing Vector-by-Matrix Multiplication (VMM) are implemented with analog circuitry including non-volatile memory devices (flash transistors) and using in-memory computation. In one example, improved performance and more accurate VMM is achieved in arrays including multi-gate flash transistors when computation uses a control gate or the combination of control gate and word line (instead of using the word line alone). In another example, very fast weight programming of the arrays is achieved using a novel programming protocol. In yet another example, higher density and faster array programming is achieved when the gate(s) responsible for erasing devices, or the source line, are re-routed across different rows, e.g., in a zigzag form. In yet another embodiment a neural network is provided with nonlinear synaptic weights implemented with nonvolatile memory devices.