Flash VMM Arrays Using Control Gates for Accurate In-Memory Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital hardware struggles to achieve efficient performance in implementing Machine Learning or Deep Neural Networks due to high computational complexity, especially in Vector-by-Matrix Multiplication (VMM) operations, as they require large systems and numerous operations to process inputs effectively.
Innovation Solution
The implementation of Vector-by-Matrix Multiplication using analog circuitry with non-volatile memory devices, specifically multi-gate flash transistors, where the control gate or a combination of control and word line is used instead of the word line alone for improved performance, and the re-routing of gates responsible for erasing or sourcing across different rows for faster programming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If digital hardware (CPUs/GPUs) is used to implement Machine Learning and Deep Neural Networks, then the system can perform computations, but the computational complexity becomes unmanageably high and performance is insufficient
Solution Approach 1:
The patent replaces digital computational systems (CPUs/GPUs) with an analog neuromorphic system using nonvolatile memory devices. The analog circuitry directly performs Vector-by-Matrix Multiplication operations that mimic neural network computations, substituting the mechanical/digital computation process with an analog physical process that naturally performs the required mathematical operations through current flow and conductance modulation.
Solution Approach 2:
The patent changes the operational parameters from digital discrete states to analog continuous states. By utilizing the analog conductance states of nonvolatile memory devices, the system can represent and process neural network weights and activations as continuous physical quantities, enabling direct hardware acceleration of neural network operations without the overhead of digital computation.
2Measurement precision
If analog circuitry with nonvolatile memory devices is used for in-memory computation, then Vector-by-Matrix Multiplication accuracy improves, but device programming complexity increases
Solution Approach 1:
The patent segments the programming process into distinct phases: initialization phase (setting default conductance states), training phase (programming weight values during network training), and inference phase (using programmed weights for computation). This segmentation allows complex programming tasks to be broken down into manageable steps, with each phase having specific programming requirements and optimization strategies.
Solution Approach 2:
The patent implements preliminary initialization of nonvolatile memory devices to known conductance states before programming. This preliminary action ensures that devices start from a consistent baseline state, simplifying subsequent programming operations and improving the reliability of weight programming during neural network training. The initialization step prepares the memory array for efficient weight programming without requiring complex in-situ calibration.
3Productivity
If erase gates and source lines are re-routed in a zigzag form across different rows, then array programming density and speed improve, but circuit routing complexity increases
Solution Approach 1:
The patent introduces a zigzag routing pattern that utilizes both horizontal and vertical dimensions for signal distribution. Instead of simple row-by-row or column-by-column routing, the erase gates and source lines are routed in a zigzag pattern that efficiently covers multiple rows and columns, improving programming throughput by enabling parallel operations across different memory regions while distributing the routing complexity systematically.
4Measurement precision
If control gate and word line are used together instead of word line alone for computation, then VMM performance and accuracy improve, but device structure complexity increases
Solution Approach 1:
The patent makes the nonvolatile memory devices multi-functional by utilizing both the control gate and word line for computational purposes. The control gate is used for weight programming and the word line for both programming and computation operations. This multi-functionality allows the same hardware structure to perform multiple operations (programming, erasing, and analog computation) without requiring separate dedicated circuits for each function, thereby improving VMM accuracy while minimizing additional device complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enhances the accuracy and speed of VMM operations, allowing for more efficient and accurate computation in larger neural networks by leveraging the properties of non-volatile memory devices, particularly through the use of multi-gate flash transistors and innovative programming protocols.
Implementation Method 1
a second gate comprising a floating gate disposed over a second portion of the channel, the floating gate controlling a second conductivity of the second portion in response to an amount of charge (electrons or holes) stored on the floating gate
Implementation Method 2
a first gate disposed over a first portion of the channel and insulated from the first portion of the channel, the first gate controlling a first conductivity of the first portion in response to a first voltage applied to the first gate
Implementation Method 3
a third gate comprising a gate coupled to the floating gate so as to control the amount of charge transferred to the floating gate during programming
Implementation Method 4
The FETs are operated in a subthreshold regime
Data Source
AI summary
Building blocks for implementing Vector-by-Matrix Multiplication (VMM) are implemented with analog circuitry including non-volatile memory devices (flash transistors) and using in-memory computation. In one example, improved performance and more accurate VMM is achieved in arrays including multi-gate flash transistors when computation uses a control gate or the combination of control gate and word line (instead of using the word line alone). In another example, very fast weight programming of the arrays is achieved using a novel programming protocol. In yet another example, higher density and faster array programming is achieved when the gate(s) responsible for erasing devices, or the source line, are re-routed across different rows, e.g., in a zigzag form. In yet another embodiment a neural network is provided with nonlinear synaptic weights implemented with nonvolatile memory devices.


