Memristive Neural Network Engine Using CMOS Charge-Trap Transistors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural networks require significant computational resources and are limited by power and area constraints, making them inefficient for practical applications, especially with the saturation of transistor scaling in traditional digital computation architectures.
Innovation Solution
A memristive neural network computing engine based on CMOS-compatible charge-trap transistors (CTT) is developed, utilizing CTT devices as analog multipliers to achieve area and power reductions, with a scalable CTT multiplier array and energy-efficient analog-digital interfaces, enabling simplified mixed-signal interfaces through a sequential analog fabric.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional digital computation processors (CPU, GPU, DSP) are used to implement neural networks, then computational accuracy can be maintained, but power consumption and area occupation increase significantly
Solution Approach 1:
The patent replaces traditional digital computation mechanisms with analog computing mechanisms based on memristive devices. The analog multipliers use continuous voltage signals and resistive memory elements to perform multiplication operations, fundamentally substituting the digital logic-based computation approach with a physics-based analog approach that naturally performs parallel computations with lower power consumption.
Solution Approach 2:
The patent changes the computational paradigm from discrete digital values to continuous analog parameters. By using analog voltages and currents to represent data and performing computations through physical laws (Ohm's law, Kirchhoff's laws), the system achieves higher energy efficiency while maintaining computational accuracy through careful design of the analog-digital interface.
2Measurement precision
If traditional digital computation processors are used to implement neural networks, then computational accuracy can be maintained, but area occupation increases significantly
Solution Approach 1:
The patent merges multiple computational functions into unified analog circuits. The memristive crossbar array simultaneously performs multiplication and accumulation operations (MAC) for multiple neural network computations in parallel, eliminating the need for separate digital logic circuits for each operation and achieving significant area reduction.
Solution Approach 2:
The patent transitions from two-dimensional planar digital circuit layouts to a three-dimensional stacked architecture where memristive devices are vertically integrated with CMOS readout circuits. This vertical stacking enables dense packing of computational elements, dramatically increasing the number of MAC operations per unit area.
3Use of energy by moving object
If new materials or extra manufacture processes are introduced for memristor-based computing, then area and power reduction can be achieved, but manufacturing complexity and cost increase
Solution Approach 1:
The patent designs memristive devices that can serve multiple functions: they act as both the computational element (analog multiplier) and the memory element (weight storage). This multi-functionality eliminates the need for separate memory and processing components, allowing the system to be manufactured using standard CMOS processes without requiring additional specialized manufacturing steps.
Solution Approach 2:
The patent modifies existing CMOS transistor structures to create memristive devices with charge-trapping characteristics by adjusting gate dielectric parameters and applying specific voltage programming sequences. This approach leverages standard CMOS fabrication processes while achieving the desired memristive behavior through parameter optimization rather than introducing entirely new materials or processes.
4Productivity
If transistor scaling is continued in traditional digital architectures, then computational density can be increased, but physical limits are reached leading to saturation
Solution Approach 1:
The patent substitutes transistor-based digital logic with memristive analog computing elements that do not rely on transistor scaling for increased density. The memristive crossbar architecture achieves computational density through the parallelism inherent in the crossbar structure, where N×M devices can perform N×M multiplications simultaneously, bypassing the sequential limitations of transistor-based logic.
Solution Approach 2:
The patent segments the computational task into independent multiply-accumulate operations that can be executed in parallel across the memristive crossbar array. Each intersection in the crossbar performs an independent multiplication, and column summing circuits aggregate results, enabling fine-grained parallelism that scales with device count without increasing interconnect complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution achieves a 100× reduction in area and power consumption compared to digital computation, with a proof-of-concept 784×784 CTT computing engine demonstrating 76.8 TOPS at 500 MHz and 14.8 mW power consumption, while maintaining performance comparable to state-of-the-art fully connected neural networks in tasks like handwritten digit recognition.
Implementation Method 1
a gate dielectric including a trap layer... storing a charge in the trap layer of the gate dielectric of the CTT device, thereby modulating a threshold voltage of the CTT device
Implementation Method 2
CTT devices are used as analog multipliers... each CTT element representing a neuron... CTT elements perform computations of a fully connected (FC) neural network
Data Source
AI summary
A neural network computing engine having an array of charge-trap-transistor (CTT) elements which are utilized as analog multipliers with all weight values preprogrammed into each CTT element as a CTT threshold voltage, with multiplicator values received from the neural network inference mode. The CTT elements perform computations of a fully connected (FC) neural network with each CTT element representing a neuron. Row resistors for each row of CTT element sum output currents as partial summation results. Counted pulse generators write weight values under control of a pulse generator controller. A sequential analog fabric (SAF) feeds multiple drain voltages in parallel to the CTT array to enable parallel analog computations of neurons. Partial summation results are read by an analog-to-digital converter (ADC).


