Hybrid Compute-in-Memory Architecture for Neural Network Filters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Compute-in-memory architectures face challenges in achieving high-speed processing while minimizing power consumption, particularly in applications like machine learning, where traditional digital architectures are preferred due to the limitations of analog-to-digital converters (ADCs) in achieving both speed and precision.
Innovation Solution
A hybrid compute-in-memory architecture that uses a capacitor with multiple switches to perform multiplication operations in parallel phases, allowing for efficient charging of capacitor plates and subsequent accumulation, thereby reducing the need for high-resolution ADCs and minimizing power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If compute-in-memory architecture is used to reduce power consumption, then power consumption is reduced, but operating speed decreases due to ADC limitations
Solution Approach 1:
The patent segments the multiplication operation into two separate phases: first multiplying activation bits with filter weight bits to charge the first plate, then multiplying complement activation bits with complement filter weight bits to charge the second plate. This segmentation allows parallel processing of multiple bit combinations simultaneously, increasing operating speed while maintaining the analog compute-in-memory architecture's power efficiency.
Solution Approach 2:
The patent utilizes both plates of the capacitor in a dual-rail configuration, adding a dimensional aspect to the charge accumulation. By charging both the first plate and second plate simultaneously in parallel phases, the system effectively doubles the computational throughput without requiring higher-resolution ADCs, thus maintaining power savings while improving speed.
2Measurement precision
If high-resolution ADC is used to achieve same precision as traditional digital computing, then precision is improved, but operating speed decreases and power consumption increases
Solution Approach 1:
Instead of using a high-resolution ADC to capture all possible voltage combinations, the patent performs partial multiplications in parallel phases, charging each capacitor plate to represent specific bit combinations. This partial action approach achieves sufficient precision for machine learning applications without requiring the overhead of high-resolution conversion, maintaining both speed and power efficiency.
3Measurement precision
If conventional ADC methods are used to convert accumulated charge, then conversion is achieved, but operating speed slows and power consumption increases
Solution Approach 1:
The patent extracts the multiplication result directly into the capacitor charge state without requiring full ADC conversion. By using the capacitor voltage to directly represent the partial sum of multiplications, the system eliminates the need for power-consuming ADC operations while maintaining the necessary precision for subsequent digital processing, thus reducing ADC power consumption significantly.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables fast computation speeds comparable to traditional digital computers while maintaining the power savings of compute-in-memory architectures, by efficiently accumulating charge across capacitors and reducing the dynamic range required for analog-to-digital conversion.
Implementation Method 1
a capacitor including a first plate and a second plate
Implementation Method 2
a first switch configured to close responsive to a first activation bit signal; a second switch coupled in series with the first switch between the voltage source and the first plate
Data Source
AI summary
A compute-in-memory array is provided that implements a filter for a layer in a neural network. The filter multiplies a plurality of activation bits by a plurality of filter weight bits for each channel in a plurality of channels through a charge accumulation from a plurality of capacitors. The accumulated charge is digitized to provide the output of the filter.


