In-Memory Computing Engine for Low-Delay Matrix-Vector Multiplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing architectures face inefficiencies in energy and delay due to the separation of memory and compute operations, particularly in large-scale matrix-vector multiplications common in machine learning and artificial intelligence, where accessing data from memory dominates energy and delay costs.
Innovation Solution
An in-memory computing architecture that integrates a reshaping buffer, a compute-in-memory array, analog-to-digital converter circuitry, and control circuitry to perform multi-bit computing operations using single-bit internal circuits, enabling bit-parallel/bit-serial operations and efficient analog-to-digital conversion, thereby reducing energy and delay costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is accessed from memory for large-scale matrix-vector multiplications, then computation can be performed, but energy consumption and delay increase significantly
Solution Approach 1:
The patent merges memory and compute operations by implementing in-memory computing where matrix elements are stored in the memory array and vector elements are broadcast in parallel fashion over the memory array, enabling compute operations within memory bit-cells to provide results as charge using voltage-to-charge conversion via a capacitor
Solution Approach 2:
The patent introduces a reshaping buffer as an intermediary component that reshapes a sequence of received data words to form massively parallel bit-wise input signals, and uses ADC circuitry as an intermediary to convert the parallel bit-wise input signals and accumulation signals into multi-bit output words
2Productivity
If data is accessed from memory for large-scale matrix-vector multiplications, then computation can be performed, but delay increases significantly
Solution Approach 1:
The patent merges memory and compute operations by implementing in-memory computing where matrix elements are stored in the memory array and vector elements are broadcast in parallel fashion over the memory array, enabling compute operations within memory bit-cells to provide results as charge using voltage-to-charge conversion via a capacitor
Solution Approach 2:
The patent enables continuous computation by performing matrix-vector multiplication operations directly within the memory array without data movement, where bit-cell circuits involve appropriate switching of a local capacitor in a given bit-cell, where that local capacitor is also appropriately coupled to other bit-cell capacitors, to yield an aggregated compute result across the coupled bit-cells
3Use of energy by moving object
If single-bit internal circuits are used for in-memory computing, then energy efficiency is improved, but computing precision is limited
Solution Approach 1:
The patent changes the parameter representation by using ADC circuitry to convert parallel bit-wise input signals and accumulation signals into multi-bit output words, allowing single-bit internal circuits to achieve multi-bit computing precision through analog-to-digital conversion
Solution Approach 2:
The patent substitutes digital bit-wise operations with analog voltage-to-charge conversion operations in the memory bit-cells, where compute operations within memory bit-cells provide their results as charge, typically using voltage-to-charge conversion via a capacitor
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables highly efficient linear algebra computations by amortizing memory access energy and delay, supporting large-dimensional operations with controlled quantization noise, and improving signal-to-quantization noise ratio, making it suitable for broad applications in machine learning and artificial intelligence.
Implementation Method 1
compute operations within memory bit-cells provide their results as charge, typically using voltage-to-charge conversion via a capacitor
Implementation Method 2
bit-cell circuits involve appropriate switching of a local capacitor in a given bit-cell, where that local capacitor is also appropriately coupled to other bit-cell capacitors, to yield an aggregated compute result across the coupled bit-cells
Data Source
AI summary
Various embodiments comprise systems, methods, architectures, mechanisms or apparatus for providing programmable or pre-programmed in-memory computing operations.


