In-Memory Computing Array Layout for Energy-Efficient Matrix Multiplication
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional accelerators for machine learning and artificial intelligence applications face energy and delay inefficiencies due to the separation of memory and compute, which limits energy/delay reductions in matrix-vector multiplications, especially when dealing with large dimensionalities.
Innovation Solution
An in-memory computing architecture that integrates a reshaping buffer, a compute-in-memory array, analog-to-digital converter circuitry, and control circuitry to perform multi-bit computing operations using single-bit internal circuits, enabling bit-parallel/bit-serial operations and near-memory computing for efficient matrix-vector multiplications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If memory and compute are separated in conventional accelerators, then device architecture is simpler and easier to manufacture, but energy consumption increases and computing speed decreases for matrix-vector multiplications
Solution Approach 1:
The patent merges memory and compute operations by implementing in-memory computing where matrix elements are stored in memory arrays and computing operations are performed directly within the memory structure. This integration eliminates the need for separate memory and compute units, reducing data movement and associated energy consumption while accepting increased architectural complexity.
Solution Approach 2:
The memory array is designed to serve dual purposes: storing data and performing compute operations. The same memory structure that stores matrix elements also executes matrix-vector multiplication through in-memory computing operations, making the system more energy-efficient by eliminating redundant data transfer between separate memory and compute units.
2Speed
If memory and compute are separated in conventional accelerators, then ease of operation is improved, but computing speed decreases for large-dimensional matrix-vector multiplications
Solution Approach 1:
By combining memory storage and compute operations into a single in-memory computing structure, the patent enables faster execution of matrix-vector multiplications. The merging allows direct computation on stored data without intermediate data movement, significantly improving computing speed for large-dimensional operations despite increased operational complexity.
3Loss of energy
If in-memory computing is implemented, then energy consumption is reduced for matrix-vector multiplications, but device complexity increases
Solution Approach 1:
The patent implements in-memory computing by merging storage and compute functions within the memory array. This approach reduces energy loss by eliminating data movement between separate memory and compute units, while the integrated structure inherently increases circuit complexity compared to conventional separated architectures.
4Productivity
If in-memory computing is implemented, then productivity is improved for linear algebra computations, but device complexity increases
Solution Approach 1:
The patent achieves improved productivity for linear algebra computations by merging memory and compute operations in an in-memory computing architecture. This integration enables parallel processing of matrix-vector multiplications directly within the memory structure, significantly increasing computing throughput despite the increased system complexity required to implement the integrated architecture.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables highly efficient and energy-proportional linear algebra computations, reducing energy consumption and improving throughput by integrating compute operations directly within memory, thereby addressing the inefficiencies of conventional accelerators.
Implementation Method 1
charge, typically using voltage-to-charge conversion via a capacitor
Implementation Method 2
analog-to-digital converter (ADC) circuitry configured to process the plurality of CIM channel output signals to provide thereby a sequence of multi-bit output words
Data Source
AI summary
Various embodiments comprise systems, methods, architectures, mechanisms or apparatus for providing programmable or pre-programmed in-memory computing operations.


