In-Memory Neural VMM Circuit for Lower Data Transfer Power
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neuromorphic processors face challenges in efficiently performing neural network operations due to high power consumption and large data transmission times, especially in systems with separated computational units and memory, which limits their performance and efficiency.
Innovation Solution
An in-memory computing circuit is introduced, integrating computational units and memory to perform vector-by-matrix multiplication operations using crossbar arrays, sharing read-write circuits, decoders, and analog-to-digital converters, and incorporating a control mechanism to optimize spatial efficiency and reduce power consumption.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If computational units and memory are separated in neuromorphic processors, then device architecture is simpler and easier to manufacture, but data transmission time increases and power consumption rises
Solution Approach 1:
The patent merges computational units and memory into an integrated in-memory computing circuit where crossbar arrays perform vector-by-matrix multiplication operations directly within the memory structure, eliminating the need for separate data transmission between computational units and memory, thereby reducing data transmission time while maintaining manufacturing feasibility
Solution Approach 2:
The crossbar array structure serves multiple functions simultaneously: it acts as memory for storing weight data, as computational unit for performing multiplication operations, and as data transmission medium, thereby eliminating the need for separate dedicated memory and computational components while reducing overall data transmission time
2Ease of manufacture
If computational units and memory are separated in neuromorphic processors, then device architecture is simpler and easier to manufacture, but power consumption increases
Solution Approach 1:
The patent merges computational units and memory into an integrated in-memory computing circuit where crossbar arrays perform vector-by-matrix multiplication operations directly within the memory structure, eliminating the need for separate data transmission between computational units and memory, thereby reducing power consumption while maintaining manufacturing feasibility
Solution Approach 2:
The patent extracts the computational function from separate processing units and integrates it directly into the memory structure through crossbar arrays, eliminating unnecessary data movement and associated power consumption while keeping the device architecture manufacturable
3Productivity
If multiple processing circuits are used to perform neural network operations, then processing capacity and productivity increase, but data transmission time and power consumption increase
Solution Approach 1:
The patent merges multiple processing circuits into a unified in-memory computing architecture where crossbar arrays perform parallel vector-by-matrix multiplication operations, increasing processing capacity while eliminating the need for data transmission between separate processing units, thereby reducing data transmission time
4Productivity
If multiple processing circuits are used to perform neural network operations, then processing capacity and productivity increase, but power consumption increases
Solution Approach 1:
The patent merges multiple processing circuits into a unified in-memory computing architecture where crossbar arrays perform parallel vector-by-matrix multiplication operations, increasing processing capacity while eliminating the need for data transmission between separate processing units, thereby reducing power consumption
Data Source
AI summary
A neural network apparatus includes: a first processing circuit and a second processing circuit each configured to perform a vector-by-matrix multiplication (VMM) operation on a weight and an input activation; a first register configured to store an output of the first processing circuit; an adder configured to add an output of the first register and an output of the second processing circuit; a second register configured to store an output of the adder; and an input circuit configured to input a same input activation to the first processing circuit and the second processing circuit and control the first processing circuit and the second processing circuit.


