Processing-in-Memory Dot Product Circuits for Low-Latency AI Compute
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computer-based computations for operations like forward and backward propagation in neural networks are processor and memory intensive, requiring extensive data transfer between compute cores and memory arrays, which can be inefficient and power-consuming.
Innovation Solution
Implementing processing-in-memory (PIM) operations within a memory device, where a processor is integrated near or on the same chip as the memory array, allowing for dot product operations to be performed internally without external data transfer, using a memory array with bit lines, word lines, and summing circuits to generate products of numbers stored along these lines.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional computer-based computations are used with separate processors and memory arrays, then data transfer between compute cores and memory arrays can be performed, but the system becomes inefficient and power-consuming due to extensive external communication
Solution Approach 1:
The patent merges the processor and memory array into a single integrated device, allowing computational operations to be performed directly within the memory device. This integration eliminates the need for extensive data transfer between separate compute cores and memory arrays, thereby improving computational efficiency and reducing power consumption associated with external communication.
Solution Approach 2:
The patent introduces summing circuits as intermediary components within the memory device that enable arithmetic operations (such as dot product computations) to be performed directly on data stored in the memory array. These summing circuits act as mediators that facilitate processing-in-memory operations, reducing the need for data to be transferred to external processors and thereby improving efficiency while reducing power consumption.
2Productivity
If data is transferred between external processors and memory arrays, then computations can be executed, but latency increases due to external communication requirements
Solution Approach 1:
By merging the processor functionality directly into the memory device, the patent eliminates the time required for data transfer between external processors and memory arrays. Computational operations are executed in-place within the memory device, significantly reducing latency and improving processing performance for tasks such as neural network computations.
Solution Approach 2:
The patent performs preliminary setup of computational operations within the memory device by configuring summing circuits and storing data in appropriate locations within the memory array. This preliminary arrangement enables computations to be executed immediately without requiring data to be fetched from external sources, thereby reducing latency.
3Adaptability or versatility
If processing operations are performed externally to the memory device, then standard computational workflows can be maintained, but the system requires extensive data transfer infrastructure
Solution Approach 1:
The patent combines processing capabilities directly within the memory device by integrating summing circuits and control logic with the memory array. This merger enables the memory device to perform computational operations such as dot product computations, thereby improving adaptability and versatility without requiring extensive external data transfer infrastructure.
Solution Approach 2:
The patent designs the memory device to serve multiple functions: it can store data in the memory array and simultaneously perform computational operations using the integrated summing circuits. This multi-functionality allows the same device to handle both memory operations and processing tasks, reducing the need for separate processing infrastructure and simplifying the overall system architecture.
Data Source
AI summary
Methods, apparatuses, and systems for in- or near-memory processing are described. Bits of a first number may be stored on a number of memory elements, wherein each memory element of the number of memory elements intersects a digit line and an access line of a number of access lines. A number of signals corresponding to bits of a second number may be driven on the number of access lines to generate a number of output signals. A value equal to a product of the first number and the second number may be generated based on the number of output signals.


