Parallel Processor Associative Memory In-Memory Computing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional von Neumann-based computing systems face a 'memory wall' due to the growing gap between computing performance and memory access times, exacerbated by applications like machine learning and 5G wireless systems that require high data processing rates and energy efficiency.
Innovation Solution
A parallel processor in associative content-addressable memory (PPAC) is introduced, which performs computation directly in memory using a fully-digital standard-cell-based CMOS architecture, supporting matrix-vector-product operations and other tasks like low-precision neural networks and cryptography, thereby improving throughput and energy efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional von Neumann architecture is used with separate processor and memory, then system flexibility is maintained, but memory access time and energy consumption increase significantly
Solution Approach 1:
The patent merges the processor and memory into a unified architecture where processing elements are directly integrated with memory cells. Each memory bit cell contains logic operators that can perform computations directly on stored data, eliminating the need to transfer data between separate processor and memory components. This integration resolves the memory wall problem by reducing memory access time while maintaining system versatility through configurable logic operators.
Solution Approach 2:
The patent introduces sense amplifiers as intermediary components that bridge the gap between memory cells and processing elements. These sense amplifiers not only read data from memory but also facilitate in-memory computations by enabling logic operations directly at the memory location, thereby reducing the time and energy required for data access and processing.
2Device complexity
If traditional von Neumann architecture is used with separate processor and memory, then system control and logic circuitry are centralized, but energy efficiency deteriorates
Solution Approach 1:
The patent segments the centralized processor into distributed processing elements, with each memory bit cell containing its own logic operators. This segmentation allows computations to be performed locally at the memory location rather than requiring data to be moved to a centralized processor, significantly reducing energy consumption associated with data transfer and centralized processing operations.
Solution Approach 2:
The patent enables memory cells to serve themselves by performing computations directly on the data stored within them. Each bit cell is equipped with logic operators that can execute operations on the stored bits without requiring external processing, thereby eliminating the energy overhead of moving data between memory and processor and improving overall system energy efficiency.
3Productivity
If computation is moved into memory for processing in memory (PIM), then throughput and energy efficiency improve, but device complexity increases
Solution Approach 1:
The patent implements universal logic operators within each memory bit cell that can perform multiple functions including read operations, various logic operations (AND, OR, XOR, XNOR), and arithmetic operations. This multi-functionality allows a single in-memory processing structure to handle diverse computational tasks, improving throughput without proportionally increasing device complexity, as the same hardware resources serve multiple purposes.
Data Source
AI summary
A parallel processor in associative content-addressable memory (PPAC) is provided. Processing in memory (PIM) moves computation into memories with the goal of improving throughput and energy-efficiency compared to traditional von Neumann-based architectures. Most existing PIM architectures are either general-purpose but only support atomistic operations, or are specialized to accelerate a single task. The PPAC described herein provides a novel in-memory accelerator that supports a range of matrix-vector-product (MVP)-like operations that find use in traditional and emerging applications. PPAC is, for example, able to accelerate low-precision neural networks, exact/approximate hash lookups, cryptography, and forward error correction. The fully-digital nature of PPAC enables its implementation with standard-cell-based complementary metal-oxide-semiconductor (CMOS), which facilitates automated design and portability among technology nodes. A comparison with recent digital and mixed-signal PIM accelerators reveals that PPAC is competitive in terms of throughput and energy-efficiency, while accelerating a wide range of applications and simplifying development.


