Digital Bit-Serial Vector Computing Architecture in DRAM Subarrays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer systems are bottlenecked by memory access bandwidth for applications with large datasets and low computational intensity, motivating the need to integrate computational capabilities within Dynamic Random-Access Memory (DRAM).
Innovation Solution
A digital bit-serial vector computing architecture is embedded in the DRAM subarray, featuring a bit-serial logic unit per subarray column, bank-level bit-serial control logic, and rank-level processing units. This architecture decouples execution from memory row access, enabling faster bit-serial logic operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If computational capabilities are placed inside DRAM to overcome memory access bandwidth bottlenecks, then processing speed and energy efficiency improve, but device complexity increases
Solution Approach 1:
The patent segments the computational architecture into hierarchical levels: bit-serial logic units at the subarray level, bank-level control logic, and rank-level processing units. This segmentation allows computation to be distributed across multiple granularities, improving processing throughput while managing complexity through modular organization. Each segment handles specific operations, enabling parallel processing across DRAM subarrays without requiring a monolithic complex processor.
Solution Approach 2:
The patent introduces bit-serial logic units as intermediary components between the DRAM storage array and traditional processing units. These logic units serve as mediators that perform computational operations directly on data residing in DRAM, eliminating the need for continuous data transfer to external processors. This intermediary approach enables in-memory computation, improving processing speed while keeping the overall system architecture manageable through clearly defined interface boundaries.
2Loss of time
If bit-serial logic operations are integrated into DRAM subarray, then execution time is reduced compared to memory row cycles, but manufacturing precision requirements increase
Solution Approach 1:
The patent changes the operational parameters of DRAM by implementing bit-serial logic units that operate on individual bits or small bit groups rather than traditional word-level operations. This parameter change enables finer-grained control over data processing, reducing execution time by performing multiple bit-serial operations within a single memory row cycle. The approach manages manufacturing precision requirements by leveraging standard DRAM cell structures and using control logic that operates at relaxed timing compared to high-speed serial interfaces.
Solution Approach 2:
The patent introduces a temporal dimension to DRAM operations by implementing bit-serial processing that unfolds computation over multiple clock cycles within a row active period. Instead of requiring all bits to be processed simultaneously (spatial parallelism), the bit-serial approach processes bits sequentially over time, reducing the need for complex simultaneous switching and thereby lowering manufacturing precision requirements while still achieving reduced execution time through efficient utilization of the row active period.
3Productivity
If massive parallelism of DRAM row operations is leveraged, then computational throughput increases, but energy consumption increases
Solution Approach 1:
The patent applies partial action by implementing bit-serial logic units that perform computation on a subset of bits within each DRAM row rather than requiring full row activation for every operation. This allows the system to leverage DRAM's inherent parallelism by activating only the necessary portions of memory rows, reducing capacitive loading and dynamic power consumption while maintaining high computational throughput through selective row activation and bit-serial processing of relevant data elements.
Data Source
AI summary
Disclosed are various approaches for bit-serial computing embedded in the DRAM subarray, leveraging the massive parallelism of DRAM row operations. The present disclosure discloses digital techniques that can outperform analog charge-sharing techniques. Digital techniques can use more area but support a wider range of computing primitives, and allow a sequence of logic operations to be performed at higher clock speeds, be-tween slower subarray row reads/writes. The present disclosure describes a range of bit-serial architectures, and evaluate raw performance as well as area and energy efficiency. Results show that the digital architecture demonstrates 20× speedup over CPU, 5× over GPU, and 1.7× over SIMDRAM, an analog architecture.


