Digital Bit-Serial Vector Computing Architecture in DRAM Subarrays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer systems are bottlenecked by memory access bandwidth for applications with large datasets and low computational intensity, motivating the need to integrate computational capabilities within Dynamic Random-Access Memory (DRAM).

Innovation Solution

A digital bit-serial vector computing architecture is embedded in the DRAM subarray, featuring a bit-serial logic unit per subarray column, bank-level bit-serial control logic, and rank-level processing units. This architecture decouples execution from memory row access, enabling faster bit-serial logic operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If computational capabilities are placed inside DRAM to overcome memory access bandwidth bottlenecks, then processing speed and energy efficiency improve, but device complexity increases

Engineering Contradiction:
Improveprocessing speedVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the computational architecture into hierarchical levels: bit-serial logic units at the subarray level, bank-level control logic, and rank-level processing units. This segmentation allows computation to be distributed across multiple granularities, improving processing throughput while managing complexity through modular organization. Each segment handles specific operations, enabling parallel processing across DRAM subarrays without requiring a monolithic complex processor.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces bit-serial logic units as intermediary components between the DRAM storage array and traditional processing units. These logic units serve as mediators that perform computational operations directly on data residing in DRAM, eliminating the need for continuous data transfer to external processors. This intermediary approach enables in-memory computation, improving processing speed while keeping the overall system architecture manageable through clearly defined interface boundaries.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If bit-serial logic operations are integrated into DRAM subarray, then execution time is reduced compared to memory row cycles, but manufacturing precision requirements increase

Engineering Contradiction:
Improveexecution timeVSAvoidmanufacturing precision
Core Design Contradiction:
Loss of timeVSManufacturing precision

Solution Approach 1:

The patent changes the operational parameters of DRAM by implementing bit-serial logic units that operate on individual bits or small bit groups rather than traditional word-level operations. This parameter change enables finer-grained control over data processing, reducing execution time by performing multiple bit-serial operations within a single memory row cycle. The approach manages manufacturing precision requirements by leveraging standard DRAM cell structures and using control logic that operates at relaxed timing compared to high-speed serial interfaces.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a temporal dimension to DRAM operations by implementing bit-serial processing that unfolds computation over multiple clock cycles within a row active period. Instead of requiring all bits to be processed simultaneously (spatial parallelism), the bit-serial approach processes bits sequentially over time, reducing the need for complex simultaneous switching and thereby lowering manufacturing precision requirements while still achieving reduced execution time through efficient utilization of the row active period.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If massive parallelism of DRAM row operations is leveraged, then computational throughput increases, but energy consumption increases

Engineering Contradiction:
Improvecomputational throughputVSAvoidenergy consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by implementing bit-serial logic units that perform computation on a subset of bits within each DRAM row rather than requiring full row activation for every operation. This allows the system to leverage DRAM's inherent parallelism by activating only the necessary portions of memory rows, reducing capacitive loading and dynamic power consumption while maintaining high computational throughput through selective row activation and bit-serial processing of relevant data elements.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250117223A1Systems, circuits, methods, and articles of manufacturer for dram-based digital bit-serial vector computing architecture
Publication Date: 2025.04.10 UNIV OF VIRGINIA PATENT FOUND
  • US20250117223A1 patent drawing
  • US20250117223A1 patent drawing
  • US20250117223A1 patent drawing

AI summary

Disclosed are various approaches for bit-serial computing embedded in the DRAM subarray, leveraging the massive parallelism of DRAM row operations. The present disclosure discloses digital techniques that can outperform analog charge-sharing techniques. Digital techniques can use more area but support a wider range of computing primitives, and allow a sequence of logic operations to be performed at higher clock speeds, be-tween slower subarray row reads/writes. The present disclosure describes a range of bit-serial architectures, and evaluate raw performance as well as area and energy efficiency. Results show that the digital architecture demonstrates 20× speedup over CPU, 5× over GPU, and 1.7× over SIMDRAM, an analog architecture.