Memory Device In-Memory Calculation to Overcome the Memory Wall

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The memory bottleneck in high-performance computing systems, particularly in AI systems using transformer models, limits performance due to the lag in memory access speed compared to processor computation speed, leading to a constraint known as the memory wall.

Innovation Solution

A memory device with an array of memory cells and a peripheral circuit that includes page buffers, process units, and control logic to perform calculations within the memory device, distributing calculation tasks and eliminating the need for large data transfer to the processor.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in memory and transferred to processor for computation, then data storage capacity is improved, but computation speed deteriorates due to memory access lag

Engineering Contradiction:
Improvedata storage capacityVSAvoidcomputation speed
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent combines memory storage function with computation function into a single integrated memory device. Process units are embedded within the memory device to perform calculations directly on stored data, eliminating the need to transfer data between memory and external processor, thus resolving the speed bottleneck caused by memory access lag

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transitions from traditional von Neumann architecture where computation and memory are separate dimensions to an in-memory computing architecture where computation is embedded within the memory dimension. This dimensional integration allows data to be processed in-place without physical transfer, overcoming the memory wall constraint

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If large amount of data is transferred between memory and processor, then computation accuracy is improved, but energy consumption increases

Engineering Contradiction:
Improvecomputation accuracyVSAvoidenergy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts the computation function from the external processor and embeds it directly within the memory device. This extraction eliminates the energy-intensive data transfer process between memory and processor while maintaining computational accuracy through integrated process units that operate directly on stored data

Inventive Principle:
Principle #2Taking out (Extraction)

3Speed

If memory access speed is increased to match processor computation speed, then system performance is improved, but device complexity increases

Engineering Contradiction:
Improvememory access speedVSAvoiddevice complexity
Core Design Contradiction:
SpeedVSDevice complexity

Solution Approach 1:

Instead of trying to increase memory access speed to match processor speed, the patent inverts the approach by bringing computation to memory. This inversion eliminates the speed mismatch problem by making the data stationary and the computation mobile, avoiding the need for complex high-speed memory interfaces

Inventive Principle:
Principle #13The other way round (Inversion)

Data Source

PatentUS20250335122A1Memory device, memory system, and method for data calculation with the memory device
Publication Date: 2025.10.30 YANGTZE MEMORY TECH CO LTD
  • US20250335122A1 patent drawing
  • US20250335122A1 patent drawing
  • US20250335122A1 patent drawing

AI summary

A memory device, a memory system, and a method for data calculation with the memory device are provided. The memory device includes an array of memory cells and a peripheral circuit coupled to the memory cells is provided. The peripheral circuit includes page buffers configured to store first data transmitted from a data interface of the memory device and to sense second data from the array of memory cells. The peripheral circuit further includes at least one process unit coupled to the page buffers via a data-path bus of the peripheral circuit and configured to perform calculation based on the first data and the second data. The peripheral circuit further includes a control logic configured to program the second data into the array of memory cells.