Parallel Memory Compute Matrix for Lower-Power Data Transfer

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computation models, such as artificial neural networks, face inefficiencies when dealing with large datasets that exceed the storage capacity of processing devices like SoC or CPU, leading to increased power and bandwidth usage due to frequent data transfers between processing devices and memory devices.

Innovation Solution

Incorporating an arithmetic logic unit matrix within memory devices to perform computations on data before transferring results, reducing the amount of data transferred over memory buses and enhancing communication bandwidth through separate IC dies connected via through-silicon vias or wire bonding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is stored in memory devices and transferred to processing devices for computation, then data storage capacity is improved, but power consumption and bandwidth usage increase due to frequent data transfers

Engineering Contradiction:
Improvedata storage capacityVSAvoidpower consumption
Core Design Contradiction:
Quantity of substanceVSUse of energy by moving object

Solution Approach 1:

The patent combines memory storage functions with arithmetic computation functions into a single memory device. The arithmetic logic unit matrix is integrated within the memory device, allowing data to be stored and computed upon without being transferred to a separate processing device. This merging eliminates the need for frequent data transfers between memory and processing devices, thereby reducing power consumption while maintaining large data storage capacity.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If data is transferred between memory devices and processing devices, then computation capability is improved, but communication bandwidth is exceeded due to large data volumes

Engineering Contradiction:
Improvecomputation capabilityVSAvoiddata volume
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent implements preliminary computation action within the memory device before data transfer. The arithmetic logic unit matrix performs computations on data while it is still stored in the memory device, generating intermediate or final results that require less data transfer volume. This preliminary action reduces the quantity of data that needs to be transferred over the memory bus, preventing bandwidth exhaustion while maintaining high computation capability.

Inventive Principle:
Principle #10Preliminary action

3Loss of substance

If arithmetic computation is performed in memory devices, then data transfer volume is reduced, but device complexity increases

Engineering Contradiction:
Improvedata transfer volumeVSAvoiddevice complexity
Core Design Contradiction:
Loss of substanceVSDevice complexity

Solution Approach 1:

The patent makes the memory device universal by giving it multiple functions: data storage and arithmetic computation. The arithmetic logic unit matrix is designed to work with the memory structure, allowing the same device to perform both storage and computation tasks. This multi-functionality reduces data transfer volume while the complexity increase is managed through integrated design where the computation units are tightly coupled with the memory architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12399655B2Parallel memory access and computation in memory devices
Publication Date: 2025.08.26 MICRON TECHNOLOGY INC
  • US12399655B2 patent drawing
  • US12399655B2 patent drawing
  • US12399655B2 patent drawing

AI summary

An integrated circuit (IC) memory device encapsulated within an IC package. The memory device includes first memory regions configured to store lists of operands; a second memory region configured to store a list of results generated from the lists of operands; and at least one third memory region. A communication interface of the memory device can receive requests from an external processing device; and an arithmetic compute element matrix can access memory regions of the memory device in parallel. When the arithmetic compute element matrix is processing the lists of operands in the first memory regions and generating the list of results in the second memory region, the external processing device can simultaneously access the third memory region through the communication interface to load data into the third memory region, or retrieve results that have been previously generated by the arithmetic compute element matrix.