In-Memory Compute Architecture for Parallel Data Access

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing computation models, such as artificial neural networks, face inefficiencies when dealing with large datasets that exceed the storage capacity of processing devices like SoC or CPU, leading to increased power consumption and bandwidth usage due to frequent data transfers between processing devices and memory devices.

Innovation Solution

Incorporating an arithmetic logic unit matrix within memory devices to perform computations on data before transferring results, reducing the amount of data transferred and enhancing data throughput and system performance by allowing parallel access and pre-processing within the memory device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If data is stored in external memory devices and transferred to processing devices for computation, then storage capacity is sufficient, but power consumption and bandwidth usage increase due to frequent data transfers

Engineering Contradiction:
Improvepower consumptionVSAvoidsystem architecture
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The patent combines memory storage functionality with arithmetic computation functionality into a single integrated memory device. The arithmetic logic unit matrix is embedded within the memory device, allowing computations to be performed on data while it remains stored in memory, eliminating the need for separate processing devices and reducing data transfer requirements.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The arithmetic logic unit matrix acts as an intermediary between the external processing device and the stored data. It performs computations directly on the data within the memory device, serving as a computational bridge that reduces the need for data to be transferred to external processing devices while still enabling complex operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If data is transferred frequently between processing devices and memory devices, then computation can be performed, but data throughput efficiency decreases

Engineering Contradiction:
Improvedata throughputVSAvoiddata transfer time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent enables preliminary computations to be performed on data while it remains stored in memory. The arithmetic logic unit matrix can perform arithmetic operations, logic operations, and other computations on data before it needs to be transferred to external processing devices, reducing the amount of data that needs to be moved and the time required for data preparation.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If computations are performed on large datasets exceeding processing device storage capacity, then comprehensive processing is possible, but power consumption increases

Engineering Contradiction:
Improveprocessing capabilityVSAvoidpower usage
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent merges memory storage and computation capabilities within the same device, enabling large datasets to be processed directly in memory without requiring transfer to external processing devices. This integration maintains comprehensive processing capability while significantly reducing the energy associated with data movement and external computation.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250390253A1Parallel memory access and computation in memory devices
Publication Date: 2025.12.25 LODESTAR LICENSING GROUP LLC
  • US20250390253A1 patent drawing
  • US20250390253A1 patent drawing
  • US20250390253A1 patent drawing

AI summary

An integrated circuit (IC) memory device encapsulated within an IC package. The memory device includes first memory regions configured to store lists of operands; a second memory region configured to store a list of results generated from the lists of operands; and at least one third memory region. A communication interface of the memory device can receive requests from an external processing device; and an arithmetic compute element matrix can access memory regions of the memory device in parallel. When the arithmetic compute element matrix is processing the lists of operands in the first memory regions and generating the list of results in the second memory region, the external processing device can simultaneously access the third memory region through the communication interface to load data into the third memory region, or retrieve results that have been previously generated by the arithmetic compute element matrix.