Parallel Memory Compute Matrix for Lower-Power Data Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computation models, such as artificial neural networks, face inefficiencies when dealing with large datasets that exceed the storage capacity of processing devices like SoC or CPU, leading to increased power and bandwidth usage due to frequent data transfers between processing devices and memory devices.
Innovation Solution
Incorporating an arithmetic logic unit matrix within memory devices to perform computations on data before transferring results, reducing the amount of data transferred over memory buses and enhancing communication bandwidth through separate IC dies connected via through-silicon vias or wire bonding.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data is stored in memory devices and transferred to processing devices for computation, then data storage capacity is improved, but power consumption and bandwidth usage increase due to frequent data transfers
Solution Approach 1:
The patent combines memory storage functions with arithmetic computation functions into a single memory device. The arithmetic logic unit matrix is integrated within the memory device, allowing data to be stored and computed upon without being transferred to a separate processing device. This merging eliminates the need for frequent data transfers between memory and processing devices, thereby reducing power consumption while maintaining large data storage capacity.
2Productivity
If data is transferred between memory devices and processing devices, then computation capability is improved, but communication bandwidth is exceeded due to large data volumes
Solution Approach 1:
The patent implements preliminary computation action within the memory device before data transfer. The arithmetic logic unit matrix performs computations on data while it is still stored in the memory device, generating intermediate or final results that require less data transfer volume. This preliminary action reduces the quantity of data that needs to be transferred over the memory bus, preventing bandwidth exhaustion while maintaining high computation capability.
3Loss of substance
If arithmetic computation is performed in memory devices, then data transfer volume is reduced, but device complexity increases
Solution Approach 1:
The patent makes the memory device universal by giving it multiple functions: data storage and arithmetic computation. The arithmetic logic unit matrix is designed to work with the memory structure, allowing the same device to perform both storage and computation tasks. This multi-functionality reduces data transfer volume while the complexity increase is managed through integrated design where the computation units are tightly coupled with the memory architecture.
Data Source
AI summary
An integrated circuit (IC) memory device encapsulated within an IC package. The memory device includes first memory regions configured to store lists of operands; a second memory region configured to store a list of results generated from the lists of operands; and at least one third memory region. A communication interface of the memory device can receive requests from an external processing device; and an arithmetic compute element matrix can access memory regions of the memory device in parallel. When the arithmetic compute element matrix is processing the lists of operands in the first memory regions and generating the list of results in the second memory region, the external processing device can simultaneously access the third memory region through the communication interface to load data into the third memory region, or retrieve results that have been previously generated by the arithmetic compute element matrix.


