Non-Volatile Memory Processing Units for Von Neumann Bottleneck
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing memory devices face limitations due to the Von Neumann bottleneck, where data transfer from non-volatile memory to volatile memory for processing is constrained by hardware bandwidth, hindering overall speed and efficiency in data manipulation operations.
Innovation Solution
A computing system with processing units integrated on the same chip as non-volatile memory, allowing in-place data processing and storage, where each processing unit is associated with a data line and can independently program and erase data at selectable locations without affecting other locations, thereby overcoming bandwidth limitations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If data is transferred from non-volatile memory to volatile memory for processing, then data manipulation can be performed, but the overall speed is limited by hardware bandwidth (Von Neumann bottleneck)
Solution Approach 1:
The patent merges processing units directly with non-volatile memory locations, creating a unified architecture where computation and storage coexist on the same chip. This eliminates the need for separate volatile memory and the data transfer bottleneck, as processing occurs in-place within the non-volatile memory array.
Solution Approach 2:
The patent introduces a new spatial dimension by distributing processing units across multiple data lines within the memory array. Each processing unit is associated with specific data lines, enabling parallel processing operations directly at the memory location level, thereby overcoming the traditional sequential data transfer limitation.
2Loss of time
If processing units are integrated on the same chip as non-volatile memory, then latency is reduced and processing speed is enhanced, but device complexity increases
Solution Approach 1:
The patent segments the non-volatile memory array into multiple data lines, with each data line associated with a dedicated processing unit. This segmentation enables independent parallel processing operations on different data lines simultaneously, reducing overall access latency while managing complexity through modular organization.
Solution Approach 2:
The processing units serve multiple functions: they read data from their associated data lines, perform computational operations, and write results back to the same data lines. This multi-functionality reduces the need for separate read/write circuits and buffers, thereby reducing overall device complexity despite the integration.
3Productivity
If each processing unit is associated with a data line for independent computation, then processing parallelism is increased, but control complexity increases
Solution Approach 1:
The patent introduces control circuits as intermediary components that manage the operation of multiple processing units. These control circuits receive operation commands, decode them, and distribute appropriate control signals to the relevant processing units and memory circuits, thereby coordinating parallel operations without requiring complex direct interconnections between all components.
4Productivity
If non-volatile memory is used for both storage and processing, then bandwidth limitations are overcome, but power consumption management becomes more complex
Solution Approach 1:
The patent implements dynamic power management where processing units and memory circuits can be selectively activated or deactivated based on operational needs. The control circuits monitor usage patterns and power down inactive processing units or memory blocks, thereby reducing overall power consumption while maintaining the ability to perform parallel operations when needed.
Data Source
AI summary
In one example, a computing system includes a device, the device including: a non-volatile memory divided into a plurality of selectable locations, each bit in the non-volatile memory configured to have corresponding data independently altered, wherein the selectable locations are grouped into a plurality of data lines; and one or more processing units coupled to the non-volatile memory, each of the processing units associated with a data line of the plurality of data lines, and each of the processing units configured to compute, based on data in an associated data line of the plurality of data lines, corresponding results, wherein the non-volatile memory is configured to selectively write, based on the corresponding results, data in selectable locations of the associated data line reserved to store results of the computation from the process unit associated with the associated data line.


