On-Chip MRAM Storage for Neural Network Weight Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional artificial neural networks face high power consumption and latency due to the need for frequent access and modification of weight and bias values stored in non-volatile memory, which are typically located off-chip, leading to increased energy usage and processing delays.
Innovation Solution
Implementing a distributed storage architecture using magnetoresistive random-access memory (MRAM) bits that are physically proximate to hardware neurons, allowing for on-chip access and reducing the need for off-chip memory access, thereby minimizing power consumption and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If weight values and bias values are stored in off-chip non-volatile memory, then storage capacity is improved, but power consumption and latency increase due to frequent off-chip access
Solution Approach 1:
The patent segments the memory storage system into two parts: off-chip non-volatile memory for bulk weight and bias value storage, and on-chip storage circuitry for frequently accessed values. This segmentation allows the system to maintain large storage capacity while reducing power consumption by keeping hot data on-chip.
Solution Approach 2:
The patent implements local quality by placing storage circuitry physically close to the hardware neurons that require the weight and bias values. This local storage approach minimizes access latency and power consumption for frequently used parameters while maintaining overall system storage capacity.
2Speed
If weight values and bias values are loaded from off-chip non-volatile memory into on-chip RAM registers, then access speed is improved, but power consumption increases due to frequent loading operations
Solution Approach 1:
The patent applies preliminary action by pre-loading weight and bias values into on-chip storage circuitry before they are needed by the hardware neurons. This allows the values to be readily available when required, improving access speed while reducing the frequency of power-consuming load operations from off-chip memory.
Solution Approach 2:
The patent implements dynamic management of the storage circuitry, where weight and bias values are loaded into on-chip storage only when needed by specific hardware neurons, and the system dynamically manages which values reside on-chip versus off-chip based on access patterns.
3Area of stationary object
If off-chip memory access is used for weight values and bias values, then chip space is conserved, but latency increases in hardware neuron operations
Solution Approach 1:
The patent segments the storage architecture into off-chip non-volatile memory for capacity and on-chip storage circuitry for speed, allowing the system to conserve overall chip space while reducing latency for critical operations through local caching of weight and bias values.
Solution Approach 2:
The patent implements a nested storage hierarchy where on-chip storage circuitry is embedded within or adjacent to the hardware neuron circuitry, creating a nested architecture that provides fast access to frequently used parameters without significantly increasing the overall chip footprint.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach reduces power consumption, computational resources, and latency by enabling on-chip access to neural network components, conserving chip space, and enhancing operational efficiency during both training and inference processes.
Implementation Method 1
Implementing a distributed storage architecture using magnetoresistive random-access memory (MRAM) bits that are physically proximate to hardware neurons
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present disclosure is drawn to, among other things, a device comprising input circuitry; weight operation circuitry electrically connected to the input circuitry; bias operation circuitry electrically connected to the weight operation circuitry; storage circuitry electrically connected to the weight operation circuitry and the bias operation circuitry; and activation function circuitry electrically connected to the bias operation circuitry, wherein at least the weight operation circuitry, the bias operation circuitry, and the storage circuitry are located on a same chip.