In-Storage Machine Learning Operations via Persistent Memory Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning systems face challenges in efficiently performing inference operations in large neural networks due to high data and processing resource requirements, leading to power consumption and delay issues.
Innovation Solution
A system comprising a first persistent memory connected to a control and inference circuit via a wideband data connection, allowing for arithmetic operations and enabling in-storage inference by reading weights from persistent memory into random-access memory for neural network operations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If inference operations are performed using traditional host-based processing, then processing power and flexibility are maintained, but power consumption increases and processing delay occurs due to frequent data transfers between host and storage
Solution Approach 1:
The patent merges the persistent memory storage function with the inference processing function by integrating a control and inference circuit directly into the persistent memory device. This allows inference operations to be performed at the storage location, eliminating the need for frequent data transfers between host and storage, thereby reducing power consumption while maintaining system functionality.
Solution Approach 2:
The patent introduces a new dimension to the storage system by adding inference processing capabilities to the persistent memory device. This transforms the storage device from a passive data repository into an active processing node, enabling computations to occur at the edge of the storage hierarchy and reducing the energy cost of data movement.
2Speed
If inference operations are performed with frequent data transfers between host and storage, then processing flexibility is maintained, but processing speed decreases due to transfer delays
Solution Approach 1:
The patent extracts the inference processing function from the host system and places it directly within the persistent memory device. By taking out the computation workload from the host and performing it locally at the storage device, the system eliminates time-consuming data transfers while accelerating inference processing speed.
Solution Approach 2:
The persistent memory device performs inference operations on its own stored data without requiring continuous host intervention. The control and inference circuit within the memory device autonomously executes neural network operations using weights and input data already present in the persistent memory, significantly reducing processing time.
3Reliability
If large amounts of data are processed in inference operations, then model accuracy is maintained, but resource requirements and power consumption increase
Solution Approach 1:
The patent copies the neural network weights from volatile memory to persistent memory, enabling inference operations to be performed using data already stored in the persistent memory. This eliminates the need for repeated loading of weight data from the host, reducing power consumption while maintaining the ability to process large datasets with high model accuracy.
Data Source
AI summary
A system and method for in-storage machine learning operations. In some embodiments, a system includes a first persistent memory, and a control and inference circuit. The first persistent memory may be connected to the control and inference circuit by a wideband data connection, and the control and inference circuit may be configured to perform arithmetic operations.


