In-Storage Machine Learning Operations via Persistent Memory Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning systems face challenges in efficiently performing inference operations in large neural networks due to high data and processing resource requirements, leading to power consumption and delay issues.

Innovation Solution

A system comprising a first persistent memory connected to a control and inference circuit via a wideband data connection, allowing for arithmetic operations and enabling in-storage inference by reading weights from persistent memory into random-access memory for neural network operations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If inference operations are performed using traditional host-based processing, then processing power and flexibility are maintained, but power consumption increases and processing delay occurs due to frequent data transfers between host and storage

Engineering Contradiction:
Improvepower consumptionVSAvoidsystem architecture complexity
Core Design Contradiction:
Use of energy by moving objectVSDevice complexity

Solution Approach 1:

The patent merges the persistent memory storage function with the inference processing function by integrating a control and inference circuit directly into the persistent memory device. This allows inference operations to be performed at the storage location, eliminating the need for frequent data transfers between host and storage, thereby reducing power consumption while maintaining system functionality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a new dimension to the storage system by adding inference processing capabilities to the persistent memory device. This transforms the storage device from a passive data repository into an active processing node, enabling computations to occur at the edge of the storage hierarchy and reducing the energy cost of data movement.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If inference operations are performed with frequent data transfers between host and storage, then processing flexibility is maintained, but processing speed decreases due to transfer delays

Engineering Contradiction:
Improveinference processing speedVSAvoiddata transfer time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent extracts the inference processing function from the host system and places it directly within the persistent memory device. By taking out the computation workload from the host and performing it locally at the storage device, the system eliminates time-consuming data transfers while accelerating inference processing speed.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The persistent memory device performs inference operations on its own stored data without requiring continuous host intervention. The control and inference circuit within the memory device autonomously executes neural network operations using weights and input data already present in the persistent memory, significantly reducing processing time.

Inventive Principle:
Principle #25Self-service

3Reliability

If large amounts of data are processed in inference operations, then model accuracy is maintained, but resource requirements and power consumption increase

Engineering Contradiction:
Improveinference accuracyVSAvoidpower consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent copies the neural network weights from volatile memory to persistent memory, enabling inference operations to be performed using data already stored in the persistent memory. This eliminates the need for repeated loading of weight data from the host, reducing power consumption while maintaining the ability to process large datasets with high model accuracy.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250053796A1In-storage machine learning operations
Publication Date: 2025.02.13 SAMSUNG ELECTRONICS CO LTD
  • US20250053796A1 patent drawing
  • US20250053796A1 patent drawing
  • US20250053796A1 patent drawing

AI summary

A system and method for in-storage machine learning operations. In some embodiments, a system includes a first persistent memory, and a control and inference circuit. The first persistent memory may be connected to the control and inference circuit by a wideband data connection, and the control and inference circuit may be configured to perform arithmetic operations.