Smart Storage AI Training With Local Memory and Reduced I/O

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems face challenges with low bandwidth and high latency in connections between host devices and semiconductor devices, leading to inefficiencies in memory sharing and coherency, particularly when training large artificial intelligence models that require significant data movement.

Innovation Solution

Implementing a smart storage device with an accelerator and non-volatile memory that learns AI models through cache coherency, using pre-set weights and biases, and updating them based on output values from the host device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If data is repeatedly moved between host device and smart storage device during AI model learning, then model learning can be performed, but input/output traffic increases and processing efficiency decreases

Engineering Contradiction:
Improvemodel learning efficiencyVSAvoidinput/output traffic
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent combines the AI accelerator and non-volatile memory into a single smart storage device, allowing the model learning process to occur locally within the device using stored learning data, thereby eliminating repeated data transfers between host and storage devices and reducing I/O traffic while improving learning efficiency

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The smart storage device acts as an intermediary between the host device and the learning data, performing model learning operations locally and only exchanging necessary results with the host, thereby reducing the volume of data traffic while maintaining productive model training

Inventive Principle:
Principle #24Intermediary (Mediator)

2Speed

If learning data is stored in non-volatile memory of smart storage device, then model learning can be performed locally, but connection bandwidth limitations still affect data transfer speed

Engineering Contradiction:
Improvedata processing speedVSAvoidbandwidth
Core Design Contradiction:
SpeedVSQuantity of substance

Solution Approach 1:

The patent segments the AI processing system into distinct functional components: the host device handles high-level coordination and result collection, while the smart storage device handles local model learning operations using its embedded accelerator and memory, thereby optimizing data processing speed within each segment without requiring high bandwidth between them

Inventive Principle:
Principle #1Segmentation

3Productivity

If AI model learning is performed with frequent data transfers, then model can be trained, but latency increases

Engineering Contradiction:
Improvemodel training capabilityVSAvoidlatency
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-loading learning data into the non-volatile memory of the smart storage device and pre-positioning the AI accelerator in a ready state, enabling model learning to begin immediately without repeated data transfer delays and reducing overall latency while maintaining full model training capability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12619559B2Smart storage devices
Publication Date: 2026.05.05 SAMSUNG ELECTRONICS CO LTD
  • US12619559B2 patent drawing
  • US12619559B2 patent drawing
  • US12619559B2 patent drawing

AI summary

There is provided a smart storage device. The smart storage device comprises an accelerator which is connected to a host device through a smart interface, and includes a first model distributed from the host device, and a non-volatile memory which includes learning data used for learning the first model, wherein the accelerator learns the first model by using the learning data on the basis of a first weight and a first bias that are set in advance, provides the host device with a first output value which is output by inputting the learning data into the first model, and learns the first model by using the learning data, on the basis of a second weight and a second bias calculated from the host device with reference to the first output value.