Smart Storage AI Training With Local Memory and Reduced I/O
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing systems face challenges with low bandwidth and high latency in connections between host devices and semiconductor devices, leading to inefficiencies in memory sharing and coherency, particularly when training large artificial intelligence models that require significant data movement.
Innovation Solution
Implementing a smart storage device with an accelerator and non-volatile memory that learns AI models through cache coherency, using pre-set weights and biases, and updating them based on output values from the host device.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is repeatedly moved between host device and smart storage device during AI model learning, then model learning can be performed, but input/output traffic increases and processing efficiency decreases
Solution Approach 1:
The patent combines the AI accelerator and non-volatile memory into a single smart storage device, allowing the model learning process to occur locally within the device using stored learning data, thereby eliminating repeated data transfers between host and storage devices and reducing I/O traffic while improving learning efficiency
Solution Approach 2:
The smart storage device acts as an intermediary between the host device and the learning data, performing model learning operations locally and only exchanging necessary results with the host, thereby reducing the volume of data traffic while maintaining productive model training
2Speed
If learning data is stored in non-volatile memory of smart storage device, then model learning can be performed locally, but connection bandwidth limitations still affect data transfer speed
Solution Approach 1:
The patent segments the AI processing system into distinct functional components: the host device handles high-level coordination and result collection, while the smart storage device handles local model learning operations using its embedded accelerator and memory, thereby optimizing data processing speed within each segment without requiring high bandwidth between them
3Productivity
If AI model learning is performed with frequent data transfers, then model can be trained, but latency increases
Solution Approach 1:
The patent performs preliminary actions by pre-loading learning data into the non-volatile memory of the smart storage device and pre-positioning the AI accelerator in a ready state, enabling model learning to begin immediately without repeated data transfer delays and reducing overall latency while maintaining full model training capability
Data Source
AI summary
There is provided a smart storage device. The smart storage device comprises an accelerator which is connected to a host device through a smart interface, and includes a first model distributed from the host device, and a non-volatile memory which includes learning data used for learning the first model, wherein the accelerator learns the first model by using the learning data on the basis of a first weight and a first bias that are set in advance, provides the host device with a first output value which is output by inputting the learning data into the first model, and learns the first model by using the learning data, on the basis of a second weight and a second bias calculated from the host device with reference to the first output value.


