Memory device-based accelerated deep learning system

KR103022371B1Active Publication Date: 2026-09-21SANDISK TECHNOLOGIES LLC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
KR1020247016204
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-02-04
Filing Date
2022-05-21
Publication Date
2026-09-21
Estimated Expiration
2042-05-21

Smart Images

  • Figure 112024052749412-PCT00003_ABST
    Figure 112024052749412-PCT00003_ABST
Patent Text Reader

Abstract

The data storage device includes memory and a controller coupled to the memory device. The controller is configured to be coupled to a host device. The controller receives a plurality of commands, generates a logical block address (LBA) to physical block address (PBA) (L2P) mapping for each of the plurality of commands, and is also configured to store the data of the plurality of commands in their respective PBAs according to the generated L2P mapping. Each L2P mapping is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure. The controller includes an NN command interpretation unit and an L2P mapping generator coupled to the NN command interpretation unit. The controller is configured to fetch training data and NN parameters from the memory device.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Cross-reference of related application(s)

[0002] This application claims the benefit of the full contents of U.S. Regular Application No. 17 / 592,953, filed February 4, 2022, titled “MEMORY DEVICE BASED ACCELERATED DEEP-LEARNING SYSTEM”, for all purposes, and incorporates it by reference into this specification.

[0003] Technology field

[0004] The embodiments of the present disclosure generally relate to data storage devices such as solid state drives (SSDs), and more specifically, to using a deep learning training model stored in non-volatile memory to improve the read and write performance of a data storage device. Background Technology

[0005] Deep learning (DL) systems are advanced technologies capable of functioning in various fields. However, as the functionality of DL systems increases, the corresponding consumption of hardware resources for DL ​​systems also increases. Due to the size of data sets and DL models, DL systems may require very large capacities of high-speed memory. This memory can be random access memory (RAM). However, non-volatile memories, such as NAND memory devices, can be interlaced in DL hardware computation.

[0006] Typically, DL models are held in the dynamic RAM (DRAM) of data storage devices. As the size of the DL model increases, more DRAM may be required, which can increase the cost of the data storage device. However, non-volatile memory, such as NAND memory, may not be as cost-effective per capacity as DRAM. Nevertheless, NAND memory may not be comparable to the performance output of DRAM. For example, the size of a dataset can be approximately 100 GB or more. A dataset is a collection of data samples and labels used to tune the DL model.

[0007] Therefore, in the relevant technical field, an improved DL system that uses non-volatile memory to train DL models is required.

[0008] The present disclosure generally relates to a data storage device such as a solid-state drive (SSD), and more specifically to using a deep learning training model stored in non-volatile memory to improve the read and write performance of the data storage device. The data storage device includes memory and a controller coupled to the memory device. The controller is configured to be coupled to a host device. The controller is also configured to receive a plurality of commands, generate a logical block address (LBA) to physical block address (PBA) (L2P) mapping for each of the plurality of commands, and store the data of the plurality of commands in their respective PBAs according to the generated L2P mapping. Each L2P mapping is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure. The controller includes an NN command interpretation unit and an L2P mapping generator coupled to the NN command interpretation unit. The controller is configured to fetch training data and NN parameters from the memory device.

[0009] In one embodiment, the data storage device includes a memory and a controller coupled to the memory device. The controller is configured to be coupled to a host device. The controller receives a plurality of commands, generates a logical block address (LBA) to physical block address (PBA) (L2P) mapping for each of the plurality of commands, and is also configured to store data of the plurality of commands in their respective PBAs according to the generated L2P mapping. Each L2P mapping is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure.

[0010] In another embodiment, the data storage device includes a memory and a controller coupled to the memory device. The controller includes an NN command resolution unit and a logical block address (LBA) to physical block address (PBA) (L2P) mapping generator coupled to the NN command resolution unit. The controller is configured to fetch training data and NN parameters from the memory device.

[0011] In another embodiment, the data storage device comprises a non-volatile memory means and a controller coupled to the non-volatile memory means. The controller stores neural network (NN) parameters and one or more hyperparameter values ​​in the non-volatile memory means, performs a fully autonomous deep learning (DL) training model or a semi-autonomous DL training model; and is configured to store data according to the performed DL training model. Brief explanation of the drawing

[0012] In a manner that allows the features of the present disclosure mentioned above to be understood in detail, a more specific description of the present disclosure, briefly summarized above, may be made with reference to embodiments, some of which are illustrated in the accompanying drawings. However, it should be noted that the accompanying drawings are merely illustrative of typical embodiments of the present disclosure and should not be construed as limiting the scope of the present disclosure, as the present disclosure may allow for other equally effective embodiments. FIG. 1 is a schematic block diagram illustrating a storage system in which a data storage device can function as a storage device for a host device, according to specific embodiments. FIG. 2 is an exemplary example of a deep neural network according to specific embodiments. FIG. 3 is a schematic block diagram illustrating an LBA / PBA addressing system according to specific embodiments. FIG. 4 is a schematic block diagram illustrating an LBA / PBA addressing system according to specific embodiments. FIG. 5 is a flowchart illustrating a method of operating a fully autonomous data storage device during deep learning training according to specific embodiments. FIG. 6 is a flowchart illustrating a method of operating a semi-autonomous data storage device during deep learning training according to specific embodiments. To facilitate understanding, the same reference numerals have been used as much as possible to denote identical elements common to the drawings. It is considered that the elements disclosed in one embodiment may be advantageously utilized in other embodiments without special reference. Specific details for implementing the invention

[0013] Refer to the embodiments of the present disclosure below. However, it should be understood that the present disclosure is not limited to the embodiments specifically described. Instead, any combination of the following features and elements, whether or not related to other embodiments, is considered in implementing and practicing the present disclosure. Furthermore, while the embodiments of the present disclosure may achieve advantages over other possible solutions and / or prior art, whether a particular advantage is achieved by a given embodiment is not a limitation of the present disclosure. Accordingly, the following aspects, features, embodiments, and advantages are merely illustrative and are not to be construed as elements or limitations of the appended claims, except where expressly stated in the claim(s). Likewise, the reference to “the present disclosure” should not be interpreted as a generalization of any inventive subject matter disclosed herein, nor should it be construed as elements or limitations of the appended claims, except where expressly stated in the claim(s).

[0014] The present disclosure generally relates to a data storage device such as a solid-state drive (SSD), and more specifically to using a deep learning training model stored in non-volatile memory to improve the read and write performance of the data storage device. The data storage device includes memory and a controller coupled to the memory device. The controller is configured to be coupled to a host device. The controller is also configured to receive a plurality of commands, generate a logical block address (LBA) to physical block address (PBA) (L2P) mapping for each of the plurality of commands, and store data of the plurality of commands in their respective PBAs according to the generated L2P mapping. Each L2P mapping is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure. The controller includes an NN command interpretation unit and an L2P mapping generator coupled to the NN command interpretation unit. The controller is configured to fetch training data and NN parameters from the memory device.

[0015] FIG. 1 is a schematic block diagram illustrating a storage system (100) in which a host device (104) communicates with a data storage device (106) according to specific embodiments. For example, the host device (104) may store and retrieve data by utilizing non-volatile memory (NVM) (110) included in the data storage device (106). The host device (104) includes a host DRAM (138). In some examples, the storage system (100) may include a plurality of storage devices, such as a data storage device (106), which can operate as a storage array. For example, the storage system (100) may include a plurality of data storage devices (106) composed of redundant arrays of inexpensive / independent disks (RAID) that collectively function as a large-capacity storage device for the host device (104).

[0016] A host device (104) can store and / or retrieve data from one or more storage devices, such as a data storage device (106). As illustrated in FIG. 1, the host device (104) can communicate with the data storage device (106) through an interface (114). The host device (104) may include any of a wide range of devices, including a computer server, a NAS (network-attached storage) unit, a desktop computer, a notebook (i.e., laptop) computer, a tablet computer, a set-top box, a telephone handset such as a so-called "smart" phone or so-called "smart" pad, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, or other devices capable of transmitting and receiving data from a data storage device.

[0017] The data storage device (106) includes a controller (108), an NVM (110), a power supply (111), a volatile memory (112), an interface (114), and a write buffer (116). In some examples, the data storage device (106) may include additional components not shown in FIG. 1 for clarity. For example, the data storage device (106) may include a printed circuit board (PCB) that includes electrically conductive traces to which components of the data storage device (106) are mechanically attached and which electrically interconnect components such as the data storage device (106). In some examples, the physical dimensions and connector configurations of the data storage device (106) may follow one or more standard form factors. Some exemplary standard form factors include, but are not limited to, 3.5" data storage devices (e.g., HDD or SSD), 2.5" data storage devices, 1.8" data storage devices, PCI (peripheral component interconnect), PCI-X (PCI-extended), and PCIe (PCI Express) (e.g., PCIe x1, x4, x8, x16, PCIe mini card, MiniPCI, etc.). In some examples, the data storage device (106) may be directly coupled to the motherboard of the host device (104) (e.g., directly soldered or plugged into a connector).

[0018] The interface (114) may include one or both of a data bus for exchanging data with the host device (104) and a control bus for exchanging commands with the host device (104). The interface (114) may operate according to any suitable protocol. For example, the interface (114) may operate according to one or more of the following protocols: ATA (advanced technology attachment) (e.g., Serial-ATA (SATA) and Parallel-ATA (PATA)), Fibre Channel Protocol (FCP), Small Computer System Interface (SCSI), SAS (serially attached SCSI), PCI, and PCIe, NVMe (non-volatile memory express), OpenCAPI, GenZ, Cache Coherent Interface Accelerator (CCIX), Open Channel SSD (OCSSD), etc. An interface (114) (e.g., a data bus, a control bus, or both) is electrically connected to a controller (108) and provides an electrical connection between the host device (104) and the controller (108), allowing data to be exchanged between the host device (104) and the controller (108). In some examples, the electrical connection of the interface (114) also allows a data storage device (106) to receive power from the host device (104). For example, as illustrated in FIG. 1, a power supply unit (111) can receive power from the host device (104) through the interface (114).

[0019] The NVM (110) may include a plurality of memory devices or memory units. The NVM (110) may be configured to store and / or retrieve data. For example, a memory unit of the NVM (110) may receive data and a message from the controller (108) instructing the memory unit to store data. Similarly, a memory unit may receive a message from the controller (108) instructing the memory unit to retrieve data. In some examples, each of the memory units may be referred to as a die. In some examples, the NVM (110) may include a plurality of dies (i.e., a plurality of memory units). In some examples, each memory unit may be configured to store a relatively large amount of data (e.g., 128MB, 256MB, 512MB, 1GB, 2GB, 4GB, 8GB, 16GB, 32GB, 64GB, 128GB, 256GB, 512GB, 1TB, etc.).

[0020] In some examples, each memory unit may include any type of non-volatile memory device, such as a flash memory device, a PCM (phase-change memory) device, a ReRAM (resistive random-access memory) device, a MRAM (magneto-resistive random-access memory) device, a F-RAM (ferroelectric random-access memory), a holographic memory device, and any other type of non-volatile memory device.

[0021] NVM (110) may include multiple flash memory devices or memory units. An NVM flash memory device may include a NAND or NOR-based flash memory device and may store data based on the charge contained in the floating gate of a transistor for each flash memory cell. In an NVM flash memory device, the flash memory device may be divided into multiple dies, wherein each die among the multiple dies has multiple physical blocks or logical blocks that can be further divided into multiple pages. Each block of the multiple blocks within a particular memory device may include multiple NVM cells. Rows of NVM cells may be electrically connected using word lines to define one of multiple pages. Each cell in each of the multiple pages may be electrically connected to their respective bit lines. Furthermore, an NVM flash memory device may be a 2D or 3D device and may be a single-level cell (SLC), a multi-level cell (MLC), a triple-level cell (TLC), or a quad-level cell (QLC). The controller (108) can write data to NVM flash memory devices at the page level and read data from them, and can erase data from NVM flash memory devices at the block level.

[0022] The power supply unit (111) can provide power to one or more components of the data storage device (106). When operating in standard mode, the power supply unit (111) can provide power to one or more components using power provided by an external device, such as a host device (104). For example, the power supply unit (111) can provide power to one or more components using power received from the host device (104) via an interface (114). In some examples, the power supply unit (111) may include one or more power storage components configured to provide power to one or more components when operating in shutdown mode, such as when power reception from an external device is interrupted. In this way, the power supply unit (111) can function as an onboard backup power source. Some examples of one or more power storage components include, but are not limited to, capacitors, supercapacitors, batteries, etc. In some examples, the amount of power that can be stored by one or more power storage components may be a function of the cost and / or size (e.g., area / volume) of one or more power storage components. In other words, as the amount of power stored by one or more power storage components increases, the cost and / or size of one or more power storage components also increase.

[0023] Volatile memory (112) can be used through a controller (108) to store information. Volatile memory (112) may include one or more volatile memory devices. In some examples, the controller (108) may use the volatile memory (112) as a cache. For example, the controller (108) may store cached information in the volatile memory (112) until the cached information is written to the NVM (110). As illustrated in FIG. 1, the volatile memory (112) may consume power received from the power supply unit (111). Examples of volatile memory (112) include, but are not limited to, RAM (random-access memory), DRAM (dynamic random access memory), SRAM (static RAM), and SDRAM (synchronous dynamic RAM) (e.g., DDR1, DDR2, DDR3, DDR3L, LPDDR3, DDR4, LPDDR4, etc.).

[0024] The controller (108) can manage one or more operations of the data storage device (106). For example, the controller (108) can manage reading data from the NVM (110) and / or writing data to the NVM (110). In some embodiments, when the data storage device (106) receives a write command from the host device (104), the controller (108) can initiate a data storage command to store data in the NVM (110) and monitor the progress of the data storage command. The controller (108) can determine at least one operation characteristic of the storage system (100) and store at least one operation characteristic in the NVM (110). In some embodiments, when the data storage device (106) receives a write command from the host device (104), the controller (108) temporarily stores the data associated with the write command in an internal memory or write buffer (116) before transmitting the data to the NVM (110).

[0025] FIG. 2 is an exemplary example of a deep neural network (DNN) (200) according to specific embodiments. The DNN (200) includes an input layer (202), a first hidden layer (204a), a second hidden layer (204b), a third hidden layer (204c), and an output layer (206). The number of hidden layers depicted is not intended to be limited but is intended to provide examples of possible embodiments. Additionally, each of the input layer (202), the first hidden layer (204a), the second hidden layer (204b), the third hidden layer (204c), and the output layer (206) includes a plurality of nodes. Each node of the input layer (202) may be an input node for data input. Each node of the first hidden layer (204a), the second hidden layer (204b), and the third hidden layer (204c) combines an input from the data with a set of coefficients or weights that give significance to the input in relation to the task the algorithm intends to learn by amplifying or weakening the input. The result of the third hidden layer (204c) is transmitted to a node of the output layer (206).

[0026] The basic forward computation operation (e.g., feedforward) of single-node activation in a DNN (200) can be expressed by the following equation: Multiplication-accumulation (MAC) operations are summed and an activation function is calculated, which may be a max (e.g., rectifier activation function or ReLU) or a sigmoid function. In other words, the forward computation operation is an activation sigmoid function applied to the sum of weights multiplied by the net input value plus the bias for each neuron or node. The DNN (200) learning scheme is based on the backpropagation equation used to update the neural network (NN) weights. The backpropagation equation is based on a weighted sum using the calculated delta terms given below in matrix and vector form for the nodes of the output layer (206) and the nodes of the first hidden layer (204a), the second hidden layer (204b), and the third hidden layer (204c).

[0027]

[0028] Backpropagation equations (BP1, BP2, BP3, and BP4) show that there is a fixed input (z) that can be handled in static memory without changing (e.g., NVM (110) in FIG. 1), and adjustable values ​​(C, δ, and w) that can be temporarily adjusted or calculated and handled in dynamic memory (e.g., DRAM). Another memory-consuming factor is the DL model itself (i.e., NN parameters, which may be "weights" or C, δ, and w). As the functionality of the DNN (200) increases, the size of the DL model also increases. Although a fully connected NN architecture is exemplified, it should be understood that the embodiments described herein may be applicable to other NN architectures.

[0029] FIG. 3 is a schematic block diagram illustrating a logical block address (LBA) / physical block address (PBA) addressing system (300) according to specific embodiments. The LBA / PBA addressing system (300) includes a host device (302) coupled to a data storage device (308). The data storage device (308) is coupled to an NVM storage system comprising a plurality of NVMs (316a to 316n). It should be understood that the plurality of NVMs (316a to 316n) may be placed in the data storage device (308). In some examples, the plurality of NVMs (316a to 316n) are NAND devices. The host device (302) includes a CPU / GPU unit (304) and a block-based command generator unit (306). The block-based command generator unit (306) generates commands to be programmed into the blocks of the NVMs of the plurality of NVMs (316a to 316n). The host device (302) recognizes the LBA where data is stored, and the data storage device (308) recognizes the PBA where data is stored in the plurality of NVMs (316a to 316n).

[0030] The data storage device (308) includes a command interpretation unit (310), a block-based flash translation layer (FTL) conversion unit (312), and a flash interface unit (314), all of which may be placed in a controller such as the controller (108) of FIG. 1. The command interpretation unit (310) may be configured to receive or retrieve a command from a block-based command generator unit (306). The command interpretation unit (310) may process the command and generate relevant control information for the processed command. Then, the command is transmitted to the block-based FTL conversion unit (312), where the command is converted from an LBA to a PBA. The flash interface unit (314) transmits the read / write command based on the PBA to the relevant NVMs of a plurality of NVMs (316a to 316n). In other words, the conversion layer between the LBA and the PBA is stored in the data storage device (308), so that whenever a command is transmitted from the host device (302) to the data storage device (308), the PBA corresponding to the LBA associated with the command is extracted from the conversion layer.

[0031] FIG. 4 is a schematic block diagram illustrating an LBA / PBA addressing system (400) according to specific embodiments. The LBA / PBA addressing system (400) includes a host device (402) coupled to a data storage device (408). The data storage device (408) is coupled to an NVM storage system comprising a plurality of NVMs (416a to 416n). It should be understood that the plurality of NVMs (416a to 416n) may be placed in the data storage device (408). The host device (402) includes a CPU / GPU unit (404) and an NN interface command generator unit (406). The NN interface command generator unit (406) generates commands to be programmed into blocks of the NVMs of the plurality of NVMs (416a to 416n). In some examples, the plurality of NVMs (416a to 416n) are NAND devices. The command may include an NN structure and one or more hyperparameter values. The NN structure and one or more hyperparameter values ​​are stored in one or more of the plurality of NVMs (416a to 416n). One or more hyperparameter values ​​may define the training procedure of the DL model. The host device (402) recognizes the LBA where the data is stored, and the data storage device (408) recognizes the PBA where the data is stored in the plurality of NVMs (416a to 416n).

[0032] The data storage device (408) includes an NN interface command interpretation unit (410), a schedule-based FTL conversion unit (412), and a flash interface unit (414), all of which may be placed in a controller such as the controller (108) of FIG. 1. The NN interface command interpretation unit (410) may be configured to receive or retrieve commands from the NN interface command generator unit (406). The NN interface command interpretation unit (410) may process commands and generate relevant control information for the processed commands. In some embodiments, to reduce overhead and improve storage utilization for both dynamic parameters (e.g., "weights" and cost calculations) and static parameters, such as data stored in the NVMs of a plurality of NVMs (416a to 416n), the data storage device may hold some or all of the NN structures and hyperparameter values.

[0033] Then, the command is transmitted to a schedule-based FTL conversion unit (412), where the command is converted from an LBA to a PBA based on a schedule (e.g., a DL model) transmitted from the host device (402) to the data storage device (408). The flash interface unit (414) transmits the read / write command based on the PBA to the corresponding NVMs of a plurality of NVMs (416a to 416n). In other words, the conversion layer between the LBA and the PBA is stored in the data storage device (408), so that whenever a command is transmitted from the host device (402) to the data storage device (408), the PBA corresponding to the LBA associated with the command is extracted from the conversion layer.

[0034] FIG. 5 is a flowchart illustrating a method (500) of fully autonomous data storage device operation during deep learning training according to specific embodiments. The method (500) may be implemented by the storage device (408) of FIG. 4 or the controller (108) of FIG. 1. For exemplary purposes, aspects of the LBA / PBA addressing system (400) may be referenced herein. Fully autonomous data storage device operation may omit the explicit transmission of NN parameters of specific read and write commands from the CPU / GPU unit (404) to the data storage device (408). When a GPU is used in addition to the CPU, dual read / write direct storage access may be allowed between the GPU and a plurality of NVMs (416a to 416n).

[0035] More precisely, the data storage device (408) can hold NN structure and hyperparameter values. The NN interface command interpretation unit (410) can receive NN structure and / or hyperparameter values ​​prior to the training process, or select NN structure and / or hyperparameter values ​​stored in a static configuration (i.e., stored offline). Accordingly, the training process and the placement of data within the buffer (i.e., placement of data into the NVMs of multiple NVMs (416a to 416n) based on L2P mapping) can be completed in a "fully autonomous" manner, for instance, without the need for feedback from the host device (402).

[0036] In block (502), the host device (402) selects an NN structure from a predefined configuration or explicitly transmits an NN structure. The predefined configuration may be a previously trained NN structure or a default NN structure. In block (504), the host device (402) initiates the training process by transmitting a data location through a dedicated interface. For example, the training process may be initiated by placing a value or data location at a node of the input layer (202) of FIG. 2. In block (506), the data storage device (408), or more specifically, the controller (108), performs read and write operations according to a predefined schedule. The predefined schedule may be an NN structure and / or hyperparameter value transmitted from the host device (402) to the data storage device (408) prior to the training process, or held in the data storage device (408) at an offline location (e.g., the NVM of a plurality of NVMs (416a to 416n)). In block (508), the host device (402) performs calculations by reading and placing data from a buffer directed toward the data storage device (408).

[0037] The method (500) may implement either block (506) or block (508) independently, or both block (506) and block (508) together. For example, the controller (108) may execute block (506) without executing block (508). In some examples, the result of block (506) may be transferred to a host device (402) for implementation in block (508) and / or the result of block (508) may be transferred to a data storage device (408) for implementation in block (506). As the need for random read and write is reduced, data may be addressed in full block size or partial block size. Accordingly, NN parameters may be addressed to a predefined schedule through a start point and an offset. In block (510), DL model training is terminated when a threshold number of iterations is reached (i.e., when a predefined training schedule is terminated) or, for example, because the cost calculation is maintained constant, the host device (402) terminates the training process.

[0038] In an alternative addressing scheme, a key-value (KV) pair interface may be used instead of a PBA to LBA mapping. Each data instance (e.g., a value) can be addressed using a key. NN parameters can be addressed as a structure associated with an iteration or a part of an iteration. For example, all NN parameters belonging to the first iteration (e.g., nodes 1 through 100 from a list of nodes greater than 100) can be addressed through a single key.

[0039] To reduce model overfitting (e.g., redundant calculations, unnecessary shifts, etc.), DL model training may use dropout. Dropout improves the robustness of the DL model by disabling one or some of the nodes in the hidden layer in each iteration of the algorithm, thereby improving the performance of the algorithm. However, dropout introduces a measure of uncertainty. Since network connections are effectively changed in each iteration, NN parameters may be used differently. If dropout can be applied before the training process, the modified NN connections may already be reflected in the NN hyperparameters. For example, either the controller (108) or the data storage device (408) may apply dropout to specific nodes by parsing the NN structure iteration by iteration or by indicating the nodes to be skipped in each iteration. In some examples, the data storage device (408) or the controller (108) may randomize the nodes to be dropped out in each iteration according to a predefined randomization setting.

[0040] FIG. 6 is a flowchart illustrating a method (600) of semi-autonomous data storage device operation during deep learning training according to specific embodiments. The method (600) may be implemented by the storage device (408) of FIG. 4 or the controller (108) of FIG. 1. For exemplary purposes, aspects of the LBA / PBA addressing system (400) may be referenced herein. When the data storage device (408) is operating in semi-autonomous mode, the CPU / GPU unit (404) may point to NN parameters to be read in each iteration. Accordingly, when storing data in a plurality of NVMs (416a to 416n) based on L2P mapping, the problem of synchronizing read / write operations is reduced, and dropout processing may be reduced.

[0041] The data storage device (408) or controller (108) can utilize the unique characteristics of the DL model training workload and update the NN parameters in a predefined deterministic manner after each read and loss calculation. Accordingly, the data storage device (408) or controller (108) can update the "weights" by implementing write commands in a semi-autonomous manner. In other words, each update or write to the NN parameters or "weights" is completed at the same address as the previous read. Accordingly, there may be no need to send specific write commands. More precisely, the CPU / GPU unit (404) will send a list of NN parameter "weights" to the data storage device (408) for updating after each iteration.

[0042] In block (602), the host device (402) selects an NN structure from a predefined configuration for one iteration or explicitly transmits an NN structure. The predefined configuration may be a previously trained NN structure or a default NN structure. In block (604), the host device (402) initiates the training process by transmitting a data location through a dedicated interface. For example, the training process may be initiated by placing a value or data location at a node of the input layer (202) of FIG. 2. In block (606), the data storage device (408), or more specifically, the controller (108), performs read and write operations according to a predefined schedule during one training iteration. The predefined schedule may be an NN structure and / or hyperparameter value transmitted from the host device (402) to the data storage device (408) prior to the training process or held in the data storage device (408) at an offline location (e.g., the NVM of a plurality of NVMs (416a to 416n)). In block (608), the host device (402) performs calculations by reading and placing data from a buffer directed toward the data storage device (408).

[0043] The method (600) may implement either block (606) or block (608) independently, or both block (606) and block (608) together. For example, the controller (108) may execute block (606) without executing block (608). In some examples, the result of block (606) may be passed to a host device (402) to be implemented in block (608) and / or the result of block (608) may be passed to a data storage device (408) to be implemented in block (606). As the need for random read and write is reduced, data may be addressed in full block size or partial block size. Accordingly, NN parameters may be addressed to a predefined schedule through a start point and an offset. In block (610), the data storage device (408) or the controller (108) determines whether DL model training has ended. For example, training is terminated when the host device (402) terminates the training process, such as when a critical number of iterations is reached (i.e., the predefined training schedule is terminated) or when, for instance, the cost calculation remains constant. If training is not terminated at block (610), the method (600) returns to block (602). However, if training is terminated at block (610), the method (600) terminates at block (612).

[0044] By reducing the overhead of command transmission and interpretation between the host device running machine learning applications and the flash memory of the data storage device, power consumption can be reduced and throughput improved.

[0045] In one embodiment, the data storage device includes a memory and a controller coupled to the memory device. The controller is configured to be coupled to a host device. The controller receives a plurality of commands, generates a logical block address (LBA) to physical block address (PBA) (L2P) mapping for each of the plurality of commands, and is also configured to store data of the plurality of commands in their respective PBAs according to the generated L2P mapping. Each L2P mapping is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure.

[0046] The controller receives an NN structure and one or more hyperparameter values, and is also configured to store the NN structure and hyperparameter values ​​in a memory device. The NN structure is received from a host device. The memory device is a non-volatile memory device. One or more hyperparameter values ​​define the training procedure of the DL training model. The NN structure and one or more hyperparameter values ​​are provided to the DL training model at the start of the training procedure. The DL training model uses predefined hyperparameter values ​​of one or more predefined parameter sets. The DL training model is updated after generating each L2P mapping. The controller is also configured to read weights according to the NN structure. The weights are updated after generating each L2P mapping. The controller is also configured to place data of multiple commands into a designated buffer. Placement is completed without intervention from the host device.

[0047] In another embodiment, the data storage device includes a memory and a controller coupled to the memory device. The controller includes an NN command resolution unit and a logical block address (LBA) to physical block address (PBA) (L2P) mapping generator coupled to the NN command resolution unit. The controller is configured to fetch training data and NN parameters from the memory device.

[0048] The NN command resolution unit is configured to interface with an NN interface command generator deployed on the host device. NN parameters are KV pair data. Training data and NN parameters are used for a Deep Learning (DL) training model. One or more parts of the DL training model are disabled. The controller is configured to perform autonomous fetching of training data and NN parameters from the memory device. The controller is also configured to update one or more weights associated with the Deep Learning (DL) training model. The update is performed to the same address as the previous read of one or more weights.

[0049] In another embodiment, the data storage device comprises a non-volatile memory means and a controller coupled to the non-volatile memory means. The controller stores neural network (NN) parameters and one or more hyperparameter values ​​in the non-volatile memory means, performs a fully autonomous deep learning (DL) training model or a semi-autonomous DL training model; and is configured to store data according to the performed DL training model.

[0050] The non-volatile memory means is a NAND-based memory means. The operation includes performing reads and writes according to a predefined training schedule.

[0051] Although the foregoing describes embodiments of the present disclosure, other embodiments and additional embodiments of the present disclosure may be devised without departing from the basic scope of the present disclosure, and the scope of the present disclosure is determined by the following claims.

Claims

Claim 1 As a data storage device, the memory device; and a controller coupled to the memory device, wherein the controller is configured to be coupled to a host device, and the controller also: receives a plurality of commands; Logical block address (LBA) to physical block address (PBA) (L2P) mappings are generated for each of the plurality of commands above — each of the L2P mappings is generated based on the results of a deep learning (DL) training model using a neural network (NN) structure, said NN structure includes NN parameters, said NN parameters are updated after each read so that each update or write to the NN parameters is completed at the same address as the previous read, said DL training model uses dropout to disable nodes in each iteration of the algorithm, said dropout is applied to specific nodes by parsing the NN structure iteration by iteration or by indicating nodes that should be skipped in each iteration, said NN parameters include fixed inputs and adjustable values ​​for the backpropagation equation, said fixed inputs are stored in non-volatile memory (NVM) and said adjustable values ​​are dynamic random access A data storage device configured to store in memory (DRAM); store data of the plurality of commands in their respective PBAs according to the generated L2P mappings; fetch the NN structure from the memory device; perform read and write according to the fetched NN structure; and determine whether the threshold number of iterations of the DL training module has been reached. Claim 2 In claim 1, the controller also: receives the NN structure and one or more hyperparameter values; and is configured to store the NN structure and the hyperparameter values ​​in the memory device, a data storage device. Claim 3 In paragraph 2, the NN structure is a data storage device received from a host device. Claim 4 In paragraph 2, the memory device is a data storage device that is a non-volatile memory device. Claim 5 In paragraph 2, the data storage device wherein the one or more hyperparameter values ​​define the training procedure of the DL training model. Claim 6 A data storage device according to claim 5, wherein the NN structure and the one or more hyperparameter values ​​are provided to the DL training model at the start of the training procedure. Claim 7 In claim 5, the DL training model is a data storage device that uses predefined hyperparameter values ​​of one or more predefined parameter sets. Claim 8 A data storage device, wherein the DL training model is updated after generating each of the L2P mappings in claim 1. Claim 9 A data storage device according to claim 1, wherein the controller is also configured to read weights according to the NN structure, and the weights are updated after each of the L2P mappings is generated. Claim 10 A data storage device according to claim 1, wherein the controller is also configured to place the data of the plurality of commands into a designated buffer, and the placement is completed without the intervention of a host device. Claim 11 As a data storage device, the device comprises: a memory device; and a controller coupled to the memory device, wherein the controller comprises: a neural network (NN) command interpretation unit; and a logical block address (LBA) to physical block address (PBA) (L2P) mapping generator coupled to the NN command interpretation unit, wherein the controller is configured to fetch training data and NN parameters from the memory device, and the controller comprises: performing read and write operations according to the fetched NN parameters, wherein the NN parameters are updated after each read so that each update or write to the NN parameters is completed at the same address as the previous read, and the NN parameters include a fixed input and an adjustable value for a backpropagation equation, wherein the fixed input is stored in non-volatile memory (NVM) and the adjustable value is stored in dynamic random access memory (DRAM). A data storage device configured to determine whether a threshold number of iterations of a deep learning (DL) training module has been reached — said DL training model uses dropout to disable a node in each iteration of the algorithm, said dropout is applied to a specific node by parsing the NN structure iteration by iteration or by indicating a node that should be skipped in each iteration. Claim 12 In claim 11, the data storage device is configured such that the NN command interpretation unit interfaces with an NN interface command generator placed on a host device. Claim 13 delete Claim 14 In claim 11, the training data and the NN parameters are a data storage device used in the deep learning (DL) training model. Claim 15 A data storage device in which one or more parts of the above-mentioned DL training model are disabled, in accordance with claim 14. Claim 16 In claim 11, the data storage device is configured such that the controller performs autonomous fetching of the training data and the NN parameters from the memory device. Claim 17 In claim 11, the controller is also configured to update one or more weights associated with the deep learning (DL) training model, and the update is performed for the same address as the previous read of the one or more weights, a data storage device. Claim 18 A data storage device comprises: a non-volatile memory means; and a controller coupled to the non-volatile memory means, wherein the controller: stores neural network (NN) parameters and one or more hyperparameter values ​​in the non-volatile memory means — the NN parameters are updated after each read so that each update or write to the NN parameters is completed at the same address as the previous read, the NN parameters include a fixed input and an adjustable value for a backpropagation equation, the fixed input is stored in the non-volatile memory means and the adjustable value is stored in dynamic random access memory (DRAM) —; performs any one of the following: a fully autonomous deep learning (DL) training model; or a semi-autonomous DL training model; stores data according to the performed DL training model; fetches the NN parameters from the non-volatile memory means; performs reads and writes according to the fetched NN parameters; A data storage device configured to determine whether a threshold number of iterations of the above-mentioned DL training module has been reached — the above-mentioned DL training model uses dropout to disable a node in each iteration of the algorithm, and the dropout is applied to a specific node by indicating a node that should be parsed by the iteration of the NN structure iteration or skipped in each iteration — Claim 19 In paragraph 18, the above non-volatile memory means is a data storage device that is a NAND-based memory means. Claim 20 A data storage device according to claim 18, wherein the above-mentioned operation includes performing reading and writing according to a predefined training schedule. Claim 21 In paragraph 18, the controller comprises a logical block address (LBA) to physical block address (PBA) (L2P) mapping generator coupled to an NN command interpretation unit, a data storage device.

Citation Information

Patent Citations

  • Deep solid state device and neural network based persistent data storage

    KR1020200059151A

  • Teacher data generation apparatus and method, and object detection system

    US20180342077A1

  • Non-volatile memory die with deep learning neural network

    US20200185027A1