Artificial intelligence storage device and storage system comprising the same
By employing a separate processor and memory architecture in the storage device and independently managing the clock/power domain, the AI function of the storage device itself can be realized, solving the problems of high power consumption and large data traffic in the existing technology and improving efficiency.
Patent Information
- Application Number
- CN202011224896.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-18
- Filing Date
- 2020-11-05
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-11-05
AI Technical Summary
Existing storage devices consume a lot of power and handle large amounts of data when performing artificial intelligence calculations, resulting in low efficiency.
It adopts a separate processor and memory architecture, which are used for data storage and AI computing respectively. Through independent clock/power domain management, the AI function of the storage device itself is realized, reducing the dependence on the host device.
It reduces the power consumption of storage devices, reduces the data traffic between host devices and storage devices, and improves work efficiency.
Smart Images

Figure CN112820339B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to Korean Patent Application No. 10-2019-0147552, filed on November 18, 2019, with the Korean Intellectual Property Office (KIPO), the entire contents of which are incorporated herein by reference. Technical Field
[0003] The example embodiments generally relate to semiconductor integrated circuits, and more specifically to artificial intelligence (AI) storage devices and storage systems including such storage devices. Background Technology
[0004] A storage system comprises host devices and storage devices. These storage devices can be memory systems that include both memory controllers and storage devices, or memory systems that consist only of storage devices. In a storage system, host devices and storage devices are interconnected via various interface standards such as Universal Flash Memory (UFS), Serial Advanced Technology Attached (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), and embedded multimedia cards (eMMC).
[0005] In computer science, artificial intelligence (AI), sometimes referred to as machine intelligence, is intelligence exhibited by machines, in contrast to natural intelligence, such as that displayed by humans. In layman's terms, the term "AI" is often used to describe machines (e.g., computers) that mimic "cognitive" functions (e.g., "learning" and "problem-solving") associated with human thought. AI can be implemented, for example, based on machine learning, neural networks, artificial neural networks (ANNs), etc. ANNs are obtained by engineering a model of the cellular structure of the human brain in which pattern recognition processes are performed. An ANN is a software and / or hardware-based computational model designed to mimic biological computational capabilities by applying numerous artificial neurons interconnected by wires. The human brain is composed of neurons, which are the basic units of nerves, and encrypts or decrypts information based on the different types of dense connections between these neurons. The artificial neurons in an ANN are obtained by simplifying the functions of biological neurons. An ANN performs cognitive or learning processes by interconnecting artificial neurons with strong connections. Recently, data processing based on AI and / or ANNs has been studied. Summary of the Invention
[0006] At least one example embodiment of this disclosure provides a storage system that includes an artificial intelligence (AI) storage device capable of improving or enhancing operational efficiency and reducing power consumption.
[0007] At least one example embodiment of the disclosure provides an AI storage device capable of improving or enhancing work efficiency and reducing power consumption.
[0008] According to an example embodiment, a storage system includes a host device and a storage device. The host device provides first input data and second input data. The storage device is configured to store the first input data and perform an AI computation based on the second input data. The storage device includes a first processor, a first non-volatile memory, a second processor, and a second non-volatile memory. The first processor controls an operation of the storage device. The first non-volatile memory stores the first input data. The second processor performs the AI computation and is different from the first processor. The second non-volatile memory stores weight data associated with the AI computation and is different from the first non-volatile memory.
[0009] According to an example embodiment, a storage device includes a first processor, a first non-volatile memory, a second non-volatile memory, and a second processor. The first processor is configured to control an operation of the storage device. The first non-volatile memory is configured to store first input data for a data storage function. The second non-volatile memory is configured to store weight data associated with an artificial intelligence (AI) computation and is different from the first non-volatile memory. The second processor is configured to perform an AI function and is different from the first processor. The second processor is configured to load the weight data stored in the second non-volatile memory, perform the AI computation based on second input data and the weight data, and output computation result data.
[0010] According to an example embodiment, a storage system includes a first clock / power domain including a first processor and a first non-volatile memory, a second clock / power domain including a second processor and a second non-volatile memory, and a third clock / power domain including a trigger unit. The first non-volatile memory is configured to store first input data in a first operation mode. The second non-volatile memory is configured to store weight data associated with artificial intelligence (AI) computation, the second non-volatile memory being different from the first non-volatile memory, and the second processor is configured to perform the AI computation based on second input data in a second operation mode, the second processor being different from the first processor. The third clock / power domain includes a trigger unit, the third clock / power domain being different from the first clock / power domain and the second clock / power domain. The trigger unit is configured to enable the first operation mode and the second operation mode. The trigger unit is configured to enable the first processor and the first non-volatile memory to store the first input data in the first operation mode, and put the second processor and the second non-volatile memory in an idle state, and enable the second processor to load the weight data stored in the second non-volatile memory to perform the AI computation based on the second input data and the weight data, and output computation result data in the second operation mode.
[0011] The AI storage device and the storage system according to example embodiments can further include a second processor to perform an AI function and an AI computation. The second processor can be independent of or different from a first processor to control an operation of the storage device. The AI function of the storage device can be performed without a control of a host device, and can be performed by the storage device itself. The storage device can receive only second input data as a target of the AI computation, and can output only computation result data as a result of the AI computation. Accordingly, data traffic between the host device and the storage device can be reduced. In addition, the first processor and the second processor can be independent of each other, and non-volatile memories accessed by the first processor and the second processor can also be independent of each other, thereby reducing power consumption. BRIEF DESCRIPTION OF DRAWINGS
[0012] Illustrative, non-limiting example embodiments will be more clearly understood from the following detailed description taken in conjunction with the accompanying drawings.
[0013] Figure 1 is a block diagram illustrating a storage device and a storage system including the same according to some example embodiments.
[0014] Figure 2 is a block diagram illustrating an example of a storage controller included in a storage device according to some example embodiments.
[0015] Figure 3 is a block diagram illustrating an example of a non-volatile memory included in a storage device according to some example embodiments.
[0016] Figure 4 Figure 5A Figure 5B Figure 5C are diagrams for describing operations of a storage system of Figure 1 .
[0017] Figure 6A Figure 6B Figure 6C are diagrams for describing examples of a network structure driven by an AI function implemented in a storage device according to some example embodiments.
[0018] Figure 7 Figure 8A Figure 8B are diagrams for describing operations of a storage system of Figure 1 .
[0019] Figure 9 Figure 10 are diagrams for describing operations of switching operation modes in a storage system according to some example embodiments.
[0020] Figure 11A Figure 11B Figure 11C are diagrams for describing operations of transferring data in a storage system according to some example embodiments.
[0021] Figure 12 is a block diagram illustrating a storage device and a storage system including the same according to some example embodiments.
[0022] Figure 13 is a flowchart illustrating a method of operating a storage device according to some example embodiments.
[0023] Figure 14 is a block diagram illustrating an electronic system according to some example embodiments. DETAILED DESCRIPTION
[0024] Various example embodiments will be described more fully with reference to the accompanying drawings. The disclosure may, however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Like reference numerals refer to like elements throughout the application.
[0025] Figure 1 is a block diagram illustrating a storage device and a storage system including the same according to some example embodiments.
[0026] Referring toFigure 1 The storage system 100 includes a host device 200 and a storage device 300.
[0027] The host device 200 is configured to control overall operations of the storage system 100. The host device 200 can include an external interface (EXT I / F) 210, a host interface (HOST I / F) 220, a host memory (HOST MEM) 230, a neural processing unit (NPU) 240, a digital signal processor (DSP) 250, a central processing unit (CPU) 260, an image signal processor (ISP) 270, and a graphics processing unit (GPU) 280.
[0028] The external interface 210 can be configured to exchange data, signals, events, etc. with the outside of the storage system 100. For example, the external interface 210 can include an input device such as a keyboard, a keypad, a button, a microphone, a mouse, a touchpad, a touch screen, a remote controller, etc., and an output device such as a printer, a speaker, a display, etc.
[0029] The host interface 220 can be configured to provide a physical connection between the host device 200 and the storage device 300. For example, the host interface 220 can provide an interface corresponding to a bus format of the host device 200 to communicate between the host device 200 and the storage device 300. In some example embodiments, the bus format of the host device 200 can be Universal Flash Storage (UFS) and / or Non-Volatile Memory Express (NVMe). In other example embodiments, the bus format of the host device 200 can be Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), Universal Serial Bus (USB), Peripheral Component Interconnect Express (PCIe), Advanced Technology Attachment (ATA), Parallel ATA (PATA), Serial ATA (SATA), etc.
[0030] The NPU 240, the DSP 250, the CPU 260, the ISP 270, and the GPU 280 can be configured to control operations of the host device 200, and can process data associated with the operations of the host device 200.
[0031] For example, the CPU 260 can be configured to control overall operations of the host device 200, and can run an operating system (OS). For example, the OS run by the CPU 260 can include a file system for file management, and a device driver configured to control peripheral devices including the storage device 300 at an OS level. The DSP 250 can be configured to process a digital signal. The ISP 270 can be configured to process an image signal. The GPU 280 can be configured to process various data associated with graphics.
[0032] The NPU 240 can be configured to run and drive a neural network system, and can process corresponding data. At least one of the DSP 250, the CPU 260, the ISP 270, and the GPU 280, in addition to the NPU 240, can also be configured to run and drive a neural network system. Accordingly, the NPU 240, the DSP 250, the CPU 260, the ISP 270, and the GPU 280 can be referred to as a plurality of processing elements (PEs), a plurality of resources, or a plurality of accelerators for driving a neural network system, and can include processing circuitry such as hardware including logic circuitry, a hardware / software combination such as a processor running software, or a combination thereof. For example, the processing circuitry can more specifically include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), and a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), etc.
[0033] The host memory 230 can be configured to store instructions and / or data run and / or processed by the NPU 240, the DSP 250, the CPU 260, the ISP 270, and / or the GPU 280. For example, the host memory 230 can include at least one of various volatile memories such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.
[0034] In some example embodiments, the host device 200 can be an application processor (AP). For example, the host device 200 can be implemented in the form of a system on chip (SoC).
[0035] The host device 200 can be configured to access the storage device 300. The storage device 300 can include a storage controller 310 and a plurality of non-volatile memories (NVMs) 320a, 320b, 320c, and 320d. Although it is illustrated as including four NVMs, example embodiments are not limited thereto, and can include more or less NVMs, for example, five or more.
[0036] The storage controller 310 can be configured to control operations of the storage device 300 and / or the plurality of non-volatile memories 320a, 320b, 320c, and 320d based on commands, addresses, and data received from the host device 200. The configuration of the storage controller 310 will be described with reference to FIG. 3. Figure 2 The configuration of the storage controller 310 will be described.
[0037] The plurality of non-volatile memories 320a, 320b, 320c, and 320d can be configured to store a plurality of data. For example, the plurality of non-volatile memories 320a, 320b, 320c, and 320d can store metadata, various user data, etc.
[0038] In some example embodiments, the plurality of non-volatile memories 320a, 320b, 320c, and 320d can each include NAND flash memory. In other example embodiments, the plurality of non-volatile memories 320a, 320b, 320c, and 320d can each include one of electrically erasable programmable read-only memory (EEPROM), phase-change random access memory (PRAM), resistive random access memory (RRAM), nano floating-gate memory (NFGM), polymer random access memory (PoRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), etc.
[0039] In some example embodiments, the storage device 300 can include universal flash storage (UFS), a multimedia card (MMC), or an embedded multimedia card (eMMC). In other example embodiments, the storage device 300 can include one of a solid state drive (SSD), a secure digital (SD) card, a micro-SD card, a memory stick, a chip card, a universal serial bus (USB) card, a smart card, a compact flash (CF) card, etc.
[0040] In some example embodiments, the storage device 300 can be connected to the host device 200 via a block accessible interface, which can include, for example, UFS, eMMC, an NVMe bus, a SATA bus, a SCSI bus, a SAS bus, etc. The storage device 300 can be configured to provide a block accessible interface to the host device 200 using a block accessible address space corresponding to an access size of the plurality of non-volatile memories 320a, 320b, 320c, and 320d, to allow access to data stored in the plurality of non-volatile memories 320a, 320b, 320c, and 320d in units of storage blocks.
[0041] In some example embodiments, the storage system 100 can be included in at least one of various mobile systems such as a mobile phone, a smart phone, a tablet computer, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a digital camera, a portable game console, a music player, a camcorder, a video player, a navigation device, a wearable device, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, an electronic book reader, a virtual reality (VR) device, an augmented reality (AR) device, a robot device, a drone, etc. In other example embodiments, the storage system 100 can be included in at least one of various computing systems such as a personal computer (PC), a server computer, a data center, a workstation, a digital television, a set-top box, a navigation system, etc.
[0042] The storage device 300 according to some example embodiments is implemented with or equipped with an artificial intelligence (AI) function. For example, the storage device 300 can be configured to function as a storage medium that performs a data storage function, and can also be configured to function as a computing device that runs a neural network system to perform an AI function.
[0043] For example, the storage device 300 can be configured to operate in one of a first operation mode and a second operation mode. In the first operation mode, the storage device 300 can perform a data storage function, such as a write operation for storing first input data UDAT received from the host device 200, a read operation for outputting stored data to the host device 200, and the like. In the second operation mode, the storage device 300 can perform an AI function such as AI computation and / or operation (e.g., arithmetic operation for AI) based on second input data IDAT received from the host device 200 to generate computation result data RDAT, an operation of outputting the computation result data RDAT to the host device 200, and the like. Although Figure 1 Only data is shown to be transmitted, but a command, an address, and the like corresponding to the data can also be transmitted.
[0044] The storage controller 310 can include a first processor 312 and a second processor 314. The first processor 312 can control overall operations of the storage device 300, and can control operations associated with a data storage function in the first operation mode. The second processor 314 can control operations, running, or computation associated with an AI function in the second operation mode. In Figure 1 In an example, the first processor 312 and the second processor 314 can be formed or implemented as one chip, or can include a processing circuit, such as hardware including a logic circuit, a hardware / software combination, such as a processor running software, or a combination thereof. For example, the processing circuit can more specifically include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), and a programmable logic unit, a microprocessor, an application specific integrated circuit (ASIC), and the like.
[0045] The plurality of non-volatile memories 320a, 320b, 320c, and 320d can include (e.g., can be divided or classified into) at least one first non-volatile memory and at least one second non-volatile memory. The first non-volatile memory can be configured to be accessed by the first processor 312, and can be designated or allocated to perform a data storage function. The second non-volatile memory can be configured to be accessed by the second processor 314, and can be designated or allocated to perform an AI function. For example, as will be described with reference to FIG. 4, the first non-volatile memory can be configured to store data, and the second non-volatile memory can be configured to store a neural network system. Figure 4As described, the first non-volatile memory can store the first input data UDAT, and the second non-volatile memory can store weight data associated with the AI computation.
[0046] According to some example embodiments, the first non-volatile memory and the second non-volatile memory can be formed or implemented as one chip, or can be formed or implemented as two independent chips. In some example embodiments, when only the first non-volatile memory is accessed by the first processor 312 and only the second non-volatile memory is accessed by the second processor 314, the plurality of non-volatile memories 320a, 320b, 320c, and 320d can further include a third non-volatile memory accessed by both the first processor 312 and the second processor 314.
[0047] The AI function of the storage device 300 can be performed without dependence on the control of the host device 200, and / or can be performed separately / independently from the control of the host device 200, can be performed by the inside of the storage device 300 and / or by the storage device 300 itself. For example, the neural network system operated by the inside of the storage device 300 can be implemented and / or driven independently of the neural network system operated by the host device 200. The storage device 300 can perform the AI function based on the neural network system operated by the inside of the storage device 300, without the control of the host device 200. According to example embodiments, the neural network system operated by the inside of the storage device 300 and the neural network system operated by the host device 200 can be the same type or different types.
[0048] A conventional storage device does not implement or equip an AI function, but performs the AI function using resources or accelerators included in the host device. When an AI function including a relatively small amount of computation is to be performed, a relatively large resource included in the host device will be used to perform a relatively small amount of computation. In this case, power consumption can increase, and data traffic between the host device and the storage device can exceed the amount of computation, and thus a bottleneck can occur. As a result, when an AI function including a relatively small amount of computation is to be performed, it can be very inefficient to perform the AI function on the host device.
[0049] The storage device 300 according to some example embodiments can be configured to implement an AI function, and the storage device 300 can further include a second processor 314 that performs the AI function and AI computation. The second processor 314 can be independent of or different from the first processor 312 that controls the operation of the storage device 300. The AI function of the storage device 300 can be performed without the control of the host device 200, and can be performed by the storage device 300 itself. The storage device 300 may, for example, be configured to receive only the second input data IDAT that is a target of AI computation, and can be configured to output only the computation result data RDAT that is a result of AI computation. Accordingly, data traffic between the host device 200 and the storage device 300 can be reduced. In addition, the first processor 312 and the second processor 314 can be independent of each other, and the non-volatile memories accessed by the first processor 312 and the second processor 314 can also be independent of each other, thereby reducing power consumption.
[0050] Figure 2 is a block diagram illustrating an example of a storage controller included in a storage device according to some example embodiments.
[0051] Referring to Figure 2 , the storage controller 400 can include a first processor 410, a second processor 420, a buffer memory 430, a host interface 440, an error correction code (ECC) block 450, and a memory interface 460.
[0052] The first processor 410 can be configured to control the operation of the storage controller 400 in response to a command received from a host device (e.g., the host device 200 in Figure 1 ) via the host interface 440. In some example embodiments, the first processor 410 can control the respective components by employing firmware for operating the storage device (e.g., the storage device 300 in Figure 1 ).
[0053] The first processor 410 can be configured to control operations associated with a data storage function, and the second processor 420 can control operations associated with an AI function and AI computation. Figure 2 The first processor 410 and the second processor 420 in Figure 1 may be the same as or similar to the first processor 312 and the second processor 314 in , respectively. For example, the first processor 410 can be a CPU, the second processor 420 can be an NPU, or can include a processing circuit, such as hardware including a logic circuit; a hardware / software combination, such as a processor running software; or a combination thereof.
[0054] In some example embodiments, the second processor 420 can be an NPU, and the NPU included in the storage controller 400 can be smaller than the NPU included in the host device 200 (e.g., the NPU 240 in the host device 200). For example, the second processor 420 can have a smaller data throughput, arithmetic capacity, power consumption, etc. than the NPU 240. Figure 1
[0055] The buffer memory 430 can be configured to store instructions and data executed and processed by the first processor 410 and the second processor 420. For example, the buffer memory 430 can be implemented by a volatile memory (e.g., static random access memory (SRAM), cache memory, etc.) having a relatively small capacity and high speed.
[0056] The ECC block 450 is configured to correct errors and can be configured to perform encoding modulation, for example, by using at least one of a Bose-Chaudhuri-Hocquenghem (BCH) code, a low-density parity-check (LDPC) code, a turbo code, a Reed-Solomon code, a convolutional code, a recursive systematic code (RSC), a trellis coded modulation (TCM), a block coded modulation (BCM), etc., or can perform ECC encoding and ECC decoding using the above-mentioned codes and / or other error correction codes.
[0057] The host interface 440 can be configured to provide a physical connection between the host device 200 and the storage device 300. For example, the host interface 440 can provide an interface corresponding to a bus format of the host for communication between the host device 200 and the storage device 300. The bus format of the host interface 440 can be the same as or similar to the bus format of the host interface 220 in the host device 200. Figure 1
[0058] The memory interface 460 can be configured to exchange data with non-volatile memories (e.g., the non-volatile memories 320a, 320b, 320c, and 320d in the storage device 300). The memory interface 460 can be configured to transmit data to the non-volatile memories 320a, 320b, 320c, and 320d, and / or can be configured to receive data read from the non-volatile memories 320a, 320b, 320c, and 320d. In some example embodiments, the memory interface 460 can be connected to the non-volatile memories 320a, 320b, 320c, and 320d via one channel. In other example embodiments, the memory interface 460 can be connected to the non-volatile memories 320a, 320b, 320c, and 320d via two or more channels. Figure 1
[0059] Figure 3 This is a block diagram illustrating an example of a non-volatile memory included in a storage device according to some example embodiments.
[0060] Reference Figure 3 The non-volatile memory 500 includes a memory cell array 510, a row decoder 520, a page buffer circuit 530, a data input / output (I / O) circuit 540, a voltage generator 550, and a control circuit 560.
[0061] The memory cell array 510 can be connected to the row decoder 520 via multiple string select lines (SSL), multiple word lines (WL), and multiple ground select lines (GSL). The memory cell array 510 can also be connected to the page buffer circuit 530 via multiple bit lines (BL). The memory cell array 510 can include multiple memory cells (e.g., multiple non-volatile memory cells) connected to the multiple word lines (WL) and multiple bit lines (BL). The memory cell array 510 can be divided into multiple memory blocks BLK1, BLK2, ..., BLKz, each containing memory cells. Furthermore, each of the multiple memory blocks BLK1, BLK2, ..., BLKz can be divided into multiple pages.
[0062] In some example embodiments, multiple memory cells may be arranged in a two-dimensional (2D) array structure and / or a three-dimensional (3D) vertical array structure. A three-dimensional vertical array structure may include a vertical string of cells vertically oriented such that at least one memory cell is located above another memory cell. At least one memory cell may include a charge trapping layer. The following patent documents, incorporated herein by reference in their entirety: U.S. Patent Nos. 7,679,133; 8,553,466; 8,654,587; 8,559,235; and U.S. Patent Publication No. 2011 / 0233648, describe suitable configurations of memory cell arrays including 3D vertical array structures, wherein the three-dimensional memory array is configured as multiple levels and shares word lines and / or bit lines between these levels.
[0063] The control circuit 560 can be configured to receive signals from an external source (e.g., Figure 1 The host device 200 and / or storage controller 310 receive commands CMD and addresses ADDR, and can be configured to control erase, program, and read operations on the non-volatile memory 500 based on the commands CMD and addresses ADDR. Erasing operations may include a series of erase cycles, and programming operations may include a series of programming cycles. Each programming cycle may include a programming operation and a programming verification operation. Each erase cycle may include an erase operation and an erase verification operation. Read operations may include normal read operations and data recovery read operations.
[0064] For example, the control circuit 560 can be configured to generate a control signal CON for controlling the voltage generator 550 based on the command CMD, and can generate a control signal PBC for controlling the page buffer circuit 530, and can generate a row address R_ADDR and a column address C_ADDR based on the address ADDR. The control circuit 560 can provide the row address R_ADDR to the row decoder 520, and can provide the column address C_ADDR to the data I / O circuit 540.
[0065] The row decoder 520 can be connected to the memory cell array 510 via a plurality of string select lines SSL, a plurality of word lines WL, and a plurality of ground select lines GSL.
[0066] For example, in a data erase / write / read operation, the row decoder 520 can determine at least one of the plurality of word lines WL as a selected word line based on the row address R_ADDR, and can determine the remaining and / or the rest of the plurality of word lines WL other than the selected word line as unselected word lines.
[0067] In addition, in the data erase / write / read operation, the row decoder 520 can determine at least one of the plurality of string select lines SSL as a selected string select line based on the row address R_ADDR, and can determine the remaining or the rest of the plurality of string select lines SSL other than the selected string select line as unselected string select lines.
[0068] Further, in the data erase / write / read operation, the row decoder 520 can determine at least one of the plurality of ground select lines GSL as a selected ground select line based on the row address R_ADDR, and can determine the remaining or the rest of the plurality of ground select lines GSL other than the selected ground select line as unselected ground select lines.
[0069] The voltage generator 550 can be configured to generate a voltage VS for an operation of the non-volatile memory 500 based on the power PWR and the control signal CON. The voltage VS can be applied to the plurality of string select lines SSL, the plurality of word lines WL, and the plurality of ground select lines GSL via the row decoder 520. In addition, the voltage generator 550 can be configured to generate an erase voltage VERS for a data erase operation based on the power PWR and the control signal CON. The erase voltage VERS can be applied to the memory cell array 510 directly or via the bit line BL.
[0070] For example, during an erase operation, the voltage generator 550 can apply an erase voltage VERS to the common source line and / or the bit line BL of a storage block (e.g., a selected storage block), and can apply an erase pass voltage (e.g., a ground voltage) to all or part of the word lines of the storage block via the row decoder 520. In addition, during an erase verify operation, the voltage generator 550 can apply an erase verify voltage to all of the word lines of the storage block simultaneously, or sequentially to the word lines one by one.
[0071] For example, during a program operation, the voltage generator 550 can apply a program voltage to a selected word line via the row decoder 520, and can apply a program pass voltage to unselected word lines. In addition, during a program verify operation, the voltage generator 550 can apply a program verify voltage to a selected word line via the row decoder 520, and can apply a verify pass voltage to unselected word lines.
[0072] In addition, during a normal read operation, the voltage generator 550 can apply a read voltage to a selected word line via the row decoder 520 and can apply a read pass voltage to unselected word lines. During a data recovery read operation, the voltage generator 550 can apply a read voltage to a word line adjacent to a selected word line via the row decoder 520, and can apply a recovery read voltage to the selected word line.
[0073] The page buffer circuit 530 can be connected to the storage unit array 510 via a plurality of bit lines BL. The page buffer circuit 530 can include a plurality of page buffers. In some example embodiments, each page buffer can be connected to one bit line. In other example embodiments, each page buffer can be connected to two or more bit lines.
[0074] The page buffer circuit 530 can store data DAT to be programmed into the storage unit array 510, or can read data DAT sensed from the storage unit array 510. For example, the page buffer circuit 530 can function as a write driver or a sense amplifier according to an operation mode of the nonvolatile memory 500.
[0075] The data I / O circuit 540 can be connected to the page buffer circuit 530 via a data line DL. The data I / O circuit 540 can be configured to provide data DAT from outside of the nonvolatile memory 500 to the storage unit array 510 via the page buffer circuit 530 based on a column address C_ADDR, or can provide data DAT from the storage unit array 510 to outside of the nonvolatile memory 500.
[0076] Figure 4 、 Figure 5A 、 Figure 5B and Figure 5C are used to describeFigure 1 A diagram illustrating the operation of the storage system. Figure 4 The operation of the storage system 100 in the first operating mode is shown. Figure 5A , Figure 5B and Figure 5C The operation of the storage system 100 in a second operating mode is illustrated. For ease of explanation, components of the storage device that are not particularly relevant to the description of the illustrated example embodiment are omitted.
[0077] Reference Figure 4 , Figure 4 The host device 200 in the middle can be with Figure 1 The host device 200 is the same as or similar to the host device 200. Figure 4 The storage device 300a may include a host interface 440, a first processor 410, a first memory interface 462, a first non-volatile memory 322, a second processor 420, a second memory interface 464, and a second non-volatile memory 324.
[0078] Figure 4 The host interface 440, the first processor 410, and the second processor 420 can be respectively connected to... Figure 2 The host interface 440, the first processor 410, and the second processor 420 are the same or similar. The first memory interface 462 and the second memory interface 464 may be included. Figure 2 The memory interface 460. The first non-volatile memory 322 and the second non-volatile memory 324 can be included. Figure 1 Multiple non-volatile memories 320a, 320b, 320c, and 320d are included. The host interface 440 may be included in a first clock / power domain DM1. The first processor 410, the first memory interface 462, and the first non-volatile memory 322 may be included in a second clock / power domain DM2, which is different from and distinct from the first clock / power domain DM1. The second processor 420, the second memory interface 464, and the second non-volatile memory 324 may be included in a third clock / power domain DM3, which is different from and distinct from the first clock / power domain DM1 and the second clock / power domain DM2.
[0079] In the first operating mode, first input data UDAT can be provided from the host interface 220 of the host device 200, and the storage device 300a can receive the first input data UDAT. For example, the first input data UDAT can be any user data processed by at least one of NPU 240, DSP 250, CPU 260, ISP 270, and GPU 280.
[0080] The storage device 300a can perform a data storage function on the first input data UDAT. For example, the first input data UDAT can be transmitted and stored in the first nonvolatile memory 322 through the host interface 440 and the first memory interface 462. Although the data storage function is described based on a write operation, the example embodiments are not limited thereto, and a read operation can be performed to provide the data UDAT stored in the first nonvolatile memory 322 to the host device 200.
[0081] In Figure 4 In the first operation mode as illustrated, the host interface 440, the first and second processors 410 and 420, the first and second memory interfaces 462 and 464, and the first and second nonvolatile memories 322 and 324 can all be enabled or activated. However, the example embodiments are not limited thereto, and the second processor 420, the second memory interface 464, and the second nonvolatile memory 324 can be in an idle state in the first operation mode, as will be described with reference to Figure 7 .
[0082] Referring Figure 5A , in the second operation mode, the second input data IDAT can be provided from the external interface 210 and the host interface 220 of the host device 200, and the storage device 300a can receive the second input data IDAT. For example, the second input data IDAT can be any inference data that is a target of an AI computation. For example, when the AI function is voice recognition, the second input data IDAT can be voice data received from a microphone included in the external interface 210. For another example, when the AI function is image recognition, the second input data IDAT can be image data received from a camera included in the external interface 210. The second input data IDAT can be transmitted to the second processor 420 through the host interface 440.
[0083] In the second operation mode, the host interface 440, the second processor 420, the second memory interface 464, and the second nonvolatile memory 324 can be enabled or can be in an activated state, and the first processor 410, the first memory interface 462, and the first nonvolatile memory 322 can be switched or changed from the activated state to an idle state (e.g., a sleep, power saving, or power-off state). In Figure 5A and subsequent drawings, components in the idle state are illustrated with a hatched pattern. Only the first processor 410, the first memory interface 462, and the first nonvolatile memory 322 can be included in a separate clock / power domain DM2, such that only the first processor 410, the first memory interface 462, and the first nonvolatile memory 322 are switched to the idle state, and thus power consumption can be reduced in the second operation mode.
[0084] Referring to Figure 5B In the second operation mode, the second processor 420 can load weight data WDAT stored in the second non-volatile memory 324. The weight data WDAT can be transmitted to the second processor 420 through the second memory interface 464. For example, the weight data WDAT can represent a plurality of weight parameters as pre-trained parameters and included in a plurality of layers of a neural network system. The weight data WDAT can be pre-trained to be suitable for or adapted to the neural network system, and can be pre-stored in the second non-volatile memory 324.
[0085] In some example embodiments, the weight data WDAT can be continuously, sequentially, and / or sequentially stored in the second non-volatile memory 324. In this example, the second processor 420 can directly load the weight data WDAT using only a start position (e.g., a start address) at which the weight data WDAT is stored and a size of the weight data WDAT, without a flash translation layer (FTL) operation.
[0086] In some example embodiments, the neural network system includes at least one of various neural network systems and / or machine learning systems such as, for example, an artificial neural network (ANN) system, a convolutional neural network (CNN) system, a deep neural network (DNN) system, a deep learning system, etc. Such machine learning systems can include a variety of learning models such as, for example, a convolutional neural network (CNN), a deconvolutional neural network, a recurrent neural network (RNN) optionally including long short-term memory (LSTM) units and / or gated recurrent units (GRUs), a stacked neural network (SNN), a state space dynamic neural network (SSDNN), a deep belief network (DBN), a generative adversarial network (GAN), and / or a restricted Boltzmann machine (RBM). Alternatively or additionally, such machine learning systems can include other forms of machine learning models such as, for example, linear and / or logistic regression, statistical clustering, Bayesian classification, decision trees, dimensionality reduction such as principal component analysis, and expert systems; and / or combinations thereof including ensembles such as random forests. Such machine learning models can also be used to provide at least one of a variety of services and / or applications such as, for example, an image classification service, a user authentication service based on bio-information or biometric data, an advanced driver assistance system (ADAS) service, a voice assistant service, an automatic speech recognition (ASR) service, etc., and can be executed, run, or processed by the host device 200 and / or the storage device 300a. The configuration of the neural network system will be described with reference to Figure 6A 、 Figure 6B and Figure 6C .
[0087] Referring to Figure 5C In the second operation mode, the second processor 420 can perform AI computation based on the second input data IDAT received in the memory device 300 and the weight data WDAT loaded in the memory device 300 to generate result data RDAT, and can transmit the result data RDAT to the host device 200. The result data RDAT can be transmitted to the host device 200 through the host interface 440. For example, the result data RDAT can represent a result of a multiply-accumulate (MAC) operation performed by a neural network system. Figure 5C Figure 5B In the second operation mode, the second processor 420 can perform AI computation based on the second input data IDAT received in the memory device 300 and the weight data WDAT loaded in the memory device 300 to generate result data RDAT, and can transmit the result data RDAT to the host device 200. The result data RDAT can be transmitted to the host device 200 through the host interface 440. For example, the result data RDAT can represent a result of a multiply-accumulate (MAC) operation performed by a neural network system.
[0088] As described with reference to Figure 5A , Figure 5B and Figure 5C , the second processor 420 can be configured to perform AI functions separately and / or independently in the memory device 300a, and the weight data WDAT can be used only inside the memory device 300a and can not be transmitted to the host device 200. For example, the memory device 300a can exchange only the second input data IDAT as an AI computation target and the result data RDAT as an AI computation result with the host device 200. Generally, the size of the weight data WDAT can be much larger than the size of the second input data IDAT and the size of the result data RDAT. Therefore, since the AI functions are implemented in the memory device 300a, the data traffic between the host device 200 and the memory device 300a can be reduced, and the amount of computation of the host device 200 and the usage rate of the host memory 230 can also be reduced.
[0089] Figure 6A , Figure 6B and Figure 6C are diagrams for describing examples of network structures driven by AI functions implemented in a memory device according to some example embodiments.
[0090] Referring to Figure 6A , a general neural network (e.g., an ANN) can include an input layer IL, a plurality of hidden layers HL1, HL2, …, HLn, and an output layer OL.
[0091] The input layer IL can include i input nodes x1, x2, …, xi, where i is a natural number. Input data (e.g., vector input data) IDAT of length i can be input to the input nodes x1, x2, …, xi, such that each element of the input data IDAT is input to a corresponding input node among the input nodes x1, x2, …, xi. i where i is a natural number. Input data (e.g., vector input data) IDAT of length i can be input to the input nodes x1, x2, …, xi, such that each element of the input data IDAT is input to a corresponding input node among the input nodes x1, x2, …, xi. i where i is a natural number. Input data (e.g., vector input data) IDAT of length i can be input to the input nodes x1, x2, …, xi, such that each element of the input data IDAT is input to a corresponding input node among the input nodes x1, x2, …, xi. i where i is a natural number. Input data (e.g., vector input data) IDAT of length i can be input to the input nodes x1, x2, …, xi, such that each element of the input data IDAT is input to a corresponding input node among the input nodes x1, x2, …, xi.
[0092] The plurality of hidden layers HL1, HL2,..., HLn can include n hidden layers, where n is a natural number, and can include a plurality of hidden nodes h 1 1, h 1 2, h 1 3,..., h 1 m , h 2 1, h 2 2, h 2 3,..., h 2 m , h n 1, h n 2, h n 3,..., h n m For example, the hidden layer HL1 can include m hidden nodes h 1 1, h 1 2, h 1 3,..., h 1 m The hidden layer HL2 can include m hidden nodes h 2 1, h 2 2, h 2 3,..., h 2 m The hidden layer HLn can include m hidden nodes h n 1, h n 2, h n 3,..., h n m where m is a natural number.
[0093] The output layer OL can include j output nodes y1, y2,..., y j where j is a natural number. Each of the output nodes y1, y2,..., y j may correspond to a respective one of the classes to be classified. The output layer OL can output, for each class, an output value ODAT (e.g., a class score or a simple score) associated with the input data IDAT. The output layer OL can be referred to as a fully connected layer and can indicate, for example, a probability that the input data IDAT corresponds to a car.
[0094] Figure 6A The structure of the illustrated neural network can be represented by information about branches (or connections) between nodes that are illustrated as lines and weight values (not shown) assigned to each branch. Nodes within a layer can not be directly connected to each other, but nodes of different layers can be fully or partially connected to each other.
[0095] Each node (e.g., the node h 11) can be configured to receive an output of a previous node (e.g., node x1), to perform a calculation operation, operation, and / or calculation on the received output, and to output a result of the calculation operation, operation, and / or calculation to a subsequent node (e.g., node h 2 1). Each node can calculate a value to be output by applying an input to a specific function (e.g., a non-linear function).
[0096] Generally, a structure of a neural network is set in advance, and a weight value for a connection between nodes is properly set using data having a known answer for a category to which it belongs. The data having a known answer is referred to as "training data", and a process of determining a weight value is referred to as "training". The neural network "learns" during the training process. A set of structures and weight values that can be independently trained is referred to as a "model", and a process of predicting which category input data belongs to by a model having a determined weight value and then outputting a predicted value is referred to as a "test" process.
[0097] Figure 6A The general neural network illustrated can not be suitable for processing input image data (or input sound data) because each node (e.g., node h 1 1) is connected to all nodes (e.g., nodes x1, x2, …, x i ) of the previous layer, and the number of weight values sharply increases as the size of the input image data increases. Therefore, a CNN has been studied by combining a filtering technique with a general neural network, thereby effectively training a two-dimensional image (e.g., input image data) through the CNN.
[0098] Referring to Figure 6B , the CNN can include a plurality of layers CONV1, RELU1, CONV2, RELU2, POOL1, CONV3, RELU3, CONV4, RELU4, POOL2, CONV5, RELU5, CONV6, RELU6, POOL3, and FC.
[0099] Unlike a general neural network, each layer of the CNN can have three dimensions of width, height, and depth, and thus data input to each layer can be volume data having three dimensions of width, height, and depth. For example, if Figure 6B the size of an input image in is 32 width (e.g., 32 pixels) and 32 height and three color channels R, G, and B, the size of input data IDAT corresponding to the input image can be 32*32*3. Figure 6B The input data IDAT in can be referred to as input volume data or input activation volume.
[0100] The convolution layers CONV1, CONV2, CONV3, CONV4, CONV5, and CONV6 can each be configured to perform a convolution operation on the input volume data. For example, in image processing, the convolution operation represents an operation of processing image data based on a mask having weight values, and obtaining an output value by multiplying an input value with a weight value and adding all the multiplied values. The mask can be referred to as a filter, a window, or a kernel.
[0101] The parameters of each convolution layer can consist of a set of learnable filters. Each filter can be small in spatial dimensions (along width and height), but can extend across the entire depth of the input volume. For example, during a forward pass, each filter can be slid over width and height of the input volume (e.g., convolved), and the dot product of the filter’s entries and the input at each position can be computed. As the filter is slid over the width and height of the input volume, a two-dimensional activation map is produced that gives the response at each spatial location of this filter. The result can be an output volume where the activation maps are stacked along the depth dimension. For example, if input volume data of size 32*32*3 is passed through the convolution layer CONV1 having four filters with zero-padding, the output volume data of the convolution layer CONV1 can have a size of 32*32*12 (e.g., the depth of the volume data is increased).
[0102] The RELU layers RELU1, RELU2, RELU3, RELU4, RELU5, and RELU6 can each be configured to perform a rectified linear unit (RELU) operation, which corresponds to an activation function defined by, for example, the function f(x) = max(0, x) (e.g., the output is zero for all negative inputs x). For example, if input volume data of size 32*32*12 is passed through the RELU layer RELU1 to perform a rectified linear unit operation, the output volume data of the RELU layer RELU1 can have a size of 32*32*12 (e.g., the size of the volume data remains).
[0103] The pooling layers POOL1, POOL2, and POOL3 can each be configured to perform a down-sampling operation on the input volume data along the spatial dimensions of width and height. For example, four input values arranged in a 2*2 matrix can be converted into one output value based on a 2*2 filter. For example, a maximum value of the four input values arranged in a 2*2 matrix can be selected based on a 2*2 max-pooling, or an average value of the four input values arranged in a 2*2 matrix can be obtained based on a 2*2 average-pooling. For example, if the input volume data having a size of 32*32*12 is passed through the pooling layer POOL1 having a 2*2 filter, the output volume data of the pooling layer POOL1 can have a size of 16*16*12 (e.g., the width and height of the volume data are reduced, and the depth of the volume data is maintained).
[0104] Generally, one convolution layer (e.g., CONV1) and one RELU layer (e.g., RELU1) can form a pair of CONV / RELU layers in the CNN, the pairs of CONV / RELU layers can be repeatedly arranged in the CNN, and a pooling layer can be periodically inserted in the CNN, thereby reducing the spatial size of the image and extracting features of the image.
[0105] The output layer or fully connected layer FC can output a result (e.g., a class score) of the input volume data IDAT for each class. For example, when the convolution operation and the down-sampling operation are repeatedly performed, the input volume data IDAT corresponding to a two-dimensional image can be converted into a one-dimensional matrix or a vector. For example, the fully connected layer FC can represent probabilities that the input volume data IDAT corresponds to a car, a truck, an airplane, a ship, and a horse.
[0106] The types and numbers of layers included in the CNN can not be limited to those described in the examples, but can be changed according to example embodiments. In addition, although not shown in Figure 6B , the CNN can further include other layers, for example, a soft-max layer for converting a score value corresponding to a prediction result into a probability value, a bias addition layer for adding at least one bias, etc. Figure 6B
[0107] Referring to Figure 6C , the RNN can include a recurrent structure using a specific node or unit N shown on the left side of Figure 6C .
[0108] In Figure 6C The structure shown on the right can represent the recurrent connections of the RNN shown on the left being unrolled (or unfolded). The term "unrolled" refers to writing or showing the network for the complete sequence or entire sequence including all nodes NA, NB, and NC. For example, if the sequence of interest is a 3-word sentence, the RNN can be unrolled into a 3-layer neural network, one layer for each word (e.g., no recurrent connections or no loops).
[0109] In an RNN of the form Figure 6C X represents the input to the RNN. For example, X t can be the input at time step t, and X t-1 and X t+1 can be the inputs at time steps t-1 and t+1, respectively.
[0110] In an RNN of the form Figure 6C S represents the hidden state. For example, S t can be the hidden state at time step t, and S t-1 and S t+1 can be the hidden states at time steps t-1 and t+1, respectively. The hidden state can be computed based on the previous hidden state and the input at the current step. For example, S t = f(UX t + WS t-1 ). For example, the function f can generally be a non-linear function such as tanh or RELU. The S -1 needed to compute the first hidden state can generally be initialized to all zeros.
[0111] In an RNN of the form Figure 6C O represents the output of the RNN. For example, O t can be the output at time step t, and O t-1 and O t+1 can be the outputs at time steps t-1 and t+1, respectively. For example, if the next word in a sentence is to be predicted, it will be a vector of probabilities over the entire vocabulary. For example, O t = softmax(VS t ).
[0112] In an RNN of the form Figure 6C the hidden state can be the "memory" of the network. For example, the RNN can have a "memory" that captures information about the results that have been computed so far. The hidden state S t can capture information about what happened in all previous time steps. The output O t can be computed based only on the memory at the current time step t. In addition, unlike traditional neural networks that use different parameters at each layer, the RNN can share the same parameters at all time steps (e.g.,Figure 6C U, V, and W in FIG. 1). This can represent the fact that the same task can be performed at each step, just using different inputs. This can greatly reduce the total number of parameters that need to be trained or learned.
[0113] In some example embodiments, the neural network system described with reference to Figure 6A , Figure 6B and Figure 6C may perform, run, or process at least one of various services and / or applications (e.g., an image classification service, a user authentication service based on biometric or biometric feature data, an advanced driver assistance system (ADAS) service, a voice assistant service, an automatic speech recognition (ASR) service, etc.).
[0114] Figure 7 , Figure 8A and Figure 8B are diagrams for describing operations of the storage system of Figure 1 . Repetitive descriptions will be omitted with Figure 4 , Figure 5A , Figure 5B and Figure 5C .
[0115] The host device 200 in Figure 7 , Figure 7 may be substantially the same or similar to the host device 200 in Figure 1 . The storage device 300b in Figure 7 may be substantially the same or similar to the storage device 300a in Figure 4 , except that the storage device 300b further includes a trigger unit (TRG) 470.
[0116] The trigger unit 470 can be configured to enable and / or activate the second processor 420, the second memory interface 464, and / or the second non-volatile memory 324 when the operating mode of the storage device 300b changes from the first operating mode to the second operating mode. The trigger unit 470 can be included in a fourth clock / power domain DM4 that is different and distinguished from the first clock / power domain DM1, the second clock / power domain DM2, and the third clock / power domain DM3. According to some example embodiments, the trigger unit 470 and the second processor 420 can be formed or implemented as one chip or two independent chips. In one example, the trigger unit 470 can be configured to enable / activate / trigger the first operating mode and the second operating mode.
[0117] In the first operating mode, first input data UDAT can be provided from the host interface 220 of the host device 200, and the storage device 300b can be configured to receive the first input data UDAT. The storage device 300b can be configured to perform data storage functions on the first input data UDAT.
[0118] exist Figure 7 In the first operating mode shown, the host interface 440, the first processor 410, the first memory interface 462, the first non-volatile memory 322, and the trigger unit 470 can be active, while the second processor 420, the second memory interface 464, and the second non-volatile memory 324 can be idle. Only the second processor 420, the second memory interface 464, and the second non-volatile memory 324 can be included in a separate clock / power domain DM3, so that only the second processor 420, the second memory interface 464, and the second non-volatile memory 324 are idle, thus reducing power consumption in the first operating mode.
[0119] Reference Figure 8A In the second operating mode, the second input data IDAT can be provided from the external interface 210 and the host interface 220 of the host device 200, and the storage device 300b can be configured to receive the second input data IDAT. The second input data IDAT can be sent to the trigger unit 470 through the host interface 440.
[0120] Reference Figure 8B The trigger unit 470 can generate a wake-up signal WK to enable the second processor 420, the second memory interface 464, and the second non-volatile memory 324, and can provide the second input data IDAT to the second processor 420 after the second processor 420, the second memory interface 464, and the second non-volatile memory 324 are enabled.
[0121] In the second operating mode, the host interface 440 and the trigger unit 470 can be in an active state, and the second processor 420, the second memory interface 464, and the second non-volatile memory 324 can switch from an idle state to an active state. Additionally, compared with reference to... Figure 5A The description is similar, stating that the first processor 410, the first memory interface 462, and the first non-volatile memory 322 can switch from an active state to an idle state.
[0122] exist Figure 8B After the operation, the operation of loading the weight data WDAT and generating / sending the calculation result data RDAT can be performed in conjunction with the reference. Figure 5B and Figure 5C The operations described are similar.
[0123] Figure 9 and Figure 10 are diagrams for describing operations of switching operation modes in a storage system according to example embodiments.
[0124] Referring to Figure 9 , an operation mode of the storage device 302 can be switched or a specific operation mode of the storage device 302 can be enabled or activated based on a mode setting signal MSS provided from the host device 202 to the storage device 302.
[0125] For example, the host device 202 can include a plurality of first pins P1 and a second pin P2 different from the plurality of first pins P1, and the storage device 302 can include a plurality of third pins P3 and a fourth pin P4 different from the plurality of third pins P3. A plurality of first signal lines SL1 for connecting the plurality of first pins P1 and the plurality of third pins P3 can be formed between the plurality of first pins P1 and the plurality of third pins P3, and a second signal line SL2 for connecting the second pin P2 and the fourth pin P4 can be formed between the second pin P2 and the fourth pin P4. For example, the pins can be contact pads or contact pins, but example embodiments are not limited thereto.
[0126] The plurality of first pins P1, the plurality of third pins P3, and the plurality of first signal lines SL1 can form a general interface between the host device 202 and the storage device 302, and can be configured to exchange Figure 1 the first input data UDAT, the second input data IDAT, and the calculation result data RDAT shown.
[0127] The second pin P2, the fourth pin P4, and the second signal line SL2 can be formed independently and additionally from the general interface between the host device 202 and the storage device 302, and can be physically added for the mode setting signal MSS. For example, the mode setting signal MSS can be transmitted through the second pin P2 and the fourth pin P4, and the second pin P2 and the fourth pin P4 can be used only for transmitting the mode setting signal MSS. For example, the second pin P2 and the fourth pin P4 can each be a general purpose input / output (GPIO) pin.
[0128] Referring to Figure 10 , a specific storage space S4 of the storage device 304 can be designated or allocated to a special function register (SFR) area for enabling or activating an AI function, and an operation mode of the storage device 304 can be switched or a specific operation mode of the storage device 304 can be enabled or activated based on an address SADDR and setting data SDAT provided from the host device 204 to the storage device 304.
[0129] For example, the host device 204 can include a plurality of first pins P1, the storage device 304 can include a plurality of third pins P3, and a plurality of first signal lines SL1 can be formed between the plurality of first pins P1 and the plurality of third pins P3. Figure 10 The plurality of first pins P1, the plurality of third pins P3, and the plurality of first signal lines SL1 in the storage system 300 can be the same as or similar to the plurality of first pins P1, the plurality of third pins P3, and the plurality of first signal lines SL1 in the storage system 200, respectively. Figure 9 The plurality of first pins P1, the plurality of third pins P3, and the plurality of first signal lines SL1 in the storage system 300 can be the same as or similar to the plurality of first pins P1, the plurality of third pins P3, and the plurality of first signal lines SL1 in the storage system 200, respectively.
[0130] The storage space of the storage device 304 can include or can be divided into a first storage space S1 for an OS, a second storage space S2 for user data, a third storage space S3 for weight data, a fourth storage space S4 for an AI function, and the like. When an address SADDR of the fourth storage space S4 and setting data SDAT for enabling or disabling the AI function are provided to the storage device 304, the second operation mode can be enabled or disabled.
[0131] For example, when the operation mode of the storage device according to an example embodiment is changed or switched, pins P2 and P4 for a mode setting signal MSS can be physically added to a general-purpose interface (as shown in Figure 9 ), or a specific address and storage space can be designated and used to switch the operation mode when using the general-purpose interface (as shown in Figure 10 ).
[0132] Although not shown in Figure 9 and Figure 10 , an unused command field in a command field used in the storage device can be designated and used to switch the operation mode.
[0133] Figure 11A , Figure 11B and Figure 11C are diagrams for describing operations of transmitting data in a storage system according to some example embodiments.
[0134] Referring to Figure 11A , an example in which a voice recognition service is executed based on a neural network service is illustrated, for example, the second input data IDAT is an example of voice data VDAT received from a microphone included in the external interface 210.
[0135] Referring to Figure 11B, shows the interface IF1 between the host device 200 and the storage device 300 when the voice data VDAT is sampled and transmitted in real time. There is initially an idle interval TID1, and then the sampling data D1, D2, D3, D4, D5, D6, D7, D8, D9, D10, D11, and D12 are sequentially transmitted. In this case, the sampling data D1 to D12, which are relatively small in size, are transmitted at a relatively slow rate (for example, about 24 kHz), and the time interval TA in which no data is transmitted is shorter than the reference time, so the interface (for example, the host interfaces 220 and 440) between the host device 200 and the storage device 300 does not enter the sleep state during the time interval TA.
[0136] Referring to Figure 11C , shows the interface IF2 between the host device 200 and the storage device 300 when the voice data VDAT is sampled and the sampling data D1 to D12 are collected in a predetermined number and transmitted. There can be initially an idle interval TID2, and then three of the sampling data D1 to D12 can be collected and transmitted at a time. Compared to the operation of Figure 11B , the time interval TH in which no data is transmitted can be longer than the reference time, so the interface (for example, the host interfaces 220 and 440) between the host device 200 and the storage device 300 can enter the sleep state during the time interval TH.
[0137] Figure 12 is a block diagram showing a storage device and a storage system including the same according to an example embodiment. The description overlapping with Figure 1 will be omitted.
[0138] Referring to Figure 12 , the storage system 100c includes the host device 200 and a storage device 300c.
[0139] Figure 12 The storage system 100c of Figure 1 may be the same as or similar to the storage system 100 except that the configuration of the storage controller 310c included in the storage device 300c is changed.
[0140] The storage controller 310c includes a first processor 312. A second processor 314 is located or disposed outside the storage controller 310c. In Figure 12 example, the first processor 312 and the second processor 314 can be formed or implemented as two independent chips.
[0141] In some example embodiments, when the storage system 100c further includes a trigger unit (for example, the trigger unit 470 in Figure 7 , the trigger unit 470 and the second processor 420 can be formed or implemented as one chip or two independent chips.
[0142] Figure 13 is a flowchart illustrating a method of operating a storage device according to some example embodiments.
[0143] Referring to Figure 1 and Figure 13 In the method of operating a storage device according to some example embodiments, the storage device 300 is configured to perform a data storage function in a first operation mode (step S100). For example, the first processor 312 can store user data in the first non-volatile memory 322, or can read user data stored in the first non-volatile memory 322.
[0144] The storage device 300 performs an AI function in a second operation mode. For example, the second processor 314 can receive inference data from the host device 200 (step S210), can load weight data from the second non-volatile memory 324 (step S220), can perform AI computation based on the inference data and the weight data to generate computation result data (step S230), and can transmit the computation result data to the host device 200 (step S240).
[0145] In some example embodiments, the host device 200 and the storage device 300 can run or drive a neural network system, respectively, and according to example embodiments, the running results can be integrated in the storage system 100. For example, the storage system 100 can be used to drive and compute two or more neural network systems simultaneously to perform complex inference operations. For example, when performing sound or speech recognition, the host device 200 can drive a CNN to recognize lip movements, while the storage device 300 can drive an RNN to recognize the sound or speech itself, and then the recognition results can be integrated to improve recognition accuracy.
[0146] Figure 14 is a block diagram illustrating an electronic system according to some example embodiments.
[0147] Referring to Figure 14 , the electronic system 4000 includes at least one processor 4100, a communication module 4200, a display / touch module 4300, a storage device 4400, and a storage 4500. For example, the electronic system 4000 can be any mobile system or any computing system.
[0148] The processor 4100 is configured to control an operation of the electronic system 4000. For example, the processor 4100 can execute an OS and at least one application program to provide an Internet browser, a game, a video, etc. The communication module 4200 is configured to perform wireless or wired communication with an external system. The display / touch module 4300 is configured to display data processed by the processor 4100 and / or receive data through a touch panel. The storage 4400 is configured to store user data. The storage 4500 temporarily stores data used for processing an operation of the electronic system 4000. The processor 4100 and the storage 4400 can correspond to the host device 200 and the storage device 300 in FIG. 4, respectively. Figure 1
[0149] The inventive concept can be applied to various electronic devices and / or systems including a storage device and a storage system. For example, the inventive concept can be applied to systems such as a mobile phone, a smart phone, a tablet computer, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a digital camera, a portable game console, a music player, a camcorder, a video player, a navigation device, a wearable device, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, an electronic book reader, a virtual reality (VR) device, an augmented reality (AR) device, a robot device, a drone, etc.
[0150] The foregoing is a summary of some example embodiments and should not be interpreted as a limitation on the scope of such embodiments. Although some example embodiments have been described, those skilled in the art will readily understand that many modifications can be made to the example embodiments without materially departing from the novel teachings and advantages of the example embodiments. Accordingly, all such variations are intended to be included within the scope of the example embodiments as defined in the claims. Accordingly, it should be understood that the foregoing is a summary of various example embodiments and should not be interpreted as a limitation on the scope of the disclosed specific example embodiments, and that modifications to the disclosed example embodiments, as well as other example embodiments, are intended to be included within the scope of the claims.
Claims
1. A storage system, comprising: A host device configured to provide first input data and second input data; as well as A storage device configured to: store the first input data in a first operating mode, and generate computation result data by performing the artificial intelligence computation based on the second input data and weight data associated with the artificial intelligence computation in a second operating mode, the storage device comprising: A first processor is configured to control operations of the storage device associated with its data storage function. A first non-volatile memory, configured to store the first input data. A second processor, configured to perform the artificial intelligence calculations, is different from the first processor. A second non-volatile memory, configured to store the weight data, is different from the first non-volatile memory.
2. The storage system according to claim 1, wherein, The second processor is configured to respond to the receipt of the second input data. Load the weight data stored in the second non-volatile memory. The artificial intelligence calculation is performed based on the second input data and the weight data to generate the calculation result data, and The calculation result data is sent to the host device.
3. The storage system according to claim 2, wherein, The weight data is used by the storage device and is not sent to the host device.
4. The storage system according to claim 2, wherein, The weight data represents multiple weight parameters that are used as pre-training parameters and are included in multiple layers of the neural network system, and The calculation result data represents the result of the multiplication and accumulation operation performed by the neural network system.
5. The storage system according to claim 4, wherein, The second processor is a neural processing unit configured to drive the neural network system.
6. The storage system according to claim 4, wherein, The neural network system includes at least one of artificial neural network systems, convolutional neural network systems, recurrent neural network systems, and deep neural network systems.
7. The storage system according to claim 1, wherein, The second operating mode is enabled based on a mode setting signal provided from the host device to the storage device.
8. The storage system according to claim 7, wherein, The host device includes a plurality of first pins and second pins. The plurality of first pins are configured to exchange the first input data, the second input data, and the calculation result data with the storage device. The second pins are configured to exchange the mode setting signal with the storage device. The storage device further includes a plurality of third pins and a fourth pin, wherein the plurality of third pins are configured to exchange the first input data, the second input data and the calculation result data with the host device, and the fourth pin is configured to exchange the mode setting signal with the host device.
9. The storage system according to claim 1, wherein, The first address of the storage device and the first storage space in the storage device corresponding to the first address are designated as a special function register region, and The second operating mode is enabled in response to the first address and first setting data being provided from the host device.
10. The storage system according to claim 1, wherein, The first processor and the first non-volatile memory are enabled in the first operating mode and switched to an idle state in the second operating mode.
11. The storage system according to claim 10, wherein, The second processor and the second non-volatile memory are in the idle state in the first operating mode and are enabled in the second operating mode.
12. The storage system according to claim 11, further comprising: A triggering unit is configured to enable the second processor and the second non-volatile memory in response to a change in the operating mode of the storage device from the first operating mode to the second operating mode.
13. The storage system according to claim 10, wherein, The first processor and the first non-volatile memory are included in the first clock / power domain, and The second processor and the second non-volatile memory are included in the second clock / power domain.
14. The storage system according to claim 1, wherein, The host device is configured to collect a plurality of the second input data in the second operating mode in order to send the collected second input data to the storage device.
15. The storage system according to claim 14, wherein, The interface between the host device and the storage device is configured to enter a sleep state in the second operating mode when the collected second input data is not sent.
16. A storage device, comprising: A first processor is configured to control operations of the storage device associated with data storage functions; A first non-volatile memory, configured to store first input data in a first operating mode; A second non-volatile memory is configured to store weight data associated with artificial intelligence computation; as well as The second processor is configured to load the weight data stored in the second non-volatile memory in a second operating mode, generate calculation result data by performing the artificial intelligence calculation based on the second input data and the weight data, and output the calculation result data.
17. The storage device according to claim 16, wherein, The first processor and the second processor are combined into a single chip.
18. The storage device according to claim 16, wherein, The first processor and the second processor are formed as two independent chips.
19. A storage device, comprising: A first clock / power domain, the first clock / power domain including a first processor and a first non-volatile memory, the first non-volatile memory being configured to store first input data in a first operating mode, and the first processor being configured to access the first non-volatile memory. A second clock / power domain, comprising a second processor and a second non-volatile memory, the second non-volatile memory being configured to store weight data associated with artificial intelligence computation, the second non-volatile memory being different from the first non-volatile memory, and the second processor being configured to access the second non-volatile memory and perform the artificial intelligence computation based on second input data in a second operating mode, the second processor being different from the first processor; as well as A third clock / power domain, comprising a trigger unit, is distinct from both the first and second clock / power domains. The trigger unit is configured to enable both the first and second operating modes. Specifically, in the first operating mode, the trigger unit is configured to: enable the first processor and the first non-volatile memory to store the first input data, and to leave the second processor and the second non-volatile memory idle; and in the second operating mode, enable the second processor to load the weight data stored in the second non-volatile memory, thereby performing the artificial intelligence calculation based on the second input data and the weight data, and outputting the artificial intelligence calculation result data.
Citation Information
Patent Citations
Three-Dimensional Semiconductor Memory Devices And Methods Of Fabricating The Same
US20110233648A1
Nonvolatile memory device, operating method thereof and memory system including the same
US8559235B2
Neural network processing method, computer system, and storage medium
WO2019128752A1
KR20190115303A