Storage device, storage controller, and method of operating a neural processor

By introducing reconfigurable neural processors and FPGAs into electronic devices, the problem of complex and time-consuming artificial intelligence calculations has been solved, enabling efficient and low-cost local computing and reducing reliance on remote data centers.

CN112748876BActive Publication Date: 2026-01-20SAMSUNG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011179725.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-10-30
Filing Date
2020-10-29
Publication Date
2026-01-20
Estimated Expiration
2040-10-29

AI Technical Summary

Technical Problem

In existing technologies, the computational operations of artificial intelligence functions are complex and time-consuming, requiring high-performance computing circuits, resulting in high costs and large space occupation, and low access efficiency to remote data centers.

Method used

Employing a reconfigurable neural processor and a field-programmable gate array (FPGA), the system receives application information from the host via an interface circuit, selects a hardware image, and reconfigures the FPGA. Combined with the neural processor and memory controller, it achieves efficient computational operations.

Benefits of technology

It enables efficient and low-cost local AI computing, reducing reliance on remote data centers and improving computing efficiency and space utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112748876B_ABST
    Figure CN112748876B_ABST
Patent Text Reader

Abstract

A storage device, a storage controller, and a method of operating a neural processor are provided. The storage device includes an interface circuit configured to receive application information from a host, a field programmable gate array (FPGA), a neural processor (NPU), and a central processing unit (CPU) configured to select a hardware image from among a plurality of hardware images stored in a memory using the application information and reconfigure the FPGA using the selected hardware image. The NPU is configured to perform an operation using the reconfigured FPGA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent application claims priority to Korean Patent Application No. 10-2019-0136595, filed on October 30, 2019, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein in its entirety by reference. TECHNICAL FIELD

[0002] The disclosure relates to an electronic device, and more particularly, to an electronic device including a neural processor. BACKGROUND

[0003] Artificial intelligence (AI) capabilities have been utilized in various fields. For example, artificial intelligence capabilities can be used to perform speech recognition and image classification in various electronic devices, such as personal computers, laptop computers, tablets, smart phones, and digital cameras. AI capabilities typically perform complex and time-consuming computing operations. These computing operations can require computing circuits having high operating speeds and large-capacity high-speed memories for processing data to implement. However, computing circuits are typically very expensive and occupy a large amount of physical space. Therefore, artificial intelligence capabilities are typically performed using a remote data center that accesses one or more of these computing circuits. SUMMARY

[0004] At least one embodiment of the disclosure provides a reconfigurable neural processor and an electronic device including the same.

[0005] According to an exemplary embodiment of the inventive concept, there is provided a storage apparatus including: an interface circuit configured to receive application information from a host; a field programmable gate array (FPGA); a neural processing unit (NPU); and a central processing unit (CPU) configured to: select a hardware image from among a plurality of hardware images stored in a memory using the application information, and reconfigure the FPGA using the selected hardware image. The NPU is configured to perform an operation using the reconfigured FPGA.

[0006] According to an exemplary embodiment of the inventive concept, there is provided a storage controller including: a neural processing unit (NPU); a field programmable gate array (FPGA); a non-volatile memory (NVM) controller connected to an NVM apparatus located outside the storage controller via a dedicated channel, the NVM apparatus storing a plurality of hardware images; and a data bus connecting the NPU, the FPGA, and the NVM controller. The NVM controller selects one hardware image corresponding to received application information from among the plurality of hardware images, and reconfigures the FPGA using the selected hardware image.

[0007] According to exemplary embodiments of the inventive concept, a mobile device is provided that includes a host device, a storage device having a device controller and a non-volatile memory (NVM) storing a plurality of hardware images, the device controller having a first neural processing unit (NPU), and an interface circuit configured to enable the host device and the storage device to communicate with each other. The device controller is configured to select a hardware image from among the plurality of hardware images using application information received from the host device and reconfigure a field programmable gate array (FPGA) of the first NPU using the selected hardware image.

[0008] According to exemplary embodiments of the inventive concept, a method of operating a neural processing unit (NPU) includes receiving, by the NPU, application information and data from a host device, selecting, by the NPU, one of a plurality of hardware images from a memory using the application information, loading, by the NPU, the selected hardware image to a field programmable gate array (FPGA) within the NPU to configure the FPGA, and executing, by the NPU, a machine learning algorithm associated with the application information on the data using the configured FPGA to generate a result. BRIEF DESCRIPTION OF DRAWINGS

[0009] Aspects and features of the present disclosure will become more fully apparent from the following detailed description, taken in conjunction with the accompanying drawings, wherein:

[0010] Figure 1 is a block diagram illustrating a computing system according to exemplary embodiments of the inventive concept;

[0011] Figure 2 is a diagram illustrating an operation of a neural processing unit according to exemplary embodiments of the inventive concept;

[0012] Figure 3 is a diagram illustrating a configuration of a field programmable gate array (FPGA) according to exemplary embodiments of the inventive concept;

[0013] Figure 4 is a flowchart illustrating a method of operating a neural processing unit according to exemplary embodiments of the inventive concept;

[0014] Figure 5 is a flowchart illustrating a method of operating a neural processing unit according to exemplary embodiments of the inventive concept;

[0015] Figure 6 is a block diagram illustrating a method of operating a neural processing unit according to exemplary embodiments of the inventive concept; and

[0016] Figures 7 to 19 is a block diagram illustrating an electronic device according to some exemplary embodiments of the inventive concept. DETAILED DESCRIPTION

[0017] The terms "unit", "module", etc. used herein or the functional blocks shown in the drawings can be implemented in the form of hardware, software, or a combination thereof configured to perform a specific function.

[0018] In the disclosure, a hardware image I is a hardware / software image for a specific operation performed by a field programmable gate array (FPGA), and can be referred to as a bitstream, a core, or a lookup table according to various embodiments.

[0019] In the disclosure, an image indicates a still image being displayed or a dynamic image changing over time.

[0020] Figure 1 is a block diagram illustrating a computing system according to an exemplary embodiment of the inventive concept.

[0021] Referring to Figure 1 , the electronic device 1 includes a neural processing unit (NPU) 100 and a plurality of intellectual property (IP) (e.g., IP1, IP2, etc.). The electronic device 1 can further include a storage 200. The IP can be a semiconductor intellectual property core, an IP core, or an IP block. The IP can be a reusable unit of an integrated circuit layout design, a unit, or logic that is intellectual property of one party.

[0022] The electronic device 1 can be designed to perform various functions in a semiconductor system. For example, the electronic device 1 can include an application processor. The electronic device 1 can analyze input data in real time based on a neural network and extract effective information. For example, the electronic device 1 can apply input data to a neural network to generate a prediction. The electronic device 1 can perform situation determination based on the extracted information, or can control elements of an electronic apparatus in communication with the electronic device 1. In one example embodiment, the electronic device 1 can be applied to one or more computing devices performing various computing functions. For example, the electronic device 1 can be applied to a robot device such as a drone, an advanced driver assistance system (ADAS), a smart TV, a smart phone, a medical device, a mobile device, an image display device, a measuring device, or an Internet of Things (IoT) device. Furthermore, the electronic device 1 can be installed in at least one of various electronic apparatuses.

[0023] The electronic device 1 can include various IPs. For example, the IP can include a processing unit, a plurality of cores included in the processing unit, a multi-format codec (MFC), a video module (e.g., a camera interface, a joint photographic experts group (JPEG) processor, a video processor, or a mixer), a 3D graphics core, an audio system, a driver, a display driver, a volatile memory, a non-volatile memory, a memory controller, an input / output interface block, and a cache memory.

[0024] As a technology for connecting IPs, a connection scheme based on a system bus 50 is mainly utilized. For example, an Advanced Microcontroller Bus Architecture (AMBA) protocol of ARM Holdings PLC can be applied as a standard bus specification. Bus types of the AMBA protocol can include an Advanced High-Performance Bus (AHB), an Advanced Peripheral Bus (APB), an Advanced Extensible Interface (AXI), AXI4, and an AXI Coherency Extension (ACE). Among the above bus types, the AXI can provide a multiple outstanding address function, a data interleave function, etc. as an interface protocol between IPs. In addition, other types of protocols such as uNetwork of SONICs Inc., CoreConnect of IBM, and an Open Core Protocol of OCP-IP can be applied to the system bus 50.

[0025] The neural processor 100 can generate a neural network, train (or learn) the neural network, operate the neural network based on input data to generate a prediction, or retrain the neural network. That is, the neural processor 100 can perform complex calculations required for deep learning or machine learning.

[0026] In one exemplary embodiment of the inventive concept, the neural processor 100 includes an FPGA (e.g., 300, 140, 25, etc.). The FPGA 300 can perform additional calculations for complex calculations performed by the neural processor 100 (e.g., a neural processor). For example, the neural processor 100 can offload one of its tasks or a portion of a task to the FPGA. The FPGA can include a plurality of programmable gate arrays. The FPGA can reconfigure one or more of the gate arrays by loading a hardware image I. The hardware image I can correspond to a particular application. Additional calculations required for the corresponding application can be processed according to the reconfigured gate arrays. This will be described in more detail below. The additional calculations can be pre-computations, intermediate computations, or post-computations required for the complex calculations performed by the neural processor 100. For example, during speech recognition, pre-computations can be used to extract features from an audio signal being analyzed so that the features can be applied to a neural network later during intermediate computations. For example, during a text-to-speech operation, post-computations can be used to synthesize speech waveforms from text.

[0027] The neural processor 100 can receive various application data from the IPs through the system bus 50 and load a hardware image I suitable for an application into the FPGA based on the application data.

[0028] The neural processor 100 can perform complex calculations associated with generation, training, or retraining of a neural network. However, additional calculations required for the corresponding complex calculations can be performed by the FPGA, so the neural processor 100 can process various calculations independently without the help of an application processor.

[0029] Figure 2 This is a diagram illustrating the operation of a neural processor according to an exemplary embodiment of the inventive concept. Figure 3 This is a diagram illustrating the configuration of an FPGA according to an exemplary embodiment of the inventive concept.

[0030] Reference Figure 2 An exemplary embodiment of the electronic device 1 according to the inventive concept includes a neural processor 100 and a memory 200.

[0031] The neural processor 100 can access the data from external devices (such as...) Figure 1 The IP shown in the diagram receives application information N and data D, processes the computation requested by the application of the IP to generate a computation result Out, and outputs the computation result Out to a corresponding external device. In an exemplary embodiment, the neural processor 100 loads a hardware image I corresponding to the application information N stored in memory 200 into the FPGA 300 to reconfigure one or more gate arrays of the FPGA 300. For example, the hardware image I may be an electronic file (such as an FPGA bitstream containing programming information for the FPGA 300). The FPGA 300 with the reconfigured one or more gate arrays performs computation on the data D to generate a computation result Out, and then outputs the result Out. The execution of the computation may include the use of weight information W. For example, if the computation involves a summation operation on a set of inputs performed by nodes (or neurons) of a neural network, some of these inputs may be modified by weights in the weight information W (e.g., multiplied) before the summation is calculated. For example, if one of the multiple inputs should be considered more than the others when making a prediction, that input may receive a higher weight. According to an exemplary embodiment of the inventive concept, FPGA 300 performs pre-computation on data D to generate a pre-computation result, neural processor 100 performs complex computation on the pre-computation result using weight information W to generate a result Out, and then outputs the result Out. For example, the complex computation can be an operation performed by nodes of a neural network. According to an exemplary embodiment of the inventive concept, neural processor 100 performs complex computation on data D using weight information W to generate intermediate data, FPGA 300 performs post-computation on the intermediate data to generate a result Out, and then outputs the result Out. Since the result Out is based on the previous computation taking into account the weight information W, the result Out reflects the weight information W.

[0032] According to exemplary embodiments of the inventive concept, complex computations requested by the application are associated with artificial intelligence functions or machine learning algorithms, and may be, for example, speech recognition, speech synthesis, image recognition, image classification, or natural language processing computations.

[0033] According to an exemplary embodiment of the inventive concept, when the data D is a user's voice (e.g., audio recorded during a human speaking) and the external device is an application using a speech recognition function (application information N), the neural processor 100 performs a calculation for extracting features from the user's voice received from the external device, and outputs an analysis result based on the extracted features to the application having the speech recognition function. In this case, the FPGA 300 can perform an additional calculation for extracting features from the voice.

[0034] According to an exemplary embodiment of the inventive concept, when the data D is a user's voice (e.g., audio) or a script (e.g., text) and the external device is an application using a natural language processing function (application information N), the neural processor 100 analyzes the voice or the script received from the external device using a word vector, performs a calculation for natural language processing to generate a processed expression, and outputs the processed expression to the application using the natural language processing function. According to some embodiments, the application using the natural language processing function can be used for speech recognition, summarization, translation, user sentiment analysis, a text classification task, a question and answer (Q&A) system, or a chatbot. In this case, the FPGA 300 can perform an additional calculation for natural language processing calculation.

[0035] According to an exemplary embodiment of the inventive concept, when the data D is a script and the external device is an application using a text-to-speech (TTS) function (application information N), the neural processor 100 performs a text-to-phoneme (TTP) conversion, a grapheme-to-phoneme (GTP) conversion calculation, or a prosody adjustment on the script received from the external device to generate a speech synthesis result, and outputs the speech synthesis result to the application using the TTS function. In this case, the FPGA 300 can perform an additional calculation for speech synthesis calculation.

[0036] According to an exemplary embodiment of the inventive concept, when the data D is an image and the external device is an application using an image processing function (application information N), the neural processor 100 performs an activation calculation on the image received from the external device to generate a result, and outputs the result to the application using the image processing function. According to an exemplary embodiment of the inventive concept, the activation calculation can be a non-linear function calculation. Examples of the non-linear function include a tangent hyperbolic (Tanh) function, a sigmoid function, a Gaussian error linear unit (GELU) function, an exponential function, and a logarithmic function. In this case, the FPGA 300 can perform an additional calculation for the activation calculation.

[0037] As Figure 2As illustrated in FIG. 1, the neural processor 100 includes an FPGA 300. The FPGA 300 is a programmable logic device configured based on a hardware image I. According to some embodiments, the FPGA 300 can include at least one of a configurable logic block (CLB), an input output block (IOB), and a configurable connection circuit for connecting two components. For example, the CLB can be configured to perform a complex combination function or to be a simple logic gate such as AND, NAND, OR, NOR, XOR, etc. For example, the CLB can also be configured as a memory element such as a flip-flop.

[0038] Referring to Figure 2 and Figure 3 Loading the hardware image I to the FPGA 300 can cause the FPGA 300 to be reconfigured according to the surrounding environment. That is, the FPGA 300 can be reconfigured by being programmed according to the hardware image I.

[0039] According to an exemplary embodiment of the inventive concept, the FPGA 300 includes a dynamic region 310 and a static region 320.

[0040] In the dynamic region 310, a specific operation can be performed according to the hardware image I loaded based on the application information N. The FPGA 300 in the dynamic region 310 can be a programmable logic device widely used in designing a digital circuit, which is dynamically reconfigured by the hardware image I to perform a specific operation.

[0041] In the static region 320, a specific operation is performed without loading the hardware image I. The FPGA 300 in the static region 320 can perform an operation corresponding to an operation or a simple calculation that must be frequently performed regardless of an application. In one example embodiment, a nonlinear function calculation for an operation frequently performed in artificial intelligence and deep learning regardless of an application can be included in the static region 320. Some examples of the nonlinear function can include a tangent hyperbolic (Tanh) function, a sigmoid function, a Gaussian error linear unit (GELU) function, an exponential function, a logarithmic function, etc.

[0042] For example, when the hardware image I associated with an application is loaded to the FPGA 300, blocks of the dynamic region 310 are reconfigured, and blocks of the static region 320 are not reconfigured. However, the static region 320 can be initially configured or reconfigured to perform a simple calculation at system power-up.

[0043] As Figure 2The memory 200 stores data required for the operation of the neural processor 100, as illustrated in FIG. 2. According to an exemplary embodiment of the inventive concept, the memory 200 includes a reconfiguration information database 210 configured to store information required for computation for artificial intelligence functions. In one exemplary embodiment, the memory 200 stores a plurality of different machine learning algorithms, and the application information N indicates a machine learning algorithm among the plurality of machine learning algorithms to be selected for the neural processor 100 or for the processor 120. For example, one of the machine learning algorithms can be used for speech recognition, another of the machine learning algorithms can be used for image recognition, etc. Each of the machine learning algorithms can include a different type of neural network. For example, the hardware image I can be used to implement one of the plurality of machine learning algorithms.

[0044] According to an exemplary embodiment of the inventive concept, the reconfiguration information database (DB) 210 includes an image database (DB) 211 including a plurality of hardware images I for reconfiguring the FPGA 300, and a weight DB 212 including a plurality of weight information W. According to an exemplary embodiment, the image DB 211 and the weight DB 212 can be independently included in the memory 200. Alternatively, the weight information W can be stored in association with the hardware image I, or the hardware image I and the weight information W can be stored in association with an application.

[0045] According to an exemplary embodiment, the reconfiguration information DB 210 stores at least one hardware image I and at least one weight information W by mapping to the application information N. According to an exemplary embodiment, the application information N, the hardware image I, and the weight information W are stored in the form of a mapping table, and can be stored by associating the hardware image I and the weight information W with the application information N using a pointer. For example, the application information N can be a unique integer identifying one among a plurality of different applications, and an entry of the mapping table is accessed using the integer. For example, if N=3 indicates a third application among a plurality of applications, a third entry of the mapping table can be retrieved, and the retrieved entry identifies a hardware image and weight information to be used for the third application.

[0046] Figure 4 is a flowchart illustrating an operation method of a neural processor according to an exemplary embodiment of the inventive concept, Figure 5 is a flowchart illustrating an operation method of a neural processor according to an exemplary embodiment of the inventive concept.

[0047] Referring to Figure 4, an external device accesses the neural processor to perform an artificial intelligence function (e.g., deep learning), and transmits application information N and data D. The neural processor receives the application information N and the data D (S10). The neural processor determines (or checks) the application information N (S11), and checks whether the FPGA needs to be reconfigured in order to output a result requested by the corresponding application (S12). That is, the neural processor checks whether the hardware image I needs to be loaded. For example, if the next operation to be performed only needs a calculation provided by the static region 320, the hardware image I does not need to be loaded.

[0048] According to exemplary embodiments of the inventive concept, when a calculation requested by the corresponding application needs to reconfigure the FPGA 300, the neural processor 100 operates the FPGA 300 in the dynamic region 310 to generate data (S13). When the FPGA 300 is used without reconfiguration, the neural processor operates the FPGA 300 in the static region 320 to generate data (S14). The neural processor 100 processes the data to generate a result (S15), and then outputs the result (S17). For example, the neural processor 100 can output the result to an IP (e.g., IP1) or the storage 200 through the bus 50.

[0049] The operated region of the FPGA 300 performs data calculation to generate a processing result, and outputs the processing result to the neural processor 100.

[0050] If the NPU 100 receives the same application information N again, the neural processor 100 can operate the FPGA 300 in the dynamic region 310 or the static region 320.

[0051] In Figure 5 , steps S20 and S21 are the same as S10 and S11, respectively, and thus the description of steps S20 and S21 will be omitted.

[0052] Referring to Figure 5 , when the FPGA 300 in the dynamic region 310 is operated, the neural processor 100 accesses a hardware image I corresponding to the application information N among a plurality of hardware images I stored in the memory (S22), and loads the hardware image I into the FPGA (S23). For example, when the dynamic region 310 needs to be reconfigured based on the application information N, a hardware image I associated with the application information N is loaded into the dynamic region 310.

[0053] The FPGA 300 is reconfigured according to the loaded hardware image I. The reconfigured FPGA 300 processes data received from the external device to generate processed data (S24). For example, the processing can be pre-computation. The FPGA 300 can output the processed data to the neural processor 100. The neural processor 100 reads the weight information W corresponding to the application information N from the memory 200. The neural processor 100 applies the weight information W to the processed data output by the FPGA to generate a result Out, and outputs the result Out to the external device (S26). The result Out can include weighted data. For example, the processed data can be multiplied by one or more weights in the weight information W. For example, if the processed data includes several features, the features can be multiplied by weights to generate weighted features, and then operations can be performed on the weighted features to generate the result Out. For example, the operations can be operations performed by nodes or neurons of a neural network.

[0054] The neural processor 100 can more adaptively and efficiently process pre-computation or post-computation for each of various applications for performing artificial intelligence functions. In addition, the neural processor 100 can adaptively reconfigure the FPGA 300 according to the FPGA application to perform additional computation required to perform complex computation. Accordingly, when the neural processor 100 performs complex computation on a neural network, dependence on an application processor or an external processing device can be reduced.

[0055] Figures 6 to 17 is a block diagram of an electronic device illustrating an example embodiment according to the inventive concept.

[0056] Figure 6 is a block diagram of an electronic device illustrating an example embodiment according to the inventive concept, Figure 7 is a block diagram of an electronic device illustrating an example embodiment according to the inventive concept, Figure 6 is a block diagram of a neural processor illustrated in FIG. 1, Figure 8 is a block diagram of an electronic device illustrating an example embodiment according to the inventive concept, Figure 6 is a block diagram of a neural processor illustrated in FIG. 1.

[0057] Referring to Figure 6 , the electronic device 1 includes a host 10 (e.g., a host device), a neural processor 100, and a storage device 500.

[0058] The host 10 can include a processor 10a, a memory 10b, and a communication interface 10c as illustrated in FIG. 1. Figure 1One of the IPs shown in the middle. For example, the host 10 can be a central processing unit (CPU), an application processor (AP), or a graphics processor (GPU) for performing overall operations of the electronic device 1. The storage 500 can include a plurality of non-volatile memory devices. As an example, the non-volatile memory devices can include a flash memory or a resistive memory such as a resistive random access memory (ReRAM), a phase change RAM (PRAM), and a magnetic RAM (MRAM). Also, the non-volatile memory devices can include an integrated circuit including a processor and a RAM (e.g., a processor-in-memory (PIM)).

[0059] In some embodiments, the flash memory included in the non-volatile memory device can be a two-dimensional (2D) memory array or a three-dimensional (3D) memory array. The 3D memory array can be monolithically formed in at least one of a plurality of memory cell arrays having an active region disposed on a silicon base and a circuit associated with an operation of a memory cell and formed on or in the base in physical levels. The term "monolithic" indicates that the layers on the levels of the array are each directly stacked above the layers on the lower levels. The 3D memory array includes vertical NAND strings disposed in a vertical direction such that at least one memory cell is placed above other memory cells. The memory cell can include a charge-trapping layer.

[0060] Referring to Figure 7 The neural processor 100 according to an example embodiment of the inventive concept includes an interface circuit (I / F) 110, a processor 120, a multiply-accumulate calculator (MAC) 130, an FPGA 140, and a memory (or internal memory) 150. In one example embodiment, the MAC 130 is implemented by a floating point unit coprocessor or a mathematical coprocessor.

[0061] In one example embodiment, each of the interface circuit 110, the processor 120, the MAC 130, the FPGA 140, and the memory 150 can be implemented as a separate semiconductor device, a separate semiconductor chip, a separate semiconductor die, or a separate semiconductor package. Alternatively, some of the interface circuit 110, the processor 120, the MAC 130, the FPGA 140, and the memory 150 can be implemented as one semiconductor device, one semiconductor chip, one semiconductor die, or one semiconductor package. For example, the processor 120 and the MAC 130 can be implemented as one semiconductor package, and the FPGA 140 and the memory 150 can be connected to each other through a high-speed line (e.g., a through silicon via (TSV), also referred to as a through hole via) in the semiconductor package.

[0062] Data bus 101 can connect channels between interface circuit 110, processor 120, MAC 130, FPGA 140, and memory 150. Data bus 101 may include multiple channels. In one exemplary embodiment, the multiple channels may indicate multiple communication paths that are driven independently, and the multiple channels may communicate with devices connected to them based on the same communication scheme.

[0063] The neural processor 100 can connect to, via interface circuit 110, such as... Figure 1 or Figure 2 The external device shown communicates. The interface circuit 110 may be based on at least one of various interfaces, such as Double Data Rate (DDR), Low Power DDR (LPDDR), Universal Serial Bus (USB), Multimedia Card (MMC), Peripheral Component Interconnect (PCI), PCI Fast (PCI-E), Advanced Technology Attachment (ATA), Serial ATA (SATA), Parallel ATA (PATA), Small Computer Small Interface (SCSI), Enhanced Small Disk Interface (ESDI), Integrated Drive Electronics (IDE), Mobile Industry Processor Interface (MIPI), Non-Volatile Memory Fast (NVM-e), and Universal Flash (UFS)).

[0064] Processor 120 (e.g., a central processing unit) can control the operation of other components 101, 110, 130, 140, and 150 of neural processor 100 based on application information N and data D received from host 10. For example, processor 120 can be configured to select hardware image I from a plurality of hardware images stored in memory using application information N, and to reconfigure FPGA 140 using the selected hardware image. The selected hardware image can be associated with a selected machine learning algorithm among a plurality of different machine learning algorithms, and application information N can indicate the machine learning algorithm to be selected.

[0065] MAC 130 can apply weight information W to data obtained through computation by FPGA 140 to generate a result Out requested by an external device, and output the result Out to the external device. According to an exemplary embodiment, MAC 130 performs MAC computation to apply weight information W corresponding to application information N to data computed by FPGA 140 to generate a result Out, and then outputs the result Out of the MAC computation.

[0066] According to an exemplary embodiment, FPGA 140 can perform computations based on dynamic region 310 (i.e., a gate array reconfigured based on hardware mirror I loaded from internal memory 150) and can also perform computations on data based on static region 320 (i.e., a fixed gate array).

[0067] According to an example embodiment, the internal memory 150 stores a plurality of hardware images I and a plurality of weight information W. In addition, the internal memory 150 can store preset information, programs, or commands related to the operation or state of the neural processor 100. According to an example embodiment, the internal memory 150 is a non-volatile RAM.

[0068] According to some embodiments, the internal memory 150 is a working memory, and can temporarily store received data and application information, or can temporarily store results obtained through calculations in the processor 120, the MAC 130, and the FPGA 140. According to an example embodiment, the internal memory 150 can be a buffer memory. According to some embodiments, the internal memory 150 can include a cache, a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), a phase-change RAM (PRAM), a flash memory, a static RAM (SRAM), or a dynamic RAM (DRAM).

[0069] According to an example embodiment, the internal memory 150 includes a first memory configured to store a plurality of hardware images I and a plurality of weight information W, and a second memory configured to store other data. According to an example embodiment, the first memory is a non-volatile memory, and the second memory as a working memory is a buffer memory.

[0070] Referring to Figure 8 , the neural processor 100 according to an example embodiment of the inventive concept includes the interface circuit 110, the processor 120, the MAC 130, the FPGA 140, and the external memory controller 160 connected to the external memory 165. For convenience of description, the following description focuses on the differences from Figure 7 , and the description of elements identical to the above-described elements will be omitted.

[0071] In Figure 8 , the neural processor 100 further includes the external memory controller 160 connected to the external memory 165.

[0072] According to some embodiments, the external memory 165 can be a non-volatile memory, a buffer memory, a register, or a static random access memory (SRAM). According to an example embodiment, the external memory 165 is a non-volatile memory, and stores a plurality of hardware images I and a plurality of weight information W. A corresponding one of the plurality of hardware images I can be loaded into the FPGA 140.

[0073] The external memory controller 160 can load any one of the hardware images I from the external memory 165 into the FPGA 140. The external memory controller 160 according to some embodiments can operate under the control of the interface circuit 110, under the control of the processor 120, or under the control of the FPGA 140.

[0074] Figure 9 is a block diagram of an electronic device illustrating an example embodiment according to the inventive concept, Figures 10 to 16 is a block diagram of a storage controller according to an example embodiment of the inventive concept. Figure 9 The following description focuses on the differences from Figure 6 and the description of elements identical to those described above will be omitted.

[0075] Referring to Figure 9 , the electronic device 1 includes a host 10, a storage controller 20, and a storage device 500. The storage device 500 can be connected to the storage controller 20 via a dedicated channel. The storage controller 20 includes a neural processor 100. Figure 2 The neural processor 100 of Figure 10 The neural processor 100 of Figure 2 The FPGA 300 of Figure 10 The FPGA 25 of Figure 2 The memory 200 of Figure 10 The storage device 500 of

[0076] The storage controller 20 can control the operation of the storage device 500. In one example embodiment, the storage controller 20 is connected to the storage device 500 through at least one channel to write or read data. According to some embodiments, the storage controller 20 can be an element provided in the storage device 500, such as a solid state drive (SSD) or a memory card.

[0077] In Figure 10 , the storage controller 20 according to an example embodiment includes a host interface 21, an internal memory 22, a processor 23, a non-volatile memory controller 24, a neural processor 100, and an FPGA 25 connected to each other through a data bus 29.

[0078] The data bus 29 can include a plurality of channels. In one example embodiment, the plurality of channels can indicate a plurality of communication paths driven independently, and the plurality of channels can communicate with a device connected thereto based on the same communication scheme.

[0079] The host interface 21 can be based on at least one of various interfaces such as DDR, LPDDR, USB, MMC, PCI, PCI-E, ATA, SATA, PATA, SCSI, ESDI, IDE, MIPI, NVM-e, and UFS.

[0080] The internal memory 22, which is a working memory, can be a buffer memory. According to some embodiments, the internal memory 22 can include a cache, a ROM, a PROM, an EPROM, an EEPROM, a PRAM, a flash memory, an SRAM, or a DRAM.

[0081] According to some embodiments, the internal memory 22 can be a working memory of the storage controller 20 or a working memory of the neural processor 100.

[0082] The processor 23 can control the overall operation of the storage controller 20.

[0083] The non-volatile memory controller 24 accesses the storage device 500 to control the operation of a plurality of non-volatile memories. In one exemplary embodiment, the non-volatile memory controller 24 is connected to the non-volatile memory device through at least one channel to write data, read data, or delete data.

[0084] The storage device 500 includes a plurality of non-volatile memory devices. The storage device 500 can include a first non-volatile memory 510 configured to store a plurality of hardware images I and a plurality of weight information W, and a second non-volatile memory 520 configured to store other data (e.g., data for operations other than machine learning) about writing, reading, or deletion. According to some embodiments, the first non-volatile memory 510 can allow only reading, and can allow writing or deletion only when a user has a predetermined authority. In one exemplary embodiment, the predetermined authority can be a setting authority or an update authority for the electronic device 1.

[0085] According to an exemplary embodiment of the inventive concept, the non-volatile memory controller 24 is connected to the first non-volatile memory 510 through a first channel CH_A and to the second non-volatile memory 520 through a second channel CH_B. The first channel CH_A and the second channel CH_B can be independent channels. When accessing the plurality of hardware images I and the plurality of weight information W, the neural processor 100 can independently perform efficient computation using the first channel CH_A. The first channel CH_A and the second channel CH_B can be included in a dedicated channel that connects only the non-volatile memory controller 24 to the storage device 500.

[0086] According to some embodiments, the non-volatile memory controller 24 can be connected to the first non-volatile memory 510 through the first channel CH_A and to the second non-volatile memory 520 through the second channel CH_B. Figure 10The different nonvolatile memory controllers 24 can be connected to the first nonvolatile memory 510 and the second nonvolatile memory 520 through a common channel.

[0087] According to Figure 10 In an embodiment shown in FIG. 1, the neural processor 100 includes the MAC 130 and the register 170. In an example embodiment, the register 170 is replaced with an SRAM. In an example embodiment, the register (or SRAM) 170 can store a plurality of hardware images I and a plurality of weight information W.

[0088] The MAC 130 can apply the weight information W to data obtained through the calculation of the FPGA 25, and can output a result requested by an external device.

[0089] The register 170 can store data obtained through the calculation of the FPGA 25, and can store a result to which the weight information W is applied. That is, the register 170 can be a working memory of the MAC 130.

[0090] The FPGA 25 can be connected to an element connected to the data bus 29 through the first channel CH1, and the FPGA 25 can be connected to the neural processor 100 through the second channel CH2. That is, the FPGA 25 can be shared by the neural processor 100 and the storage controller 20. In an example embodiment, the neural processor 100 is directly connected to the FPGA 25 through the second channel CH2. For example, a result of an operation performed by the FPGA 25 can be directly output to the neural processor 100 through the second channel CH2. For example, the neural processor 100 can perform an operation, and can directly provide a result to the FPGA 25 through the second channel CH2 (e.g., a signal line different from a signal line used for the bus 29) for further processing of the FPGA 25.

[0091] According to an example embodiment, when the first channel CH1 and the second channel CH2 are independent, the FPGA 25 can efficiently perform a calculation requested by the neural processor 100 through the second channel CH2 directly connected to the neural processor 100.

[0092] According to an example embodiment, the FPGA 25 can load a hardware image I, and can perform a calculation based on a request of the storage controller 20 or the host 10 received through the first channel CH1.

[0093] In Figure 11 In an embodiment shown in FIG. 1, the storage controller 20 according to an example embodiment of the inventive concept includes a host interface 21, an internal memory 22, a processor 23, a nonvolatile memory controller 24, a neural processor 100, and an FPGA 25 connected to each other through a data bus 29. Figure 2The neural processor 100 of the present disclosure can be Figure 11 The neural processor 100 of the present disclosure, Figure 2 The FPGA 300 of the present disclosure can be Figure 11 The FPGA 25 of the present disclosure, Figure 2 The memory 200 of the present disclosure can be Figure 11 The storage device 500 of the present disclosure. For convenience of description, the following description focuses on the difference from the above-described elements, and the description of the same elements as the above-described elements will be omitted. Figure 10

[0094] In the embodiment shown in Figure 11 The neural processor 100 includes the MAC 130 and the internal memory 170.

[0095] According to an exemplary embodiment of the present disclosure, the internal memory 170 includes storage information P indicating a location in which a plurality of hardware images I and a plurality of weight information W are stored. According to an exemplary embodiment of the present disclosure, the storage information P can be a pointer to a location in which the plurality of hardware images I and the plurality of weight information W are stored. That is, the storage information P can indicate that the plurality of hardware images 501 and the plurality of weight information 502 are stored in a storage location in the storage device 500. The non-volatile memory controller 24 can access the storage device 500 based on the storage information P, and can read at least one hardware image I and at least one weight information W.

[0096] In addition, according to some embodiments, the internal memory 170 can be a working memory for the MAC 130.

[0097] According to some embodiments, the internal memory 170 can include a cache, a ROM, a PROM, an EPROM, an EEPROM, a PRAM, a flash memory, an SRAM, or a DRAM.

[0098] According to an exemplary embodiment, the internal memory 170 can include a first memory configured to store the storage information P, and include a second memory as a working memory. The first memory can be a non-volatile memory.

[0099] In Figure 12 According to an exemplary embodiment of the present disclosure, the storage controller 20 includes a host interface 21, an internal memory 22, a processor 23, a non-volatile memory controller 24, a neural processor 100, and an FPGA 25 connected to each other through a data bus 29. Figure 2 The neural processor 100 of the present disclosure can be Figure 12 The neural processor 100 of the present disclosure, Figure 2 The FPGA 300 of the present disclosure can be Figure 12 The FPGA 25 of the present disclosure, Figure 2 The memory 200 of the present disclosure can be​Figure 12 external memory 165. For convenience of description, the following description focuses on the differences from Figure 10 and Figure 11 and the description of elements identical to the above-described elements will be omitted.

[0100] As shown in Figure 12 , the neural processor 100 includes the MAC 130 and the external memory controller 160.

[0101] The external memory controller 160 is connected to the external memory 165 to control the overall operation of the external memory 165. According to an example embodiment, the external memory controller 160 can read or delete data stored in the external memory 165, or can write new data into the external memory 165.

[0102] The external memory 165 can store data related to the neural processor 100. In one example embodiment, the external memory 165 stores a plurality of hardware images I and a plurality of weight information W. The external memory 165 can also store one or more neural networks executable by the neural processor 100 or functions executed by nodes of the neural network.

[0103] The FPGA 25 can load at least one hardware image I among the plurality of hardware images I stored in the external memory 165. The FPGA 25 can receive the hardware image I to be loaded through the second channel CH2.

[0104] In Figure 13 , the storage controller 20 according to an example embodiment of the inventive concept includes a host interface 21, an internal memory 22, a processor 23, a non-volatile memory controller 24, a neural processor 100, and an FPGA 25 connected to each other through a data bus 29. Figure 2 The neural processor 100 of Figure 13 , the neural processor 100 of Figure 2 , the FPGA 300 of Figure 13 , the FPGA 25 of Figure 2 , the memory 200 of Figure 13 , the internal memory 150 of. For convenience of description, the following description focuses on the differences from Figures 10 to 12 and the description of elements identical to the above-described elements will be omitted.

[0105] The neural processor 100 includes the MAC 130 and the internal memory 150. According to an exemplary embodiment of the inventive concept, the internal memory 150 stores the plurality of hardware images I and the plurality of weight information W. According to some embodiments, the internal memory 150 can be a working memory of the neural processor 100. Alternatively, according to some embodiments, the internal memory 150 can include a first memory including the plurality of hardware images I and the plurality of weight information W, and a second memory configured to store operation data.

[0106] The FPGA 25 can be connected to the elements through a plurality of common channels, and can be accessed by the elements through the data bus 29. According to an exemplary embodiment, the gate array of the FPGA 25 can be reconfigured by connecting to the neural processor 100 through a selected one of the plurality of channels, receiving the hardware image I through the selected channel, and loading the received hardware image I into the FPGA 25. The FPGA 25 having the reconfigured gate array can process the data D received from the host interface 21 through any one of the plurality of channels to generate a computation result, and output the computation result to the neural processor 100 through any one of the plurality of channels.

[0107] In Figure 14 , the storage controller 20 according to an exemplary embodiment of the inventive concept includes the host interface 21, the internal memory 22, the processor 23, the non-volatile memory controller 24, and the neural processor 100 connected to each other through the data bus 29. Figure 2 The neural processor 100 of Figure 14 The neural processor 100 of Figure 2 The FPGA 300 of Figure 14 The FPGA 300 of Figure 2 The memory 200 of Figure 14 The external memory 165 of Figures 10 to 13 For ease of description, the following description focuses on the differences from

[0108] The neural processor 100 according to an exemplary embodiment of the inventive concept includes the MAC 130, the external memory controller 160 connected to the external memory 165, and the FPGA 300.

[0109] The external memory 165 includes the plurality of hardware images I and the plurality of weight information W. In one exemplary embodiment, the FPGA 300 is used only for the operation of the neural processor 100, and is not shared by other elements of the storage controller 20 as shown in Figures 10 to 13

[0110] ​The FPGA 300 is reconfigured by loading the hardware image I via the external memory controller 165 of the neural processor 100, and the reconfigured FPGA 300 performs a computation on data received by the neural processor 100.

[0111] In Figure 15 , the storage controller 20 according to an exemplary embodiment of the inventive concept includes a host interface 21, an internal memory 22, a processor 23, a non-volatile memory controller 24, and a neural processor 100 connected to each other through a data bus 29. Figure 2 The neural processor 100 can be Figure 15 The neural processor 100, Figure 2 The FPGA 300 can be Figure 15 The FPGA 300, Figure 2 The memory 200 can be Figure 15 The internal memory 150. For ease of description, the following description focuses on the differences from Figure 10 and Figure 11 , and the description of elements identical to the above-described elements will be omitted.

[0112] According to an exemplary embodiment of the inventive concept, the internal memory 150 stores a plurality of hardware images I and a plurality of weight information W. According to some embodiments, the internal memory 150 can be a working memory of the neural processor 100. Alternatively, according to some embodiments, the internal memory 150 can include a first memory including the plurality of hardware images I and the plurality of weight information W, and a second memory configured to store operation data.

[0113] In one exemplary embodiment, the FPGA 300 is used only for the operation of the neural processor 100, and is not shared by other elements of the storage controller 20 as shown in Figure 14

[0114] In Figure 16 , the storage controller 20 according to an exemplary embodiment of the inventive concept includes a host interface 21, an internal memory 22, a processor 23, a non-volatile memory controller 24, a neural processor 100, and an FPGA 25 connected to each other through a data bus 29. Figure 2 The neural processor 100 can be Figure 16 The neural processor 100, Figure 2 The FPGA 300 can be Figure 16 The FPGA 25, Figure 2 The memory 200 can be Figure 16 The storage device 500. For ease of description, the following description focuses on the differences from Figures 10 to 15 ​differences, and the description of elements identical to those described above will be omitted.

[0115] In Figure 16 the embodiment shown in FIG. 5, the storage device 500 includes a first nonvolatile memory 521 including the storage information P and a plurality of second nonvolatile memories 525-1, 525-2, and 525-3 configured to store pairs of the hardware image I and the weight information W (e.g., I1and W1, I2and W2, …, and I N and W N ). Each second nonvolatile memory can include at least one pair of the hardware image and the weight information. Although the number of the plurality of second nonvolatile memories is three, the inventive concept is not limited to any particular number of second nonvolatile memories. Figure 16

[0116] The nonvolatile memory controller 24 can determine the locations where the hardware image I and the weight information W are stored from the storage information P stored in the first nonvolatile memory 521, access any one of the plurality of second nonvolatile memories, and read the hardware image I and the weight information W corresponding to the application information. For example, if the storage information P indicates that the hardware image I and the weight information W are stored in the memory 525-2, the nonvolatile memory controller 24 can retrieve the corresponding information from the memory 525-2.

[0117] Figure 17 is a block diagram of an electronic device according to an exemplary embodiment of the inventive concept.

[0118] Referring to Figure 17 , the electronic device 1 includes an application processor 15 and a storage device 200.

[0119] The application processor 15 includes a neural processor 100. The neural processor 100 includes an FPGA 300. The neural processor 100 is configured to perform machine learning (i.e., complex calculations related to a neural network), and the application processor 15 performs calculations for other operations that are not related to a neural network.

[0120] The application processor 15 can directly access the storage device 200. For example, the application processor 15 can directly access the storage device 200 using a channel connecting the application processor 15 to the storage device 200.

[0121] The storage device 200 can be Figure 2 the storage device 200. The description duplicated with Figure 2 will be omitted.

[0122] ​In one example embodiment, the neural processor 100 receives application information directly from the application processor 15. According to some embodiments, the application processor 15 can determine a workload that activates and operates only the neural processor 100 or both the neural processor 100 and the FPGA 300. For example, if the workload of the neural processor 100 is above a certain threshold, the application processor can activate the FPGA 300, otherwise the FPGA is deactivated to reduce power consumption.

[0123] Figure 18 is a block diagram illustrating an electronic device according to an example embodiment of the inventive concept.

[0124] Referring to Figure 18 , the electronic device can be a universal flash storage (UFS) system. The UFS system 1000 includes a UFS host 1100 (e.g., a host device) and a UFS device 1200. The UFS host 1100 and the UFS device 1200 can be connected to each other through a UFS interface 1300. The UFS system 1000 is based on a flash memory as a non-volatile memory device. For example, the UFS system 100 can be used in a mobile device such as a smart phone.

[0125] The UFS host 1100 includes an application 1120, a device driver 1140, a host controller 1160, and a host interface 1180.

[0126] The application 1120 can include various application programs to be executed by the UFS host 1100. The device driver 1140 can be used to drive a peripheral device connected to the UFS host 1100, and can drive the UFS device 1200. The application 1120 and the device driver 1140 can be implemented in software or firmware.

[0127] The host controller 1160 (e.g., a control circuit) can generate a protocol or a command to be provided to the UFS device 1200 according to a request from the application 1120 and the device driver 1140, and can provide the generated command to the UFS device 1200 through the host interface 1180 (e.g., an interface circuit). When a write request is received from the device driver 1140, the host controller 1160 provides a write command and data to the UFS device 1200 through the host interface 1180. When a read request is received, the host controller 1160 provides a read command to the UFS device 1200 through the host interface 1180 and receives data from the UFS device 1200.

[0128] In one example embodiment, the UFS interface 1300 uses a serial advanced technology attachment (SATA) interface. In one embodiment, the SATA interface is divided into a physical layer, a link layer, and a transport layer according to functions.

[0129] The host-side SATA interface 1180 includes a transmitter and a receiver, and the UFS device-side SATA interface 1210 includes a receiver and a transmitter. The transmitter and the receiver correspond to a physical layer of the SATA interface. The transmission part of the host-side SATA interface 1180 is connected to the reception part of the UFS device-side SATA interface 1210, and the transmission part of the UFS device-side SATA interface 1210 can be connected to the reception part of the host-side SATA interface 1180.

[0130] The UFS device 1200 can be connected to the UFS host 1100 through a device device 1210 (e.g., an interface circuit). The host interface 1180 and the device interface 1210 can be connected to each other through a data line for transmitting and receiving data or a power line for providing power.

[0131] The UFS device 1200 includes a device controller 1220 (e.g., a control circuit), a buffer memory 1240, and a nonvolatile memory device 1260. The device controller 1220 can control overall operations (such as writing, reading, and erasing) of the nonvolatile memory device 1260. The device controller 1220 can transmit and receive data to and from the buffer memory 1240 or the nonvolatile memory device 1260 through an address and a data bus. The device controller 1220 can include at least one of a central processing unit (CPU), a direct memory access (DMA) controller, a flash DMA controller, a command manager, a buffer manager, a flash translation layer (FTL), and a flash manager.

[0132] The UFS device 1200 can provide a command received from the UFS host 1100 to the device DMA controller and the command manager through the device interface 1210, the command manager can allocate the buffer memory 1240 so that data is input through the buffer manager, and then can transmit a response signal to the UFS host 1100 when the data is ready to be transferred.

[0133] The UFS host 1100 can transfer data to the UFS device 1200 in response to the response signal. The UFS device 1200 can store the transferred data in the buffer memory 1240 through the device DMA controller and the buffer manager. The data stored in the buffer memory 1240 can be provided to the flash manager through the flash DMA controller, and the flash manager can store the data in a selected address of the nonvolatile memory device 1260 by referring to address mapping information of the flash translation layer (FTL).

[0134] When the data transfer and the program required by the command of the UFS host 1100 are completed, the UFS device 1200 transmits a response signal to the UFS host 1100 through the device interface 1210 and informs the UFS host 1100 that the command is completed. The UFS host 1100 can inform the device driver 1140 and the application 1120 whether the command corresponding to the received response signal has been completed, and can terminate the corresponding command.

[0135] The device controller 1220 and the non-volatile memory device 1260 in the UFS system 1000 can include the neural processor 100, the storage controller, and the memory 200 that have been described with reference to Figures 1 to 17 The device controller 1220 can include the neural processor 100.

[0136] In some embodiments, the non-volatile memory device 1260 can include the memory 200. Alternatively, in some embodiments, the device controller 1220 can include the memory 200. Alternatively, in some embodiments, the memory 200 can be an external memory connected to the neural processor 100 in the device controller 1220. In some embodiments, the FPGA 300 can be included in the neural processor 100. Alternatively, in some embodiments, the FPGA 300 can be included in the device controller 1220 and connected to the neural processor 100.

[0137] Figure 19 Exemplary embodiments of the apparatus including the host device 10 and Figure 7 or Figure 8 the NPU 100 depicted in FIG. 1A. Alternatively, Figure 19 the NPU 100 depicted in FIG. 1A can be used Figure 10 , Figure 11 , Figure 12 , Figure 13 , Figure 14 , Figure 15 or Figure 16The host device 10 also includes its own NPU 100'. The NPU 100' can be referred to as a master NPU, while the NPU 100 can be referred to as a slave NPU or a sub-NPU. In this embodiment, work can be split between the master NPU and the sub-NPU. For example, the host device 10 can use the master NPU to perform a portion of a neural processing operation, and then transfer the remaining portion of the neural processing operation to the sub-NPU. For example, the host device 10 can output commands and data to the interface circuit 110, which can forward the commands and data to the processor 120 to perform the commands on the data itself or using the FPGA 140 and / or MAC 130. If the execution of the commands requires a new hardware image, the processor 120 can reconfigure the FPGA 140 as described above. In another example embodiment, the master NPU is configured to perform a machine learning operation or algorithm of a particular type (e.g., for image recognition, speech recognition, etc.) at a first level of accuracy, and the sub-NPU is configured to perform the same type of machine learning operation or algorithm at a second level of accuracy that is lower than the first level of accuracy. For example, the machine learning operation or algorithm at the first level of accuracy takes longer to perform than the machine learning operation or algorithm at the second level of accuracy. For example, when the workload of the host device 10 is below a certain threshold, the host device 10 can use the master NPU to perform the machine learning algorithm. However, when the workload of the host device 10 exceeds the certain threshold, then the host device 10 offloads some of the operations of the machine learning algorithm or all of the operations of the machine learning algorithm to the sub-NPU. In one example embodiment, the master NPU is implemented to have the same structure as the NPU 100 or the storage controller 20.

[0138] In concluding the detailed description, it should be noted that many variations and modifications will be apparent to those skilled in the art, once given the benefit of the present description.

Claims

1. A storage apparatus comprising: an interface circuit configured to receive application information from a host; a field programmable gate array including a dynamic region and a static region; a neural processor; and a central processing unit configured to select a hardware image from among a plurality of hardware images stored in a memory using the application information, and reconfigure the dynamic region of the field programmable gate array using the selected hardware image, wherein the neural processor is configured to perform an operation using the reconfigured field programmable gate array, wherein a first operation is performed in the dynamic region according to the hardware image loaded based on the application information, wherein a second operation is performed in the static region without loading the hardware image, the second operation including a frequently performed non-linear function calculation irrespective of the application.

2. The memory device of claim 1, wherein, the selected hardware image is associated with a selected machine learning algorithm among a plurality of different machine learning algorithms, and the application information indicates the selected machine learning algorithm to be selected.

3. The memory device of claim 2, wherein, the reconfigured field programmable gate array performs pre-computation on data input to the neural processor to generate a value, and the neural processor performs the selected machine learning algorithm on the value using weight data stored in the memory to generate a result.

4. The storage device of claim 3, further comprising: a product-sum calculator configured to perform the selected machine learning algorithm on the value using the weight data to generate the result.

5. The memory device of claim 2, wherein, the neural processor performs the selected machine learning algorithm on input data using weight data stored in the memory to generate a value, and the reconfigured field programmable gate array performs post-computation on the value to generate a result.

6. The memory device of claim 2, wherein, the reconfigured field programmable gate array performs the selected machine learning algorithm on input data using weight data stored in the memory to generate a result.

7. The memory device of claim 1, wherein, the memory is a static random access memory or a register located inside the neural processor.

8. The storage device of claim 1, further comprising: a non-volatile memory controller connected to the memory, and the memory is located outside the controller including the neural processor, the central processing unit, and the field programmable gate array. 9.A storage controller comprising: a neural processor; a field programmable gate array including a dynamic region and a static region; a non-volatile memory controller connected to a non-volatile memory device located outside the storage controller via a dedicated channel, the non-volatile memory device storing a plurality of hardware images; and a data bus connecting the neural processor, the field programmable gate array, and the non-volatile memory controller, wherein the non-volatile memory controller selects a hardware image corresponding to received application information from among the plurality of hardware images, and reconfigures the dynamic region of the field programmable gate array using the selected hardware image, wherein the neural processor is configured to perform an operation using the reconfigured field programmable gate array, wherein a first operation is performed in the dynamic region according to the hardware image loaded based on the application information, wherein a second operation is performed in the static region without loading the hardware image, the second operation including a frequently performed non-linear function calculation irrespective of the application. the selected hardware image is associated with a selected machine learning algorithm among a plurality of different machine learning algorithms, and the application information indicates the selected machine learning algorithm to be selected.

10. The storage controller of claim 9, wherein, The dedicated channels include a first channel to receive the selected hardware image from a first non-volatile memory of the non-volatile memory device and a second channel to receive data for operations other than machine learning from a second non-volatile memory of the non-volatile memory device.

11. The storage controller of claim 10, wherein, The selected hardware image is associated with a selected machine learning algorithm of a plurality of different machine learning algorithms and the application information indicates the selected machine learning algorithm that is to be selected.

12. The storage controller of claim 11, wherein, The first channel is also to receive weight data stored in the first non-volatile memory for use in executing the selected machine learning algorithm.

13. The storage controller of claim 9, wherein, The neural processor is configured to request the field programmable gate array to perform the operation through a signal line connecting the neural processor to the field programmable gate array that is different from the data bus.

14. The storage controller of claim 12, wherein, The neural processor further includes a multiply-accumulate calculator configured to execute the selected machine learning algorithm on a result of the operation using the weight data.

15. A mobile device comprising: a host device; a storage device including a storage controller and a non-volatile memory storing a plurality of hardware images, the storage controller including a first neural processor; and an interface circuit configured to enable the host device and the storage device to communicate with each other, wherein a field programmable gate array of the first neural processor includes a dynamic region and a static region, wherein the storage controller is configured to select a hardware image from among the plurality of hardware images using application information received from the host device and to reconfigure the dynamic region of the field programmable gate array of the first neural processor using the selected hardware image, wherein the first neural processor is configured to perform an operation using the reconfigured field programmable gate array, wherein a first operation is performed in the dynamic region according to the hardware image loaded based on the application information, wherein a second operation is performed in the static region without loading the hardware image, the second operation including a frequently executed non-linear function calculation that is independent of the application.

16. The mobile device of claim 15, wherein, The selected hardware image is associated with a selected machine learning algorithm of a plurality of different machine learning algorithms and the application information indicates the selected machine learning algorithm that is to be selected.

17. The mobile device of claim 16, wherein, The selected machine learning algorithm is to perform one of speech recognition and image classification.

18. The mobile device of claim 16, wherein, The reconfigured field programmable gate array performs a pre-computation on data input to the storage device from the host device to generate a value for the selected machine learning algorithm and the first neural processor performs the selected machine learning algorithm on the value using weight data stored in the non-volatile memory to generate a result.

19. The mobile device of claim 18, wherein, The first neural processor further includes a multiply-accumulate calculator configured to perform the selected machine learning algorithm on the value using the weight data to generate the result.

20. The mobile device of claim 18, wherein, The host device further includes a second neural processor, wherein the host device is configured to perform a first portion of the operation using the first neural processor and to perform a second portion of the operation using the second neural processor.

21. A method of operating a neural processor, the method comprising: receiving, by the neural processor, application information and data from a host device; selecting, by the neural processor, one of a plurality of hardware images from a memory using the application information; loading, by the neural processor, the selected hardware image into a field programmable gate array within the neural processor to configure the field programmable gate array; and executing, by the neural processor, a machine learning algorithm associated with the application information on the data using the configured field programmable gate array to generate a result, wherein the loading includes configuring a dynamic region of the field programmable gate array using the selected hardware image and leaving a static region of the field programmable gate array, wherein a first operation is performed in the dynamic region according to the hardware image loaded based on the application information, wherein a second operation is performed in the static region without loading a hardware image, the second operation including a frequently executed non-linear function calculation that is application independent.

22. The method of claim 21, wherein, the executing includes: loading weight data from a memory; executing the machine learning algorithm on the data using the loaded weight data to generate a value; and commanding the configured field programmable gate array to perform a post-computation on the value to generate the result.

23. The method of claim 21, wherein, the executing includes: loading weight data from a memory; commanding the configured field programmable gate array to perform a pre-computation on the data to generate a value; and executing the machine learning algorithm on the value using the loaded weight data to generate the result.

Citation Information

Patent Citations

  • A Torsional Damper

    KR1020190136595A

  • Hardware accelerated neural network subgraphs

    US20190286972A1