Method of operating artificial neural network model and storage device performing same
By dividing the artificial neural network model into multiple node groups and allocating it to multiple hardware accelerators, the problem of low operation efficiency of the artificial neural network model under limited resources is solved, and the effect of shortening the running time and reducing power consumption is achieved.
Patent Information
- Application Number
- CN202411163913.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-08-23
- Publication Date
- 2025-06-27
AI Technical Summary
As the scale of artificial neural network models grows, resource management becomes more complex, especially when resources are limited, requiring efficient operation of these models.
By dividing the artificial neural network model into multiple node groups and assigning these node groups to multiple hardware accelerators, the allocation of node groups and model divisions are optimized to improve the efficiency of inference operations.
By optimizing the allocation of node groups and the partition of models, the running time of the artificial neural network model is shortened, the power consumption of equipment that performs inference operations is reduced, and memory space and computing limitations are overcome without significantly affecting the running time.
Smart Images

Figure CN120218145A_ABST
Abstract
Description
Technical Field
[0001] The example embodiments relate to semiconductor integrated circuits and, more particularly, to methods of operating an artificial neural network model. Background Art
[0002] Artificial Intelligence (AI) is a branch of computer science that focuses on creating systems capable of performing tasks that typically require human intelligence. The human brain consists of many nerve cells called neurons. An Artificial Neural Network (ANN) model is a computational model inspired by the structure and function of biological neural networks. The ANN model includes neurons (e.g., also referred to as nodes) organized into several layers, including an input layer, hidden layers, and an output layer.
[0003] Recently, due to the development of AI-related technologies, the provision of systems and services using AI is increasing. For example, as the performance of systems or services using AI increases, ANN models are becoming larger. As ANN models become larger, a large amount of resources are required to operate and manage the ANN models. Therefore, when resources are limited, systems and methods for efficiently running ANN models are needed. Summary of the Invention
[0004] At least one example embodiment of the present disclosure provides a method of operating an artificial neural network model for effectively performing an inference operation using an artificial neural network model provided with a plurality of node groups.
[0005] At least one example embodiment of the present disclosure provides a storage device for performing a method of operating an artificial neural network model.
[0006] According to an example embodiment, a method of operating an artificial neural network model including a plurality of nodes is provided. The method includes: dividing the artificial neural network model into a divided artificial neural network including a plurality of node groups using a first grouping manner, each node group in the plurality of node groups including at least one node of the plurality of nodes; assigning a first subset of the plurality of node groups to a plurality of first hardware accelerators and assigning a second other subset of the plurality of node groups to a plurality of second hardware accelerators using a first corresponding manner to generate an assignment, wherein the operation speed of the plurality of second hardware accelerators is faster than the operation speed of the plurality of first hardware accelerators; running the divided artificial neural network model on the plurality of input values using the plurality of first hardware accelerators and the plurality of second hardware accelerators to generate a plurality of inference result values; for each of the plurality of inference result values, recording activation region information and call counts of the plurality of node groups; and performing at least one of a first operation for changing the assignment and a second operation for changing the divided artificial neural network model based on the activation region information and the call counts.
[0007] According to an example embodiment, a storage device includes: a plurality of non-volatile memories configured to store an artificial neural network model including a plurality of nodes; a plurality of hardware accelerators configured to calculate a plurality of inference result values based on the plurality of input values and the artificial neural network model; and a storage controller configured to control the plurality of non-volatile memories and the plurality of hardware accelerators. The storage controller includes: a model splitting module configured to divide the artificial neural network model into a plurality of node groups, each node group in the plurality of node groups including at least one node of the plurality of nodes; a node group assignment module configured to assign each node group in the plurality of node groups to a corresponding hardware accelerator in the plurality of hardware accelerators to generate an assignment; and a recording module configured to record activation region information and call counts of the plurality of node groups for each of the plurality of inference result values. The node group assignment module is configured to perform a first operation for changing the assignment based on the activation region information and the call counts. The model splitting module is configured to perform a second operation for changing the divided artificial neural network based on the activation region information and the call counts.
[0008] According to an example embodiment, a method of operating an artificial neural network model including a plurality of nodes is provided. The method includes: dividing the artificial neural network model into a divided artificial neural network including a plurality of node groups using a first grouping method, where each node group in the plurality of node groups includes at least one node among the plurality of nodes; assigning a first subset of the plurality of node groups to a plurality of first hardware accelerators and assigning a second other subset of the plurality of node groups to a plurality of second hardware accelerators using a first corresponding method, where the operating speed of the plurality of second hardware accelerators is faster than the operating speed of the plurality of first hardware accelerators; running the divided artificial neural network model on the plurality of input values using the plurality of first hardware accelerators and the plurality of second hardware accelerators to generate a plurality of inference result values; for each inference result value among the plurality of inference result values, recording activation region information and call counts of the plurality of node groups; selecting a first reference inference result value from among the plurality of inference result values; reassigning the plurality of node groups to the plurality of first hardware accelerators and the plurality of second hardware accelerators using a second corresponding method different from the first corresponding method based on the activation region information of the plurality of node groups for the first reference inference result value; and dividing the artificial neural network model into a new divided artificial neural network using a second grouping method different from the first grouping method based on the activation region information.
[0009] In the method of operating an artificial neural network model according to an example embodiment and a storage device for executing the method, node groups with a high activation level can be preferentially reassigned to a plurality of second hardware accelerators having an operating speed faster than that of the plurality of first hardware accelerators. Therefore, the running time of the artificial neural network model can be shortened, and the power consumption of the device performing the inference operation can be reduced. Additionally, the plurality of node groups can be adjusted such that the number of nodes included in the frequently called node groups is reduced based on the activation region information and call counts. Therefore, memory space limitations and computational limitations can be overcome without a significant difference in the running time of the artificial neural network model before and after the adjustment. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 is a flowchart illustrating a method of operating an artificial neural network model according to an example embodiment.
[0011] Figure 2 and Figure 3 is a block diagram illustrating a method of operating an artificial neural network model according to an example embodiment and / or a system for executing the method of operating an artificial neural network model.
[0012] Figure 4 is a diagram illustrating activation region information and setting a plurality of node groups in a method of operating an artificial neural network model according to an example embodiment.
[0013] Figure 5 is a diagram for describing the call count and inference operations in the method of operating an artificial neural network model according to an exemplary embodiment.
[0014] Figure 6 is a flowchart showing a first operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0015] Figure 7 and Figure 8 is a diagram for describing a first operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0016] Figure 9 is a flowchart showing a second operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0017] Figure 10 and Figure 11 is a diagram for describing a second operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0018] Figure 12 and Figure 13 is a diagram showing an example of the way of grouping nodes in the method of operating an artificial neural network model according to an exemplary embodiment.
[0019] Figure 14 is a block diagram showing a storage device according to an exemplary embodiment and a storage system including the storage device.
[0020] Figure 15 is a block diagram showing a storage controller included in the storage device according to an exemplary embodiment.
[0021] Figure 16 is a diagram showing an example of a storage device according to an exemplary embodiment. Detailed Description
[0022] Various exemplary embodiments will be described more fully with reference to the accompanying drawings. However, the present disclosure may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein. Throughout this application, like reference numerals represent like elements.
[0023] Figure 1 is a flowchart showing a method of operating an artificial neural network model according to an exemplary embodiment.
[0024] Figure 1The method can be executed on a device that performs inference operations by dividing an artificial neural network model into multiple parts and allocating these parts to multiple hardware accelerators. Examples of hardware accelerators include graphics processing units (GPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), and digital signal processors (DSPs). For example, the device that performs inference operations can be a storage device, but the example embodiments are not limited thereto, and the device that performs inference operations can be at least one of various electronic devices. For example, the storage device can include a hardware accelerator.
[0025] The artificial neural network model includes multiple layers, where each layer includes multiple nodes. A node group can include some or all of the nodes of one or more layers of the artificial neural network model. For example, in an artificial neural network model including an input layer, a hidden layer, and an output layer, the model can be divided into a first node group including the first half of the nodes of the input layer, the first half of the nodes of the hidden layer, and the first half of the nodes of the output layer; and a second node group including the second half of the nodes of the input layer, the second half of the nodes of the hidden layer, and the second half of the nodes of the output layer.
[0026] Figure 1 The method includes dividing the artificial neural network model into multiple node groups (operation S100). For example, the artificial neural network model can represent a prediction technique based on a mathematical brain model. For example, the artificial neural network model can include multiple nodes, multiple layers, multiple weights, etc. For example, the artificial neural network model can be divided using at least one of various grouping methods. For example, the grouping method at the start of the operation can be referred to as the first grouping method. Reference will be made to Figure 4 , Figure 12 and Figure 13 to describe the grouping methods.
[0027] For example, the artificial neural network model can include multiple nodes and can be divided into multiple node groups. For example, each node group in the multiple node groups can include at least one of the multiple nodes. For example, if the device that performs inference operations is a storage device and the storage device includes a total of N (N is a positive integer) hardware accelerators for running the artificial neural network model, the artificial neural network model can be divided into N node groups.
[0028] Figure 1The method further includes allocating a plurality of node groups to a plurality of first hardware accelerators and a plurality of second hardware accelerators (operation S200). For example, when N is 5, the first node group to the second node group can be respectively allocated to two hardware accelerators of the first type (i.e., the first hardware accelerator), and the third node group to the fifth node group can be respectively allocated to three hardware accelerators of the second type different from the first type (i.e., the second hardware accelerator). For example, the plurality of first hardware accelerators and the plurality of second hardware accelerators can be included in a storage device. For example, the plurality of first hardware accelerators and the plurality of second hardware accelerators can represent devices that perform certain functions in computing faster than a central processing unit (CPU). For example, the plurality of first hardware accelerators and the plurality of second hardware accelerators can represent devices that perform inference operations faster than the central processing unit included in the storage device.
[0029] For example, the operation speed of the plurality of second hardware accelerators can be faster than that of the plurality of first hardware accelerators. For example, various corresponding methods can be used to allocate the plurality of node groups to the plurality of first hardware accelerators and the plurality of second hardware accelerators. For example, the corresponding method at the start of the operation can be referred to as the first corresponding method. For example, a single node group among the plurality of node groups can be assigned to a single hardware accelerator among the first hardware accelerator and the second hardware accelerator.
[0030] Figure 1 The method includes receiving a plurality of input values (operation S300). For example, the plurality of input values can represent data inputs from one or more users for inference operations rather than preset test data.
[0031] Figure 1 The method includes calculating a plurality of inference result values from the plurality of input values by running an artificial neural network model (operation S400). For example, the plurality of first hardware accelerators and the plurality of second hardware accelerators can be used to run an artificial neural network model provided with a plurality of node groups. For example, each of the plurality of first hardware accelerators and the plurality of second hardware accelerators can use a single allocated node group to perform a sub-inference operation on the input value to generate a plurality of sub-inference result values, and the inference result value can be calculated from the plurality of sub-inference result values.
[0032] Figure 1 The method includes recording the activation region information of the plurality of node groups and the call count of each of the plurality of inference result values (i.e., sub-inference result values) (operation S500). For example, the activation region information can include information indicating which node among the plurality of nodes is activated when running the artificial neural network model. For example, a node in a neural network is activated when the node has applied an activation function to its input. For example, the activation region information can be recorded while performing operation S400. Reference will be made to Figure 4Describe the activation area information. For example, the call count may indicate the number of times each of the multiple inference result values is calculated. For example, the call count may be recorded after operation S400 ends. Figure 5 Describes the call count.
[0033] For example, operations S300, S400, and S500 may be repeatedly performed. For example, operations S300, S400, and S500 may represent a single reasoning operation, and when the single reasoning operation is performed M (M is a positive integer) times, multiple activation region information and multiple call counts of multiple reasoning result values may be generated. For example, operation S600 may be performed based on multiple activation region information and multiple call counts of multiple reasoning result values.
[0034] Figure 1 The method includes performing at least one of a first operation and a second operation based on activation region information and a call count (operation S600). For example, the allocation between a plurality of node groups and a plurality of first hardware accelerators and a plurality of second hardware accelerators may be changed through the first operation. For example, the partitioning of the artificial neural network model may be changed through the second operation.
[0035] For example, the first operation may represent an operation of allocating a plurality of node groups using a corresponding manner different from the corresponding manner (eg, the first corresponding manner) in operation S200. Figure 6 , Figure 7 and Figure 8 The first operation is described. For example, the second operation may represent an operation of setting a plurality of node groups using a grouping method different from the grouping method (eg, the first grouping method) in operation S100. Figure 9 , Figure 10 and Figure 11 The second operation is described.
[0036] In an embodiment, a node group having a high activation level is preferentially reassigned to a plurality of second hardware accelerators having an operating speed faster than that of the plurality of first hardware accelerators. Accordingly, the running time of the artificial neural network model can be shortened, and the power consumption of the device performing the inference operation can be reduced. In an embodiment, if a first node group has a higher activation level than a second node group, in the calculation of the result of the neural network model, more nodes in the first node group than in the second node group have been activated. For example, a plurality of node groups can be reset such that the number of nodes included in the frequently called node groups is reduced, and thus, memory space limitations and computational limitations can be overcome without a significant difference in the running time of the artificial neural network model before and after the reset. For example, when it is determined that a first node group in a node group runs more frequently than a specific threshold, the node group can be reset such that the first node group includes fewer nodes or a smaller portion of the neural network than before.
[0037] Figure 2 And Figure 3 is a block diagram illustrating a method of operating an artificial neural network model and / or a system for operating a method of operating an artificial neural network model according to an example embodiment.
[0038] Referring Figure 2 , system 1000 includes a processor 1100, a storage device 1200, an inference optimization module 1300, and an inference module 1400.
[0039] Here, the term "module" may refer to, but is not limited to, software and / or hardware components, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC) that performs certain tasks. A module can be configured to reside in a tangible addressable storage medium and be configured to run on one or more processors. For example, a "module" can include components such as software components, object-oriented software components, class components, and task components, as well as procedures, functions, routines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. A "module" can be divided into multiple "modules" that perform detailed functions.
[0040] In some example embodiments, system 1000 can be a computing system and can be provided as a dedicated system for a method of operating an artificial neural network model according to an example embodiment.
[0041] The processor 1100 may control the operation of the system 1000 and may be utilized when the inference optimization module 1300 and the inference module 1400 perform operations or calculations. For example, the processor 1100 may include a microprocessor, an application processor (AP), a central processing unit (CPU), a digital signal processor (DSP), a graphics processing unit (GPU), or a neural processing unit (NPU). Although Figure 2 the system 1000 is shown to include one processor 1100, the example embodiments are not limited thereto. For example, the system 1000 may include multiple processors. In addition, the processor 1100 may include a cache memory to increase the computing power.
[0042] The storage device 1200 may store data for the operation of the system 1000 and / or the operation of the inference optimization module 1300 and the inference module 1400. For example, the storage device 1200 may store a deep learning model (or data related to the deep learning model) DLM, multiple data DAT, activation region information AA of multiple node groups NDG1, NDG2, ……, and NDGN, and a call count CNT. For example, the multiple data DAT may include sample data, simulation data, real data, and various other data. The real data may also be referred to herein as actual data or measurement data from a manufactured semiconductor device and / or a manufacturing process. The deep learning model DLM may be provided from the storage device 1200 to the inference optimization module 1300. The inference optimization module 1300 may divide the deep learning model DLM to generate a divided deep learning model DLM_D including multiple node groups NDG1, NDG2, ……, and NDGN, and may provide the divided deep learning model DLM_D to the inference module 1400. The deep learning model DLM may include learning training data and generate a generative model of similar data following the distribution of the training data. Hereinafter, the deep learning model DLM may be used with substantially the same meaning as an inference model or an artificial neural network model.
[0043] In some example embodiments, the storage device (or storage medium) 1200 may include any non-transitory computer-readable storage medium for providing commands and / or data to a computer. For example, the non-transitory computer-readable storage medium may include: volatile memories such as static random access memory (SRAM), dynamic random access memory (DRAM), etc.; and non-volatile memories such as flash memory, magnetic random access memory (MRAM), phase change random access memory (PRAM), resistive random access memory (RRAM), etc. The non-transitory computer-readable storage medium may be inserted into a computer, may be integrated in a computer, or may be coupled to the computer through a communication medium such as a network and / or a wireless link.
[0044] The inference optimization module 1300 may generate a partitioned deep learning model DLM_D of the deep learning model DLM with multiple node groups NDG1, NDG2, ……, and NDGN set based on the deep learning model DLM.
[0045] For example, the inference optimization module 1300 may set multiple node groups NDG1, NDG2, ……, and NDGN based on the deep learning model DLM. For example, the inference optimization module 1300 may allocate multiple node groups NDG1, NDG2, ……, and NDGN to multiple hardware accelerators. For example, when the inference module 1400 performs an inference operation using the deep learning model DLM with multiple node groups NDG1, NDG2, ……, and NDGN set, the activation region information AA and the call count CNT of the multiple node groups NDG1, NDG2, ……, and NDGN may be recorded.
[0046] In addition, based on the activation region information AA and the call count CNT of the multiple node groups NDG1, NDG2, ……, and NDGN, the inference optimization module 1300 may perform a first operation of changing the way of allocating the multiple node groups NDG1, NDG2, ……, and NDGN to multiple hardware accelerators and a second operation of changing the way of partitioning the artificial neural network model. In this case, the deep learning model DLM may be Figure 1 the artificial neural network model in Figure 1 the operations S100, S200, S500, and S600 in
[0047] The inference module 1400 may generate an inference result value IFRV based on the input value INV. The inference module 1400 may use the deep learning model DLM with multiple node groups NDG1, NDG2, ……, and NDGN set to perform an inference operation.
[0048] For example, the inference module 1400 may receive the input value INV and may perform an inference operation on the input value INV using (e.g., running) the deep learning model DLM. In this case, the deep learning model DLM may be Figure 1 the artificial neural network model in Figure 1 the operations S300 and S400 in
[0049] In some example embodiments, the inference optimization module 1300 and the inference module 1400 may be implemented in the form of instructions or program code run by the processor 1100. For example, the inference optimization module 1300 and the inference module 1400 may be stored in a computer-readable recording medium. At this time, the processor 1100 may load the instructions or program code of the inference optimization module 1300 and the inference module 1400 into a working memory (e.g., DRAM, etc.).
[0050] In other example embodiments, the processor 1100 may be manufactured to execute the functions of the inference optimization module 1300 and the inference module 1400. For example, the processor 1100 may implement the inference optimization module 1300 and the inference module 1400 by receiving information corresponding to the inference optimization module 1300 and the inference module 1400.
[0051] Reference Figure 3 , the system 2000 includes a processor 2100, an input / output (I / O) device 2200, a network interface 2300, a random access memory (RAM) 2400, a read-only memory (ROM) 2500, and / or a storage device 2600. Figure 3 Shows Figure 2 An example in which all components of the inference optimization module 1300 and the inference module 1400 in are implemented in software.
[0052] The system 2000 may be a computing system. For example, the computing system may be a fixed computing system, such as a desktop computer, a workstation, or a server, or may be a portable computing system, such as a laptop computer.
[0053] The processor 2100 may be substantially the same as Figure 2 the processor 1100 in. For example, the processor 2100 may include cores or processor cores for running any instruction set (e.g., Intel architecture - 32 (IA - 32), 64 - bit extended IA - 32, x86 - 64, PowerPC, Sparc, MIPS, ARM, IA - 64, etc.). For example, the processor 2100 may access a memory (e.g., RAM 2400 or ROM 2500) through a bus and may run instructions stored in the RAM 2400 or ROM 2500. As Figure 3 shown, the RAM 2400 may store a program PR or at least some elements of the program PR corresponding to Figure 2 the inference optimization module 1300 and the inference module 1400 in, and the program PR may enable the processor 2100 to perform operations for partitioning into multiple node groups (e.g., Figure 1Part of operations S100 and S600 in) and / or operations for assigning multiple node groups to multiple hardware accelerators (e.g., Figure 1 Part of operations S200 and S600 in) and / or operations for recording activation region information and call counts of multiple node groups (e.g., Figure 1 Operations S500 in) and / or operations for performing inference on input values (e.g., Figure 1 Operations S300 and S400 in).
[0054] In other words, program PR may include multiple instructions and / or procedures that can be run by processor 2100, and the multiple instructions and / or procedures included in program PR may enable processor 2100 to execute the method of operating an artificial neural network model according to the example embodiments. Each procedure may represent a series of instructions for performing a specific task. A procedure may be referred to as a function, routine, subroutine, or subprogram. Each procedure may process data provided from the outside and / or data generated by another procedure.
[0055] In some example embodiments, RAM 2400 may include any volatile memory, such as SRAM or DRAM.
[0056] Storage device 2600 may store program PR. Program PR or at least some elements of program PR may be loaded from storage device 2600 into RAM 2400 before being run by processor 2100. Storage device 2600 may store files written in a programming language, and program PR or at least some elements of program PR generated by a compiler may be loaded into RAM 2400.
[0057] Storage device 2600 may store data to be processed by processor 2100 or data obtained through the processing of processor 2100. Processor 2100 may process the data stored in storage device 2600 based on program PR to generate new data, and may store the generated data in storage device 2600.
[0058] I / O device 2200 may include input devices, such as a keyboard or a pointing device, and may include output devices, such as a display device or a printer. For example, a user may trigger processor 2100 to run program PR through I / O device 2200, and may provide or check various inputs, outputs, and / or data, etc.
[0059] The network interface 2300 may provide access to a network external to the system 2000. For example, the network may include multiple computing systems and communication links, and the communication links may include wired links, optical links, wireless links, or any other type of link. Various inputs may be provided to the system 2000 through the network interface 2300, and various outputs may be provided to another computing system through the network interface 2300.
[0060] In some example embodiments, the computer program code and / or the inference module 1400 may be stored in a transient or non-transient computer-readable medium. In some example embodiments, the result value from the inference operations performed by the processor 2100 or the value obtained from the arithmetic processing performed by the processor 2100 may be stored in a transient or non-transient computer-readable medium. In some example embodiments, the intermediate values during the inference operations and / or various data generated by the inference operations may be stored in a transient or non-transient computer-readable medium. However, the example embodiments are not limited thereto.
[0061] Figure 4 It is a diagram showing activation region information and setting multiple node groups in a method of operating an artificial neural network model according to an example embodiment.
[0062] Referring to Figure 4 , multiple node groups NG1, NG2, and NG3 may be set based on an artificial neural network model including multiple nodes ND1_A, ND2_A, ND3_A, ND4_A, ND5_DA, ND6_DA, ND7_DA, ND8_DA, and ND9_DA, and activation region information AA may be obtained when performing the inference operation IF. The inference operation IF may be an operation of running the artificial neural network model on a first input value INV1 to calculate a first inference result value IFRV1.
[0063] For example, the first node group NG1 may be set to include the first node ND1_A, the second node ND2_A, the fourth node ND4_A, and the fifth node ND5_DA. For example, the second node group NG2 may be set to include the sixth node ND6_DA. For example, the third node group NG3 may be set to include the third node ND3_A, the seventh node ND7_DA, the eighth node ND8_DA, and the ninth node ND9_DA. For example, before running the artificial neural network model, the first to third node groups NG1, NG2, and NG3 may be set as Figure 4 shown.
[0064] For example, the first to fourth nodes ND1_A, ND2_A, ND3_A, and ND4_A may represent activated nodes that are activated when performing the inference operation IF. For example, an activated node may represent a node that has been accessed or has routed data at least once when performing the inference operation IF. For example, the fifth to ninth nodes ND5_DA, ND6_DA, ND7_DA, ND8_DA, and ND9_DA may represent deactivated nodes that are deactivated when performing the inference operation IF. For example, a deactivated node may represent a node that has never been accessed or has never routed data when performing the inference operation IF.
[0065] For example, the activation region information AA corresponding to the first inference result value IFRV1 may include the activation levels of multiple node groups NG1, NG2, and NG3. For example, the activation levels of the multiple node groups NG1, NG2, and NG3 can be calculated by dividing the number of activated nodes by the total number of nodes and converting the value to a percentile. For example, since the first node ND1_A, the second node ND2_A, and the fourth node ND4_A in the first node group NG1 are activated, the activation level of the first node group NG1 can be calculated as 3 / 4 * 100% = 75%. For example, since there are no activated nodes in the second node group NG2, the activation level of the second node group NG2 can be calculated as 0%. For example, since the third node ND3_A in the third node group NG3 is activated, the activation level of the third node group NG3 can be calculated as 1 / 4 * 100% = 25%.
[0066] For example, the activation region information AA corresponding to the first inference result value IFRV1 may further include information indicating the activated nodes. For example, the information indicating the activated nodes may include the addresses or locations of the activated nodes. For example, the activation region information AA can indicate which nodes are activated and in which node group they are located.
[0067] Figure 5 is a diagram for describing the call count and inference operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0068] Reference Figure 5 , shows the process of performing the inference operation and recording the activation region information AA and the call count CNT, and shows the process of repeatedly executing Figure 1 the operations S300, S400, and S500 in
[0069] For example, an artificial neural network model can be started, and a first inference result value IFRV1 can be calculated through a first inference operation IF1 on a first input value. For example, the activation region information AA corresponding to the first inference result value IFRV1 can include the activation levels of the first to fifth node groups NG1, …, NG5. For example, the call count CNT of the first inference result value IFRV1 can be recorded as “1”.
[0070] For example, after the first inference operation IF1, a second inference operation IF2 can be performed. For example, a second inference result value IFRV2 can be calculated by performing the second inference operation IF2 on a second input value. For example, the activation region information AA corresponding to the second inference result value IFRV2 can be different from the activation region information AA corresponding to the first inference result value IFRV1. For example, the call count CNT of the second inference result value IFRV2 can be recorded as “1”.
[0071] For example, after the second inference operation IF2, a third inference operation IF3 can be performed. For example, a first inference result value IFRV1 can be calculated by performing the third inference operation IF3 on a third input value. For example, the same inference result value can be calculated for different input values (e.g., the first inference result value IFRV1 for the first input value and the third input value). For example, the activation region information AA can be determined by the inference result value, regardless of the chronological order of the inference operations. For example, the activation region information AA recorded through the first inference operation IF1 and the activation region information AA recorded through the third inference operation IF3 can be the same. For example, the call count CNT of the first inference result value IFRV1 can be recorded as “2”.
[0072] In some example embodiments, a total of 10,000 inference operations can be performed. For example, the call count CNT of the first inference result value IFRV1 can be 4,000, the call count CNT of the second inference result value IFRV2 can be 300, the call count CNT of the third inference result value IFRV3 can be 2,000, the call count CNT of the fourth inference result value IFRV4 can be 1,000, and the call count CNT of the fifth inference result value IFRV5 can be 2,700. In this case, since the inference result value with the largest call count CNT is the first inference result value IFRV1, the first operation and the second operation can be performed based on the activation region information AA corresponding to the first inference result value IFRV1. In other example embodiments, the total number of inference operations performed can be determined differently.
[0073] Figure 6 is a flowchart showing a first operation in a method of operating an artificial neural network model according to an example embodiment.
[0074] Reference Figure 6 , operations S611 and S612 can be examples of the first operation of performing the operation S600 in Figure 1 .
[0075] Figure 6 The method of includes selecting a first reference inference result value from among a plurality of inference result values (operation S611). For example, the first reference inference result value can represent the inference result value having the largest call count among the plurality of inference result values (i.e., high activation level). For example, the first operation and the second operation can be performed based on the activation region information corresponding to the first reference inference result value.
[0076] Figure 6 The method of includes reassigning a plurality of node groups to a plurality of first hardware accelerators and a plurality of second hardware accelerators using a second corresponding method different from the first corresponding method (operation S612). For example, node groups with a high activation level can be preferentially reassigned to a plurality of second hardware accelerators having an operation speed faster than that of the plurality of first hardware accelerators, and thus, the running time of the artificial neural network model can be shortened, and the power consumption of the device performing the inference operation can be reduced. For example, if there are a total of B (B is a positive integer) second hardware accelerators, a total of B node groups with a high activation level can be preferentially assigned to a total of B second hardware accelerators.
[0077] Figure 7 and Figure 8 are diagrams for describing the first operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0078] Reference Figure 7 , shows the selection of the first reference inference result value from among a plurality of inference result values ( Figure 6 operation S611 in ). For example, a total of 1000 inference operations are performed. For example, the call count CNT of the first inference result value IFRV1 is 500, the call count CNT of the second inference result value IFRV2 is 100, the call count CNT of the third inference result value IFRV3 is 200, the call count CNT of the fourth inference result value IFRV4 is 150, and the call count CNT of the fifth inference result value IFRV5 is 50.
[0079] For example, since the call count CNT of the first inference result value IFRV1 is the largest, the first inference result value IFRV1 can be selected as the first reference inference result value. Thus, as will be described with reference to Figure 8 , the first operation can be performed based on the activation region information AA corresponding to the first inference result value IFRV1.
[0080] ReferenceFigure 8 , shows the first operation OP1 and Figure 6 the operation S612 in
[0081] The first operation may be an operation of reassigning multiple node groups NG1, NG2, NG3, NG4, and NG5 to multiple first hardware accelerators HA1_1, HA1_2, and HA1_3 and multiple second hardware accelerators HA2_1 and HA2_2 using a second corresponding method CM2. In an embodiment, the second corresponding method CM2 is different from the first corresponding method CM1. Figure 7 For example, the first corresponding method CM1 may represent a way of assigning multiple node groups NG1, NG2, NG3, NG4, and NG5 to multiple first hardware accelerators HA1_1, HA1_2, and HA1_3 and multiple second hardware accelerators HA2_1 and HA2_2 before running an artificial neural network model. For example, as described above in
[0082] The first operation OP1 may be performed after a total of 1000 inference operations have been executed. Figure 7 and Figure 8 may show a case where there are two second hardware accelerators HA2_1 and HA2_2. In this case, the second node group NG2 (about 95%) and the fifth node group NG5 (about 82%) among the multiple node groups NG1, NG2, NG3, NG4, and NG5 with high activation levels are preferentially assigned to the multiple second hardware accelerators HA2_1 and HA2_2.
[0083] For example, the second node group NG2 and the fifth node group NG5 may be preferentially reassigned to the multiple second hardware accelerators HA2_1 and HA2_2 with an operation speed faster than that of the multiple first hardware accelerators HA1_1, HA1_2, and HA1_3, and thus, the running time of the artificial neural network model can be shortened, and the power consumption of the device performing the inference operations can be reduced.
[0084] Figure 9 is a flowchart showing a second operation in a method of operating an artificial neural network model according to an exemplary embodiment.
[0085] Referring to Figure 9 , operations S621 and S622 are examples of the second operation of performing Figure 1 the operation S600 in Figure 6The operation S611 in is basically the same. In the following, the descriptions repeated with Figure 6 and Figure 7 will be omitted.
[0086] Figure 9 The method of includes resetting a plurality of node groups (operation S622) by dividing an artificial neural network model using a second grouping method different from the first grouping method. For example, by resetting the plurality of node groups such that the number of nodes included in the node groups assigned to a plurality of second hardware accelerators is reduced, it is possible to overcome memory space and computational limitations without a significant difference in the running time of the artificial neural network model before and after the reset. For example, the activation level of B (B is a positive integer) node groups can be increased by including the deactivated nodes assigned to the plurality of second hardware accelerators in the B node groups in adjacent node groups. In this case, the total number of nodes included in each of the B node groups can be reduced, and the activation level of each of the B node groups can be increased. For example, the activation level of the first node group can be increased by moving the deactivated nodes of the first node group to the second node group.
[0087] Figure 10 and Figure 11 are diagrams for describing a second operation in the method of operating an artificial neural network model according to an exemplary embodiment.
[0088] Referring to Figure 10 , the second operation OP2 and Figure 9 the operation S622 in are shown. The second operation may be an operation of resetting a plurality of node groups NG1, NG2, NG3, NG4, and NG5 using a second grouping method GM2 different from the first grouping method GM1.
[0089] In the following, it will be assumed that the artificial neural network model includes a total of 500 nodes, and the artificial neural network model is divided such that each of the plurality of node groups NG1, NG2, NG3, NG4, and NG5 includes 100 nodes. For example, the addresses and positions of the activated nodes and deactivated nodes included in the plurality of node groups NG1, NG2, NG3, NG4, and NG5 may be included in the activation region information AA associated with the first inference result value IFRV1.
[0090] For example, multiple node groups can be reset such that the number of deactivated nodes included in the second node group NG2 and the fifth node group NG5 having a high activation level is reduced. For example, the deactivated nodes included in the second node group NG2 and the fifth node group NG5 can be included in the adjacent node groups NG3 and NG4. For example, one or more deactivated nodes in the second node group NG2 can be moved to the third node group NG3, and one or more deactivated nodes in the fifth node group NG5 can be moved to the fourth node group NG4.
[0091] For example, the second node group NG2 can include 95 activated nodes and 5 deactivated nodes. For example, 2 of the 5 deactivated nodes in the second node group NG2 can be included in the third node group NG3. In this case, the new second node group NNG2 can include 95 activated nodes and 3 deactivated nodes, and the new third node group NNG3 can include 7 activated nodes and 95 deactivated nodes. Thus, the activation level of the new second node group NNG2 can be 95 / 98 = approximately 96.9%, and the activation level of the new third node group NNG3 can be 7 / 102 = approximately 6.9%.
[0092] For example, the fifth node group NG5 can include 82 activated nodes and 18 deactivated nodes. For example, 10 of the 18 deactivated nodes in the fifth node group NG5 can be included in the fourth node group NG4. In this case, the new fifth node group NNG5 can include 82 activated nodes and 8 deactivated nodes, and the new fourth node group NNG4 can include 1 activated node and 109 deactivated nodes. Thus, the activation level of the new fifth node group NNG5 can be 82 / 90 = approximately 91%, and the activation level of the new fourth node group NNG4 can be 1 / 110 = approximately 0.9%.
[0093] As referred to above Figure 10 described, when only the number of deactivated nodes included in the second node group NG2 and the fifth node group NG5 is reduced, the activation levels of the second node group NG2 and the fifth node group NG5 can actually increase. Thus, the memory space and computational limitations can be overcome without a significant difference in the running time of the artificial neural network model before and after the second operation OP2.
[0094] Refer to Figure 11, which shows an example of the second operation OP2. For example, before the second operation OP2 is executed, the first node group NG1 can be set to include the first node ND1_A, the second node ND2_A, the fourth node ND4_A, and the fifth node ND5_DA. For example, the second node group NG2 can be set to include the sixth node ND6_DA. For example, the third node group NG3 can be set to include the third node ND3_A, the seventh node ND7_DA, the eighth node ND8_DA, and the ninth node ND9_DA.
[0095] For example, the first node group NG1 can include a total of 4 nodes. For example, the activation level of the first node group NG1 can be 75%, the activation level of the second node group NG2 can be 0%, and the activation level of the third node group NG3 can be 25%.
[0096] For example, the second operation OP2 can be executed to reduce the number of deactivated nodes included in the first node group NG1. For example, the fifth node ND5_DA included in the first node group NG1 can be included in the second node group NG2. In this case, the new first node group NNG1 can include a total of three nodes. For example, the activation level of the new first node group NNG1 can be 100%, which can be greater than the activation level of the first node group NG1.
[0097] Figure 12 and Figure 13 is a diagram showing an example of a node grouping method in a method of operating an artificial neural network model according to an exemplary embodiment.
[0098] Reference Figure 12 , which shows a case where nodes included in one layer among multiple layers (e.g., IL, HL1, HL2, …… HLn, OL) are set to one node group among multiple node groups (e.g., NG1a, NG2a, NG3a, …… NGk-1a, NGka). A general neural network (or artificial neural network) can include an input layer IL, multiple hidden layers HL1, HL2, ……, HLn, and an output layer OL.
[0099] The input layer IL can include i input nodes x1, x2, …… xi, where i is a natural number. An input value INV of length i can be input to the input nodes x1 to xi such that each element of the input value INV is input to a corresponding one of the input nodes x1 to xi. The input value INV can include information associated with various features of different classes to be classified. For example, the input value INV can be an array of features, where each feature is output to a corresponding one input node.
[0100] Multiple hidden layers HL1, HL2, ……, HLn may include n hidden layers, where n is a natural number, and may include multiple hidden nodes h11, h12, h13, ……, h1m, h21, h22, h23, ……, h2m, hn1, hn2, hn3, ……, hnm. For example, the hidden layer HL1 may include m hidden nodes h11 to h1m, the hidden layer HL2 may include m hidden nodes h21 to h2m, and the hidden layer HLn may include m hidden nodes hn1 to hnm, where m is a natural number.
[0101] The output layer OL may include j output nodes y1, y2, ……, yj, where j is a natural number. Each of the output nodes y1 to yj may correspond to a respective one of the multiple classes to be classified. The output layer OL may generate output values (e.g., class scores or numerical outputs such as regression variables) and / or output data associated with the input values INV for each class. In some example embodiments, the output layer OL may be a fully connected layer and may generate an inference result value IFRV corresponding to the input value INV.
[0102] Figure 12 The structure of the neural network shown can be represented by information about the branches (or connections) shown as lines between the nodes and the weighted values assigned to each branch. In some neural network models, the nodes within a layer may not be connected to each other, but the nodes of different layers may be fully or partially connected to each other. In some other neural network models (such as the unrestricted Boltzmann machine), at least some of the nodes within a layer may also be connected to other nodes within a layer in addition to (or alternatively) one or more nodes of other layers.
[0103] Each node (e.g., node h11) may receive the output of a previous node (e.g., node x1), may perform a computational operation, operation, or calculation on the received output, and may output the result of the computational operation, operation, or calculation as an output to the next node (e.g., node h21). Each node may calculate the value to be output by applying the input to a specific function (e.g., a non - linear function). This function may be referred to as the activation function of the node.
[0104] In an example embodiment, the structure of a neural network is preset, and by using sample data with sample answers (also referred to as "labels"), the weighted values for connections between nodes are appropriately set, where the sample answers indicate the categories of data corresponding to the sample input values. The data with sample answers can be referred to as "training data", and the process of determining the weighted values can be referred to as "training". During the training process, the neural network "learns" to associate data with the corresponding labels. A neural network structure that can be independently trained and a cluster of weighted values that have been trained using an algorithm can be referred to as a "model", and the process of predicting which category new input data belongs to by the model with the determined weighted values and then outputting the predicted value can be referred to as the "testing" process or operating the neural network in an inference mode. Additionally, the process of generating results for new input data based on the patterns trained with the model is referred to as "inference". In other words, inference refers to the operation of predicting results for unknown data after the model has been trained.
[0105] In some example embodiments, the neural network can be set such that the first through kth (k is a positive integer) node groups NG1a, NG2a, NG3a, …… NGk-1a, and NGka respectively correspond to the input layer IL, multiple hidden layers HL1, HL2, ……, and HLn, and the output layer OL.
[0106] For example, the first node group NG1a may include input nodes x1, x2, ……, and xi. For example, the second through k-1th node groups NG2a, NG3a, ……, and NGk-1a may include hidden nodes h11, h12, h13, …… h1m, h21, h22, h23, …… h2m, hn1, hn2, hn3, ……, and hnm. For example, the kth node group NGka may include output nodes y1, y2, ……, and yj.
[0107] Reference Figure 13 , at least some of the nodes included in two or more of the multiple layers (e.g., IL, HL1, HL2, ……, HLn, and OL) are set to one of the multiple node groups (e.g., NG1b, NG2b, …… NGk-1b, and NGkb). For example, in Figure 13 shows a different grouping method from Figure 12 . Hereinafter, descriptions that are repeated with Figure 12 will be omitted.
[0108] In some example embodiments, the first through kth node groups NG1b, NG2b, ……, NGk-1b, and NGkb of the neural network can be set independently of the input layer IL, multiple hidden layers HL1, HL2, ……, and HLn, and the output layer OL.
[0109] For example, the first node group NG1b may include some of the input nodes x1 and x2, and some of the hidden nodes h11, h12, h21, and h22. For example, the second node group NG2b may include some of the input nodes xi, and some of the hidden nodes h13, …, h1m, h23, …, and h2m.
[0110] In other example embodiments, the neural network may be arranged such that the first through kth node groups NG1b, NG2b, …, NGk-1b, and NGkb include the same number of nodes, but the grouping method for partitioning the neural network is not limited to the above description.
[0111] Figure 14 is a block diagram showing a storage device and a storage system including the storage device according to an example embodiment.
[0112] Reference Figure 14 , the storage system 100 includes a host device 200 and a storage device 300.
[0113] The host device 200 may control the overall operation of the storage system 100. For example, the host device 200 may include a host processor and a host memory. For example, the host processor may control the operation of the host device 200 and may run an operating system (OS). For example, the host memory may store instructions and data that are run and processed by the host processor. For example, the OS run by the host processor may include a file system for file management and a device driver for controlling peripheral devices including the storage device 300 at the OS level.
[0114] The storage device 300 may be accessed by the host device 200. The storage device 300 may include a storage controller 310 (e.g., a control circuit), a plurality of non-volatile memories (NVM) 320, a buffer memory 330, a plurality of first hardware accelerators 312, and a plurality of second hardware accelerators 313. The storage controller 310 may include an inference optimization module 311. The storage device 300 may store program code for running an artificial neural network model.
[0115] The storage controller 310 may control the operation of the storage device 300. For example, the storage controller 310 may control the operation of the plurality of non-volatile memories 320, the plurality of first hardware accelerators 312, and the plurality of second hardware accelerators 313 based on commands and data received from the host device 200. For example, the storage controller 310 may receive an input value INV from the host device 200 and send an inference result value IFRV corresponding to the input value (INV) to the host device 200.
[0116] The storage controller 310 may include an inference optimization module 311 for performing a method of operating an artificial neural network model according to an example embodiment. For example, as will be described with reference to Figure 15 it, the inference optimization module 311 may be implemented to include a model splitting module, a node group allocation module, and a recording module. The model splitting module sets a plurality of node groups by partitioning the artificial neural network model, and each of the plurality of node groups includes at least one of the plurality of nodes. The node group allocation module allocates the plurality of node groups to a plurality of hardware accelerators. The recording module records each of a plurality of inference result values, activation region information of the plurality of node groups, and a call count.
[0117] For example, the model splitting module may perform a second operation (e.g., OP2 in Figure 10 ), and the node group allocation module may perform a first operation (e.g., OP1 in Figure 8 ). The plurality of non-volatile memories 320 may store data. For example, the data may include an artificial neural network model that includes a plurality of nodes.
[0118] In some example embodiments, each of the plurality of non-volatile memories 320 may include a NAND flash memory. In other example embodiments, each of the plurality of non-volatile memories 320 may include one of an electrically erasable programmable read-only memory (EEPROM), a phase change random access memory (PRAM), a resistive random access memory (RRAM), a nano floating gate memory (NFGM), a polymer random access memory (PoRAM), a magnetic random access memory (MRAM), a ferroelectric random access memory (FRAM), etc.
[0119] The buffer memory 330 may store instructions and / or data that are run and / or processed by the storage controller 310, and may temporarily store data that is stored in or to be stored in the plurality of non-volatile memories 320. For example, the artificial neural network model may be stored in the buffer memory 330. For example, the buffer memory 330 may include at least one of various volatile memories, such as a dynamic random access memory (DRAM) or a static random access memory (SRAM).
[0120] The plurality of first hardware accelerators 312 and the plurality of second hardware accelerators 313 may be included in the storage device 300. For example, the plurality of first hardware accelerators 312 and the plurality of second hardware accelerators 313 may represent devices that perform certain functions in computing faster than a central processing unit. For example, the plurality of first hardware accelerators 312 and the second hardware accelerators 313 may represent devices that perform inference operations faster than the central processing unit included in the storage device 300.
[0121] For example, multiple first hardware accelerators 312 and multiple second hardware accelerators 313 may be implemented by a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or an application specific standard part (ASSP). For example, the first hardware accelerator 312 and the second hardware accelerator 313 may have a computing device and a memory space separate from the storage device 300.
[0122] For example, the multiple second hardware accelerators 313 may have an operating speed faster than that of the multiple first hardware accelerators 312. In some example embodiments, the multiple first hardware accelerators 312 may be included in the storage controller 310, and the multiple second hardware accelerators 313 may be disposed outside the storage controller 310. For example, the multiple first hardware accelerators 312 may be embedded field programmable gate arrays (eFPGAs), and the multiple second hardware accelerators 313 may be field programmable gate arrays (FPGAs).
[0123] For example, a single node group among the multiple node groups may be assigned a single hardware accelerator among the first hardware accelerator 312 and the second hardware accelerator 313. For example, each of the first hardware accelerator 312 and the second hardware accelerator 313 may perform a sub-inference operation on the single assigned node group to generate multiple sub-inference result values, and the multiple sub-inference result values may be sent to the storage controller 311. For example, the storage controller 311 may calculate an inference result value IFRV from the multiple sub-inference result values.
[0124] In some example embodiments, the storage device 300 may be a solid state drive (SSD). In other example embodiments, the storage device 300 may be one of a universal flash storage (UFS), a multimedia card (MMC), an embedded multimedia card (eMMC), a secure digital (SD) card, a micro SD card, a memory stick, a chip card, a universal serial bus (USB) card, a smart card, a compact flash (CF) card, etc.
[0125] In some example embodiments, the storage device 300 may be connected to the host device 200 through a block accessible interface, which may include, for example, a UFS, an eMMC, a serial advanced technology attachment (SATA) bus, a non-volatile memory express (NVMe) bus, a serial attached SCSI (SAS) bus, etc. The storage device 300 may use a block accessible address space corresponding to the access size of the multiple non-volatile memories 320 to provide a block accessible interface to the host device 200 to allow access to the data stored in the multiple non-volatile memories 320 through the cells of the memory blocks.
[0126] In some example embodiments, the storage system 100 can be any mobile system, such as a mobile phone, a smartphone, a tablet computer, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a digital camera, a portable game console, a music player, a video camera recorder, a video player, a navigation device, a wearable device, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, an e - book reader, a virtual reality (VR) device, an augmented reality (AR) device, a robotic device, etc. In other example embodiments, the storage system 100 can be any computing system, such as a personal computer (PC), a server computer, a workstation, a digital TV, a set - top box, or a navigation system.
[0127] Figure 15 is a block diagram showing a storage controller included in a storage device according to an example embodiment.
[0128] Referring to Figure 15 , the storage controller 400 includes a host interface 410, a processor 420, a memory 430, an error - correcting code (ECC) module 440, a memory interface 450, a model splitting module 460, a node group allocation module 465, and a recording module 470.
[0129] The processor 420 can control the operation of the storage controller 400 in response to commands received via the host interface 410 from a host (e.g., Figure 14 the host device 200 in Figure 14 ). In some example embodiments, the processor 420 can control the corresponding components by employing firmware for operating a storage device (e.g.,
[0130] the storage device 300 in
[0131] The ECC module 440 for error correction can perform encoding modulation using Bose-Chaudhuri-Hocquenghem (BCH) codes, low density parity check (LDPC) codes, turbo codes, Reed-Solomon codes, convolutional codes, recursive systematic codes (RSCs), trellis-coded modulation (TCM), block coded modulation (BCM), etc., or can perform ECC encoding and ECC decoding using the above codes or other error correction codes.
[0132] The host interface 410 can provide a physical connection between the host device 200 and the storage device 300. The host interface 410 can provide an interface corresponding to the bus format of the host for communication between the host device 200 and the storage device 300. In some example embodiments, the bus format of the host device 200 can be a Small Computer System Interface (SCSI) or a Serial Attached SCSI (SAS) interface. In other example embodiments, the bus format of the host device 200 can be formats such as USB, Peripheral Component Interconnect (PCI) Express (PCIe), Advanced Technology Attachment (ATA), Parallel ATA (PATA), Serial ATA (SATA), Non-Volatile Memory (NVM) Express (NVMe), etc.
[0133] The memory interface 450 can exchange data with non-volatile memory (e.g., Figure 14 the non-volatile memory 320 therein). The memory interface 450 can transfer data to the non-volatile memory 320, or can receive data read from the non-volatile memory 320. In some example embodiments, the memory interface 450 can be connected to the non-volatile memory 320 via one channel. In other example embodiments, the memory interface 450 can be connected to the non-volatile memory 320 via two or more channels.
[0134] The model splitting module 460 can perform Figure 1 the operation S100 in. For example, the model splitting module 460 can set multiple node groups by partitioning an artificial neural network model, where each node group in the multiple node groups includes at least one node among the multiple nodes. The model splitting module 460 can perform Figure 1 a part of the operation S600 in. For example, the model splitting module 460 can perform a second operation for changing the partitioning of the artificial neural network model (e.g., Figure 10 the OP2 in).
[0135] The node group allocation module 465 may perform Figure 1 the operation S200 in Figure 14 . For example, the node group allocation module 465 may allocate a plurality of node groups to a plurality of first hardware accelerators and a plurality of second hardware accelerators (e.g., Figure 1 312 and 313 in Figure 14 ). The node group allocation module 465 may perform Figure 8 a part of the operation S600 in
[0136] . For example, the node group allocation module 465 may perform a first operation (e.g., Figure 1 OP1 in
[0137] to change the allocation of a plurality of node groups to a plurality of first hardware accelerators and a plurality of second hardware accelerators (e.g.,
[0138] Figure 16 is a diagram showing an example of a storage device according to an example embodiment.
[0139] Referring to Figure 16 , an example in which a plurality of second hardware accelerators 313a are provided outside the storage device 300a is shown. Hereinafter, descriptions that are Figure 14 repeated will be omitted.
[0140] For example, multiple first hardware accelerators 312a may be included in the storage device 300a, and multiple second hardware accelerators 313a may be provided outside the storage device 300a. For example, the storage device 300a may use the multiple second hardware accelerators 313a that are not included in the storage device 300a to run an artificial neural network model. Additionally, the multiple second hardware accelerators 313a may have an operating speed faster than that of the multiple first hardware accelerators 312a. For example, the running time of the artificial neural network model in the multiple second hardware accelerators 313a may be shorter than that in the multiple first hardware accelerators 312a.
[0141] The inventive concept can be applied to various electronic devices and systems including a storage device. For example, the inventive concept can be applied to systems such as a personal computer (PC), a server computer, a data center, a workstation, a mobile phone, a smartphone, a tablet computer, a laptop computer, a personal digital assistant (PDA), a portable multimedia player (PMP), a digital camera, a portable game console, a music player, a camcorder, a video player, a navigation device, a wearable device, an Internet of Things (IoT) device, an Internet of Everything (IoE) device, an e - book reader, a virtual reality (VR) device, an augmented reality (AR) device, a robotic device, a drone, etc.
[0142] The foregoing illustrates example embodiments and should not be construed as limiting them. Although some example embodiments have been described, those skilled in the art will readily understand that many modifications are possible in the example embodiments without materially departing from the teachings of the example embodiments. Accordingly, all such modifications are intended to be included within the scope of the example embodiments.
Claims
1. A method for operating an artificial neural network model comprising a plurality of nodes, the method comprising: Using a first grouping method, the artificial neural network model is divided into divided artificial neural network models including a plurality of node groups, each of the plurality of node groups including at least one node of the plurality of nodes; allocating a first subset of the plurality of node groups to a plurality of first hardware accelerators and allocating a second other subset of the plurality of node groups to a plurality of second hardware accelerators using a first corresponding manner to generate an allocation, wherein an operating speed of the plurality of second hardware accelerators is faster than an operating speed of the plurality of first hardware accelerators; Using the plurality of first hardware accelerators and the plurality of second hardware accelerators to run the partitioned artificial neural network model on a plurality of input values to generate a plurality of inference result values; For each of the plurality of inference result values, recording activation region information and call counts of the plurality of node groups; and At least one of a first operation for changing the allocation and a second operation for changing the partitioned artificial neural network model is performed based on the activation region information and the call count.
2. The method according to claim 1, wherein: The activation region information of the plurality of node groups includes information indicating which node among the plurality of nodes is activated when the artificial neural network model is executed.
3. The method according to claim 1, wherein: The call count indicates a number of times each of the plurality of inference result values is calculated.
4. The method according to claim 1, wherein: Performing at least one of the first operation and the second operation includes: The first operation is performed based on a first reference inference result value having a maximum call count among the plurality of inference result values.
5. The method according to claim 4, wherein: Executing the first operation includes: selecting the first reference inference result value from among the plurality of inference result values; and Based on the activation region information of the plurality of node groups for the first reference inference result value, the plurality of node groups are reallocated to the plurality of first hardware accelerators and the plurality of second hardware accelerators using a second corresponding manner different from the first corresponding manner.
6. The method according to claim 5, wherein: When the plurality of node groups are reallocated to the plurality of first hardware accelerators and the plurality of second hardware accelerators in the second corresponding manner, N of the plurality of node groups sorted by high activation levels are allocated to N of the second hardware accelerators, where N is a positive integer.
7. The method according to claim 1, wherein: Performing at least one of the first operation and the second operation includes: The second operation is performed based on a first reference inference result value having a maximum call count among the plurality of inference result values.
8. The method according to claim 7, wherein: Executing the second operation includes: selecting the first reference inference result value from among the plurality of inference result values; and Based on the activation area information of the plurality of node groups for the first reference inference result value, the artificial neural network model is divided into new divided artificial neural network models using a second grouping method different from the first grouping method.
9. The method according to claim 8, wherein: Dividing the artificial neural network model into new divided artificial neural network models reduces the number of deactivated nodes included in one of the node groups having an activation level greater than a specific threshold.
10. The method according to claim 1, in, The program code for running the artificial neural network model is stored in a storage device, wherein the storage device includes a storage controller and a plurality of non-volatile memories controlled by the storage controller. wherein the plurality of first hardware accelerators and the plurality of second hardware accelerators are included in the storage device, and Wherein, the artificial neural network model is run by the storage device.
11. The method according to claim 10, wherein: The plurality of first hardware accelerators are included in the storage controller, and the plurality of second hardware accelerators are disposed outside the storage controller.
12. The method according to claim 11, wherein: The first plurality of hardware accelerators are embedded field programmable gate arrays, and the second plurality of hardware accelerators are field programmable gate arrays.
13. The method according to claim 11, wherein: The plurality of first hardware accelerators and the plurality of second hardware accelerators are graphics processing units.
14. The method according to claim 1, in, The program code for running the artificial neural network model is stored in a storage device, The plurality of first hardware accelerators are included in the storage device, and the plurality of second hardware accelerators are arranged outside the storage device, and Wherein, the artificial neural network model is run by the storage device.
15. The method according to claim 1, in, The artificial neural network model includes multiple layers, and The nodes included in one layer among the plurality of layers are assigned to one node group among the plurality of node groups.
16. The method according to claim 1, in, The artificial neural network model includes multiple layers, and At least some of the nodes included in two or more layers among the plurality of layers are assigned to one of the plurality of node groups.
17. A storage device, comprising: a plurality of non-volatile memories configured to store an artificial neural network model including a plurality of nodes; a plurality of hardware accelerators configured to calculate a plurality of inference result values based on a plurality of input values and the artificial neural network model; as well as a storage controller configured to control the plurality of non-volatile memories and the plurality of hardware accelerators, Wherein, the storage controller comprises: A model splitting module, wherein the model splitting module is configured to divide the artificial neural network model into a plurality of node groups, each of the plurality of node groups including at least one node of the plurality of nodes; a node group allocation module configured to allocate each of the plurality of node groups to a corresponding hardware accelerator of the plurality of hardware accelerators to generate an allocation; and a recording module configured to record, for each of the plurality of inference result values, activation region information and call counts of the plurality of node groups; The node group allocation module is further configured to perform a first operation for changing the allocation based on the activation area information and the call count, and The model splitting module is further configured to perform a second operation for changing the divided artificial neural network based on the activation area information and the call count.
18. The storage device according to claim 17, in, The plurality of hardware accelerators include a plurality of first hardware accelerators and a plurality of second hardware accelerators, the plurality of second hardware accelerators having an operation speed faster than an operation speed of the plurality of first hardware accelerators, wherein the plurality of first hardware accelerators are included in the storage controller, and The plurality of second hardware accelerators are arranged outside the storage controller.
19. The storage device according to claim 17, further comprising: a buffer memory configured to temporarily store the artificial neural network model, and Wherein, the storage controller is further configured to control the buffer memory.
20. A method of operating an artificial neural network model comprising a plurality of nodes, the method comprising: Using a first grouping method, the artificial neural network model is divided into divided artificial neural network models including a plurality of node groups, each of the plurality of node groups including at least one node of the plurality of nodes; Allocating a first subset of the plurality of node groups to a plurality of first hardware accelerators and allocating a second other subset of the plurality of node groups to a plurality of second hardware accelerators using a first corresponding manner, wherein an operation speed of the plurality of second hardware accelerators is faster than an operation speed of the plurality of first hardware accelerators; Using the plurality of first hardware accelerators and the plurality of second hardware accelerators to run the partitioned artificial neural network model on a plurality of input values to generate a plurality of inference result values; For each of the plurality of inference result values, record activation region information and call counts of the plurality of node groups; selecting a first reference inference result value from among the plurality of inference result values; reallocating the plurality of node groups to the plurality of first hardware accelerators and the plurality of second hardware accelerators using a second corresponding manner different from the first corresponding manner based on the activation region information of the plurality of node groups for the first reference inference result value; and Based on the activation area information, the artificial neural network model is divided into new divided artificial neural network models using a second grouping method different from the first grouping method.