Programming method of activation function and activation function programming unit
Machine learning is carried out through artificial neural networks, and the activation function is approximateed to a programmatic activation function and programmed in hardware, solving the problems of high complexity of activation function processing, high power consumption and inability to independently process new activation functions in the prior art, and achieving efficient and flexible activation function processing.
Patent Information
- Application Number
- CN202411573902.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-08
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-09
AI Technical Summary
When the prior art processes activation functions in hardware, there are problems such as high computational complexity, large power consumption, and increased gate count, and the new activation functions cannot be processed independently.
Machine learning is performed through artificial neural networks, the target activation function is approximateed to a programmable activation function, and converted into programmable segments, stored in a lookup table to achieve efficient and flexible programming of activation functions in hardware.
It realizes efficient processing of nonlinear activation functions in hardware, reduces computational complexity and power consumption, and can independently handle unpredefined activation functions.
Smart Images

Figure CN119962591A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the priority of Korean Patent Application No. 10-2023-0153485 filed in the Korean Intellectual Property Office on November 8, 2023, the disclosure of which is incorporated herein by reference. Technical Field
[0002] The present disclosure relates to an activation function programming method and an activation function conversion program unit. Background Art
[0003] Humans have intelligence that can identify, classify, reason, predict, and control / decide. Artificial intelligence (AI) refers to artificial imitation of human intelligence.
[0004] The human brain is composed of many nerve cells called neurons. Each neuron is connected to hundreds to thousands of other neurons through connections called synapses. In order to imitate human intelligence, the operation principle of biological neurons and the connection relationship between neurons are simulated, which is called an artificial neural network (ANN) model. In other words, ANN is a system in which the nodes that imitate neurons are connected in a layered structure.
[0005] The ANN-specific processor developed to accelerate the calculation of ANN is the Neural Processing Unit (NPU). Summary of the invention [Technical solution]
[0006] ANN is divided into "single-layer neural network" and "multi-layer neural network" according to the number of layers. A typical multi-layer neural network consists of an input layer, a hidden layer, and an output layer. (1) The input layer is a layer that receives input values. The number of input layers is the same as the number of input variables. (2) The hidden layer is located between the input layer and the output layer. It is a layer that receives signals from the input layer, extracts features, and transmits them to the output layer. (3) The output layer is a layer that receives signals from the hidden layer and outputs them to the outside.
[0007] When a signal is transmitted between neurons in the human brain, the transmission strength of the signal changes. By simulating this, the transmission strength of the signal transmitted between layers, that is, the activation, is determined by the activation function in the ANN.
[0008] Depending on the characteristics of the activation function implemented in the NPU, the inference accuracy of the ANN may vary. That is, the performance and efficiency of the ANN are determined by the hardware implementation characteristics of the activation function processing circuit of the NPU. In addition, ANNs that process complex mathematical activation functions can be processed by hardware accelerators. When ANN-specific processors are implemented in hardware, the ANN-specific processors may require a large chip area (i.e., a large number of logic gates). In addition, these chips may have considerable power consumption.
[0009] In order to realize higher artificial intelligence, a deep neural network (DNN) having an increased number of hidden layers has been disclosed. The activation function of the DNN is used to determine the transmission strength of the calculated values of the applied weights and biases. The DNN is being developed in various structures.
[0010] For example, a convolutional neural network (CNN), which is an example of a DNN, is known to be easy to extract features of an input value (ie, a video or image) and recognize patterns of the extracted features. A CNN can be configured to process a convolution operation, an activation function operation, a pooling operation, etc. in a specific order.
[0011] For example, in each layer of a DNN, the input values and parameters (i.e., weights or kernels) can be matrices consisting of multiple channels. The input values and parameters can be processed in the NPU through convolution or matrix multiplication. After processing the calculations in each layer, calculated values are generated. Activation functions can be applied to these calculated values.
[0012] For example, Transformer is a DNN based on attention technology. Transformer utilizes many matrix multiplication operations. Transformer can use parameters such as input value and query (Q), key (K) and value (V) to obtain the operation value (Q, K, V) of attention. Transformer can handle various reasoning operations based on the operation value (i.e. attention (Q, K, V)). Transformer tends to show better reasoning performance than CNN.
[0013] The above neural network can be referred to as a DNN. Meanwhile, the activation function can be selectively applied to the operation value of a specific layer among the multiple layers of the DNN.
[0014] It can be configured to include an X-axis value corresponding to the input value of the activation function (i.e., the operation value of a specific layer) and a Y-axis value corresponding to the activation value of the activation function. The activation function plays the role of converting the mathematical linear combination of the input values into various types of linear combinations or nonlinear combinations. Therefore, DNNs can be designed to perform various inference functions by applying appropriate activation functions to the operation values of a specific layer.
[0015] Most of the complex functions to be solved in DNNs have nonlinearity. To address this problem, most activation functions are nonlinear.
[0016] The performance and efficiency of DNN models processed in hardware may vary depending on the nonlinearity of an activation function applied to at least one of the DNN models processed by the NPU.
[0017] An activation function can improve or reduce the inference accuracy of the activation function input values by emphasizing features in specific regions more and emphasizing features in other regions less.
[0018] The nonlinearity of at least some of the various activation functions may include logarithmic operations, exponential operations, etc. Implementing activation functions including logarithmic and exponential operations in hardware is very complex in terms of digital logic design. For example, for logarithmic and exponential operations, the configuration of the hardware operator becomes very complex. Therefore, the inventors of the present disclosure recognize that the power consumption of the hardware may increase, and the computing processing speed may slow down.
[0019] In the case of an NPU, it may be necessary to design each activation function processing module for each activation function processing. In addition, the hardwired processor may only use the respective hardwired dedicated activation function processing logic unit to process the predefined activation function. At this point, the inventors of the present disclosure recognize that there is a disadvantage that the number of gates in the hardwired processor will increase rapidly depending on the computational complexity of the activation function.
[0020] Without hardware modification, the hardwired processor cannot independently process new activation functions. Activation functions that the hardwired processor cannot handle must be calculated using separate software. For example, the hardwired processor can be an application-specific integrated circuit (ASIC) dedicated to artificial intelligence. In other words, the hardwired processor can be an NPU.
[0021] Various methods have been proposed to process various types of activation functions in hardwired processors. For example, conventionally, activation functions are processed using a method using a lookup table (LUT), a method using a nonlinear approximation equation, a method using a polynomial approximation, and the like.
[0022] However, the inventors of the present disclosure have recognized that conventional methods of approximating activation functions in hardware using polynomial approximation or the like require a processor to perform a large amount of computation to improve inference accuracy.
[0023] Therefore, the inventors of the present disclosure have recognized the need to improve the problems of deteriorating inference accuracy of DNN models applying traditional activation function approximation technology, increasing the number of gates in the activation function processing unit of the processor, and increasing the power consumption of the processor.
[0024] Furthermore, the inventors of the present disclosure have recognized that in order to enable a processor to independently process: 1) activation functions that are not included in predetermined data (e.g., a lookup table that a processor applying a conventional activation function processing method cannot process), 2) new activation functions, and / or 3) activation functions in which some conventional activation functions have been modified, a programming method capable of approximating any activation function and a hardware design for driving the activation function are required.
[0025] Furthermore, the inventors of the present disclosure have recognized that there is a need to design an NPU that is capable of driving an approximation algorithm optimized for activation function characteristics.
[0026] Furthermore, the inventors of the present disclosure have recognized that activation functions can be efficiently and flexibly programmed in hardware if hardware optimized for such programming methods is provided.
[0027] In addition, each region can be set according to the shape of the activation function to be programmed, and an approximation parameter can be programmed for each set region. The inventors of the present disclosure have recognized that by considering the characteristics of each region of the activation function, the activation function can be efficiently programmed with a low approximation error.
[0028] Furthermore, the inventors of the present disclosure have recognized that a programmable activation function (PAF) may be provided in a hardwired processor including a programmed activation function execution unit (PAFE Unit).
[0029] Therefore, an object to be solved by the present disclosure is to provide a method that is relatively superior to traditional approximation methods and is capable of programming nonlinear activation functions in hardware with various hardware options.
[0030] In addition, an object to be solved by the present disclosure is to provide a method for approximating a nonlinear activation function in a more customized manner by considering the characteristics of the activation function itself, approximation error, hardware option information, etc.
[0031] Furthermore, a problem to be solved by the present disclosure is to provide a hardwired processor comprising a PAFE unit.
[0032] Furthermore, the problem to be solved by the present disclosure is to provide a hardwired processor comprising a PAFE unit, the hardwired processor being configured to process at least one programmed activation function.
[0033] However, the tasks of the present disclosure are not limited to the above-mentioned tasks, and other tasks not mentioned will be clearly understood by those skilled in the art from the following description.
[0034] Detailed descriptions of other examples are included in the detailed description and accompanying drawings.
[0035] According to an example of the present disclosure, an activation function conversion program unit is provided. The activation function conversion program unit can be configured to approximate a target activation function to a programmed activation function through machine learning of an artificial neural network.
[0036] The artificial neural network may include a first layer including multiple neurons and a second layer including multiple neurons, and a rectified linear unit (ReLU) function may be applied to the outputs of the multiple neurons of the first layer, and the value of the ReLU function may be applied as the input of the multiple neurons of the second layer.
[0037] The artificial neural network may include a first layer including a plurality of neurons and a second layer including a plurality of neurons, each of the plurality of neurons in the first layer may include a weight and a bias, and the plurality of neurons in the second layer may include only a weight.
[0038] The artificial neural network can be allowed to perform machine learning so as to minimize the error between the target activation function and the programmed activation function.
[0039] The programmed activation function may include a plurality of segments including a programmable segment implemented as a first-order function.
[0040] The artificial neural network may include a plurality of neurons, and the programmed activation function may include a plurality of programmable segments, the plurality of programmable segments being respectively separated by breakpoints of outputs of the plurality of neurons.
[0041] The programmed activation function may include multiple programmable segments, the number of the multiple programmable segments may correspond to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and the hardware information may correspond to a comparator included in the PAFE Unit.
[0042] The artificial neural network may include a plurality of neurons, and at least one of the outputs of the plurality of neurons may be pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0043] The artificial neural network may include a plurality of neurons, and the number of the plurality of neurons may be less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0044] According to an example of the present disclosure, an activation function programming method may be provided. The activation function programming method may include: setting a target activation function; approximating the target activation function to a programmed activation function by allowing an artificial neural network to perform machine learning; and converting the programmed activation function into a slope and an offset and storing them in a lookup table.
[0045] The artificial neural network may include a first layer including multiple neurons and a second layer including multiple neurons, and a rectified linear unit (ReLU) function may be applied to the outputs of the multiple neurons of the first layer, and the value of the ReLU function may be applied as the input of the multiple neurons of the second layer.
[0046] The artificial neural network may include a first layer including a plurality of neurons and a second layer including a plurality of neurons, each of the plurality of neurons of the first layer may include a weight and a bias, and the plurality of neurons in the second layer may include only weights.
[0047] The artificial neural network can be allowed to perform machine learning so as to minimize the error between the target activation function and the programmed activation function.
[0048] The programmed activation function may include a plurality of segments including a programmable segment implemented as a first-order function.
[0049] The artificial neural network may include a plurality of neurons, and the programmed activation function may include a plurality of programmable segments, the plurality of programmable segments being respectively separated by breakpoints of outputs of the plurality of neurons.
[0050] The programmed activation function may include multiple programmable segments, the number of the multiple programmable segments may correspond to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and the hardware information may correspond to a comparator included in the PAFE Unit.
[0051] The artificial neural network may include a plurality of neurons, and at least one of the outputs of the plurality of neurons may be pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0052] The artificial neural network may include a plurality of neurons, and the number of the plurality of neurons may be less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function. [Beneficial Effects]
[0053] According to the present disclosure, the NPU may receive programming parameters of an activation function and process the activation function.
[0054] According to the present disclosure, by using segment data, various nonlinear activation functions, especially newly proposed or known activation functions with some modifications, can be programmed to be processable in hardware.
[0055] In addition, according to the present disclosure, when approximating various nonlinear activation functions, segment data including the characteristics of the activation function itself, approximation errors, hardware option information, etc. can be used. Therefore, nonlinear activation functions can be programmed in a more customized manner while ensuring high performance and high efficiency of DNN.
[0056] Furthermore, according to the present disclosure, when approximating various nonlinear activation functions, by using segment data including characteristics of the activation function itself, approximation error, hardware option information, etc., the approximation error can be minimized while minimizing the hardware cost.
[0057] In addition, according to the present disclosure, each segment of the activation function can be programmed with various algorithms. The NPU can provide a hardware option that can process the algorithm of each segment of the programmed activation function.
[0058] Furthermore, according to the present disclosure, a hardwired processor including a PAFE unit can be implemented. Thus, the processor can handle any activation function by only changing programmable parameters without hardware changes.
[0059] Furthermore, according to the present disclosure, a hardwired processor including a PAFE unit configured to process at least one programmed activation function can be implemented. Thus, the processor can process different activation functions with the PAFE unit simultaneously or sequentially without hardware changes.
[0060] In addition, through the training of artificial neural networks, various nonlinear activation functions can be converted into programmed activation functions with multiple linear forms and optimized calculations, which have the effect of optimizing the calculation speed and power consumption of the programmed activation function execution unit of the NPU.
[0061] The effects of the present invention are not limited to the above-mentioned examples, and the present disclosure also includes more various effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a schematic conceptual diagram illustrating a device for executing a method of programming an activation function according to an example of the present disclosure.
[0063] Figure 2 is a schematic flow chart illustrating a method of programming an activation function according to an example of the present disclosure.
[0064] Figures 3A-3C is a diagram illustrating a process of programming an activation function by a method of programming an activation function according to an example of the present disclosure.
[0065] Figures 4A-4D2 is a diagram illustrating various cases where an activation function is segmented into a plurality of segments by a method of programming an activation function according to an example of the present disclosure.
[0066] Figures 5A-5C is a diagram showing an example of segmenting an activation function into a linear section and a nonlinear section using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0067] Figure 6A-6B is a diagram showing an example of segmenting an activation function into a basic linear interval and a nonlinear interval using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0068] Figure 7A-7B is a diagram showing another example of segmenting an activation function into a basic linear interval and a nonlinear interval using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0069] Figures 8A-8B is a diagram showing another example of segmenting an activation function into a basic linear interval and a nonlinear interval using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0070] Fig. 9 is a diagram showing an example of converting a segment into a programmable segment using an error value in an activation function programming method according to an example of the present disclosure.
[0071] Figures 10A-10B is a diagram showing an example of converting a segment into a programmable segment using a maximum error value in an activation function programming method according to an example of the present disclosure.
[0072] Figures 11A-11B is a diagram showing an example of converting a segment into a programmable segment using an integral value of an error value in an activation function programming method according to an example of the present disclosure.
[0073] Fig.12 is a diagram showing an example of using machine learning to approximate a segment to an optimal programmable segment in an activation function programming method according to an example of the present disclosure.
[0074] Fig.13 is a diagram showing an example of segmenting an activation function using an accumulated value of a second-order derivative of an activation function and a threshold in an activation function programming method according to an example of the present disclosure.
[0075] Fig.14 and 15 is a diagram showing the ELU activation function and the Hardswish activation function.
[0076] Fig.16 is a conceptual diagram illustrating a PAFE unit configured to process a programmed activation function according to an example of the present disclosure.
[0077] Fig.17 is a conceptual diagram illustrating a PAFE unit of an NPU of a device configured to process a programmed activation function according to an example of the present disclosure.
[0078] Fig.18 is a conceptual diagram illustrating an NPU of a device for processing a programmed activation function according to another example of the present disclosure.
[0079] Fig.19 is a conceptual diagram illustrating a PAFE unit of an NPU of a device for processing a programmed activation function according to another example of the present disclosure.
[0080] Fig. 20 is a conceptual diagram illustrating an NPU of a device for processing a programmed activation function according to another example of the present disclosure.
[0081] Fig.21 is a conceptual diagram illustrating an NPU of a device for processing a programmed activation function according to another example of the present disclosure.
[0082] Fig. 22 is a conceptual diagram illustrating a PAFE unit configured to process a programmed activation function according to another example of the present disclosure.
[0083] Fig.23 is a conceptual diagram illustrating a PAFE unit of an NPU of a device for processing a programmed activation function according to another example of the present disclosure.
[0084] Fig.24 2 is a diagram showing an example in which a device for processing a programmed activation function according to another example of the present invention approximates a S-shaped activation function as a programmable activation function.
[0085] Fig.25 is a conceptual diagram illustrating a PAFE unit of an NPU of an apparatus for processing an activation function programmed according to another example of the present invention.
[0086] Fig.26 is a schematic flow chart illustrating a method of programming an activation function according to another example of the present disclosure.
[0087] Fig. 27 is a diagram illustrating an artificial neural network for approximating an activation function according to another example of the present disclosure.
[0088] Fig.28is a diagram showing the output of the first layer of Epoch 50.
[0089] Fig.29 is a diagram showing the output of the first layer when the ReLU function is applied to the output of the first layer of Epoch 50.
[0090] Fig.30 is a diagram showing the output of the second layer of Epoch 50.
[0091] Fig.31 is a graph showing the error between the programmed activation function and the target activation function at Epoch 50.
[0092] Fig.32 is a diagram showing the output of the first layer of Epoch 300.
[0093] Fig.33 is a graph showing the values of the output of the first layer when the ReLU function is applied to Epoch 300.
[0094] Fig.34 is a diagram showing the output of the second layer of Epoch 300.
[0095] Fig.35 is a graph showing the error between the programmed activation function of Epoch 50 and the target activation function of Epoch 300.
[0096] Fig.36 is a diagram showing the breakpoints of the programmed activation function of Epoch 300.
[0097] Fig.37 is an enlarged view showing the breakpoints of the programmed activation function of Epoch 300.
[0098] Fig.38 is a diagram illustrating a process of deriving a programmable segment from a first segment.
[0099] Fig.39 is a diagram illustrating a process of deriving a programmable segment from a second segment.
[0100] Fig.40 is a diagram showing a process of deriving a programmable segment from a third segment.
[0101] Fig.41 is a diagram showing a process of deriving a programmable segment from a fourth segment.
[0102] Fig.42 is a diagram showing a process of deriving a programmable segment from a fifth segment.
[0103] Fig.43 is a diagram showing a process of deriving a programmable segment from a sixth segment.
[0104] Fig.44 is a diagram showing a process of deriving a programmable segment from a seventh segment.
[0105] Fig.45 is a diagram showing a process of deriving a programmable segment from an eighth segment.
[0106] Fig.46 is a diagram showing a process of deriving a programmable segment from a ninth segment.
[0107] Fig.47 is a diagram showing a process of deriving a programmable segment from a tenth segment.
[0108] Fig.48 is a diagram showing a process of deriving a programmable segment from an eleventh segment.
[0109] Fig.49 is a diagram showing a process of deriving a programmable segment from a twelfth segment.
[0110] Fig.50 is a diagram showing a process of deriving a programmable segment from a thirteenth segment.
[0111] Fig.51 is a diagram showing a process of deriving a programmable segment from a fourteenth segment.
[0112] Fig.52 is a diagram showing a process of deriving a programmable segment from a fifteenth segment.
[0113] Fig.53 is a diagram showing a process of deriving a programmable segment from a sixteenth segment.
[0114] Fig.54 is a diagram illustrating an artificial neural network, wherein at least one of a plurality of neurons of the artificial neural network is pruned. DETAILED DESCRIPTION
[0115] The specific structures or step-by-step descriptions of the examples according to the concepts of the present disclosure disclosed in this specification or application are merely exemplified for explaining the examples according to the concepts of the present disclosure.
[0116] Examples of the concepts according to the present disclosure may be embodied in various forms. Examples of the concepts according to the present disclosure should not be interpreted as being limited to the examples described in this specification or application.
[0117] Various changes can be applied to the embodiments according to the conception of the present disclosure. The present disclosure can take various forms. Therefore, specific examples are shown in the drawings and described in detail in the present disclosure. However, this does not mean that the examples according to the conception of the present disclosure are limited to specific public forms. Therefore, it should be understood that all changes, equivalents or alternatives included in the spirit and scope of the present disclosure are included in the present disclosure.
[0118] Terms such as first and / or second may be used to describe various components. However, the present disclosure should not be limited by the above terms.
[0119] These terms are only used to distinguish one component from another. For example, without departing from the scope of the concept of the present disclosure, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element.
[0120] When it is mentioned that one element is "connected" or "contacted" with another element, it should be understood that the other element can be directly connected or contacted with the other element, but other elements can be arranged between them. On the other hand, when it is mentioned that a certain element is "directly connected" or "directly connected" with another element, it should be understood that there are no other elements between them.
[0121] Other expressions describing the relationship between elements, such as "between" and "immediately between" or "adjacent" and "directly adjacent", etc., should be similarly interpreted.
[0122] In the present disclosure, expressions such as "A or B", "at least one of A or / and B", or "one or more of A or / and B" may include all possible combinations thereof. For example, "A or B", "at least one of A and B", or "at least one of A or B" may refer to (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B.
[0123] As used herein, expressions such as "first", "second", "first or second" may modify various elements, regardless of order and / or importance. The expressions are only used to distinguish one element from other elements and do not limit the elements. For example, a first user device and a second user device may represent different user devices, regardless of order and / or importance. For example, without departing from the scope of rights described in the present disclosure, a first element may be named a second element, and similarly, a second element may be renamed a first element.
[0124] The terms used in the present disclosure are only used to describe specific examples and are not intended to limit the scope of other examples.
[0125] Unless the context clearly dictates otherwise, a singular expression may include a plural expression. Terms used herein, including technical or scientific terms, may have the same meanings as those generally understood by one of ordinary skill in the technical field described herein.
[0126] Among the terms used in the present disclosure, the terms defined in the general dictionary may be interpreted as having the same or similar meaning as that in the context of the relevant technical field. Unless clearly defined herein, they should not be interpreted in an ideal or overly formal sense. In some cases, even the terms defined in the present disclosure cannot be interpreted as excluding examples of the present disclosure.
[0127] The terminology used herein is for describing particular examples only and is not intended to be limiting of the present disclosure.
[0128] Unless the context clearly dictates otherwise, singular expressions include plural expressions. In this specification, terms such as "including" or "having" are intended to indicate the presence of the described features, numbers, steps, operations, components, intervals, or combinations thereof. Therefore, it should be understood that the presence or addition of one or more other features, numbers, steps, operations, components, intervals, or combinations thereof is not excluded.
[0129] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as those commonly understood by those of ordinary skill in the art to which the present disclosure belongs. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with that in the context of the relevant art. Unless clearly defined in the present disclosure, they should not be interpreted in an ideal or overly formal sense.
[0130] Each feature of each example of the present disclosure can be combined in intervals or in full or in combination with each other. Each example of the present disclosure can technically realize various interlocks and drives that can be fully understood by those skilled in the art. Each example of the present disclosure can be implemented independently of each other or together in an associated relationship.
[0131] When describing examples, descriptions of technical contents that are well-known in the technical field to which the present disclosure belongs and are not directly related to the present disclosure may be omitted. This is to convey the gist of the present disclosure more clearly without making the gist of the present disclosure difficult to understand due to the omission of unnecessary descriptions.
[0132] Hereinafter, examples of the present disclosure will be described in detail with reference to the accompanying drawings.
[0133] Figure 1 is a schematic conceptual diagram illustrating a device for executing an activation function programming method according to an example of the present disclosure.
[0134] refer to Figure 1, the device A for executing the activation function programming method may include a neural processing unit NPU 1000 and an activation function conversion program unit 3000. Here, the device A may represent a system. The device A may also include a processor 2000, a main memory 4000, an image sensor 5000, and a decoder 6000. Therefore, the device A may be configured to perform various artificial neural network reasoning functions.
[0135] Each of the elements that may be included in the device A may communicate through the bus 7000 to transmit and receive data.
[0136] Here, the NPU 1000, the processor 2000, the main memory 4000, the image sensor 5000, and the decoder 6000 may be configured as a circuit. The activation function conversion program unit 3000 may be a computer program, software, firmware, application, or executable code stored in a recording medium. However, the present disclosure is not limited thereto.
[0137] The activation function conversion program unit 3000 may be a computer program configured to execute instructions for converting an activation function into a PAF represented by programmable parameters. The activation function conversion program unit 3000 may be stored in a computer-readable recording medium. The computer-readable recording medium may include a ROM, RAM, SSD, HDD, CD-ROM, flash memory, magnetic tape, floppy disk, optical data storage device, etc.
[0138] NPU 1000 is a processor dedicated to deep neural network (DNN) operations separate from processor 2000. Specifically, NPU 1000 may include operators dedicated to convolution and matrix multiplication, which occupy a large area in the computational load of DNN. NPU 1000 and processor 2000 may be semiconductor chips including circuits.
[0139] NPU 1000 may include controller 100, direct memory access (DMA) 200, memory 300, at least one processing element 400, and programmed activation function execution unit (PAFE unit) 500. Hereinafter, programmed activation function execution unit 500 will be referred to as PAFE unit and will be described.
[0140] The controller 100 may be electrically connected to the DMA 200, the memory 300, the at least one processing element 400, and the PAFE unit 500. The controller 100 may be configured to control operations related to the DNN operation in the NPU 1000.
[0141] However, the present disclosure is not limited thereto, and at least one processing element 400 may be modified and implemented as a processing element array (eg, a systolic array).
[0142] The DMA 200 is configured to enable the NPU 1000 to directly access the main memory 4000 outside the NPU 1000 to perform read / write operations. The NPU 1000 can read various data related to the DNN from the main memory 4000 through the DMA 200. The DMA 200 can be configured to perform tasks such as setting, generating, and controlling addresses of the internal memory 300.
[0143] The memory 300 may be a memory provided in an on-chip area of the NPU 1000, and may be a memory for caching or storing data processed in the on-chip area. The memory 300 may read and store data required for calculating the artificial neural network model from the main memory 4000. The memory 300 may include, for example, one of memories such as ROM, SRAM, DRAM, resistive RAM, magnetoresistive RAM, phase change RAM, ferroelectric RAM, flash memory, HBM, etc. The memory 300 may include at least one memory unit. The memory 300 may be configured as a homogeneous memory unit or a heterogeneous memory unit.
[0144] At least one processing element 400 may be configured to process operations on parameters (e.g., weights, kernels, queries (Q), keys (K), values (V), etc.) corresponding to input data of the DNN. At least one processing element 400 may include a multiplication and accumulation (MAC) operator and / or an arithmetic logic unit (ALU) operator.
[0145] The PAFE unit 500 is configured to receive data of a programmable activation function (PAF) converted from an activation function (ie, programmable parameters).
[0146] For ease of explanation, the programmable activation function is referred to as PAF.
[0147] The programmable parameter may be data generated by the activation function conversion program unit 3000. The programmable parameter may be configured to have a form compatible with the circuit of the PAFE unit 500 of the NPU 1000. The programmable parameter may be configured to implement at least one PAF. That is, the PAFE unit 500 may be configured to receive a programmable parameter corresponding to at least one PAF generated by the activation function conversion program unit 3000. Specifically, the PAF programmed by the activation function conversion program unit 3000 may include at least one programmable segment. That is, the programmable parameter may implement at least one programmable segment.
[0148] The NPU 1000 may perform a DNN operation by receiving data for a PAF related to an activation function. The PAFE unit 500 may generate an activation value (e.g., an activation map) by applying the PAF generated by the activation function converter unit 3000 to a calculated value (e.g., a feature map) output from at least one processing element 400. The PAFE unit 500 uses at least one programmable parameter generated corresponding to at least one PAF. Therefore, the PAFE unit 500 enables the NPU 1000 to process various activation functions, particularly newly proposed or known but interval-modified activation functions.
[0149] The PAFE unit 500 may be pipelined with at least one processing element 400. According to the above configuration, a value calculated by at least one processing element 400 may be input through a pipeline. Therefore, at least one pipelined processing element 400 and the PAFE unit 500 may be configured to receive an operation value from at least one processing element 400 and output an activation value to which the PAF is applied. In this case, a bottleneck that may occur in at least one processing element 400 and the PAFE unit 500 may be minimized or substantially eliminated. However, examples of the present disclosure are not limited to a pipeline structure, and the PAFE unit may be implemented by merging with at least one processing element 400.
[0150] The activation function conversion program unit 3000 may be operated by the processor 2000, but is not limited thereto. The processor 2000 may be a computing device capable of executing the activation function programming method disclosed in the present disclosure, such as a central processing unit (CPU) or an application processor (AP).
[0151] The activation function conversion program unit 3000 may be stored in a computer-readable recording medium. The activation function conversion program unit 3000 may be implemented in firmware or in software included in hardware. A separate computing system and operating system may be provided to drive the activation function conversion program unit 3000. The activation function conversion program unit 3000 may be a program for operating the NPU 1000 including the PAFE unit 500. The activation function conversion program unit 3000 may be configured to execute the activation function programming method. The activation function conversion program unit 3000 may be executed by the processor 2000 or a processor outside the device A. The activation function conversion program unit 3000 may be configured separately from the compiler configured to compile the DNN in the device A. Alternatively, the activation function conversion program unit 3000 may be integrated with the compiler.
[0152] The activation function conversion program unit 3000 may be configured to program at least one activation function. The activation function conversion program unit 3000 may be configured to provide the PAFE unit 500 with programmable parameters corresponding to at least one PAF.
[0153] The activation function converter unit 3000 may be configured to receive activation function information included in the DNN to be processed by the NPU 1000. The activation function converter unit 3000 may obtain information of all activation functions to be processed by the NPU 1000 based on the provided information on at least one activation function. Therefore, the activation function converter unit 3000 may program at least one activation function required for the DNN to be processed by the NPU 1000.
[0154] In various examples, the activation function conversion program unit 3000 may generate segment data for segmenting the activation function, segment the activation function into a plurality of segments using the generated segment data, and approximate at least one of the plurality of segments to a programmable segment. When determining the value of the programmable parameter, the approximation level of the programmable segment may be determined. The activation function conversion program unit 3000 may determine the number and width of the plurality of segments according to the segment data.
[0155] The activation function conversion program unit 3000 may be configured to analyze the characteristics of the activation function. For example, the activation function conversion program unit 3000 may be configured to analyze the gradient change of the activation function. The slope change data of the activation function may refer to various data that can determine the slope change of the activation function.
[0156] The activation function conversion program unit 3000 can analyze the characteristics of the activation function based on the slope change data. In other words, in the area where the slope of the activation function changes more dramatically, the approximation error tends to increase, while in the area where the slope does not change, the approximation error may be zero. Therefore, the activation function conversion program unit 3000 can be configured to approximate the activation function to the optimal condition by analyzing the slope change data.
[0157] For example, the slope change data of the activation function may be differential data of the activation function. The slope change data may include at least one of a slope change value, a first-order derivative value, a second-order derivative value, a third-order derivative value, and the like.
[0158] For example, the activation function conversion program unit 3000 may determine the linear interval and the nonlinear interval of the PAF based on the slope change data of the activation function.
[0159] In some examples, the activation function conversion program unit 3000 may determine an interval having a substantially insignificant gradient change in the non-linear interval of the PAF as a substantially linear interval.
[0160] The activation function conversion program unit 3000 may convert at least one segment into a programmable segment approximated by a specific equation.
[0161] For example, the activation function converter unit 3000 may convert a specific segment of an activation function into a programmable segment approximated by a linear function.
[0162] In detail, the activation function converter unit 3000 may convert at least one segment into a programmable segment approximated with a specific gradient and a specific offset value. The activation function converter unit 3000 may convert at least one segment of the plurality of segments into a programmable segment using a specific nonlinear approximation equation. The activation function converter unit 3000 may determine a gradient and an offset for approximating at least one segment into a programmable segment corresponding to a linear function.
[0163] The activation function conversion program unit 3000 may search for a minimum error value while converting the gradient value and the offset value of the programmable segment. Alternatively, the activation function conversion program unit 3000 may search for a minimum error value by executing a cost function.
[0164] The activation function converter unit 3000 may calculate an error value between at least one segment of the activation function to be converted and at least one candidate segment having a candidate gradient and a candidate offset. The activation function converter unit 3000 may determine at least one candidate segment as a programmable segment based on the calculated error value. The activation function converter unit 3000 may search for at least one minimum error value between the segment of the activation function and each corresponding programmable segment. The activation function converter unit 3000 may determine a programmable parameter of the programmable segment based on the searched at least one minimum error value. Here, the determined error value may be a minimum error value. When the activation function converter unit 3000 determines the programmable parameter based on the minimum error value, the decrease in the inference accuracy of the DNN may be suppressed or minimized.
[0165] However, examples of the present disclosure are not limited to the minimum error value, and programmable parameters may be determined differently according to different priorities among the amount of calculation, the amount of power consumption, and the approximate error value.
[0166] In other words, the activation function conversion program unit 3000 can measure the approximate error value of the programmable segment obtained by converting the specific segment into a specific approximate function. For example, the activation function conversion program unit 3000 can measure the first error value of the programmable segment by approximating the specific segment as a programmable segment of a linear function. In addition, the activation function conversion program unit 3000 can measure the second error value of the programmable segment by approximating the specific segment as a programmable segment of a quadratic function. The activation function conversion program unit 3000 can compare the first error value with the second error value, and select an approximate function with a relatively small error value as the programmable segment. Through the above process, the activation function conversion program unit 3000 can select an activation function for artificial neural network operation and convert the activation function into a PAF.
[0167] That is, when determining the approximation function of the programmable segment, the format of the programmable parameter may also be determined. For example, if the specific segment is approximated as a programmable segment of a linear function, the corresponding programmable parameter may include a gradient and an offset value. For example, if the specific segment is approximated as a programmable segment of a quadratic function, the corresponding programmable parameter may include a coefficient of a quadratic term. The approximation function of each programmable segment may be selectively determined. That is, the approximation functions of the first programmable segment and the second programmable segment may be the same or different.
[0168] The criterion for determining the approximate function characteristic of each programmable segment may be determined based on any one of the amount of calculation, power consumption, and approximation error value of the PAFE unit 500 .
[0169] For example, the criteria for determining the approximate function characteristics of the programmable segment may vary according to the relative priorities of the amount of computation, the amount of power consumption, and the approximate error value. The priorities may be set in the activation function conversion program unit 3000. In other words, the activation function conversion program unit 3000 may search for programmable parameters that implement the approximate function of the programmable segment to achieve a specific performance between high-speed operation, low power consumption, and suppression of a decrease in inference accuracy. However, the examples of the present disclosure are not limited to a specific approximation criterion.
[0170] The main memory 4000 can store data required for the calculation of the artificial neural network model. The main memory 4000 can include one of memories such as ROM, SRAM, DRAM, resistive RAM, magnetoresistive RAM, phase change RAM, ferroelectric RAM, flash memory, HBM, etc. The main memory 4000 can be composed of at least one memory unit, and the main memory 4000 can be configured as a homogeneous memory unit or a heterogeneous memory unit.
[0171] The image sensor 5000 generates image or video data from light entering through a lens. The NPU 1000 may use the image or video data as input data for a DNN processed in the NPU 1000.
[0172] The decoder 6000 decodes the input data of the encoded bit stream, and the decoded input data can be used as the input of the DNN.
[0173] The bitstream may be a bitstream encoded to perform at least one task.
[0174] Tasks that may be included in the bitstream may include object detection, object segmentation, image / video reconstruction, image / video enhancement, object tracking, event recognition, event prediction, anomaly detection, density estimation, event search, measurement, etc.
[0175] The bitstream may include multiple encoding operation values that can handle multiple tasks.
[0176] The output data of the decoder 6000 may be an image, a video, a calculated value of a specific layer of a DNN, etc.
[0177] The following will refer to Figure 2 to 4 describe the activation function programming method in detail.
[0178] Figure 2 is a schematic flow chart illustrating an activation function programming method according to an example of the present disclosure.
[0179] See also Figure 2 The activation function programming method includes a step S200 of generating segment data for segmenting the activation function, a step S210 of segmenting the activation function into a plurality of segments using the generated segment data, and a step S220 of approximating at least one of the plurality of segments as a programmable segment.
[0180] In step S200, segment data is generated. The segment data is data generated to segment the activation function into multiple segments. The segment data will be described later.
[0181] In step S210, the activation function is segmented into a plurality of segments using the generated segment data. In the present disclosure, the term "segment" refers to a portion of an activation function divided into a plurality of intervals, and can be distinguished from a "candidate segment" or a "programmable segment", which are terms related to the approximation of an activation function.
[0182] In various examples, step S210 may include a step of determining the number and width of the plurality of segments based on the segment data. In step S210, the number of the plurality of segments of the activation function to be segmented and the width of each of the plurality of segments may be determined using the segment data. At least one segment of the plurality of segments may have the same width or a different width than the other segments.
[0183] In the present disclosure, a segment in the plurality of segments may be represented as the coordinates of the start point and the end point along the X-axis. At the same time, it should be understood that when determining the number and width of each segment in the plurality of segments, the number and width of the plurality of segments may be used to obtain the coordinates of the segment in the plurality of segments.
[0184] In step S220, at least one of the plurality of segments is approximated as a programmable segment. The programmable segment may be programmed according to the hardware configuration of the PAFE unit 500. That is, the activation function converter unit 3000 may be configured to program the activation function to be processed in the NPU 1000 based on the hardware configuration of the PAFE unit 500.
[0185] For example, the PAFE unit 500 may be configured to have hardware configured to calculate each segment with a specific gradient and a specific offset. The activation function converter unit 3000 may be configured to receive configuration information of the PAFE unit 500 .
[0186] In this case, the activation function converter unit 3000 may program the segments of the corresponding activation function in the form of a linear function with a slope and an offset or in the form of a higher than quadratic function. For example, the programmable segment may be approximated by a linear function according to certain criteria. In this case, the activation function converter unit 3000 may generate a programmable segment expressed in the form of '(gradient a)*(input value x)+(offset b)'. The above-mentioned specific gradient and specific offset may be programmable parameters. In the case of determining that the programmable segment is approximated by a linear function, step S220 may include the step of approximating the selected segment with a specific gradient and a specific offset value.
[0187] In detail, in some examples, steps 210 and 220 can be substantially performed in one step. This is because the step of segmenting the segment and the step of generating programmable parameters of the corresponding programmable segment can be performed simultaneously. In detail, in some examples, steps 210 and 220 can be modified to segment the activation function into multiple segments using the generated segment data and approximate at least one of the multiple segments as a programmable segment.
[0188] Figures 3A-3C is a diagram illustrating a process of approximating an activation function by an activation function programming method according to an example of the present disclosure.
[0189] Figure 3A The activation function shown in can be used as Figure 3B The segment data shown is segmented into a plurality of segments s1, s2, s3 and s4. The plurality of segments s1, s2, s3 and s4 are approximated as follows Figure 3C The programmable segments shown are a1x+b1, a2x+b2, a3x+b3, and a4x+b4. Here, an example will be described in which the activation function conversion program unit 3000 generates programmable parameters so that all programmable segments correspond to linear functions.
[0190] Each programmable segment includes corresponding programmable parameters. Figure 3C In the example, all the multiple segments are approximated as programmable segments in the form of linear functions. However, in various examples, some of the multiple segments can be approximated with other types of programmable segments. For example, the activation function conversion program unit 3000 can program each programmable segment in the form of a linear function, a quadratic function, a cubic function, a logarithmic function, etc.
[0191] For example, only segments s1, s3, and s4 are approximated as programmable segments, while segment s2 can be approximated using various methods available in the device (in which the activation function will be processed). Specifically, if there is a lookup table, a nonlinear approximation equation, etc. for the interval of segment s2 previously determined and stored in the hardware, the segment s2 can be approximated using the predetermined and stored lookup table, nonlinear approximation equation, etc.
[0192] In other words, the activation function conversion program unit 3000 may be configured to independently program each of the segments s1, s2, s3, and s4. At this time, the activation function conversion program unit 3000 receives the hardware configuration information of the PAFE unit 500. The activation function conversion program unit 3000 may be configured to independently determine the approximation method of each of the segments s1, s2, s3, and s4 based on the hardware configuration information of the PAFE unit 500.
[0193] For example, the PAFE unit 500 may be configured to include a circuit that supports linear function operations. In this case, the activation function conversion program unit 3000 may program each of the segments s1, s2, s3, and s4 in the form of a linear function.
[0194] For example, the PAFE unit 500 may be configured to include circuits supporting linear function and quadratic function operations. In this case, the activation function conversion program unit 3000 may program each of the segments s1, s2, s3, and s4 in the form of a linear function or a quadratic function.
[0195] For example, the PAFE unit 500 may be configured to include circuits supporting linear, quadratic and logarithmic functions. In this case, the activation function converter unit 3000 may selectively program each of the segments s1, s2, s3 and s4 in the form of a linear, quadratic or logarithmic function.
[0196] For example, the PAFE unit 500 may be configured to include circuits supporting linear, quadratic, logarithmic, and exponential functions. In this case, the activation function converter unit 3000 may selectively program each of the segments s1, s2, s3, and s4 in the form of a linear, quadratic, logarithmic, or exponential function.
[0197] For example, if the PAFE unit 500 is configured to include a circuit configured to support at least one specific function operation, the activation function conversion program unit 3000 may program each of the segments s1, s2, s3, and s4 in the form of a corresponding specific function.
[0198] For example, the PAFE unit 500 may be configured to include at least one of a linear function calculation circuit, a quadratic function calculation circuit, a cubic function calculation circuit, a logarithmic function calculation circuit, an exponential function calculation circuit, or a similar function calculation circuit designed as hardware.
[0199] For example, the activation function converter unit 3000 can program the same activation function in different ways.
[0200] For example, the activation function conversion program unit 3000 may program a specific activation function as only a linear function.
[0201] For example, the activation function conversion program unit 3000 may program a specific activation function as only a quadratic function.
[0202] For example, the activation function conversion program unit 3000 may program a specific activation function as only a cubic function.
[0203] For example, the activation function conversion program unit 3000 may program a specific activation function as only a logarithmic function.
[0204] For example, the activation function conversion program unit 3000 may program a specific activation function as only an exponential function.
[0205] For example, the activation function conversion program unit 3000 may program each of a plurality of segments of a specific activation function into a corresponding approximate function.
[0206] For example, the activation function conversion program unit 3000 may program multiple segments of a specific activation function into a set of approximate functions with different functions.
[0207] Figures 4A-4D 2 is a diagram illustrating various cases of segmenting an activation function into a plurality of segments by an activation function programming method according to an example of the present disclosure.
[0208] See also Figure 4A , the PAF can be segmented into 4 segments with uniform width.
[0209] See also Figure 4B , the PAF can be segmented into four segments with different widths.
[0210] See also Figure 4C , the PAF can be segmented into four segments with different widths.
[0211] See also Figure 4D , PAF can be segmented into 6 segments with different widths.
[0212] The segment data may be used to determine the number of segments and the width of each segment.
[0213] The activation function converter unit 3000 may be configured to segment a plurality of segments having different widths by analyzing the nonlinearity of the activation function. However, the present disclosure is not limited thereto.
[0214] The activation function conversion program unit 3000 may be configured to analyze the nonlinearity of the activation function so that each of the plurality of segments is segmented to have an optimal width. However, the present disclosure is not limited thereto.
[0215] In the present disclosure, the activation function may be implemented in various forms including feature segments. When the activation function is segmented into multiple segments, the number and width of the multiple segments may be determined differently according to various shapes of the activation function.
[0216] For example, various activation functions, such as swish function, Mish function, sigmoid function, hyperbolic tangent (tanh) function, SELU function, Gaussian error linear unit (GELU) function, SOFTPLUS function, ReLU function, Leaky ReLU function, Maxout function, ELU function, etc., can have various shapes, which are divided into multiple feature intervals including (basic) linear intervals and / or nonlinear intervals. Therefore, when approximating a nonlinear activation function to make it processable in hardware, these feature intervals are considered for segmentation, that is, if the number and width of the segments are determined in consideration of the (basic) linear interval and the nonlinear interval, the activation function can be approximated more effectively in response to the characteristics of each activation function.
[0217] Therefore, in the method of approximating the activation function according to the present disclosure, the concept of segmented data is proposed to segment the activation function taking into account these characteristic intervals of the activation function. The segment data may include discontinuity information of the activation function, derivative data, information about the hardware (in which the activation function is processed), etc., and may include data after processing thereof.
[0218] Next, we will refer to FIG. 5A to FIG. 7B Describes the detailed process of segmenting the activation function into multiple segments using the discontinuity information in the segment data.
[0219] Figures 5A-5C is a diagram showing an example of segmenting an activation function into linear segments or nonlinear segments by using slope change data of segment data of an activation function programming method according to an example of the present disclosure.
[0220] The gradient change point of the activation function may represent a point at which the gradient of the activation function changes. For example, the activation function conversion program unit 3000 may be configured to generate slope change data (e.g., differential data) for analyzing the gradient change point of the activation function. However, the slope change data of the present disclosure is not limited to differential data, and may also include similar data.
[0221] The slope change data according to the example of the present disclosure may include an n-order differential value of the activation function, such as a first-order derivative value, a second-order derivative value, a third-order derivative value, etc. The slope change data may represent a gradient change rate and a gradient change point associated with the activation function.
[0222] The slope change data according to the example of the present disclosure may include an n-order derivative value of the activation function, such as a linear derivative value, a second-order derivative value, and a third-order derivative value, wherein the slope change data may represent a gradient change rate and a gradient change point associated with the activation function.
[0223] The following will refer to Figures 5A-5C Describes the process of searching for gradient change points.
[0224] exist Figure 5A In the differential data of the activation function f(x) shown, Figure 5B The first-order derivative f'(x) is shown. In addition, Figure 5A In the differential data of the activation function f(x) shown, Figure 5C The second order derivative f"(x) is shown.
[0225] For example, the activation function conversion program unit 3000 can be configured to extract the start and end points of the interval where the first-order derivative value does not change. Figure 5B As shown, the activation function conversion program unit 3000 generates the slope change data corresponding to the first-order derivative value. Then, the activation function conversion program unit 3000 identifies that in interval w3 and interval w4, although the first-order derivative values are different from each other, the first-order derivative value does not change. Therefore, the activation function conversion program unit 3000 can determine the w2 interval and the w3 interval as linear intervals respectively. In other words, the slope change data corresponding to the first-order derivative value in the linear interval remains unchanged. However, since the first-order derivative values in the w2 interval and the w3 interval are different, there are discontinuous points d1 and d2 in the slope change data corresponding to the first-order derivative value at the boundary between the w2 interval and the w3 interval. In other words, since the slope change data corresponding to the first-order derivative value at the boundary of each of the w2 interval and the w3 interval is a discontinuous point, the boundary of each w2 interval and the w3 interval may correspond to a gradient change point.
[0226] In this case, the activation function conversion program unit 3000 can convert the linear segment into a programmable parameter in the form of a corresponding linear function. Therefore, the linear segment of the activation function to be programmed can be segmented into a linear function with a specific slope and a specific offset. The first-order derivative of the linear segment can be a constant value. In other words, even if the linear segment is approximated by a linear function, the approximation error value may be zero. Therefore, the activation function conversion program unit 3000 can determine that there is basically no approximation error in the w2 and w3 intervals. That is, when the activation function conversion program unit 3000 approximates each of the w2 and w3 intervals with a linear function, the calculation amount and power consumption of the PAFE unit 500 are minimized, and the approximation error value may also be zero.
[0227] The activation function conversion program unit 3000 may be configured to determine a segment where the first-order derivative of the activation function is a constant or non-zero as a segment of a quadratic function or a high-order term or a curve (non-linear function).
[0228] In the present disclosure, the term "linear segment" related to differential data refers to a segment where the first-order derivative of the activation function is an integer or zero, or a segment where the activation function is represented by a linear function, and the term "non-linear segment" may refer to a segment where the first-order derivative of the activation function is not an integer or zero. However, the determination of the linear segment of the example of the present disclosure is not determined only by the differential value. That is, the activation function conversion program unit 3000 can be configured to determine or classify the linear segment in various ways by receiving the activation function.
[0229] The activation function conversion program unit 3000 may be configured to preferentially determine whether a linear segment exists. The activation function conversion program unit 3000 may be configured to convert the linear segment into a programmable parameter in a linear function form, and convert the remaining nonlinear segment into a programmable parameter in a specific function form.
[0230] In detail, the differential data described in the examples of the present disclosure is only a mathematical calculation method for calculating the slope of the activation function. Therefore, the present disclosure is not limited to differential values, but can use a substantially similar slope calculation method.
[0231] The search for the gradient change point is not limited to the above method, and the activation function converter unit 3000 may be configured to determine the corresponding point as a gradient change point when the change in the first-order derivative of the activation function along the X-axis is greater than a specific threshold.
[0232] Then, the activation function conversion program unit 3000 can be configured to extract the start and end points of the segments where the second-order derivative value does not change. Figure 5CAs shown, the activation function conversion program unit 3000 generates the slope change data corresponding to the second-order derivative. Then, the activation function conversion program unit 3000 determines that the second-order derivative values are different but not changing in the w1-1 and w1-2 intervals. However, since the second-order derivative values in the w1-1 and w1-2 intervals are different, the slope change data corresponding to the second-order derivative at the boundary between the w1-1 and w1-2 intervals has a discontinuous point d3. That is, since the slope change data corresponding to the second-order derivative at the boundary between the w1-1 interval and the w1-2 interval is a discontinuous point d3, the boundary between the w1-1 interval and the w1-2 interval can correspond to a gradient change point.
[0233] In this case, the activation function conversion program unit 3000 can convert the nonlinear interval into a programmable parameter in the form of a corresponding quadratic function. Therefore, the nonlinear interval of the activation function to be programmed can be segmented into a quadratic function including a quadratic term coefficient and a coefficient of a linear function including a specific slope and a specific offset. The second-order derivative of the nonlinear interval can be a constant value. In other words, even if the nonlinear interval is approximated by a quadratic function, the approximation error value may also be zero. Therefore, the activation function conversion program unit 3000 can determine that there is basically no approximation error in each of the w1-1 and w1-2 intervals. That is, when the activation function conversion program unit 3000 approximates each of the w1-1 and w1-2 intervals with a quadratic function, the computational complexity and power consumption of the PAFE unit 500 are minimized, and the approximation error value may also be zero.
[0234] However, the examples of the present disclosure are not limited thereto, and a linear function may be used to approximate the w1-1 and w1-2 intervals. In this case, the approximation error value may increase, but the power consumption of the NPU 1000 may be reduced by reducing the amount of calculation of the PAFE unit 500 of the NPU 1000. That is, the activation function converter unit 3000 may determine programmable parameters differently according to different priorities among the amount of calculation, the amount of power consumption, and the approximation error value.
[0235] The second-order derivative of the above activation function can represent the rate of change of the slope of the activation function. Since the interval where the second-order derivative of the activation function is larger is the interval where the slope change rate is larger, the segment of the activation function corresponding to the interval has a larger slope change, so there is a significant increase or decrease. On the contrary, since the interval where the second-order derivative of the activation function is relatively small is the interval where the slope change rate is smaller, the segment of the activation function corresponding to the interval has a smaller slope change, so there is a smaller increase or decrease.
[0236] Specifically, the interval where the second-order derivative of the activation function is less than or equal to a certain threshold is an interval where the rate of change of the slope is very small.
[0237] Therefore, the activation function conversion program unit 3000 may be configured to determine the activation function of such an interval as a substantially linear function interval whose slope hardly changes.
[0238] For example, the activation function conversion program unit 3000 may be configured to determine an interval in which the second-order derivative of the activation function is less than or equal to a threshold as a “substantially linear interval.” The threshold of the second-order derivative of the activation function will be described later.
[0239] The differential order when the differential value of the activation function becomes zero or an integer can represent the degree of change of the slope of the activation function. Specifically, in general, since the gradient of a function changes rapidly as the number of the highest order term of the function increases, the interval with a higher number of the highest order term of the activation function is an interval with a steeper slope change, and it can be segmented into more segments by distinguishing it from other intervals.
[0240] The order of the highest order term of the activation function in a specific interval can be determined by the order of differences at which the difference value becomes zero or an integer in the specific interval.
[0241] For example, for an activation function whose highest order term in a specific interval is the third order, since the third order derivative of the activation function in the specific interval becomes an integer (i.e., the coefficient of the highest order term), and the fourth order derivative of the activation function becomes zero, it can be determined that the activation function whose third order derivative is an integer or whose fourth order derivative is zero in the specific interval has the third order as the highest order term in the specific interval.
[0242] In various examples, intervals where the highest order term of the activation function is third order or higher may be segmented into more segments than other intervals. For example, the number of segments may be determined as the maximum number of segmentable segments of the corresponding interval in the hardware to process the activation function.
[0243] The slope change data (i.e., the first-order derivative f'(x)) can be used to identify the gradient change point of the activation function. Using the slope change data (i.e., the first-order derivative f'(x)), the activation function f(x) can be segmented into three intervals w1, w2, and w3, including two linear intervals w2 and w3.
[0244] That is, the activation function conversion program unit 3000 can determine the linear intervals w2 and w3 and the nonlinear interval w3 using the slope change data of the activation function f(x) to be programmed and segment the linear intervals w2 and w3 and the nonlinear interval w3.
[0245] That is, the activation function f(x) can be segmented according to the points or intervals where the first-order derivative f'(x) is a constant (non-zero), zero, a curve (non-linear function) below a threshold, or a curve (non-linear function). In other words, the activation function f(x) can be segmented according to the points where the activation function f(x) is not differentiable or the points where the first-order derivative f'(x) is discontinuous.
[0246] Although Figure 5B The result of segmentation into three intervals is shown in , but this is for the purpose of briefly explaining the process of segmentation into linear intervals and nonlinear intervals. Therefore, it should be understood that the activation function f(x) can be segmented into four or more intervals, i.e., at least four segments, using segment data.
[0247] For example, according to the activation function programming method of the example of the present invention, the linear interval w1 can be further segmented into multiple intervals using segment data. The activation function can be segmented into a larger number of segments and approximated by additional segmentation of the linear interval w1, thereby reducing the approximation error. In the present invention, the term "approximation error" refers to the difference between a specific segment of the activation function and a programmable segment that approximates the specific segment.
[0248] Figure 6A-6B is a diagram showing an example of segmenting an activation function into a basic linear interval and a nonlinear interval using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0249] Figure 6B Shows Fig. 6A The absolute value of the second-order derivative f" (x) of the derivative data of the activation function f (x) shown in . The activation function conversion program unit 3000 can be configured to determine the basic linear interval by setting a specific threshold for the second-order derivative f" (x). Figure 6B , when the maximum value Max of the absolute value of the second-order derivative f" (x) of the activation function f(x) is 0.5, the threshold Th can be set to 0.05, which is 10% of the maximum value Max. In other words, it can be determined that the activation function has linear characteristics as the second-order derivative f" (x) becomes smaller, and has nonlinear characteristics as the second-order derivative f" (x) becomes larger.
[0250] That is, the threshold value Th can be determined as a relative ratio of the maximum value Max of the absolute value of the second-order derivative f"(x) of the activation function f(x). The threshold value Th of the basic linear interval can be determined based on whether the error occurring when the nonlinear interval is approximated as a linear interval is acceptable. For example, the threshold value of the basic linear interval can be determined based on the error value level of each segment, which determines the degree of deterioration of the inference accuracy of the DNN to which the PAF is applied.
[0251] In other words, as the threshold of the basic linear interval increases, the segments of the linear interval can be programmed to be wider. At the same time, as the width of the segment increases, the number of segments can be reduced. That is, the total number and width of the segments of the PAF can be different according to the threshold of the basic linear interval.
[0252] The search of the basic linear interval may be performed after the search of the linear interval. However, the present disclosure is not limited to the order of the linear interval search and the basic linear interval search.
[0253] exist Figure 6B In the example, the relative ratio can be determined as 10%. However, the present disclosure is not limited to this, and can be determined as 5% of the maximum value Max according to the allowable error of the DNN. Using differential data, i.e., the second-order derivative f"(x), the activation function f(x) can be segmented into intervals w1 and w3 where the second-order derivative f"(x) is less than the threshold Th of the basic linear interval, and interval w2 where the second-order derivative f"(x) is greater than or equal to the threshold Th of the basic linear interval. In the activation function f(x), the basic linear intervals w1 and w3 and the nonlinear interval w2 can be determined and segmented using slope change data. When the first to third intervals w1, w2 and w3 are determined, the first to third segments s1, s2 and s3 can be programmed as programmable segments using corresponding programmable parameters.
[0254] In 6B, the result of segmenting into three segments s1, s2, s3 corresponding to the three intervals w1, w2, w3 is shown, which is to simply explain the process of segmenting into basic linear intervals and nonlinear intervals. That is, it should be understood that the activation function f(x) can be segmented into four or more intervals, i.e., four or more segments, using segment data.
[0255] For example, according to the activation function programming method of the example of the present disclosure, the nonlinear interval w2 can be further segmented into multiple intervals using segment data. By additionally segmenting the nonlinear interval w2, the approximation error can be reduced.
[0256] 7 is a diagram showing another example of segmenting an activation function into a basic linear interval and a nonlinear interval using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0257] 7 , in the activation function f(x), the nonlinear interval can be determined according to the threshold Th of the basic linear interval of the segment data, that is, the absolute value of the second-order derivative value f”(x). That is, the interval equal to or greater than the threshold Th of the basic linear interval is determined as the nonlinear interval. Specifically, referring to Figure 7B, the activation function conversion program unit 3000 can use the differential data, that is, the second-order derivative f" (x), to segment the activation function f(x) into a basic linear interval and a nonlinear interval. Further, as an example, the activation function conversion program unit 3000 can segment the nonlinear interval of the activation function f(x) into segments s2 and s3 corresponding to two intervals w2 and w3.
[0258] That is, the activation function conversion program unit 3000 can use the slope change data of the activation function f(x) to divide the basic linear intervals w1 and w4 and the nonlinear intervals w2 and w3, and then the nonlinear intervals w2 and w3 can be segmented.
[0259] The activation function conversion program unit 3000 may be configured to search for optimal programmable parameters corresponding to each segment in various ways. For example, the activation function conversion program unit 3000 may search for optimal programmable parameters that can achieve specific performance between high-speed operation, low power consumption, and suppression of inference accuracy degradation.
[0260] exist Figure 7B In FIG. 1 , segments s1, s2, s3, and s4 are shown segmented into four intervals w1, w2, w3, and w4, however, this is to briefly explain the process of segmenting into basic linear intervals and nonlinear intervals. Therefore, it should be understood that the activation function f(x) can be segmented into five or more intervals, i.e., five or more segments, using segment data.
[0261] For example, according to the activation function programming method of the example of the present disclosure, the nonlinear intervals w2 and w3 can be further segmented into multiple intervals using segment data. Specifically, the nonlinear intervals w2 and w3 can be segmented based on the maximum value Max of the second-order derivative f" (x). That is, the area from the threshold value Th of the basic linear interval to the maximum value Max of the second-order derivative f" (x) is segmented into interval w2. Further, the area from the maximum value Max of the second-order derivative f" (x) to the threshold value Th of the basic linear interval is segmented into interval w3.
[0262] When additional segmentation is performed in the nonlinear intervals w2 and w3, the approximation error can be further reduced.
[0263] Figures 8A-8B is a diagram showing another example of segmenting an activation function into nonlinear intervals using slope change data in segment data in an activation function programming method according to an example of the present disclosure.
[0264] See also Figures 8A-8BIn the activation function f(x), the nonlinear interval can be determined according to the threshold Th of the basic linear interval of the segment data, that is, the absolute value of the second-order derivative value f"(x). That is, the area equal to or greater than the threshold Th of the basic linear interval can be determined as the nonlinear interval. Specifically, see Figure 8B , the activation function conversion program unit 3000 can use the differential data, that is, the second-order derivative value f" (x) to segment the activation function f(x) into a basic linear interval and a nonlinear interval. In addition, the activation function conversion program unit 3000 can, for example, segment the nonlinear interval of the activation function f(x) into segments s2, s3 and s4 corresponding to the three intervals w2, w3 and w4.
[0265] The activation function conversion program unit 3000 can divide the basic linear intervals w1 and w5 and the nonlinear intervals w2, w3 and w4, and then segment the nonlinear intervals w2, w3 and w4 using the slope change data of the activation function f(x).
[0266] However, the examples of the present disclosure are not limited to the basic linear interval, and the basic linear interval may also be segmented into nonlinear intervals. That is, the step of determining the basic linear interval may not be performed in some cases.
[0267] The activation function conversion program unit 3000 may be configured to search for optimal programmable parameters corresponding to each segment in various ways. For example, the activation function conversion program unit 3000 may search for optimal programmable parameters that can achieve specific performance between high-speed operation, low power consumption, and suppression of inference accuracy degradation.
[0268] exist Figure 8B In FIG. 1 , segments s1, s2, s3, s4, s5 are shown segmented into five intervals w1, w2, w3, w4, w5, but this is only for the purpose of simply explaining the process of segmenting into basic linear intervals and nonlinear intervals. Therefore, it should be understood that the activation function f(x) can be segmented into six or more intervals, i.e., six or more segments, using segmented data. However, the examples of the present disclosure are not limited to basic linear intervals, and basic linear intervals can also be segmented into nonlinear intervals.
[0269] For example, according to the activation function programming method of the example of the present disclosure, the nonlinear intervals w2, w3, and w4 can be further segmented into multiple intervals using segmentation data.
[0270] Specifically, the nonlinear intervals w2, w3, and w4 may be segmented based on the integral value (∫f″(x)dx) of the second-order derivative f″(x). In other words, the activation function conversion program unit 3000 may segment the nonlinear interval based on the integral value of the slope change data.
[0271] When the value of the integral (∫f″(x)dx) of the second-order derivative f”(x) is high, the approximation error value between the PAF and the activation function may increase. That is, when the value of the integral (∫f″(x)dx) of the second-order derivative value f”(x) is high, an error may occur, resulting in a decrease in inference accuracy. On the other hand, as the value of the integral (∫f″(x)dx) of the second-order derivative f”(x) increases, the width of the segment becomes wider. Conversely, the smaller the value of the integral (∫f″(x)dx) of the second-order derivative f”(x), the narrower the width of the segment.
[0272] Accordingly, the activation function conversion program unit 3000 may set the integral value (∫f″(x)dx) of a specific second-order derivative f”(x) as an integral threshold of the segment approximation error. For example, the activation function conversion program unit 3000 may integrate the second-order derivative f”(x) starting from the end of interval w1. Accordingly, interval w2 may start from the end of interval w1 until a preset integral threshold of the segment approximation error reaches a specific value.
[0273] More specifically, in the interval w2, the integral of the second-order derivative f”(x) can be segmented into s2 to correspond to the integral threshold of the segment approximation error. Further, in the w3 interval, the integral of the second-order derivative f”(x) It can be segmented into s3 to correspond to the integral threshold of the segmented approximation error. Further, in segment w4, the integral of the second-order derivative f”(x) It can be segmented into s4 to correspond to the integral threshold of the segmented approximation error.
[0274] That is, the integral value of the second-order derivative f”(x) in the interval w2 The integral value of the second-order derivative f”(x) in the interval w3 The integral value of the second-order derivative f′(x) in the interval w4 Both can be the same value as the integration threshold of the segment approximation error.
[0275] However, the integral threshold of the segment approximation error may be affected by hardware data, including at least one of the number of comparators of the PAFE unit 500 of the NPU 1000, the number of gates of a circuit for implementing the PAFE unit 500, and the type of arithmetic circuit implemented (linear function circuit, quadratic function circuit, cubic function circuit, exponential circuit, logarithmic circuit, antilogarithmic circuit, etc.). That is, the activation function converter unit 3000 may be configured to determine the integral threshold of the segment approximation error in consideration of the hardware data.
[0276] That is, the smaller the integral threshold of the segment approximation error is, the closer the PAF is to the activation function. In other words, when the integral threshold of the segment approximation error is reduced, the number of programmable segments increases, thereby further reducing the approximation error value of the PAF.
[0277] However, since the number of programmable segments is limited by hardware data, there is a limit to lowering the integral threshold of the segment approximation error. That is, the minimum limit of the integral threshold of the segment approximation error can be determined according to the hardware data.
[0278] When additional segmentation is performed in the above nonlinear intervals w2, w3, w4, the approximation error can be further reduced. However, the examples of the present disclosure are not limited to the basic linear intervals, and the basic linear intervals can also be segmented into nonlinear intervals. That is, the step of determining the basic linear interval may not be performed in some cases.
[0279] like FIG. 5A to FIG. 8B As shown, the activation function converter unit 3000 can segment the activation function using the slope change data, and determine a linear interval from the activation function before approximating the activation function. When the activation function converter unit 3000 segments the activation function using the slope change data, a nonlinear interval can be determined from the activation function before approximating the activation function. When the activation function converter unit 3000 segments the activation function using the slope change data, a basic linear interval can be determined from the activation function before approximating the activation function.
[0280] A segment having a clearly linear interval or a substantially linear interval can be approximated as a programmable segment expressed in the form of '(slope a)*(input value x)+(offset b)'.
[0281] At this time, the segment with a linear interval or a substantially linear interval is in the form of a linear function or a substantially linear function with a substantially constant slope. Therefore, when the activation function is compared with the programmable segment represented by the slope and the offset, the programmed segment has no approximation error or can be minimized.
[0282] Therefore, if the activation function is programmed using slope change data, the amount of computation and power consumption in the linear interval or the substantially linear interval can be greatly reduced.
[0283] Therefore, according to the examples of the present disclosure, using a linear or basic linear interval programming activation function is efficient and the approximation error is minimized, so that the operation speed of the DNN processed in the NPU 1000 can be improved, the decrease in inference accuracy can be minimized, and the power consumption of the NPU 1000 can be reduced.
[0284] In various examples, step S210 may further include determining a linear interval of the activation function based on the slope change data of the activation function.
[0285] In various examples, step S210 may further include determining a nonlinear interval of the activation function based on slope change data of the activation function.
[0286] In various examples, step S210 may further include determining a substantially linear interval of the activation function based on slope change data of the activation function.
[0287] In various examples, step S210 may further include determining a linear interval and a nonlinear interval of the activation function based on the slope change data of the activation function.
[0288] In various examples, step S210 may further include determining a basic linear interval and a nonlinear interval of the activation function based on the slope change data of the activation function.
[0289] In various examples, step S210 may further include determining a linear interval, a substantially linear interval, and a nonlinear interval of the activation function based on differential data of the activation function.
[0290] However, examples of the present disclosure are not limited to differential data of the activation function, but various mathematical analyses capable of analyzing the slope change and linearity of the activation function may also be performed.
[0291] In various examples, the segment data may include information of hardware that processes the activation function. In the activation function programming method according to the example of the present disclosure, the activation function may be segmented using the hardware information. The hardware data may include at least one of the number of comparators of the PAFE unit 500 of the NPU 1000, the number of gates for implementing the circuit of the PAFE unit 500, and the type of arithmetic circuit implemented (linear function circuit, quadratic function circuit, cubic function circuit, exponential circuit, logarithmic circuit, antilogarithmic circuit, etc.).
[0292] For example, the number of segments for segmenting the activation function may be limited according to the number of comparators of the PAFE unit 500 of the NPU 1000. Therefore, the activation function may be segmented into the maximum number of segments to be processed that can be processed by the NPU 1000 or the number of segments corresponding to the allocated resources of the NPU 1000. Therefore, the activation function converter unit 3000 may program the activation function using predetermined hardware resources more efficiently or in a more customized manner.
[0293] In various examples, step 220 may also include approximating at least one of the plurality of segments as a programmable segment based on the gradient change point.
[0294] In various examples, step 220 may also include approximating at least one of the plurality of segments as a programmable segment based on the error value.
[0295] In the present disclosure, the term "error value" or "approximate error value" refers to the difference between a specific segment of an activation function and a programmable segment that is approximated by the specific segment. The approximate error value may also include an average value, a minimum value, a maximum value, and a cumulative value. In other words, the activation function conversion program unit 3000 may be configured to calculate an average error value, a minimum error value, a maximum error value, a cumulative error value, etc. between a specific segment and an approximate programmable segment. The cumulative error value may be a value obtained by integrating the error value between a specific segment and an approximate programmable segment.
[0296] Regarding the error value, various activation functions can be divided into multiple feature segments including (basic) linear intervals and / or nonlinear intervals. If these feature segments are segmented into segments of the same width, the error value of each segment is very different. Therefore, in the activation function programming method according to the example of the present disclosure, in order to reduce the approximation error, at least one feature of these feature intervals can be considered and approximated as a programmable segment.
[0297] In various examples, step S220 may also include calculating an error value by comparing the gradient and offset of the programmable segment with the corresponding segment of the activation function.
[0298] In various examples, step S220 may also include determining a programmable parameter for converting at least one segment of the activation function into a programmable segment. In other words, step S220 may also include searching for optimal programmable parameters for converting at least one segment of the activation function into a programmable segment. Here, when the programmable segment is a linear function, the programmable parameter may include a gradient and an offset corresponding to the linear function. Here, when the programmable segment is a quadratic function, the programmable parameter may include a coefficient corresponding to a quadratic term of the quadratic function. The coefficients of the quadratic function may include a quadratic coefficient, a linear coefficient, and a constant. The approximate function of the programmable parameter may be determined by considering performance such as high-speed operation, low power consumption, and suppression of a decrease in inference accuracy. For example, as the approximate function formula becomes more complex, the calculation speed may decrease and the power consumption may increase. As the approximation error decreases, the decrease in inference accuracy may decrease.
[0299] In various examples, step S220 may further include calculating an error value between at least one segment of the activation function and at least one candidate segment having a (temporary) gradient and a (temporary) offset. As the number of candidate segments increases, the likelihood of searching for the best programmable parameter value increases, but the search time may increase.
[0300] In various examples, step S220 may include determining a parameter of at least one candidate segment as a programmable parameter of the programmable segment based on the calculated error value.
[0301] Therefore, the activation function conversion program unit 3000 may provide programmed activation function data to the NPU 1000. Here, the programmed activation function data may include at least one programmed activation function. Here, the programmed activation function data may include programmable parameters corresponding to each programmable segment of the at least one programmed activation function.
[0302] The following will refer to Figures 9 to 11B A process of approximating at least one segment of a plurality of segments as a programmable segment based on an error value is described in detail.
[0303] In the process of programming the activation function, steps may occur at the boundaries between programmable segments. In the activation function programming method according to the example of the present disclosure, the approximation error can be greatly reduced by generating predetermined steps between programmable segments or at the beginning and / or end of a programmable segment.
[0304] Therefore, in the present disclosure, by allowing steps between programmable segments in the process of segmenting the activation function into multiple segments using segment data and approximating at least one of the multiple segments as a programmable segment based on the error value, the error value can be significantly reduced.
[0305] See also Fig. 9 , multiple candidate segments Sc1, Sc2 and Sc3 of segment s of the non-linear activation function are shown.
[0306] In the examples of the present disclosure, the term “candidate segment” means that a function that can be made into a programmable segment represented by a “programmable parameter” using an activation function programming method.
[0307] For example, when the programmable segment is represented as a linear function, the programmable segment can be represented as “(gradient a)*(input value x)+(offset b)”. Here, the programmable parameters include the gradient a and the offset b.
[0308] For example, when the programmable segment is represented as a quadratic function, the programmable segment can be represented as '(quadratic coefficient a)*(input value x2)+(linear coefficient b)*(input value x)+(constant c)'. Here, the programmable parameters include quadratic coefficient a, linear coefficient b and constant c.
[0309] Therefore, the programmable parameters may be configured to have a form capable of expressing first-order functions and second-order functions. However, the present disclosure is not limited to the format of the programmable parameters.
[0310] The following description will be made by taking a linear function as an example. The candidate segment may be in the form of a linear function corresponding to a programmable segment segmented using segment data. The candidate segment of a segment may be determined by a linear function passing through a start point and an end point of a segment.
[0311] For example, a candidate for a segment may be a linear function with an adjusted offset while having the same gradient as a linear function passing through the start and end points of the segment.
[0312] For example, a candidate for a segment may be a linear function with an adjusted offset and having a different gradient than a linear function passing through a start point and an end point of a segment.
[0313] For example, a candidate segment of a segment may be determined as one of the tangent lines of the segment.
[0314] exist Fig. 9 In order to briefly describe the process of determining a programmable segment among multiple candidate segments, three candidate segments having a common gradient passing through the start and end of segment s are shown. The first candidate segment Sc1 is a linear function passing through the start and end of segment s, the second candidate segment Sc2 and the third candidate segment Sc3 are linear functions with an adjusted offset and have a common slope with the first candidate segment Sc1, and the third candidate segment Sc3 has an offset such that the candidate segment Sc3 is tangent to segment s. Fig. 9 The candidate segments shown in are used to briefly describe segments that can become approximate programmable segments, and the gradient and / or offset of the actual candidate segments can be adjusted in various ways to reduce the error value.
[0315] In various examples, at least one of the multiple segments may be approximated as a programmable segment by searching for an error value Δy. At this time, the activation function conversion program unit 3000 may determine the width of each of the multiple segments as a uniform width. Subsequently, the activation function conversion program unit 3000 may approximate at least one of the multiple segments as a programmable segment by searching for an error value Δy of at least one segment. However, the present disclosure is not limited thereto.
[0316] Figures 10A-10B is a diagram showing an example of approximating a segment as a programmable segment by searching for a maximum error value max(Δy) in an activation function programming method according to an example of the present disclosure, wherein the maximum error value max(Δy) is the maximum value among the error values Δy.
[0317] Fig. 10A The segments s1 and s2 that segment the activation function f(x), the first candidate segment sc1(x) corresponding to the first segment s1, and the second candidate segment sc2(x) corresponding to the second segment s2 are shown. Fig. 10A , each of the candidate segments sc1(x) and sc2(x) is searched for the best programmable parameters (ie, gradient and offset) of each linear function representing the start and end points of each segment s1 and s2.
[0318] like Fig. 10AIn the example shown, the activation function conversion program unit 3000 may calculate the error value Δy between the second segment s2 and the second candidate segment sc2(x), that is, the absolute value of 'f(x)-sc2(x)' |f(x)-sc2(x)|. The activation function conversion program unit 3000 may calculate the maximum error value max(Δy), which is the maximum value of the error values Δy. In order to reduce the maximum error value max(Δy) of the second segment s2, as Fig. 10B As shown, the second candidate segment obtained by adjusting the candidate segment sc2(x) in the Y-axis direction (i.e., adjusting the offset) max(Δy) / 2 (i.e., half of the maximum error value max(Δy)) can be determined as the second programmable segment sp2(x) obtained by approximating the second segment s2.
[0319] When Fig. 10B As shown, when the first programmable segment sp1(x) is obtained by approximating the first segment s1, a step may occur between the first programmable segment sp1(x) and the second programmable segment sp2(x).
[0320] exist Fig. 10B In the process of approximating the second segment s2 of the activation function f(x) as a programmable segment based on the error value |f(x)-sc2(x)|, such a step may be intentionally generated at the boundary between adjacent programmable segments on the Y-axis. That is, in the process of approximating a specific programmable segment to reduce the maximum error value within the specific programmable segment, a step may be generated at the boundary point between adjacent programmable segments.
[0321] In other words, each programmable segment can be approximated independently of each other.
[0322] In other words, as the approximation error value of the PAF increases, the decrease in the inference accuracy of the NPU 1000 using the PAF may increase. Conversely, as the approximation error value of the PAF decreases, the decrease in the inference accuracy of the NPU 1000 using the PAF may decrease.
[0323] In various examples, at least one of the plurality of segments may be approximated as a programmable segment using an integrated value of the error value ∫[sc(x)-f(x)]dx The activation function conversion program unit 3000 may be configured to integrate or accumulate the approximation error value of each segment.
[0324] In more detail, the first programmable segment sp1(x) and the second programmable segment sp2(x) can be programmed in different ways. That is, each programmable segment can be programmed by selecting a method such as a linear function, a quadratic function, a logarithmic function, an exponential function, etc. Therefore, each programmable segment can be programmed with the same function, or can be programmed with different functions.
[0325] Figures 11A-11B is a diagram showing an example of approximating a segment as a programmable segment using an integral value ∫[sc(x)-f(x)]dx with respect to an error value in an activation function programming method according to an example of the present disclosure.
[0326] Fig.11A The segments s1 and s2 that segment the activation function f(x), the first candidate segment sc1(x) corresponding to the first segment s1, and the second candidate segment sc2(x) corresponding to the second segment s2 are shown. Fig.11A In , for each of the candidate segments sc1(x) and sc2(x), the best programmable parameters (i.e., gradient and offset) representing a linear function are searched for the start and end points of each segment s1 and s2. In practice, the offset of the second candidate segment sc2(x) can be adjusted while having the same gradient as the linear function passing through the start and end points of the second segment s2. Alternatively, the offset can be adjusted while having a different gradient than the linear function passing through the start and end points of the second segment s2.
[0327] refer to Figures 10A-10B and Figures 11A-11B , the first segment s1 includes a start point x0 and an end point x1. Here, the start point x0 and the end point x1 may represent segment boundary values.
[0328] refer to Figures 10A-10B and Figures 11A-11B , the second segment s2 includes a start point x1 and an end point x2. Here, the start point x0 and the end point x1 may represent segment boundary values.
[0329] For example, the first segment s1 may be set from the start point x0 to less than the end point x1.
[0330] For example, the second segment s2 may be set from the start point x1 to less than the end point x2.
[0331] Programmable parameters may be configured to include segment boundary values.
[0332] like Fig.11A As shown, the activation function conversion program unit 3000 calculates the integral value between the second segment s2 and the candidate segment sc1(x) As the approximate error value, and in the integral value Search for points in The candidate segment with the smallest absolute value. Fig. 11B As shown, in order to reduce the error value, the integral value can be The minimum absolute value of The candidate segment is determined as the second programmable segment sp2(x).
[0333] when Fig. 11BWhen the first programmable segment sp1(x) approximates the first segment s1, a predetermined step may occur on the Y axis between the first programmable segment sp1(x) and the second programmable segment sp2(x). Fig. 11B In the case of a step, the step can occur based on the approximate error value In the process of approximating the second segment s2 of the activation function f(x) to the second programmable segment sp2(x), even if there is a step, if the approximation error value of each programmable segment is minimized, the decrease in the inference accuracy of the NPU 1000 using the PAF can be reduced.
[0334] In various examples, step S220 may further include searching for a minimum approximation error value between the programmable segment and the corresponding segment of the activation function. The approximation error value may be at least one of an average error value, a minimum error value, a maximum error value, and a cumulative error value.
[0335] For example, step S220 may further include searching for at least one minimum error value between at least one programmable segment and a corresponding segment of at least one activation function.
[0336] For example, step S220 may further include determining a slope and an offset of the programmable segment based on the searched at least one minimum error value.
[0337] For example, step S220 may further include approximating at least one segment as a programmable segment according to the determined slope and offset.
[0338] In various examples, step S220 may also include determining the programmable segments using machine learning of a loss function.
[0339] Fig.12 is a diagram showing an example of using machine learning to approximate a segment to an optimal programmable segment in an activation function programming method according to an example of the present disclosure.
[0340] See also Fig.12 , the activation function conversion program unit 3000 can set the candidate segment sc(x) of the activation function f(x) as the initial value of the loss function. The activation function conversion program unit 3000 can determine the candidate segment with the minimum value of the loss function as the optimal programmable segment sop(x) through machine learning. Therefore, the optimized programmable parameters can be explored.
[0341] For optimization parameter search, learning can be repeated. One learning can represent an epoch. As the number of learning increases, the error value will decrease. If the number of training times is too small, it will lead to underfitting. Too many training times will lead to overfitting.
[0342] The loss function may use mean square error (MSE), root mean square error (RMSE), etc., but is not limited thereto. In the present disclosure, the candidate segment used as the initial value of the loss function may be, for example, a linear function, a quadratic function, a cubic function, etc., which approximately corresponds to the segment segmented using the segment data. However, the examples of the present disclosure are not limited to the above functions. That is, the loss function may be used after segmenting the activation function f(x) into a plurality of segments using the segment data.
[0343] Therefore, machine learning using the loss function can be performed after considering the characteristics of its activation function, such as multiple characteristic intervals including the (basic) linear interval and / or nonlinear interval of the activation function, approximation errors, etc. Therefore, the amount of calculation and search time for optimizing the programmable parameter search can be reduced, and the decrease in the inference accuracy of the NPU 1000 due to the use of the PAF can be minimized.
[0344] In addition, according to the example of the present disclosure, the effect of reducing unnecessary segments can be provided. That is, according to the example of the present disclosure, the number of segments can also be minimized. In other words, if the sum of the approximate error values of two adjacent programmable segments is less than a preset threshold, the two programmable segments can be merged into one programmable segment.
[0345] In various examples, step S210 may further include segmenting the activation function into a plurality of segments using an integral (accumulated value) of a second-order derivative of the activation function. Here, the accumulated value of the second-order derivative may be used as segment data.
[0346] For example, step S210 may further include calculating an accumulated value of a second-order derivative of the activation function.
[0347] For example, step S210 may further include segmenting the activation function into a plurality of segments based on an integral threshold of the segment approximation error (ie, a threshold of the accumulated second-order derivative).
[0348] In addition, the activation function programming method according to the present disclosure may further include the following steps: first, when the number of multiple segments determined by segmenting the activation function into multiple segments using the accumulated value of the second-order derivative is greater than or less than the target number, the threshold of the accumulated value of the second-order derivative is adjusted, and the activation function is re-segmented into another number of multiple segments based on the adjusted threshold. Specifically, the adjustment may be as follows: (1) when the number of the determined multiple segments is greater than the target number, the threshold is adjusted to increase; (2) when the number of the determined multiple segments is less than the target number, the threshold is adjusted to decrease.
[0349] In various examples, the activation function conversion program unit 3000 can segment the activation function into multiple segments based on a threshold of the accumulated value of the second-order derivative. In this case, the activation function conversion program unit 3000 can segment all intervals of the activation function based on the threshold of the accumulated value of the second-order derivative, or segment a portion of the interval of the activation function based on the threshold of the accumulated value of the second-order derivative. Specifically, the activation function conversion program unit 3000 can determine that some intervals of the activation function are nonlinear intervals rather than (basically) linear intervals, and can segment only part of the nonlinear interval based on the threshold of the accumulated value of the second-order derivative. The activation function conversion program unit 3000 can segment the remaining intervals of the nonlinear interval by the activation function programming method described in various examples of the present disclosure.
[0350] Fig.13 is a diagram showing an example of segmenting an activation function using an integral threshold of a segment approximation error of an activation function in an activation function programming method according to an example of the present disclosure.
[0351] See also Fig.13 , the activation function f(x) can be segmented using the accumulated value of the second-order derivative of the activation function f(x), i.e., ∫f”(x)dx. The point of the minimum value (min) of the X-axis of the activation function f(x) can be determined as the starting point, or the point of the maximum value (max) of the X-axis can be determined as the starting point. However, the present disclosure is not limited to this, and the starting point can also be a specific point.
[0352] The PAF can be programmed to include, for example, multiple segment boundary values x1, x2, x3, x4, x5.
[0353] The PAF can be programmed to also include, for example, a minimum value (min) and a maximum value (max). In implementing the examples according to the present disclosure, the minimum value (min) and the maximum value (max) can be used to improve the programming efficiency of the activation function when implementing clipping. Values less than or equal to the minimum value can be output as minimum values. Values greater than or equal to the maximum value can be output as maximum values.
[0354] The activation function f(x) starts from the starting point, and for each interval (where the accumulated value of the second-order derivative of the activation function f(x) reaches the threshold E Th (i.e., the integral threshold of the segment approximation error)) for segmentation.
[0355] For example, the activation function conversion program unit 3000 may be When w1 is determined, When w2 is determined, When w3 is determined, When w4 is determined, When determining w5, Determine w6 when. Specifically, different Es can also be set for each segment. Th Value. That is, multiple Es can be set according to the situation. Th Value, such as E Th1 And E Th2 Value.
[0356] In addition, the programmable activation function used in the artificial neural network operation can be configured to only process input values within a limited range. For example, the minimum value (min) of the X-axis as the input value of the programmable activation function can be negative six, and the maximum value (max) can be six. According to the above configuration, there is an effect that can reduce the data size of the programmed activation function. However, the present disclosure is not limited to this.
[0357] See Fig.13 , since the accumulated value of the second derivative of the activation function is the rate of change of the slope of the activation function, it can be determined that: (1) in the activation function f(x), the widths w2, w3, and w4 of the segments corresponding to the intervals with relatively large gradient change rates are determined to be relatively narrow, and (2) in the activation function f(x), the widths w1 and w6 of the segments including the parts that are linear functions without a rate of change of the slope are determined to be relatively wide.
[0358] Fig.14 And 15 Are diagrams showing the ELU activation function and the Hardswish activation function.
[0359] The ELU activation function f(x) is x when x > 0 and α(ex - 1) when x ≤ 0 (where α is a hyperparameter).
[0360] As Fig.14 Shown, the ELU activation function has a linear interval when the x value is zero or greater, and a non-linear interval when the x value is less than zero. That is, the ELU activation function has the characteristic of being divided into a linear interval and a non-linear interval.
[0361] The Hardswish activation function f(x) is 0 when x ≤ -3, x when x ≥ +3, and x*(x + 3) / 6 when -3 < x < +3.
[0362] As Fig.14 Shown, when the value of x is less than -3 or greater than 3, the Hardswish activation function has a linear interval, otherwise it has a non-linear interval. That is, the Hardswish activation function has the characteristic of being divided into a linear interval and a non-linear interval.
[0363] However, the present disclosure is not limited to the ELU activation function and the Hardswish activation function, and there are various activation functions with the characteristic of being divided into a linear interval and a non-linear interval.
[0364] In particular, in the field of artificial neural networks, various customized activation functions have been proposed, in which various linear and nonlinear functions are combined to improve the accuracy of artificial neural networks. In this case, the activation function programming method according to the example of the present disclosure will be more effective.
[0365] In the activation function programming method according to the present disclosure, the activation function conversion program unit 3000 can distinguish between the linear interval and the nonlinear interval of the activation function, and further, can distinguish between the basic linear interval and the nonlinear interval, so that the activation function can be selectively segmented into multiple segments. Therefore, the activation function programming method according to the present disclosure is efficient and minimizes the approximation error, especially in the programming for approximating the activation function with (basic) linear and nonlinear intervals, so that the running speed of the artificial neural network model processed in the NPU 1000 can be improved, the decrease in reasoning accuracy can be minimized, and the power consumption of the NPU 1000 can be reduced. In the activation function programming method according to the present disclosure, the activation function conversion program unit 3000 can generate programmable parameters for at least one segment. The NPU 1000 can process at least one programmed activation function based on the above information. The NPU 1000 can receive the information and process at least one programmed activation function.
[0366] The coordinates of the start and end points of the intervals of the plurality of segments may be defined as segment boundary values. That is, each segment may be displayed as a segment boundary value. That is, according to the activation function programming method of the present disclosure, the programmable parameter may include a segment boundary value. In various examples, the activation function programming method according to the present disclosure may further include approximating at least one of the plurality of segments using a predetermined lookup table, a nonlinear approximation equation, etc.
[0367] In the activation function programming method according to the present disclosure, a plurality of segments are segmented using segment data, and since the segmented plurality of segments can be selectively approximated with programmable segments, there may be a segment that is determined not to be approximated with the PAF. If a stored lookup table, a nonlinear approximation, etc. for the interval is available in hardware in a predetermined manner, the predetermined and stored lookup table, the nonlinear approximation, etc. can be used to approximate the interval.
[0368] In various examples, the activation function programming method according to the present disclosure may further include determining not to approximate at least one of the plurality of segments as a programmable segment. For example, it may be determined that a segment having a very complex shape or a segment with low importance in the DNN is not approximated as a programmable segment. These segments may be processed in another predetermined manner, or if the number of these segments is large, they may be combined and processed in another predetermined manner.
[0369] In various examples, the activation function programming method according to the present disclosure may process the programming method of each segment in an individual manner.
[0370] An example activation function programming method according to the present disclosure may include: selecting an activation function for artificial neural network operation, and converting the activation function into a programmable activation function. Fig.13 As an example, the programmed activation function may include multiple segments with specific widths, and the specific width may be determined based on a specific threshold, i.e., determined for each segment where the accumulated value of the second-order derivative of the selected activation function reaches the threshold.
[0371] According to another example of the present disclosure, a device including a programmable activation function generator may be provided. The activation function conversion program may be configured to generate segment data for segmenting the activation function, segment the activation function into a plurality of segments using the generated segment data, and convert at least one of the plurality of segments into a programmable segment.
[0372] At least one of the plurality of segments may have a different width than the other segments.
[0373] The activation function conversion program may be configured to determine the number and width of the plurality of segments based on the segment data, and segment the activation function into the plurality of segments based on the determined number and width.
[0374] The segment data may include slope change data (eg, differential data) of the activation function.
[0375] The segment data may include information of hardware capable of processing the activation function. The activation function conversion program may be configured to receive the hardware information.
[0376] The activation function conversion program may be configured to determine a basic linear interval and a nonlinear interval of the activation function based on the slope change data of the activation function, and segment the activation function into a plurality of segments according to the determined basic linear interval and nonlinear interval.
[0377] The activation function conversion program searches for programmable parameters for approximating at least one segment to a programmable segment. The activation function conversion program may be configured to approximate at least one segment to a programmable segment according to the searched optimal programmable parameters.
[0378] The apparatus may further include a PAFE unit, and the PAFE unit may be configured to approximate the at least one segment using a predetermined nonlinear approximation equation.
[0379] An NPU configured to process an activation function programmed by an activation function programming method according to an example of the present disclosure will be described in detail below.
[0380] For ease of description, reference will be made to Figure 1An NPU of an apparatus for executing an activation function programming method according to an example of the present disclosure is described.
[0381] Fig.16 is a conceptual diagram illustrating a PAFE unit configured to process a programmed activation function according to an example of the present disclosure.
[0382] The PAFE unit 500 according to an example of the present disclosure is an example of a circuit configured to program an activation function as a linear function. The activation function programming method can be implemented by one of the various programming examples of the present disclosure described above. Hereinafter, the PAFE unit 500 may be referred to as the PAFE unit 500. The activation function conversion program unit 3000 may be configured to determine the type of programmable parameter based on the provided hardware information. For example, when the PAFE unit 500 includes only a linear function calculation circuit, the activation function conversion program unit 3000 may operate so that all programmable segments become linear functions. For example, when the PAFE unit 500 includes a linear function calculation circuit and a quadratic function calculation circuit, the activation function conversion program unit 3000 may operate so that all programmable segments become linear functions or quadratic functions.
[0383] The memory 300 may include a segment register 310, a first register 320, and a second register 330. For example, the at least one register may be implemented by setting an address of at least one memory or a register map. For example, the at least one register may be implemented by allocating a dedicated memory or at least one dedicated register. That is, the memory 300 of the PAFE unit 500 may be configured to store programmed activation function data.
[0384] The segment register 310 stores information on intervals of a plurality of segments.
[0385] Specifically, the coordinates of the start and end points of the X-axis of the interval of the plurality of segments determined by one of the methods proposed by the activation function conversion program unit 3000 may be stored in the segment register 310. The coordinates of the start and end points of the interval of the plurality of segments may be defined as a segment boundary value (SB). That is, the interval of the plurality of segments may be determined by the segment boundary values SB0 to SB(N-2).
[0386] For example, to define an interval of N segments, N-1 segment boundary values SB0 to SB(N-2) may be required.
[0387] For example, the first segment boundary value SB0 may be used to define an interval from negative infinity -∞ to the first segment boundary value SB0 based on the X-axis coordinate. In addition, the last segment boundary value SB(N-2) may be used to define an interval from the last segment boundary value SB(N-2) to positive infinity ∞ based on the X-axis coordinate. However, this is not limiting, and clipping may also be performed appropriately by setting the maximum and minimum values of an infinite range.
[0388] Then, the interval of N-1 segments between the first segment boundary value SB0 and the last segment boundary value SB(N-2) can be defined by using the segment boundary values SB1, SB2, ... between the first segment boundary value SB0 and the last segment boundary value SB(N-2). In addition, the segment register 310 provides the PAFE unit 500 with the plurality of segment boundary values SB0 to SB(N-2). Therefore, the PAFE unit 500 can obtain information about the interval of the plurality of segments.
[0389] PAFE unit 500 may be configured to receive data from segment register 310 .
[0390] That is, the interval of the segment of the programmed activation function may be set in the PAFE unit 500 .
[0391] In case of a first order polynomial, the first register 320 may be configured to store gradients A0 to A(N-1) of a plurality of programmable segments.
[0392] For example, in the case of a first order polynomial, the first register 320 may be used as a gradient register.
[0393] In other words, the first register 320 may be set to store a specific value, such as a gradient, according to a programming method.
[0394] For a first order polynomial, the second register 330 may be configured to store a plurality of programmable segment offsets B0 to B(N-1).
[0395] For example, in the case of a first order polynomial, the second register 330 may be used as an offset register.
[0396] In other words, the second register 330 may be set to store a specific value, such as an offset, according to a programming method.
[0397] Specifically, the N-segment interval can be approximated into N programmable segments by the activation function conversion program unit 3000. In addition, each programmable segment includes a specific gradient A and a specific offset B value. That is, a specific register of the memory 300 can selectively store a specific value.
[0398] In other words, in the example approximated by a linear function, in the interval from the minimum value to the first segment boundary value SB0, the gradient of the programmable segment can be represented as the first gradient A0, and the offset of the programmable segment is represented as the first offset B0. Here, the minimum value Min can be negative infinity -∞.
[0399] In the interval between the last segment boundary value SB(N-2) and the maximum value, the gradient of the programmable segment can be expressed as the last slope A(N-1), and the offset of the programmable segment can be expressed as the last offset B(N-1). Here, the maximum value Max can be positive infinity ∞.
[0400] Therefore, the first register 320 may store the gradient A0 to A(N-1) of each of the N programmable segments. In addition, the second register 330 may store the offset B0 to B(N-1) of each of the N programmable segments.
[0401] The activation function conversion program unit 3000 may be configured to provide programmed activation function data to be processed by the NPU to the memory 300 .
[0402]
[0403] Referring to , data for driving a programmable activation function may be configured to be generated in the activation function conversion program unit 3000 and stored in the memory 300 of the NPU, such as the segment register 310 , the first register 320 , and the second register 330 .
[0404] For example, the segment register 310 may be configured to store the segment boundary value SB of .
[0405] For example, the first register 320 may be configured to store the gradient A of . The gradient A may be referred to as a coefficient of a linear term.
[0406] For example, the second register 330 may be configured to store the offset B of . The offset B may be referred to as a bias.
[0407] The controller 100 and / or the DMA 200 may instruct the memory 300 to store the data of the programmed activation function of . However, the examples of the present disclosure are not limited thereto, and the data of the programmed activation function may be configured to be stored in at least one of a register inside the controller 100, a register inside the PAFE unit 500, a separate memory, and a separate register. That is, the storage location of the data of the programmed activation function is not limited to a specific location.
[0408] Referring to , an example of programmed activation function data is disclosed.
[0409] For example, the programmed activation function data may be configured to include a segment boundary value SB.
[0410] For example, the programmed activation function data may be configured to include an interval for each segment S.
[0411] For example, the programmed activation function data may include a gradient A for each segment S.
[0412] For example, the programmed activation function data may include an offset B for each segment S.
[0413] In addition, under the control of the controller 100, the first register 320 may output the gradient A0 to A(N-1) of each of the N programmable segments to the PAFE unit 500. In addition, under the control of the controller 100, the second register 330 may output the offset B0 to B(N-1) of each of the N programmable segments to the PAFE unit 500.
[0414] Therefore, the PAFE unit 500 may receive the gradients A0 to A(N−1) and the offsets B0 to B(N−1) of each programmable segment. That is, the PAFE unit 500 may receive information about a plurality of programmable segments through the first register 320 and the second register 330 .
[0415]
[0416] Referring to , data for driving the programmable ReLU may be configured to be generated in the activation function converter unit 3000 and stored in the memory 300 of the NPU, such as the segment register 310 , the first register 320 , and the second register 330 .
[0417] For example, the segment register 310 may be configured to store the segment boundary value SB of .
[0418] For example, the first register 320 may be configured to store the gradient A of .
[0419] For example, the second register 330 may be configured to store the offset B of .
[0420] In the case of programmed ReLU, it can be programmed to have only one segment boundary value SB. As described above, according to various examples of the present disclosure, determining to have only one segment boundary value SB can be performed by an approximation method.
[0421] In the case of programmed ReLU, since only the first segment boundary value SB1 is programmed, the operation of the PAFE unit 300 may require only one comparator. Therefore, unnecessary comparators may be disabled.
[0422] Since the comparator enable (En) signal of is input to the PAFE unit 500, unnecessary comparator power consumption can be reduced.
[0423]
[0424] Referring to , data for driving the programmed ReLU to which clipping is applied may be configured to be generated in the activation function converter unit 3000 and stored in the memory 300 of the NPU, for example, in the segment register 310 , the first register 320 , and the second register 330 .
[0425] For example, the segment register 310 may be configured to store the segment boundary value SB of .
[0426] For example, the first register 320 may be configured to store the gradient A of .
[0427] For example, the second register 330 may be configured to store the offset B of . When clipping is applied, the minimum and maximum values of the input values of the activation function may be limited.
[0428] In addition, in the PAFE unit 500, data for driving the programmed ReLU of and data for driving the programmed ReLU with clipping of can be stored in the NPU 1000. In addition, the activation function converter unit 3000 can be configured to provide the NPU 1000 with data for driving the programmed ReLU and data for driving the programmed ReLU with clipping.
[0429] The NPU 1000 may be configured to selectively input a plurality of programmed activation functions stored in the NPU 1000 to the PAFE unit 500 according to the compiled DNN information.
[0430] For example, the NPU 1000 may use the programmed activation function data of for a first artificial neural network operation, and may control the PAFE unit 500 to use the programmed activation function data of for a second artificial neural network operation.
[0431]
[0432] Referring to , data for driving a program of the NPU 1000 may be generated in the activation function conversion program unit 3000 and stored in the memory 300 of the NPU, such as the segment register 310 , the first register 320 , and the second register 330 .
[0433] For example, the segment register 310 may be configured to store the segment boundary value SB of .
[0434] For example, the first register 320 may be configured to store the slope A of .
[0435] For example, the second register 330 may be configured to store the offset B of .
[0436] In the case of a program, it can be programmed to have two segment boundary values SB. As described above, the judgment of having two segment boundary values SB can be performed by an approximation method according to various examples of the present disclosure.
[0437] In addition, the PAFE unit 500 may store all the data for driving the programmed ReLU in , the data for driving the programmed ReLU with clipping in , and the data for driving the programmed ReLU6 in in the NPU 1000. In addition, the activation function converter unit 3000 may be configured to provide all the data for driving the programmed ReLU, the programmed ReLU with clipping, and the programmed ReLU6 to the NPU 1000.
[0438] The NPU 1000 may be configured to selectively input a plurality of programmed activation functions stored in the NPU 1000 according to the compiled DNN information.
[0439] For example, the NPU 1000 may control the PAFE unit 500 to use the data from the programmed activation function of to perform the first artificial neural network operation, use the data from the programmed activation function of to perform the subsequent second artificial neural network operation, and use the data from the programmed activation function of to perform the subsequent third artificial neural network operation. In the case of programmed ReLU6, only the first segment boundary value SB1 and the second segment boundary value SB2 are programmed, and the operation of the PAFE unit 300 may only require two comparators. Therefore, unnecessary comparators can be disabled.
[0440] In summary, the NPU 1000 can store multiple programmed activation functions. The NPU 1000 can selectively input data of a specific activation function into the PAFE unit 500 to process a specific artificial neural network operation. In addition, the PAFE unit 500 can input data from the programmed activation function in real time without changing the hardware to process the artificial neural network operation.
[0441] Fig.17 is a conceptual diagram illustrating a PAFE unit of an NPU configured as a device for processing a programmed activation function according to an example of the present disclosure.
[0442] The exemplary PAFE unit 500 configured to process a programmed activation function having a linear function may be configured to include a plurality of comparators (comparator 0 to comparator (N-2)) and (510 to 51 (n-2)), a selector 520, a multiplier 530, and an adder 540. However, the examples of the present disclosure are not limited thereto, and the region of each segment may be distinguished by configuring the circuit in various ways. In addition, the PAFE unit 500 may be modified to further include additional circuit configurations to process activation functions other than linear functions with other programming methods.
[0443] In the example of the present disclosure, since the PAFE unit 500 is an example configured to process a main function, the PAFE unit 500 may be configured to process a linear function through inputs of the segment register 310, the first register 320, and the second register 330. However, the PAFE unit 500 may be modified to also include additional registers to process various approximate functions.
[0444] Each of the plurality of comparators 510 to 51(N-2) compares the input value X calculated in at least one processing element 400 with each of the plurality of segment boundary values SB0 to SB(N-2), respectively.
[0445] For example, if the input value X is greater than each of the segment boundary values SB0 to SB(N-2), each of the plurality of comparators 510 to 51(n-2) may output an output value of a first level. On the other hand, if the input value X is less than or equal to each of the segment boundary values SB0 to SB(N-2), each of the plurality of comparators 510 to 51(n-2) may output an output value of a second level.
[0446] The first level may represent a high level, and the second level may represent a low level. Alternatively, the first level may represent a low level, and the second level may represent a high level.
[0447] Therefore, the interval of the segment to which the input value X in the interval of the plurality of segments belongs can be determined by the output value output from each of the plurality of comparators 510 to 51(n-2). The output value output from each of the plurality of comparators 510 to 51(n-2) described above can be referred to as interval determination data (SDD).
[0448] For example, if the first segment boundary value SB0 is -4, the first segment boundary value SB0 is input to the first comparator 510. In the first comparator 510, the input value X calculated in the processing element is input.
[0449] For example, if the second segment boundary value SB1 is -2, the second segment boundary value SB1 is input to the second comparator 511. In the second comparator 511, the input value X calculated in the processing element is input.
[0450] In other words, the input value X calculated in the processing element can be input to multiple comparators simultaneously.
[0451] For example, when the first segment boundary value SB0 is -4, the second segment boundary value SB1 is -2, and the input value X is -3, the first interval determination data SDD1 and the output value of the first comparator (comparator 0 and 510) are output to the first stage, and the plurality of interval determination data SDD1 to SDD(N-2) except the first interval determination data SDD1 which is the output value of the remaining comparators (comparator 1 to comparator (N-2)) can be output to the second stage. Therefore, the input value X can determine that the segment boundary value SB corresponds to the segment between -4 and -2 through the interval determination data SDD, the output values outputted from each of the plurality of comparators 510 to 51(n-2).
[0452] The section determination data SDD1 to SDD(N-2) may correspond to the segment S described in Tables 1 to 4 above.
[0453] describes the determination of the segment S of the programmed activation function according to the results of the section determination data SDD1 to SDD(N-2).
[0454] scope SDD0 SDD1 SDD2 ... SDD(N-2) Segment (S0) min<X≤SB0 L L L ... L Segment (S1) SB0<X≤SB1 H L L ... L Segment (S2) SB1<X≤SB2 H H L ... L Segment (S(N-1)) SB(N-2)<X≤max H H H ... H
[0455] Referring to , the segment S illustrated in or may be determined according to the output of the interval determination data SDD0, SDD1, SDD2, and SDD(N-2). When determining a specific segment S, a corresponding gradient A and offset B may be selected. However, the examples of the present disclosure are not limited thereto, and the corresponding segment may also be determined by configuring the circuit for determining the segment in various ways. In addition, the PAFE unit 500 may be modified by configuring the circuit to process the activation function in another way other than the comparator.
[0456] On the other hand, the operation state of each of the plurality of comparators 510 to 51 (n-2) may be determined according to each of the enable signals Comp En1 to Comp En(N-2).
[0457] That is, if each of the plurality of enable signals Comp En1 to Comp En(n-2) is at the first level, each of the plurality of comparators 510 to 51(n-2) may be operated to compare the input value X with the segment boundary values SB0 to SB(N-2). On the contrary, if each of the plurality of enable signals Comp En1 to Comp En(n-2) is at the second level, each of the plurality of comparators 510 to 51(n-2) may be operated not to compare the input value X with the segment boundary values SB0 to SB(N-2). That is, each comparator may be disabled.
[0458] As described above, the number of segment boundary values SB0 to SB(N-2) is determined according to the number of segments of the programmed activation function. For example, when the number of segments is N, the number of segment boundary values SB0 to SB(N-2) is N-1.
[0459] For example, even if the activation function conversion program unit 3000 programs the same activation function, the first programmed activation function may be programmed to have ten segments, and the second programmed activation function may be programmed to have five segments. Therefore, even if the activation function is the same, the PAFE unit 500 may control the number of comparators activated in the PAFE unit 500 differently according to each programmed activation function data. Therefore, the accuracy of the artificial neural network calculation and the power consumption of the NPU 1000 may also vary according to the programming. That is, even if the same activation function is used, a high-performance activation function calculation function or a low-power activation function calculation function may be provided according to user requirements.
[0460] Meanwhile, according to the maximum number of segment boundary values SB, the number of the plurality of comparators using the segment boundary value SB as an input should also vary.
[0461] For example, when the maximum number of segment boundary values SB is ten, at least eleven or more comparators may be provided. That is, the minimum number of comparators may be the maximum number of segment boundary values.
[0462] Therefore, each of the plurality of comparators 510 to 51(N-2) can determine whether to operate based on each of the plurality of comparator enable signals CompEn1 to Comp En(N-2). Therefore, the power consumption of the NPU can be reduced by controlling unnecessary comparator operations according to the number of segments.
[0463] However, due to hardware limitations, the number of comparators may be limited. Therefore, the number of segments used to segment the activation function may be limited according to the number of comparators of the PAFE unit 500. That is, the activation function may be segmented into the maximum number of segments to be processed that can be processed by the NPU 1000 or the number of segments corresponding to the allocated resources of the NPU 1000.
[0464] At the same time, according to the programming method of the example of the present disclosure, the linear interval and the nonlinear interval of the activation function can be distinguished, and the number of segments can be minimized by providing a variable segment width while minimizing the error value. Therefore, there is an advantage that the number of hardware gates of the PAFE unit 500 of the NPU 1000 can be minimized by minimizing the number of comparators.
[0465] In addition, the activation function programming method according to the example of the present disclosure can be configured to program a specific activation function based on information of the maximum comparator that can be provided.
[0466] Then, the selector 520 outputs the gradient A of the programmable segment corresponding to the segment to which the input value X belongs among the plurality of gradients A0 to A(N-1) of the plurality of programmable segments according to the segment determination data SDD0 to SDD(N-2).
[0467] Specifically, the first register 320 provides the selector 520 with a plurality of gradients A0 to A(N-1) for each of the plurality of programmable segments. Then, the selector 520 may determine the interval of the segment to which the input value X belongs among the intervals of the plurality of segments according to the interval determination data SDD0 to SDD(N-2) output from each of the plurality of comparators 510 to 51(N-2). In addition, the selector 520 may output the gradient A of the programmable segment corresponding to the interval of the determined segment among the plurality of gradients A0 to A(N-1) of the plurality of programmable segments.
[0468] The selector 520 outputs an offset B of a programmable segment corresponding to the segment to which the input value X belongs among a plurality of offsets B0 to B(N-1) of a plurality of programmable segments according to the segment determination data SDD0 to SDD(N-2).
[0469] Specifically, the second register 330 provides the selector 520 with a plurality of offsets B0 to B(N-1) for each of the plurality of programmable segments. In addition, the selector 520 may determine the interval of the segment to which the input value X belongs among the intervals of the plurality of segments according to the interval determination data SDD0 to SDD(N-2) output from each of the plurality of comparators 510 to 51(N-2). Then, the selector 520 may output the offset B of the programmable segment corresponding to the interval of the determined segment among the plurality of offsets B0 to B(N-1) of the plurality of programmable segments.
[0470] Therefore, the selector 520 may output the gradient A and the offset B of the programmable segment corresponding to the interval of the segment to which the input value X belongs.
[0471] Meanwhile, the selector 520 may be a multiplexer composed of a plurality of switch elements controlled according to the section determination data SDD0 to SDD(N-2), but the configuration of the selector 520 may be variously changed.
[0472] The programmed activation function calculation unit of the PAFE unit 500 may refer to a circuit unit configured to receive an input value X, a gradient A, and an offset B and calculate an output value Y.
[0473] The programmed activation function calculator of the PAFE unit 500 may include at least one multiplier 530 and an adder 540 .
[0474] The programmed activation function calculator of the PAFE unit 500 may be a hard-wired circuit.
[0475] The multiplier 530 of the programmed activation function operator multiplies the input value X by the gradient A of the programmable segment corresponding to the interval of the segment to which the input value X belongs.
[0476] Specifically, the multiplier 530 multiplies the input value X calculated in at least one processing element 400 by the gradient A of the programmable segment output from the selector 520. That is, the input value X may be a calculated value of at least one processing element 400. However, the present disclosure is not limited thereto.
[0477] Therefore, the multiplier 530 may multiply the input value X by the gradient A of the programmable segment and output the result. That is, the output of the multiplier 530 may be expressed as A×X.
[0478] Then, the adder 540 of the programmable activation function operator adds the offset B of the programmable segment corresponding to the interval of the segment to which the input value X belongs to the output value of the multiplier 530 of the programmable activation function operator.
[0479] Specifically, the adder 540 adds the offset B of the programmable segment to a value obtained by multiplying the input value X by the gradient A of the programmable segment. That is, the output of the adder 540 may be expressed as A×X+B.
[0480] Therefore, the adder 540 may output the activation value to which the PAF is applied to the input value X of the calculation value.
[0481] That is, the PAFE unit 500 according to an example of the present disclosure may be a circuit configuration configured to implement an activation function programmed as a linear function.
[0482] For example, the PAFE unit 500 pipeline-connected with at least one processing element 400 according to an example of the present disclosure may also be configured as a hard-wired circuit configured to implement an activation function programmed as a linear function.
[0483] As described above, the PAFE unit 500 of the NPU of the device for executing the activation function programming method according to the example of the present disclosure is composed of only a plurality of comparators 511 to 51(N-2), a selector 520, a multiplier 530 and an adder 540, and all activation functions can be programmed and applied to the input value X.
[0484] Since each of the above-mentioned multiple comparators 511 to 51(N-2), selector 520, multiplier 530 and adder 540 is a relatively simplified hardware, the device for executing the activation function programming method according to the example of the contents of the present disclosure has the effect of processing all activation functions with only simplified hardware.
[0485] Meanwhile, conventional activation function processing devices can only process predefined activation functions. However, the apparatus for executing the activation function programming method according to the example of the present disclosure can program and apply unpredefined activation functions, so that all programmed activation functions can be applied. Specifically, since the PAFE unit 500 can adjust the number of segments according to the characteristics of various activation functions, the approximation error can be minimized using a minimum number of comparators. Specifically, since the PAFE unit 500 can adjust the width of each segment according to the characteristics of various activation functions, the approximation error can be minimized by using a minimum number of comparators. Specifically, since the PAFE unit 500 can adjust the width and number of segments according to the characteristics of various activation functions, the approximation error can be minimized by using a minimum number of comparators.
[0486] The NPU of an apparatus for executing an activation function programming method according to another example of the present disclosure will be described in detail below.
[0487] Since the NPU of the device for executing the activation function programming method according to the example of the present disclosure differs from the NPU of the device for executing the activation function programming method according to another example of the present disclosure only in the technical characteristics of the PAFE unit, this will be mainly described.
[0488] Fig.18 A conceptual diagram of an NPU of an apparatus for processing a programmed activation function is provided according to another example of the present disclosure.
[0489] Fig.19 A conceptual diagram of a PAF unit of an NPU of an apparatus for processing a programmed activation function is provided according to another example of the present disclosure.
[0490] The PAF units 500-1 to 500-N of the NPU of the device processing the programmed activation function may be divided into a plurality. Specifically, the PAF unit may include a first PAFE unit 500-1 to an Nth PAF unit 500-N. In addition, each of the first PAFE unit 500-1 to the Nth PAF unit 500-N may process a different activation function or the same activation function. That is, the activation function programmed in each of the first PAFE unit 500-1 to the Nth PAF unit 500-N may be the same or different from each other.
[0491] From the perspective of the number of processing elements 400 , the amount of data to be processed by the PAFE units 500 - 1 to 500 -N may increase, and thus, the number of the PAFE units 500 - 1 to 500 -N may be determined in consideration of the number of processing elements 400 .
[0492] That is, if the maximum data bandwidth of the processing element 400 corresponding to the input value X (i.e., the output value of the processing element 400) is greater than the maximum data bandwidth that the PAFE unit 500 can process, the number of PAFE units 500-1 to 500-N may increase. Therefore, the bottleneck of insufficient data bandwidth of the PAFE units 500-1 to 500-N can be solved.
[0493] For example, Fig.19 As shown, the PAFE unit 500 may include a demultiplexer (DEMUX) and a multiplexer (MUX) as well as a plurality of PAFE units.
[0494] The demultiplexer (DEMUX) distinguishes input values X to which a non-linear PAF should be applied and input values X to which a linear PAF should be applied.
[0495] An input value that should be applied to the nonlinear PAF is allocated to the first PAFE unit 500-1. In addition, an input value that should be applied to the linear PAF may be allocated to the second PAFE unit 500-2.
[0496] In addition, the first PAFE unit 500-1 stores a programmed activation function of a non-linear activation function. Therefore, the first PAFE unit 500-1 can process a non-linear PAF.
[0497] In addition, the second PAFE unit 500-2 stores a programmed activation function of a linear activation function. Therefore, the second PAFE unit 500-2 can process a nonlinear PAF.
[0498] In addition, since the first PAFE unit 500-1 can be configured to process a nonlinear activation function, it can be configured to have relatively more comparators than the second PAFE unit 500-2. On the other hand, since the second PAFE unit 500-2 can be configured to have a relatively smaller number of comparators than the first PAFE unit 500-1, it can operate with relatively less power consumption.
[0499] One of the first PAFE unit 500 - 1 and the second PAFE unit 500 - 2 may be selectively disabled according to the type of programmed activation function processed by the NPU 1000 .
[0500] Furthermore, the multiplexer MUX may receive an output value with the nonlinear PAF from the first PAFE unit 500 - 1 and receive an output value with the linear PAF from the second PAFE unit 500 - 2 .
[0501] In addition, the multiplexer MUX may collect and output a non-linear PAF applied output from the first PAFE unit 500 - 1 and a linear PAF applied output from the second PAFE unit 500 - 2 .
[0502] Therefore, the multiplexer MUX can output activation values regarding the linear PAF and the nonlinear PAF to the calculation value as the input value X.
[0503] According to an example of the present disclosure, the first PAFE unit 500 - 1 and the second PAFE unit 500 - 2 may be configured to process specific intervals of the activation function, respectively, to process the activation function having linear and nonlinear intervals.
[0504] For example, Fig.14 The ELU activation function shown has a linear interval when the X value is zero or greater, and has a nonlinear interval when the X value is less than zero. That is, the ELU activation function has the characteristics of a linear interval and a nonlinear interval. Here, the first PAFE unit 500-1 can be configured to process the nonlinear interval of the ELU activation function. The second PAFE unit 500-2 can be configured to process the linear interval of the ELU activation function.
[0505] The NPU of an apparatus for executing an activation function programming method according to another example of the present disclosure will be described in detail below.
[0506] Since the NPU of the device for executing the activation function programming method according to the example of the present disclosure differs from the NPU of the device for executing the activation function programming method according to another example of the present disclosure only in the technical features of the PAF library 600, this will be mainly described.
[0507] Fig. 20 is a conceptual diagram of an NPU illustrating an apparatus for processing a programmed activation function according to another example of the present disclosure.
[0508] The NPU may further include a controller 100 , a memory 300 , at least one processing element 400 , a PAFE unit 500 , and a PAF library 600 .
[0509] The PAF library 600 can store PAFs that approximate activation functions. Specifically, the PAF library 600 can store the gradient A0 to A(N-1) and offset B0 to B(N-1) information of multiple programmable segments that make up the PAF. As an explanation, the PAF library 600 can store multiple PAFs. In addition, the PAF library 600 can store the gradient A0 to A(N-1) and offset B0 to B(N-1) information of multiple programmable segments of each of the multiple PAFs. However, through the activation function conversion program, multiple PAFs are not limited to linear functions and can be approximated by selectively combining second-order polynomials, third-order polynomials, logarithmic functions, etc. For example, the PAF library 600 can be configured to store each programmed activation function data shown in Tables 2 to 4. Therefore, the PAF library 600 can be configured to store programmed ReLU, programmed ReLU with clipping, and programmed ReLU6. In addition, as needed, the controller 100 can be controlled to select a specific activation function from the PAF library 600 and input it into the PAFE unit 500.
[0510] The plurality of programmed activation functions stored in the PAF library 600 may approximate representative activation functions. For example, the representative activation function may be a Swish function, a Mish function, a sigmoid function, a hyperbolic tangent (TANH) function, a SELU function, a GELU (Gaussian error linear unit) function, a SOFTPLUS function, a ReLU function, a Leaky ReLU function, a Maxout function, an ELU function, etc.
[0511] Therefore, the PAFE unit 500 can select a desired PAF from a plurality of PAFs stored in the PAF library 600 according to the control of the controller 100. In addition, the PAFE unit 500 can also import information such as gradients A0 to A(N-1) and offsets B0 to B(N-1) from a plurality of programmable segments for the selected PAF from the PAF library 600.
[0512] As described above, the apparatus for performing the activation function programming method according to another example of the present disclosure may program commonly used activation functions and store them in the PAF library 600 .
[0513] Therefore, in an apparatus for performing an activation function programming method according to another example of the present disclosure, the PAF library 600 may store PAFs without an activation function conversion program to program all activation functions.
[0514] Therefore, there are advantages in that the processing speed of the device for executing the activation function programming method according to another example of the present disclosure can be improved and the power consumption of driving the activation function conversion program can be reduced.
[0515] The NPU of an apparatus for executing an activation function programming method according to another example of the present disclosure will be described in detail below.
[0516] Since the NPU of the device for executing the activation function programming method according to the example of the present disclosure differs from the NPU of the device for executing the activation function programming method according to another example of the present disclosure only in at least one processing element (PE array) and PAFE unit, this point will be mainly described.
[0517] Fig.21 is a conceptual diagram of an NPU illustrating an apparatus for processing a programmed activation function according to another example of the present disclosure.
[0518] like Fig.21 As shown, in the NPU of the apparatus for executing the activation function programming method according to another example of the present disclosure, a plurality of processing elements PROCESSING ELEMENTS#0 to PROCESSING ELEMENTS#N-1 may be grouped. The grouped processing elements may be referred to as at least one processing element.
[0519] In other words, the plurality of processing elements may include the zeroth processing element PROCESSING ELEMENTS#0 to the N-1th processing element PROCESSING ELEMENTS#N-1. Each of the plurality of processing elements PROCESSING ELEMENTS#0 to PROCESSING ELEMENTS#N-1 may be referred to as a PE thread (Processing Element Thread) or a PE core (PE Core). Hereinafter, at least one of the plurality of processing elements is referred to as a PE core.
[0520] On the other hand, the structure of each of the plurality of PE cores may be different from each other. For example, each of the plurality of PE cores may be one of an input fixed type, a weight fixed type, and an output fixed type.
[0521] In addition, according to the optimization of driving, each of the multiple PE cores can be driven separately. That is, each of the multiple PE cores is not driven at the same time, but can be driven in sequence according to the operation of the PAFE unit.
[0522] In addition, the number of processing elements, multiplication and accumulation (MAC) operators, and arithmetic logic unit (ALU) operators included in each of the plurality of PE cores may be different. Therefore, the size of each of the plurality of PE cores may be different.
[0523] In addition, each of the plurality of PE cores may be connected to the PAFE unit through a multiplexer (MUX). Specifically, the multiplexer (MUX) receives a plurality of calculation values output from each of the plurality of PE cores and outputs at least one of the plurality of calculation values to the PAFE unit.
[0524] It is also possible to configure a buffer memory to be provided between the PAFE unit 500 and the plurality of PE cores. However, this is not limited thereto.
[0525] Therefore, one PAFE unit can process multiple calculation values output from each of the multiple PE cores. Therefore, the number of PAFE units provided in the apparatus for performing the activation function programming method according to another example can be minimized. Ultimately, this can minimize the manufacturing cost of the apparatus for performing the activation function programming method.
[0526] Fig. 22 is a conceptual diagram illustrating a PAFE unit configured to process a programmed activation function according to another example of the present disclosure.
[0527] Fig.23 is a conceptual diagram illustrating a PAFE unit of an NPU of an apparatus for processing a programmed activation function according to another example of the present disclosure.
[0528] Figures 22 to 23 Each of the plurality of programmable segments of the PAF applied to the PAFE unit shown in FIG. 1 can operate as a linear or quadratic function. Therefore, the coefficients A, B, and C of the programmable segments may include a quadratic term coefficient A, a linear term coefficient B, and an offset C.
[0529] Thus, the activation function conversion program unit 3000 may be configured to provide programmed activation function data for processing in the PU and the memory 300 .
[0530]
[0531] Referring to , data for driving the programmable activation function may be generated in the activation function converter unit 3000 and configured to be stored in the memory 300 of the NPU, such as the segment register 310, the first register 320, the second register 330, and the third register 340.
[0532] For example, the segment register 310 may be configured to store the segment boundary value SB of .
[0533] For example, the first register 320 may be configured to store the coefficient of the quadratic term A of .
[0534] For example, the second register 330 may be configured to store the coefficient of the linear term B of . For example, the third register 340 may be configured to store the offset C of .
[0535] The controller 100 and / or the DMA 200 may instruct that the data of the programmed activation function in be stored in the memory 300. The examples of the present disclosure are not limited thereto, and the data of the programmed activation function may be configured to be stored in at least one of a register in the controller 100, a register in the PAFE unit 500', a separate memory, and a separate register. That is, the storage location of the data of the programmed activation function is not limited to a specific location.
[0536] See , which discloses an example of programmed activation function data.
[0537] For example, the programmed activation function data may be configured to include a segment boundary value SB.
[0538] For example, for each segment, the programmed activation function data may be configured to include a series of segments S.
[0539] For example, the programmed activation function data may be configured to include a quadratic term coefficient A and a linear term coefficient B for each segment.
[0540] For example, the programmed activation function data may be configured to include an offset C for each segment.
[0541] An exemplary PAFE unit configured to process a programmed activation function of a quadratic term can be configured to include a plurality of comparators: comparator 0 to comparator (N-2) or 511 to 51(N-2), a selector 520, a plurality of multipliers 531, 532 and 533, and a plurality of adders 541 and 542.
[0542] Each of the plurality of comparators 510 to 51 (N-2) compares the input value X calculated in at least one processing element 400 with each of the plurality of segment boundary values SB0 to SB (N-2). For example, when the input value X is greater than each of the plurality of segment boundary values SB0 to SB (N-2), each of the plurality of comparators 510 to 51 (N-2) may output a first level output value. Conversely, when the input value X is less than or equal to each of the plurality of segment boundary values SB0 to SB (N-2), each of the plurality of comparators 510 to 51 (N-2) may output a second level output value.
[0543] Therefore, the interval of the segment to which the input value X belongs may be determined among the intervals of the plurality of segments by the output value output from each of the plurality of comparators 510 to 51(N-2).
[0544] Meanwhile, the operation of each of the plurality of comparators 510 to 51(N-2) may be determined by each of the plurality of comparator enable signals Comp En1 to Comp En(N-2).
[0545] In addition, based on the interval determination data SDD0 to SDD(N-2), the selector 520 outputs coefficients A, B, and C of the programmable segment corresponding to the interval of the segment to which the input value X belongs, among the coefficients of the plurality of programmable segments A0 to A(N-1), B0 to B(N-1), and C0 to C(N-1).
[0546] Specifically, the first register 320 provides the selector 520 with quadratic term coefficients A0 to A(N−1), linear term coefficients B0 to B(N−1), and offsets C0 to C(N−1) of each of the plurality of programmable segments.
[0547] Furthermore, the selector 520 may determine a section to which the input value X among the sections of the plurality of sections belongs according to the section determination data SSD0 to SSD(N-2) output from each of the plurality of comparators 510 to 51(N-2).
[0548] In addition, the selector 520 outputs the quadratic term coefficient A, linear term coefficient B and offset C of the programmable segment corresponding to the interval of the determined segment among the quadratic term coefficients A0 to A(N-1), linear term coefficients B0 to B(N-1) and offsets C0 to C(N-1) of the multiple programmable segments.
[0549] Therefore, the selector 520 may output the quadratic term coefficient A, the linear term coefficient B, and the offset C of the programmable segment corresponding to the interval of the segment to which the input value X belongs.
[0550] Meanwhile, the selector 520 may be a multiplexer composed of a plurality of switch elements controlled according to the section determination data SDD, but the configuration of the selector 520 may be variously changed.
[0551] The programmed activation function calculation unit of the PAFE unit 500 ′ may represent a circuit unit configured to receive an input value X, a quadratic term coefficient A, a linear term coefficient B, and an offset C as inputs and calculate an output value Y.
[0552] The programmed activation function calculator of the PAFE unit 500 ′ may be configured to include a plurality of multipliers 531 , 532 , and 533 and a plurality of adders 541 and 542 to process a quadratic function or a linear function.
[0553] The programmed activation function calculation unit of the PAFE unit 500 ′ may be a hard-wired circuit.
[0554] The plurality of multipliers of the programmed activation function calculator may include a first multiplier 531 , a second multiplier 532 , and a third multiplier 533 .
[0555] The first multiplier 531 multiplies the input value X by a coefficient of a quadratic term A of a programmable segment corresponding to the segment to which the input value X belongs.
[0556] Specifically, the first multiplier 531 multiplies the input value X calculated in at least one processing element 400 by the coefficient of the quadratic term A of the programmable segment output from the selector 520 .
[0557] Therefore, the first multiplier 531 may multiply the input value X by the coefficient of the quadratic term A of the programmable segment and output the result. That is, the output of the first multiplier 531 may be expressed as A×X.
[0558] Then, the second multiplier 532 multiplies the output value output from the first multiplier 531 by the input value X.
[0559] Specifically, the second multiplier 532 multiplies the input value X calculated by at least one processing element 400 by the output value output from the second multiplier 532 .
[0560] Therefore, the output of the second multiplier 532 can be expressed as A×X 2 However, the above configuration is only to achieve A×X 2 The examples can also be modified by various circuit combinations.
[0561] The third multiplier 533 multiplies the input value X by a coefficient of a linear term B of a programmable segment corresponding to an interval of a segment to which the input value X belongs.
[0562] Specifically, the third multiplier 533 multiplies the input value X calculated in at least one processing element 400 by the coefficient of the linear term B of the programmable segment output from the selector 520 .
[0563] Therefore, the third multiplier 533 may multiply the input value X by the coefficient of the linear term B of the programmable segment and output the result. That is, the output of the third multiplier 533 may be expressed as B×X.
[0564] The plurality of adders may include a first adder 541 and a second adder 542 .
[0565] The first adder 541 adds the output value of the third multiplier 533 and the output value of the second multiplier 532 .
[0566] Specifically, the first adder 541 may output the sum of the quadratic term and the linear term of each of the plurality of programmable segments consisting of quadratic terms. That is, the output of the first adder 541 may be expressed as A×X 2 +B×X.
[0567] Then, the second adder 542 adds the offset C of the programmable segment corresponding to the interval of the segment to which the input value X belongs to, to the output value of the first adder 541 .
[0568] Specifically, the adder 540 adds the offset C of the programmable segment to the sum of the quadratic term and the linear term of the programmable segment composed of the quadratic term. That is, the output of the second adder 542 can be expressed as A×X 2 +B×X+C.
[0569] Therefore, the adder 540 may output an activation value to which the activation function programmed as a quadratic function is applied to the input value X as an operation value.
[0570] According to the above configuration, the PAFE unit 500 ′ can process the operation of a second-order polynomial.
[0571] Meanwhile, operations of the second multiplier 532 , the third multiplier 533 , and the second adder 542 may be controlled by the first enable signal EN1 .
[0572] Specifically, when the second multiplier 532 , the third multiplier 533 , and the second adder 542 do not operate due to the first enable signal EN1 , the operations are as follows.
[0573] The first multiplier 531 multiplies the input value X by a coefficient of a quadratic term A of a programmable segment corresponding to an interval of a segment to which the input value X belongs.
[0574] Specifically, the first multiplier 531 multiplies the input value X calculated in at least one processing element 400 by the coefficient of the quadratic term A of the programmable segment output from the selector 520 .
[0575] Therefore, the first multiplier 531 may multiply the input value X by the coefficient of the quadratic term A of the programmable segment and output the result. That is, the output of the first multiplier 531 may be expressed as A×X.
[0576] In addition, the second multiplier 532 and the third multiplier 533 do not operate, and the output of the first multiplier 531 is input as is to the first adder 541. That is, the calculator disabled by the first enable signal EN1 can be bypassed.
[0577] Then, the first adder 541 adds the coefficient of the linear term B of the programmable segment corresponding to the interval of the segment to which the input value X belongs to, to the output value of the first multiplier 531 .
[0578] Specifically, the first adder 541 adds the coefficient of the linear term B of the programmable segment to a value obtained by multiplying the input value X by the coefficient of the second-order term A of the programmable segment. That is, the output of the first adder 541 may be expressed as A×X+B.
[0579] Likewise, the second adder 542 does not operate, and the output of the first adder 541 is output as it is. That is, the calculator disabled by the first enable signal EN1 may be bypassed.
[0580] That is, the first adder 541 may output an activation value to which an activation function programmed as a linear function is applied to an operation value as an input value X.
[0581] According to the above configuration, the PAFE unit 500 ′ can process the operation of a first-order polynomial.
[0582] As described above, some components of the plurality of multipliers and the plurality of adders can be controlled by the first enable signal EN1. Therefore, according to the first enable signal EN1, the PAFE unit can be driven not only when each programmable segment is a second-order polynomial but also when each programmable segment is a first-order polynomial.
[0583] In other words, the at least one processing element 400 and the PAFE unit 500 ′ pipelined according to the example of the present disclosure may also be composed of hard-wired circuits configured to implement activation functions programmed as quadratic functions and linear functions.
[0584] Therefore, there is an advantage in that PAF can be processed in various situations with one PAFE unit.
[0585] Fig.24is a diagram showing an example in which a device for processing a programmed activation function approximates a S-shaped activation function as a programmable activation function according to another example of the present disclosure.
[0586] As described above, according to another example of the present disclosure, each of the plurality of programmable segments of the PAF applied in the PAFE unit of the device executing the activation function programming method is a second-order polynomial. In detail, at least a portion of the S-shaped function, for example, only the range of -6.0 to 2.0 can be approximated by dividing it into three segments.
[0587] For example, when the S-shaped activation function is approximated by PAF, it can be approximated as follows.
[0588] In the interval S0 where the input value X is greater than -6.0 or less than or equal to -2.6, the programmable segment can be approximately "0.07X 2 +0.08X+0.23. In addition, in the interval S1 where the input value X is greater than -2.6 or less than or equal to -0.6, the programmable segment can be approximately "0.05X 2 +0.3X+0.25. In addition, in the interval S2 where the input value X is greater than -0.6 or less than or equal to 2, the programmable segment can be approximately -0.03X 2 +0.26X+0.5.
[0589] Therefore, the programmable parameters can be mapped according to the format of .
[0590] For example, A0 in may be 0.07, B0 in may be 0.08, and C0 in may be 0.23.
[0591] For example, A1 in may be 0.05, B1 in may be 0.3, and C1 in may be 0.52.
[0592] For example, A2 in may be -0.03, B2 in may be 0.26, and C2 in may be 0.5.
[0593] For example, SB0 in may be -2.6. SB1 in may be -0.6.
[0594] For example, the Min in may be -6.0, and the Max in may be 2.0.
[0595] For example, according to Fig.12 In the activation function programming method of the example, each segment can also be approximated as an optimal programmable segment by using machine learning to obtain the segment boundary value SB, quadratic term coefficient A, linear term coefficient B and offset C of the segment.
[0596] Fig.24 The coefficients in are only examples derived through machine learning and can be modified in various ways. For example, some of the programmable segments S0 and S2 may correspond to linear intervals, while another part of the programmable segment S1 may correspond to nonlinear intervals.
[0597] Therefore, some of the programmable segments S0 and S2 can be approximated by a linear function, while another portion S1 of the programmable segment can be approximated by a quadratic function.
[0598] In some examples, the output of the PAFE unit may further include a logarithm operator. Fig.25 , the PAFE unit including the logarithm operator will be described in detail.
[0599] Fig.25 is a conceptual diagram showing a PAFE unit of an NPU of an apparatus for processing an activation function programmed according to another example of the present disclosure.
[0600] See also Fig.25 The PAFE unit 500″ may include a plurality of comparators: comparator 0 to comparator (N-2) or 511 to 51(N-2), a selector 520, a plurality of multipliers 531, 532, and 533, and a plurality of adders 541 and 542, and a logarithm operator 550.
[0601] because Fig.23 The PAFE unit shown is Fig.25 The PAFE units shown differ only in whether the logarithm operator 550 operates, so this will be described in detail.
[0602] The operation of the logarithm operator 550 may be controlled by the second enable signal EN2. When the second enable signal EN2 is applied to the logarithm operator 550, the logarithm coefficient D may be input to the logarithm operator 550. When the logarithm operator 550 is activated, operators 531, 532, 533, 541, and 542 associated with the coefficient of the second-order term A, the coefficient of the first-order term A, and the offset C may be deactivated.
[0603] That is, the output of the logarithm operator 550 can be expressed as logD.
[0604] That is, the logarithm operator 550 may output the activation value to which the PAF including the logarithm operation is applied to the input value X.
[0605] Fig.25Each of the plurality of programmable segments of the PAF applied in the PAFE unit shown in FIG. 1 can operate as a linear function, a quadratic function, or a logarithmic function. Therefore, the coefficients A, B, C, and D of the programmable segments may include the coefficient of the quadratic term A, the coefficient of the linear term A, the offset C, and the logarithm D.
[0606]
[0607] Referring to , data for driving the programmable activation function can be configured to be generated in the activation function conversion program unit 3000 and stored in the memory 300 of the NPU, such as the segment register 310, the first register 320, the second register 330, the third register 340 and the fourth register 350.
[0608] For example, the programmed activation function data may be configured to include a segment boundary value SB. The segment boundary value SB may be stored in a first register of the memory.
[0609] For example, the programmed activation function data may include a series of segments S for each segment.
[0610] For example, the programmed activation function data may include a quadratic coefficient A for each segment. The coefficient A of the quadratic term may be stored in a second register of the memory.
[0611] For example, the programmed activation function data may include coefficients of the linear term B for each segment. The coefficients of the linear term B may be stored in a third register of the memory.
[0612] For example, the programmed activation function data may include an offset C for each segment. The offset C may be stored in a fourth register of the memory.
[0613] For example, the programmed activation function data may include a logarithmic coefficient D for each segment. The logarithmic coefficient D may be stored in a fifth register of the memory.
[0614] As described above, the application of the PAF including the logarithmic operation by adding the logarithmic operator 550 to the PAFE unit has been described. However, as an operator added to the output terminal of the PAFE unit, not only the logarithmic operator 550 but also various types of operators may be added.
[0615] In other words, the programmed activation function data may be determined according to the operator circuit configuration and the supported equation of the programmed activation function calculator of the PAFE unit.
[0616] Fig.26 is a schematic flow chart illustrating a method of programming an activation function according to another example of the present disclosure.
[0617] Fig. 27 is a diagram illustrating an artificial neural network for approximating an activation function according to another example of the present disclosure.
[0618] See also Fig.26 , the activation function programming method includes the following steps: setting a target activation function S110, training an artificial neural network to approximate the target activation function to a programmed activation function S320, converting the programmed activation function into a slope and an offset and storing them in a lookup table S330.
[0619] In step S310, an activation function is set, i.e., a target activation function to be programmed. For example, the target activation function may be a Swish function, a Mish function, a S-shaped function, a hyperbolic tangent (tanh) function, a SELU function, a Gaussian error linear unit (GELU) function, a SOFTPLUS function, a square root (SQRT) function, and other nonlinear functions.
[0620] In step S320, the target activation function is approximated by using a programmed activation function through training of an artificial neural network.
[0621] See also Fig. 27 , an artificial neural network for approximating a target activation function may include two layers and a plurality of rectified linear unit (ReLU) functions between the two layers.
[0622] More specifically, the first layer includes a plurality of neurons. Each of the plurality of neurons in the first layer has a node in the input layer as an input, and has each of the plurality of nodes in the hidden layer as an output.
[0623] For example, the number of neurons in the first layer may be 15. Therefore, the number of nodes in the plurality of hidden layers may be 15. However, the number of neurons in the first layer and the number of nodes in the hidden layers may vary as desired.
[0624] In addition, the first layer may be a fully connected layer in which one node of the input layer as an input and multiple nodes of the hidden layer as outputs are fully connected. Therefore, each of the multiple neurons in the first layer may have a weight and a bias.
[0625] That is, the weight of each of the multiple neurons in the first layer can be represented by n1, n2, ..., n15, and the bias of each of the multiple neurons in the first layer can be represented by b1, b2, ..., b15.
[0626] Therefore, when an input x is input to the first layer, each of the multiple nodes in the hidden layer can output z i =n i*x+b i A rectified linear unit (ReLU) function may then be applied to the output of each of the plurality of neurons in the first layer.
[0627] Rectified Linear Unit (ReLU)(z) can be represented as max(0,z), which means that all negative values are converted to zero when the ReLU function is applied.
[0628] Therefore, the output value of the first layer to which the ReLU function is applied can be expressed as ReLU(n i *x+b i ).
[0629] The second layer also includes a plurality of neurons. Each of the plurality of neurons in the second layer has each of the plurality of nodes in the hidden layer as an input, and has a node in the output layer as an output.
[0630] For example, the number of neurons in the second layer may be 15. Therefore, the number of multiple nodes in the hidden layer may be 15. However, the number of neurons in the second layer and the number of nodes in the hidden layer may vary as desired.
[0631] In addition, the second layer may be a fully connected layer in which a plurality of nodes as inputs of the hidden layer and a node as output of the output layer are fully connected. Therefore, each of the plurality of neurons included in the second layer may have a weight.
[0632] That is, the weight of each of the multiple neurons included in the second layer can be expressed as m1, m2, ..., m 15 express.
[0633] Therefore, the second layer can be ReLU (n i *x+b i ) as input.
[0634] Therefore, the output of the second layer is the output of the first layer ReLU (n i *x+b i ) multiplied by the sum of the weights of the second layer.
[0635] Therefore, a node of the output layer as an output of the second layer can output an operation value according to [Formula 1].
[0636] [Formula 1]
[0637] By executing the above operations of the artificial neural network, the error between the approximate programmed function and the target activation function is calculated, and the training of the artificial neural network is repeatedly performed to minimize the error value. Through the above training process, the activation function conversion program unit can approximate the target activation function to the programmed activation function.
[0638] Finally, by calculating the breakpoints of the programmed activation function, the linear interval of the programmed activation function can be set. Each linear interval can then be segmented into a first-order function with a specific slope and a specific offset.
[0639] In step S320, the programmed activation function is converted into a slope and an offset and stored in a lookup table.
[0640] As described above, each programmed activation function can be segmented into a first-order function with a specific slope and a specific offset for each linear segment.
[0641] Therefore, the specific slope and the specific offset for each linear segment can be stored in a lookup table.
[0642] See also Figures 28 to 37 , the training process of an artificial neural network for approximating an activation function will be described below according to another example of the present disclosure.
[0643] Figures 28 to 31 Epoch 50 is shown. Figures 32 to 37 Epoch 300 is shown.
[0644] As used herein, epoch refers to the number of training cycles (training epochs) used to train an artificial neural network.
[0645] For example, Epoch 50 means that 50 neural network training runs have been performed to train the neural network. Also, Epoch 300 means that 300 neural network training runs have been performed to train the neural network.
[0646] Fig.28 is a diagram showing the output of the first layer of Epoch 50.
[0647] Fig.29 is a graph showing the values of the output of the first layer when the ReLU function is applied to Epoch 50.
[0648] Fig.30 is a diagram showing the output of the second layer of Epoch 50.
[0649] Fig.31 is a graph showing the error between the programmed activation function and the target activation function at Epoch 50.
[0650] First, refer to Fig.28 , for Epoch 50, the weight of each of the multiple neurons in the first layer can be expressed as n 1-50 , n 2-50 , ..., n 15-50 Indicates that the bias of each of the multiple neurons in the first layer can be expressed as b 1-50 , b 2-50 , ..., b 15-50 express.
[0651] Therefore, for Epoch 50, the first node among the multiple nodes in the hidden layer can output z 1-50 =n 1-50 *x+b 1-50 , that is, the output value of the first neuron in the first layer, Neuron 1 Output (Epoch 50).
[0652] For Epoch 50, the second node among the multiple nodes in the hidden layer can output z 2-50 =n 2-50 *x+b 2-50 , which is the output value of the second neuron in the first layer, Neuron 2 Output (Epoch 50).
[0653] For Epoch 50, the third node among the multiple nodes in the hidden layer can output z 3-50 =n 3-50 *x+b 3-50 , which is the output value of the third neuron in the first layer, Neuron 3 Output (Epoch 50).
[0654] For Epoch 50, the fourth node among the multiple nodes in the hidden layer can output z 4-50 =n 4-50 *x+b 4-50 , which is the output value of the fourth neuron in the first layer, Neuron 4 Output (Epoch 50).
[0655] For Epoch 50, the fifth node among the multiple nodes in the hidden layer can output z 5-50 =n 5-50 *x+b 5-50 , which is the output value of the fifth neuron in the first layer, Neuron 5 Output (Epoch 50).
[0656] For Epoch 50, the sixth node among the multiple nodes in the hidden layer can output z 6-50 =n 6-50 *x+b6-50 , which is the output value of the sixth neuron in the first layer, Neuron 6 Output (Epoch 50).
[0657] For Epoch 50, the seventh node among the multiple nodes in the hidden layer can output z 7-50 =n 7-50 *x+b 7-50 , which is the output value of the seventh neuron in the first layer, Neuron 7 Output (Epoch 50).
[0658] For Epoch 50, the eighth node among the multiple nodes in the hidden layer can output z 8-50 =n 8-50 *x+b 8-50 , which is the output value of the fourth neuron in the first layer, Neuron 8 Output (Epoch 50).
[0659] For Epoch 50, the ninth node among the multiple nodes in the hidden layer can output z 9-50 =n 9-50 *x+b 9-50 , that is, the output value of the ninth neuron in the first layer, Neuron 9 Output (Epoch 50).
[0660] For Epoch 50, the tenth node among the multiple nodes in the hidden layer can output z 10-50 =n 10-50 *x+b 10-50 , that is, the output value of the tenth neuron in the first layer, Neuron 10 Output (Epoch 50).
[0661] For Epoch 50, the eleventh node among the multiple nodes in the hidden layer can output z 11-50 =n 11-50 *x+b 11-50 , that is, the output value of the eleventh neuron in the first layer, Neuron 11 Output (Epoch 50).
[0662] For Epoch 50, the twelfth node among the multiple nodes in the hidden layer can output z 12-50 =n 12-50 *x+b 12-50 , which is the output value of the eleventh neuron in the first layer, Neuron 12 Output (Epoch 50).
[0663] For Epoch 50, the thirteenth node among the multiple nodes in the hidden layer can output z 13-50 =n13-50 *x+b 13-50 , which is the output value of the thirteenth neuron in the first layer, Neuron 13 Output (Epoch 50).
[0664] For Epoch 50, the fourteenth node among the multiple nodes in the hidden layer can output z 14-50 =n 14-50 *x+b 14-50 , which is the output value of the fourteenth neuron in the first layer, Neuron 14 Output (Epoch 50).
[0665] For Epoch 50, the fifteenth node among the multiple nodes in the hidden layer can output z 15-50 =n 15-50 *x+b 15-50 , which is the output value of the fifteenth neuron in the first layer, Neuron 15 Output (Epoch 50).
[0666] refer to Fig.29 , for Epoch 50, a rectified linear unit (ReLU) function may be applied to the output value of the first layer output to each of the plurality of nodes of the hidden layer.
[0667] That is, by applying the ReLU function to the output value of the first neuron in the first layer, Neuron 1 Output (Epoch 50), i.e., z 1-50 =n 1-50 *x+b 1-50 , you can output ReLU(n 1-50 *x+b 1-50 ), which is the output value Neuron 1 Activation (Epoch 50).
[0668] Then, the output value of the second neuron in the first layer, Neuron 2 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 2-50 =n 2-50 *x+b 2-50 , you can output ReLU(n 2-50 *x+b 2-50 ), which is the output value Neuron 2 Activation (Epoch 50).
[0669] Then, the output value of the third neuron in the first layer, Neuron 3 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 3-50 =n3-50 *x+b 3-50 , you can output ReLU(n 3-50 *x+b 3-50 ), which is the output value Neuron 3 Activation (Epoch 50).
[0670] Then, the output value of the fourth neuron in the first layer, Neuron 4 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 4-50 =n 4-50 *x+b 4-50 , you can output ReLU(n 4-50 *x+b 4-50 ), which is the output value Neuron 4 Activation (Epoch 50).
[0671] Then, by applying the ReLU function to the output value of the fifth neuron in the first layer, Neuron 5 Output (Epoch 50), i.e., z 5-50 =n 5-50 *x+b 5-50 , you can output ReLU(n 5-50 *x+b 5-50 ), which is the output value Neuron 5 Activation (Epoch 50).
[0672] Then, the output value of the sixth neuron in the first layer, Neuron 6 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 6-50 =n 6-50 *x+b 6-50 , you can output ReLU(n 6-50 *x+b 6-50 ), which is the output value Neuron 6 Activation (Epoch 50).
[0673] Then, by applying the ReLU function to the output value of the seventh neuron in the first layer, Neuron 7 Output (Epoch 50), i.e., z 7-50 =n 7-50 *x+b 7-50 , you can output ReLU(n 7-50 *x+b 7-50 ), which is the output value Neuron 7 Activation (Epoch 50).
[0674] Then, by applying the ReLU function to the output value of the eighth neuron in the first layer, Neuron 8 Output (Epoch 50), i.e., z 8-50 =n 8-50 *x+b 8-50 , you can output ReLU(n 8-50 *x+b 8-50 ), which is the output value Neuron 8 Activation (Epoch 50).
[0675] Then, the output value of the ninth neuron in the first layer, Neuron 9 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 9-50 =n 9-50 *x+b 9-50 , you can output ReLU(n 9-50 *x+b 9-50 ), which is the output value Neuron 9 Activation (Epoch 50).
[0676] Then, the output value of the tenth neuron in the first layer, Neuron 10 Output (Epoch 50), is applied to the rectified linear unit (ReLU) function, i.e., z 10-50 =n 10-50 *x+b 10-50 , you can output ReLU(n 10-50 *x+b 10-50 ), which is the output value Neuron 10 Activation (Epoch 50).
[0677] Then, by applying the ReLU function to the output value of the eleventh neuron in the first layer, Neuron 11 Output (Epoch 50), i.e., z 11-50 =n 11-50 *x+b 11-50 , you can output ReLU(n 11-50 *x+b 11-50 ), which is the output value Neuron 11 Activation (Epoch 50).
[0678] Then, by applying the ReLU function to the output value of the twelfth neuron in the first layer, Neuron 12 Output (Epoch 50), i.e., z 12-50 =n 12-50 *x+b 12-50 , you can output ReLU(n 12-50 *x+b12-50 ), which is the output value Neuron 12 Activation (Epoch 50).
[0679] Then, by applying the ReLU function to the output value of the thirteenth neuron in the first layer, Neuron 13 Output (Epoch 50), i.e., z 13-50 =n 13-50 *x+b 13-50 , you can output ReLU(n 13-50 *x+b 13-50 ), which is the output value Neuron 13 Activation (Epoch 50).
[0680] Then, by applying the ReLU function to the output value of the fourteenth neuron in the first layer, Neuron 14 Output (Epoch 50), i.e., z 14-50 =n 14-50 *x+b 14-50 , you can output ReLU(n 14-50 *x+b 14-50 ), which is the output value Neuron 14 Activation (Epoch 50).
[0681] Then, by applying the ReLU function to the output value of the fifteenth neuron in the first layer, Neuron 15 Output (Epoch 50), i.e., z 15-50 =n 15-50 *x+b 15-50 , you can output ReLU(n 15-50 *x+b 15-50 ), which is the output value Neuron 15 Activation (Epoch 50).
[0682] See also Fig.30 , for Epoch 50, the weight of each of the multiple neurons included in the second layer can be expressed as m 1-50 、m 2-50 ,...,m 15-50 express.
[0683] Therefore, at the first neuron in the second layer, the weight m 1-50 With the input value ReLU(n 1-50 *x+b 1-50 ) so that the output of the first neuron in the second layer, Neuron 1 Outcome, can be m 1-50 *ReLU(n 1-50 *x+b1-50 ).
[0684] Then, at the second neuron in the second layer, the weight m 2-50 With the input value ReLU(n 2-50 *x+b 2-50 ) so that the output of the second neuron in the second layer, Neuron 2 Outcome, can be m 2-50 *ReLU(n 2-50 *x+b 2-50 ).
[0685] At the third neuron in the second layer, the weight m 3-50 With the input value ReLU(n 3-50 *x+b 3-50 ) is multiplied so that the output of the third neuron in the second layer, Neuron 3 Outcome, can be m 3-50 *ReLU(n 3-50 *x+b 3-50 ).
[0686] Then, at the fourth neuron in the second layer, the weight m 4-50 With the input value ReLU(n 4-50 *x+b 4-50 ) is multiplied so that the output of the fourth neuron in the second layer, Neuron 4 Outcome, can be m 4-50 *ReLU(n 4-50 *x+b 4-50 ).
[0687] At the fifth neuron in the second layer, the weight m 5-50 With the input value ReLU(n 5-50 *x+b 5-50 ) is multiplied so that the output of the fifth neuron in the second layer, Neuron 5 Outcome, can be m 5-50 *ReLU(n 5-50 *x+b 5-50 ).
[0688] Then, at the sixth neuron in the second layer, the weight m 6-50 With the input value ReLU(n 6-50 *x+b 6-50 ) is multiplied so that the output of the sixth neuron in the second layer, Neuron 6 Outcome, can be m 6-50 *ReLU(n 6-50 *x+b 6-50 ).
[0689] At the seventh neuron in the second layer, the weight m 7-50With the input value ReLU(n 7-50 *x+b 7-50 ) so that the output of the seventh neuron in the second layer, Neuron 7 Outcome, can be m 7-50 *ReLU(n 7-50 *x+b 7-50 ).
[0690] Then, at the eighth neuron in the second layer, the weight m 8-50 With the input value ReLU(n 8-50 *x+b 8-50 ) is multiplied so that the output of the eighth neuron in the second layer, Neuron 8 Outcome, can be m 8-50 *ReLU(n 8-50 *x+b 8-50 ).
[0691] At the ninth neuron in the second layer, the weight m 9-50 With the input value ReLU(n 9-50 *x+b 9-50 ) is multiplied so that the output of the ninth neuron in the second layer, Neuron 9 Outcome, can be m 9-50 *ReLU(n 9-50 *x+b 9-50 ).
[0692] Then, at the tenth neuron in the second layer, the weight m 10-50 With the input value ReLU(n 10-50 *x+b 10-50 ) is multiplied so that the output of the tenth neuron in the second layer, Neuron 10 Outcome, can be m 10-50 *ReLU(n 10-50 *x+b 10-50 ).
[0693] At the eleventh neuron in the second layer, the weight m 11-50 With the input value ReLU(n 11-50 *x+b 11-50 ) is multiplied so that the output of the eleventh neuron in the second layer, Neuron 11 Outcome, can be m 11-50 *ReLU(n 11-50 *x+b 11-50 ).
[0694] Then, at the twelfth neuron in the second layer, the weight m 12-50 With the input value ReLU(n 12-50 *x+b 12-50) so that the output of the twelfth neuron in the second layer, Neuron 12 Outcome, can be m 12-50 *ReLU(n 12-50 *x+b 12-50 ).
[0695] Then, at the thirteenth neuron in the second layer, the weight m 13-50 With the input value ReLU(n 13-50 *x+b 13-50 ) is multiplied so that the output of the thirteenth neuron in the second layer, Neuron 13 Outcome, can be m 13-50 *ReLU(n 13-50 *x+b 13-50 ).
[0696] At the fourteenth neuron in the second layer, the weight m 14-50 With the input value ReLU(n 14-50 *x+b 14-50 ) is multiplied so that the output of the fourteenth neuron in the second layer, Neuron 14 Outcome, can be m 14-50 *ReLU(n 14-50 *x+b 14-50 ).
[0697] Then, at the fifteenth neuron in the second layer, the weight m 15-50 With the input value ReLU(n 15-50 *x+b 15-50 ) so that the output of the fifteenth neuron in the second layer, Neuron 15 Outcome, can be m 15-50 *ReLU(n 15-50 *x+b 15-50 ).
[0698] Then, by adding all the outputs of the plurality of neurons of the second layer (Neuron 1 Outcome to Neuron 15 Outcome) as described above, the final output of the second layer (ie, the final output of the artificial neural network) can be expressed as [Formula 2].
[0699] [Formula 2]
[0700] refer to Fig.31 ,For Epoch 50, calculate the final output of the second layer, the error of the programmed function and the target activation function.
[0701] That is, the loss function is used to calculate the error between the final output of the second layer of Epoch 50, the programmed function, and the target activation function.
[0702] For example, the loss function can be derived by calculating the mean squared error (MSE) of the error between the final output of the second layer, the programmed function, and the target activation function.
[0703] However, the loss function used to calculate the error of the final output of the second layer, the programmed function, and the target activation function is not limited to the mean square error (MSE), but can also be the root mean square error (RMSE), the cross entropy error (CEE), the binary cross entropy error (BCEE), and the categorical cross entropy error (BCEE).
[0704] Fig.32 is a diagram showing the output of the first layer of Epoch 300.
[0705] Fig.33 is a graph showing the values of the output of the first layer when the ReLU function is applied to Epoch 300.
[0706] Fig.34 is a diagram showing the output of the second layer of Epoch 300.
[0707] Fig.35 is a graph showing the error between the programmed activation function and the target activation function for Epoch 300.
[0708] First, refer to Fig.32 , for Epoch 300, the weight of each of the multiple neurons in the first layer can be expressed as n 1-300 , n 2-300 , ..., n 15-300 Represented by, and the bias of each of the multiple neurons in the first layer can be represented by b 1-300 , b 2-300 , ..., b 15-300 express.
[0709] Accordingly, for Epoch 300, the first node among the multiple nodes in the hidden layer can output z 1-300 =n 1-300 *x+b 1-300 , that is, the output value of the first neuron in the first layer Neuron 1 Output (Epoch 300).
[0710] For Epoch 300, the second node among the multiple nodes in the hidden layer can output z 2-300 =n 2-300 *x+b 2-300 , which is the output value of the second neuron in the first layer, Neuron 2 Output (Epoch 300).
[0711] For Epoch 300, the third node among the multiple nodes in the hidden layer can output z 3-300 =n 3-300 *x+b 3-300 , which is the output value of the third neuron in the first layer, Neuron 3 Output (Epoch 300).
[0712] For Epoch 300, the fourth node among the multiple nodes in the hidden layer can output z 4-300 =n 4-300 *x+b 4-300 , which is the output value of the fourth neuron in the first layer, Neuron 4 Output (Epoch 300).
[0713] For Epoch 300, the fifth node among the multiple nodes in the hidden layer can output z 5-300 =n 5-300 *x+b 5-300 , which is the output value of the fifth neuron in the first layer, Neuron 5 Output (Epoch 300).
[0714] For Epoch 300, the sixth node among the multiple nodes in the hidden layer can output z 6-300 =n 6-300 *x+b 6-300 , which is the output value of the sixth neuron in the first layer, Neuron 6 Output (Epoch 300).
[0715] For Epoch 300, the seventh node among the multiple nodes in the hidden layer can output z 7-300 =n 7-300 *x+b 7-300 , which is the output value of the seventh neuron in the first layer, Neuron 7 Output (Epoch 300).
[0716] For Epoch 300, the eighth node among the multiple nodes in the hidden layer can output z 8-300 =n 8-300 *x+b 8-300 , which is the output value of the fourth neuron in the first layer, Neuron 8 Output (Epoch 300).
[0717] For Epoch 300, the ninth node among the multiple nodes in the hidden layer can output z 9-300 =n 9-300 *x+b 9-300 , which is the output value of the ninth neuron in the first layer, Neuron 9 Output (Epoch 300).
[0718] For Epoch 300, the tenth node among the multiple nodes in the hidden layer can output z 10-300 =n 10-300 *x+b 10-300 , which is the output value of the tenth neuron in the first layer, Neuron 10 Output (Epoch 300).
[0719] For Epoch 300, the eleventh node among the multiple nodes in the hidden layer can output z 11-300 =n 11-300 *x+b 11-300 , that is, the output value of the eleventh neuron in the first layer, Neuron 11 Output (Epoch 300).
[0720] For Epoch 300, the twelfth node among the multiple nodes in the hidden layer can output z 12-300 =n 12-300 *x+b 12-300 , that is, the output value of the eleventh neuron in the first layer, Neuron 12 Output (Epoch 300).
[0721] For Epoch 300, the thirteenth node among the multiple nodes in the hidden layer can output z 13-300 =n 13-300 *x+b 13-300 , that is, the output value of the thirteenth neuron in the first layer, Neuron 13 Output (Epoch 300).
[0722] For Epoch 300, the fourteenth node among the multiple nodes in the hidden layer can output z 14-300 =n 14-300 *x+b 14-300 , which is the output value of the fourteenth neuron in the first layer, Neuron 14 Output (Epoch 300).
[0723] For Epoch 300, the fifteenth node among the multiple nodes in the hidden layer can output z 15-300 =n 15-300 *x+b 15-300 , which is the output value of the fifteenth neuron in the first layer, Neuron 15 Output (Epoch 300).
[0724] refer to Fig.33 , for Epoch 50, a rectified linear unit (ReLU) function may be applied to the output value of the first layer output to each of the plurality of nodes of the hidden layer.
[0725] That is, by applying the ReLU function to the output value of the first neuron in the first layer, Neuron 1 Output (Epoch 300), i.e., z 1-300 =n 1-300 *x+b 1-300 , you can output ReLU(n 1-300 *x+b 1-300 ), which is the output value Neuron 1 Activation (Epoch 300).
[0726] Then, the output value of the second neuron in the first layer, Neuron 2 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 2-300 =n 2-300 *x+b 2-300 , you can output ReLU(n 2-300 *x+b 2-300 ), which is the output value Neuron 2 Activation (Epoch 300).
[0727] Then, by applying the ReLU function to the output value of the third neuron in the first layer, Neuron 3 Output (Epoch 300), i.e., z 3-300 =n 3-300 *x+b 3-300 , you can output ReLU(n 3-300 *x+b 3-300 ), which is the output value Neuron 3 Activation (Epoch 300).
[0728] Then, the output value of the fourth neuron in the first layer, Neuron 4 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 4-300 =n 4-300 *x+b 4-300 , you can output ReLU(n 4-300 *x+b 4-300 ), which is the output value Neuron 4 Activation (Epoch 300).
[0729] Then, the output value of the fifth neuron in the first layer, Neuron 5 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 5-300 =n 5-300 *x+b 5-300, you can output ReLU(n 5-300 *x+b 5-300 ), which is the output value Neuron 5 Activation (Epoch 300).
[0730] Then, the output value of the sixth neuron in the first layer, Neuron 6 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 6-300 =n 6-300 *x+b 6-300 , you can output ReLU(n 6-300 *x+b 6-300 ), which is the output value Neuron 6 Activation (Epoch 300).
[0731] Then, the output value of the seventh neuron in the first layer, Neuron 7 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 7-300 =n 7-300 *x+b 7-300 , you can output ReLU(n 7-300 *x+b 7-300 ), which is the output value Neuron 7 Activation (Epoch 300).
[0732] Then, the output value of the eighth neuron in the first layer, Neuron 8 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 8-300 =n 8-300 *x+b 8-300 , you can output ReLU(n 8-300 *x+b 8-300 ), which is the output value Neuron 8 Activation (Epoch 300).
[0733] Then, the output value of the ninth neuron in the first layer, Neuron 9 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 9-300 =n 9-300 *x+b 9-300 , you can output ReLU(n 9-300 *x+b 9-300 ), which is the output value Neuron 9 Activation (Epoch 300).
[0734] Then, the output value of the tenth neuron in the first layer, Neuron 10 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 10-300 =n 10-300 *x+b 10-300 , you can output ReLU(n 10-300 *x+b 10-300 ), which is the output value Neuron 10 Activation (Epoch 300).
[0735] Then, the output value of the eleventh neuron in the first layer, Neuron 11 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 11-300 =n 11-300 *x+b 11-300 , you can output ReLU(n 11-300 *x+b 11-300 ), which is the output value Neuron 11 Activation (Epoch 300).
[0736] Then, the output value of the twelfth neuron in the first layer, Neuron 12 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 12-300 =n 12-300 *x+b 12-300 , you can output ReLU(n 12-300 *x+b 12-300 ), which is the output value Neuron 12 Activation (Epoch 300).
[0737] Then, by applying the ReLU function to the output value of the thirteenth neuron in the first layer, Neuron 13 Output (Epoch 300), i.e., z 13-300 =n 13-300 *x+b 13-300 , you can output ReLU(n 13-300 *x+b 13-300 ), which is the output value Neuron 13 Activation (Epoch 300).
[0738] Then, the output value of the fourteenth neuron in the first layer, Neuron 14 Output (Epoch 300), is applied to the rectified linear unit (ReLU) function, i.e., z 14-300 =n 14-300 *x+b 14-300 , you can output ReLU(n 14-300*x+b 14-300 ), which is the output value Neuron 14 Activation (Epoch 300).
[0739] Then, by applying the ReLU function to the output value of the fifteenth neuron in the first layer, Neuron 15 Output (Epoch 300), i.e., z 15-300 =n 15-300 *x+b 15-300 , you can output ReLU(n 15-300 *x+b 15-300 ), which is the output value Neuron 15 Activation (Epoch 300).
[0740] See also Fig.33 , for Epoch 300, the weight of each of the multiple neurons included in the second layer can be expressed as m 1-300 、m 2-300 ,...,m 15-300 express.
[0741] Therefore, at the first neuron in the second layer, the weight m 1-300 With the input value ReLU(n 1-300 *x+b 1-300 ) so that the output of the first neuron in the second layer, Neuron 1 Outcome, can be m 1-300 *ReLU(n 1-300 *x+b 1-300 ).
[0742] Then, at the second neuron in the second layer, the weight m 2-300 With the input value ReLU(n 2-300 *x+b 2-300 ) so that the output of the second neuron in the second layer, Neuron 2 Outcome, can be m 2-300 *ReLU(n 2-300 *x+b 2-300 ).
[0743] At the third neuron in the second layer, the weight m 3-300 With the input value ReLU(n 3-300 *x+b 3-300 ) is multiplied so that the output of the third neuron in the second layer, Neuron 3 Outcome, can be m 3-300 *ReLU(n 3-300 *x+b 3-300 ).
[0744] Then, at the fourth neuron in the second layer, the weight m 4-300 With the input value ReLU(n 4-300 *x+b 4-300 ) is multiplied so that the output of the fourth neuron in the second layer, Neuron 4 Outcome, can be m 4-300 *ReLU(n 4-300 *x+b 4-300 ).
[0745] At the fifth neuron in the second layer, the weight m 5-300 With the input value ReLU(n 5-300 *x+b 5-300 ) is multiplied so that the output of the fifth neuron in the second layer, Neuron 5 Outcome, can be m 5-300 *ReLU(n 5-300 *x+b 5-300 ).
[0746] Then, at the sixth neuron in the second layer, the weight m 6-300 With the input value ReLU(n 6-300 *x+b 6-300 ) is multiplied so that the output of the sixth neuron in the second layer, Neuron 6 Outcome, can be m 6-300 *ReLU(n 6-300 *x+b 6-300 ).
[0747] At the seventh neuron in the second layer, the weight m 7-300 With the input value ReLU(n 7-300 *x+b 7-300 ) so that the output of the seventh neuron in the second layer, Neuron 7 Outcome, can be m 7-300 *ReLU(n 7-300 *x+b 7-300 ).
[0748] Then, at the eighth neuron in the second layer, the weight m 8-300 With the input value ReLU(n 8-300 *x+b 8-300 ) is multiplied so that the output of the eighth neuron in the second layer, Neuron 8 Outcome, can be m 8-300 *ReLU(n 8-300 *x+b 8-300 ).
[0749] At the ninth neuron in the second layer, the weight m 9-300 With the input value ReLU(n 9-300 *x+b 9-300) is multiplied so that the output of the ninth neuron in the second layer, Neuron 9 Outcome, can be m 9-300 *ReLU(n 9-300 *x+b 9-300 ).
[0750] Then, at the tenth neuron in the second layer, the weight m 10-300 With the input value ReLU(n 10-300 *x+b 10-300 ) is multiplied so that the output of the tenth neuron in the second layer, Neuron 10 Outcome, can be m 10-300 *ReLU(n 10-300 *x+b 10-300 ).
[0751] At the eleventh neuron in the second layer, the weight m 11-300 With the input value ReLU(n 11-300 *x+b 11-300 ) is multiplied so that the output of the eleventh neuron in the second layer, Neuron 11 Outcome, can be m 11-300 *ReLU(n 11-300 *x+b 11-300 ).
[0752] Then, at the twelfth neuron in the second layer, the weight m 12-300 With the input value ReLU(n 12-300 *x+b 12-300 ) so that the output of the twelfth neuron in the second layer, Neuron 12 Outcome, can be m 12-300 *ReLU(n 12-300 *x+b 12-300 ).
[0753] Then, at the thirteenth neuron in the second layer, the weight m 13-300 With the input value ReLU(n 13-300 *x+b 13-300 ) is multiplied so that the output of the thirteenth neuron in the second layer, Neuron 13 Outcome, can be m 13-300 *ReLU(n 13-300 *x+b 13-300 ).
[0754] At the fourteenth neuron in the second layer, the weight m 14-300 With the input value ReLU(n 14-300 *x+b 14-300 ) is multiplied so that the output of the fourteenth neuron in the second layer, Neuron 14 Outcome, can be m 14-300 *ReLU(n14-300 *x+b 14-300 ).
[0755] Then, at the fifteenth neuron in the second layer, the weight m 15-300 With the input value ReLU(n 15-300 *x+b 15-300 ) so that the output of the fifteenth neuron in the second layer, Neuron 15 Outcome, can be m 15-300 *ReLU(n 15-300 *x+b 15-300 ).
[0756] Then, the final output of the second layer, that is, the final output of the artificial neural network, can be expressed as [Formula 3], which is the sum of the outputs of multiple neurons (Neuron 1 Outcome to Neuron 15 Outcome) of the second layer as described above.
[0757] [Formula 3]
[0758] refer to Fig.35 ,For Epoch 300, calculate the final output of the second layer, the error of the programmed function and the target activation function.
[0759] That is, the loss function is used to calculate the error between the final output of the second layer of Epoch 300, the programmed function, and the target activation function.
[0760] For example, the loss function can be derived by calculating the mean squared error (MSE) of the error between the final output of the second layer, the programmed function, and the target activation function.
[0761] However, the loss function for calculating the error of the final output of the second layer, the programmed function, and the target activation function is not limited to the mean square error (MSE), but also includes the root mean square error (RMSE), the cross entropy error (CEE), the binary cross entropy error (BCEE), and the categorical cross entropy error (BCEE).
[0762] The following takes Epoch 300 as an example to describe the process of deriving multiple programmable segments of a programmed activation function.
[0763] Fig.36 is a diagram showing the breakpoints of the programmed activation function of Epoch 300.
[0764] Fig.37 is an enlarged view showing the breakpoints of the programmed activation function of Epoch 300.
[0765] That is to say, Fig.36 and37 In , a ReLU function is applied to the output of the first layer to derive the breakpoints of the programmed activation function.
[0766] Specifically, Fig.37 The breakpoints of the programmed activation function are shown.
[0767] That is to say, Fig.36 and 37 In , a ReLU function is applied to the output of the first layer to derive the breakpoints of the programmed activation function.
[0768] Specifically, Fig.37 The breakpoints of the programmed activation function are shown.
[0769] That is to say, Fig.36 and 37 In , a ReLU function is applied to the output of the first layer to derive the breakpoints of the programmed activation function.
[0770] Specifically, Fig.36 shows the value of the output of the first layer when the value of x is between -10 and 10. In addition, Fig.37 A zoomed-in view of the result of applying the ReLU function to the output of the first layer when the value of x is between -1 and 1 is shown.
[0771] refer to Fig.36 and 37 , the boundaries of multiple programmable segments of the programmed activation function can be set by computing breakpoints for each value of the ReLU function applied to the output of the first layer.
[0772] Specifically, the first breakpoint bp1, ie, the breakpoint of the ReLU function value applied to the output Neuron 1Activation (Epoch 300) of the second neuron of the first layer, may be calculated.
[0773] Then, the second breakpoint bp2, ie, the breakpoint of the ReLU function value applied to the output Neuron 2Activation (Epoch 300) of the second neuron of the first layer, can be calculated.
[0774] Then, the third breakpoint bp3, ie, the breakpoint of the ReLU function value applied to the output Neuron 3Activation (Epoch 300) of the third neuron of the first layer, can be calculated.
[0775] Then, the fourth breakpoint bp4, ie, the breakpoint of the ReLU function value applied to the output Neuron 4Activation (Epoch 300) of the fourth neuron of the first layer, can be calculated.
[0776] Then, the fifth breakpoint bp5, ie, the breakpoint of the ReLU function value applied to the output Neuron 5Activation (Epoch 300) of the fifth neuron of the first layer, can be calculated.
[0777] Then, the sixth breakpoint bp6, ie, the breakpoint of the ReLU function value applied to the output Neuron 6Activation (Epoch 300) of the sixth neuron of the first layer, can be calculated.
[0778] Then, the seventh breakpoint bp7, ie, the breakpoint of the ReLU function value applied to the output Neuron 7Activation (Epoch 300) of the seventh neuron of the first layer, can be calculated.
[0779] Then, the eighth breakpoint bp8, ie, the breakpoint of the ReLU function value applied to the output Neuron 8Activation (Epoch 300) of the eighth neuron of the first layer, can be calculated.
[0780] Then, the ninth breakpoint bp9, ie, the breakpoint of the ReLU function value applied to the output Neuron 9Activation (Epoch 300) of the ninth neuron of the first layer, can be calculated.
[0781] Then, the tenth breakpoint bp can be calculated 10 , which is the breakpoint of the ReLU function value applied to the output Neuron 10Activation (Epoch 300) of the tenth neuron in the first layer.
[0782] Then, the eleventh breakpoint bp can be calculated 11 , which is the breakpoint of the ReLU function value applied to the output Neuron 5Activation (Epoch 300) of the eleventh neuron in the first layer.
[0783] Then, the twelfth breakpoint bp can be calculated 12 , which is the breakpoint of the ReLU function value applied to the output Neuron12 Activation (Epoch 300) of the twelfth neuron in the first layer.
[0784] Then, the seventh breakpoint bp7, ie, the breakpoint of the ReLU function value applied to the output Neuron 7Activation (Epoch 300) of the seventh neuron of the first layer, can be calculated.
[0785] Then, the thirteenth breakpoint bp can be calculated 13, which is the breakpoint of the ReLU function value applied to the output of the thirteenth neuron in the first layer, Neuron13 Activation (Epoch 300).
[0786] Then, the fourteenth breakpoint bp can be calculated 14 , which is the breakpoint of the ReLU function value applied to the output Neuron14 Activation (Epoch 300) of the fourteenth neuron in the first layer.
[0787] Then, the fifteenth breakpoint bp can be calculated 15 , which is the breakpoint of the ReLU function value applied to the output of the fifteenth neuron in the first layer, Neuron15 Activation (Epoch 300).
[0788] More specifically, if the weights of the plurality of neurons in the first layer are all n, and the biases of the plurality of neurons in the first layer are all b, then each of the plurality of breakpoints bp may be calculated as -b / n.
[0789] For example, each of the plurality of breakpoints according to the weight and bias of each of the plurality of neurons may be expressed as shown in Table 8.
[0790] Neurons n b bp 1 0.494768 -1.0633781 2.149245909 2 0.867136 -0.0400035 0.046132902 3 -0.2581 -0.9045406 -3.504612941 4 -0.27437 -0.83794 -3.054051099 5 0.271104 0.05046014 -0.186128349 6 0.589796 0.41333011 -0.70080182 7 -0.42154 0.05490794 0.130255587 8 0.853946 -0.3598031 0.421341748 9 0.537062 0.03628266 -0.067557675 10 -0.22686 0.76596254 3.376366658 11 0.433543 0.3996537 -0.921831744 12 -0.30503 -0.1894374 -0.621045143 13 0.240677 0.11094899 -0.460987091 14 0.019641 -0.3599817 18.32807393 15 0.847085 -0.4981247 0.58804571
[0791] That is to say, reference Fig.36 and Fig.37 , the first breakpoint bp1 is 2.149245909, the second breakpoint bp2 is 0.046132902, the third breakpoint bp3 is -3.504612941, the fourth breakpoint bp4 is -3.054051099, the fifth breakpoint bp5 is -0.186128349, the sixth breakpoint bp6 is -0.70080182, the seventh breakpoint bp7 is 0.130255587, the eighth breakpoint bp8 is 0.421341748, the ninth breakpoint bp9 is -0.067557675, and the tenth breakpoint bp 10 3.376366658, the eleventh breakpoint bp 11 is -0.921831744, the twelfth breakpoint bp 12 is -0.621045143, the thirteenth breakpoint bp 13 is -0.460987091, the fourteenth breakpoint bp 14 is 18.32807393, and the fifteenth breakpoint bp 15 It is 0.58804571.
[0792] Using the breakpoints described above as boundaries, one can set the linear interval for the programmed activation function.
[0793] For example, the first interval may be a segment from negative infinity to the third breakpoint bp3, -3.504612941.
[0794] The second interval can be from the third breakpoint bp3, -3.504612941 to the fourth breakpoint bp4, -3.054051099.
[0795] The third interval can be from the fourth breakpoint bp4, -3.054051099 to the eleventh breakpoint bp 11 , -0.921831744.
[0796] The fourth interval can be from the eleventh breakpoint bp 11 , -0.921831744, to the sixth breakpoint bp6, -0.70080182.
[0797] The fifth interval can be from the sixth breakpoint bp6, -0.70080182, to the twelfth breakpoint bp 12 , -0.621045143.
[0798] The sixth interval can be from the twelfth breakpoint bp 12 , -0.621045143, to the thirteenth breakpoint bp 13 , -0.460987091.
[0799] The seventh interval can be from the thirteenth breakpoint bp 13 , -0.460987091 to the fifth breakpoint bp5, -0.186128349.
[0800] The eighth interval may be the segment from the fifth breakpoint bp5, -0.186128349 to the ninth breakpoint bp9, -0.067557675.
[0801] The ninth interval may be a segment from the ninth breakpoint bp9, -0.067557675 to the second breakpoint bp2, 0.046132902.
[0802] The tenth interval may be the segment from the second breakpoint bp2, 0.046132902 to the seventh breakpoint bp7, 0.130255587.
[0803] The eleventh interval can be from the seventh breakpoint bp7, 0.130255587, to the eighth breakpoint bp8, 0.421341748.
[0804] The twelfth interval can be from the eighth breakpoint bp8, 0.421341748, to the fifteenth breakpoint bp 15, 0.58804571.
[0805] The thirteenth interval can be from the fifteenth breakpoint bp 15 , 0.58804571, to the first breakpoint bp1, 2.149245909.
[0806] The fourteenth interval can be from the first breakpoint bp1,2.149245909 to the tenth breakpoint bp 10 , 3.376366658.
[0807] The fifteenth segment can be from the tenth breakpoint bp 10 , 3.376366658 to the 14th breakpoint bp 14 , 18.32807393.
[0808] The sixteenth segment can be from the fourteenth breakpoint bp 14 (not shown), a segment from 18.32807393 to positive infinity.
[0809] refer to Figures 38 to 52 , the following describes a process of deriving a programmable segment of a programmable activation function by adding the output of each of a plurality of neurons in each of a plurality of segments.
[0810] Each of the plurality of neurons of the artificial neural network includes each of the plurality of neurons of the first layer and each of the plurality of neurons of the second layer connected through each of the plurality of nodes of the hidden layer.
[0811] Fig.38 is a diagram illustrating a process of deriving a programmable segment from a first segment.
[0812] Fig.39 is a diagram illustrating a process of deriving a programmable segment from a second segment.
[0813] Fig.40 is a diagram showing a process of deriving a programmable segment from a third segment.
[0814] Fig.41 is a diagram showing a process of deriving a programmable segment from a fourth segment.
[0815] Fig.42 is a diagram showing a process of deriving a programmable segment from a fifth segment.
[0816] Fig.43 is a diagram showing a process of deriving a programmable segment from a sixth segment.
[0817] Fig.44 is a diagram showing a process of deriving a programmable segment from a seventh segment.
[0818] Fig.45is a diagram showing a process of deriving a programmable segment from an eighth segment.
[0819] Fig.46 is a diagram showing a process of deriving a programmable segment from a ninth segment.
[0820] Fig.47 is a diagram showing a process of deriving a programmable segment from a tenth segment.
[0821] Fig.48 is a diagram showing a process of deriving a programmable segment from an eleventh segment.
[0822] Fig.49 is a diagram showing a process of deriving a programmable segment from a twelfth segment.
[0823] Fig.50 is a diagram showing a process of deriving a programmable segment from a thirteenth segment.
[0824] Fig.51 is a diagram showing a process of deriving a programmable segment from a fourteenth segment.
[0825] Fig.52 is a diagram showing a process of deriving a programmable segment from a fifteenth segment.
[0826] Fig.53 is a diagram showing a process of deriving a programmable segment from a sixteenth segment.
[0827] The activation function converter unit may add the output of each of the plurality of neurons in each of the plurality of segments to generate a programmable segment.
[0828] Neurons n b m 1 0.494768 -1.0633781 -0.11528 2 0.867136 -0.0400035 0.250981 3 -0.2581 -0.9045406 -0.04777 4 -0.27437 -0.83794 -0.14026 5 0.271104 0.05046014 0.236343 6 0.589796 0.41333011 0.136813 7 -0.42154 0.05490794 0.119756 8 0.853946 -0.3598031 0.21637 9 0.537062 0.03628266 0.291831 10 -0.22686 0.76596254 -0.18697 11 0.433543 0.3996537 0.079995 12 -0.30503 -0.1894374 0.146994 13 0.240677 0.11094899 0.225464 14 0.019641 -0.3599817 0.001106 15 0.847085 -0.4981247 0.305554
[0829] In Table 9, the weight of each of the multiple neurons in the first layer is n, the bias of each of the multiple neurons in the first layer is b, and the weight of each of the multiple neurons in the second layer is m.
[0830] refer to Fig.38 In the first segment Segment_w1, the output of the third neuron Neuron 3 Contribution, the output of the fourth neuron Neuron 4 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution and the output of the twelfth neuron Neuron12 Contribution can be added to generate the programmable segment Model's Output in the first segment Segment_w1.
[0831] In this regard, the output of the third neuron can be expressed as m3*(n3*x+b3), the output of the fourth neuron can be expressed as m4*(n4*x+b4), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the twelfth neuron can be expressed as m 12 *(n 12 *x+b 12 ).
[0832] Therefore, in the first segment Segment_w1, the programmable segment can be expressed as m3*(n3*x+b3)+m4*(n4*x+b4)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 12 *(n 12 *x+b 12 ).
[0833] In other words, in the first segment Segment_w1, the programmable segment can be expressed as (m3*n3+m4*n4+m7*n7+m 10 *n 10 +m 12 *n 12 )*x+(m3*b3+m4*b4+m7*b7+m 10 *b 10 +m 12 *b 12 ), which is a first-order function.
[0834] See also Fig.39 , in the second segment Segment_w2, the output of the fourth neuron Neuron 4 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, and the output of the twelfth neuron Neuron 12 Contribution may be added to generate the programmable segment Model's Output in the second segment Segment_w2.
[0835] In this regard, the output of the fourth neuron can be expressed as m4*(n4*x+b4), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10), the output of the twelfth neuron can be expressed as m 12 *(n 12 *x+b 12 ).
[0836] Therefore, in the second segment Segment_w2, the programmable segment can be expressed as m4*(n4*x+b4)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 12 *(n 12 *x+b 12 ).
[0837] In other words, in the second segment Segment_w2, the programmable segment can be expressed as (m4*n4+m7*n7+m 10 *n 10 +m 12 *n 12 )*x+(m4*b4+m7*b7+m 10 *b 10 +m 12 *b 12 ), which is a first-order function.
[0838] See also Fig.40 , in the third segment Segment_w3, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, and the output of the twelfth neuron Neuron 12 Contribution may be added to generate the programmable segment Model's Output in the third segment Segment_w3.
[0839] In this regard, the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the twelfth neuron can be expressed as m 12 *(n 12 *x+b 12 ).
[0840] Therefore, in the third segment Segment_w3, the programmable segment can be expressed as m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 12 *(n 12 *x+b 12 ).
[0841] In other words, in the third segment Segment_w3, the programmable segment can be expressed as (m7*n7+m 10 *n 10 +m 12 *n 12 )*x+(m7*b7+m 10 *b 10 +m 12 *b 12 ), which is a first-order function.
[0842] See also Fig.41 In the fourth segment Segment_w4, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, and the output of the twelfth neuron Neuron 12 Contribution may be added to generate the programmable segment Model's Output in the fourth segment Segment_w4.
[0843] In this regard, the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the twelfth neuron can be expressed as m 12 *(n 12 *x+b 12 ).
[0844] Therefore, in the fourth segment Segment_w4, the programmable segment can be expressed as m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 12 *(n 12 *x+b 12 ).
[0845] In other words, in the fourth segment Segment_w4, the programmable segment can be expressed as (m7*n7+m 10 *n 10 +m 11 *n11 +m 12 *n 12 )*x+(m7*b7+m 10 *b 10 +m 11 *b 11 +m 12 *b 12 ), which is a first-order function.
[0846] See also Fig.42 In the fifth segment Segment_w5, the output of the sixth neuron Neuron 6 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, and the output of the twelfth neuron Neuron 12 Contribution may be added to generate the programmable segment Model's Output in the fifth segment Segment_w5.
[0847] In this regard, the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the twelfth neuron can be expressed as m 12 *(n 12 *x+b 12 ).
[0848] Therefore, in the fifth segment Segment_w5, the programmable segment can be expressed as m6*(n6*x+b6)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 12 *(n 12 *x+b 12 ).
[0849] In other words, in the fifth segment Segment_w5, the programmable segment can be expressed as (m6*n6+m7*n7+m 10 *n 10+m 11 *n 11 +m 12 *n 12 )*x+(m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 12 *b 12 ), which is a first-order function.
[0850] See also Fig.43 , in the sixth segment Segment_w6, the output Neuron 6 Contribution of the sixth neuron, the output Neuron 7 Contribution of the seventh neuron, the output Neuron 10 Contribution of the tenth neuron, and the output Neuron 11 Contribution of the eleventh neuron may be added to generate the programmable segment Model's Output in the sixth segment Segment_w6.
[0851] In this regard, the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ).
[0852] Therefore, in the sixth segment Segment_w6, the programmable segment can be expressed as m6*(n6*x+b6)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 ).
[0853] In other words, in the sixth segment Segment_w6, the programmable segment can be expressed as (m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 )*x+(m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 ), which is a first-order function.
[0854] See also Fig.44 In the seventh segment Segment_w7, the output of the sixth neuron Neuron 6 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, and the output of the thirteenth neuron Neuron 13 Contribution can be added to generate the programmable segment Model's Output in the seventh segment Segment_w7.
[0855] In this regard, the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ).
[0856] Therefore, in the seventh segment Segment_w7, the programmable segment can be expressed as m6*(n6*x+b6)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 ).
[0857] In other words, in the seventh segment Segment_w7, the programmable segment can be expressed as (m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 )*x+(m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 13 *b13 ), which is a first-order function.
[0858] See also Fig.45 In the eighth segment Segment_w8, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron11 Contribution, and the output of the thirteenth neuron Neuron 13 Contribution can be added to generate a programmable segment Model's Output in the eighth segment Segment_w8.
[0859] In this regard, the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ).
[0860] Therefore, in the eighth segment Segment_w8, the programmable segment can be expressed as m5*(n5*x+b5)+m6*(n6*x+b6)+m7*(n7*x+b7)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 ).
[0861] In other words, in the eighth segment Segment_w8, the programmable segment can be expressed as (m5*n5+m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 +m 13 *n13 )*x+(m5*b5+m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 ), which is a first-order function.
[0862] See also Fig.46 In the ninth segment Segment_w9, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, and the output of the thirteenth neuron Neuron13 Contribution can be added to generate a programmable segment Model's Output in the ninth segment Segment_w9.
[0863] In this regard, the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ).
[0864] Therefore, in the ninth segment Segment_w9, the programmable segment can be expressed as m5*(n5*x+b5)+m6*(n6*x+b6)+m7*(n7*x+b7)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n13 *x+b 13 ).
[0865] In other words, in the ninth segment Segment_w9, the programmable segment can be expressed as (m5*n5+m6*n6+m7*n7+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 )*x+(m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 ), which is a first-order function.
[0866] See also Fig.47 In the tenth segment Segment_w10, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the seventh neuron Neuron 7 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron11 Contribution, and the output of the thirteenth neuron Neuron 13 Contribution can be added to generate the programmable segment Model's Output in the tenth segment Segment_w10.
[0867] In this regard, the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the seventh neuron can be expressed as m7*(n7*x+b7), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13*x+b 13 ).
[0868] Therefore, in the tenth segment Segment_w10, the programmable segment can be expressed as m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m7*(n7*x+b7)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 ).
[0869] In other words, in the tenth segment Segment_w10, the programmable segment can be expressed as (m2*n2+m5*n5+m6*n6+m7*n7+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 )*x+(m2*b2+m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 ), which is a first-order function.
[0870] See also Fig.48 In the eleventh segment Segment_w11, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, and the output of the thirteenth neuron Neuron 13 Contribution can be added to generate the programmable segment Model's Output in the eleventh segment Segment_w11.
[0871] In this regard, the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ).
[0872] Therefore, in the eleventh segment Segment_w11, the programmable segment can be expressed as m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 ).
[0873] In other words, in the 11th segment Segment_w11, the programmable segment can be expressed as (m2*n2+m5*n5+m6*n6+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 )*x+(m2*b2+m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 ), which is a first-order function.
[0874] See also Fig.49In the twelfth segment Segment_w12, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the eighth neuron Neuron 8 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron11 Contribution, and the output of the thirteenth neuron Neuron 13 Contribution can be added to generate the programmable segment Model's Output in the twelfth segment Segment_w12.
[0875] In this regard, the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the eighth neuron can be expressed as m8*(n8*x+b8), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ).
[0876] Therefore, in the twelfth segment Segment_w12, the programmable segment can be expressed as m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m8*(n8*x+b8)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 ).
[0877] In other words, in the twelfth segment Segment_w12, the programmable segment has a first-order function form (m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 )*x+(m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 ).
[0878] See also Fig.50 In the thirteenth segment Segment_w13, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the eighth neuron Neuron 8 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, the output of the thirteenth neuron Neuron 13 Contribution, and the output of the fifteenth neuron Neuron 15 Contribution can be added to generate the programmable segment Model's Output in the thirteenth segment Segment_w13.
[0879] In this regard, the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the eighth neuron can be expressed as m8*(n8*x+b8), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b13 ), the output of the fifteenth neuron can be expressed as m 15 *(n 15 *x+b 15 ).
[0880] Therefore, in the thirteenth segment Segment_w13, the programmable segment can be expressed as m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m8*(n8*x+b8)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 )+m 15 *(n 15 *x+b 15 ).
[0881] In other words, in the thirteenth segment Segment_w13, the programmable segment can be expressed as (m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 +m 15 *n 15 )*x+(m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 +m 15 *b 15 ), which is a first-order function.
[0882] See also Fig.51In the fourteenth segment Segment_w14, the output of the first neuron Neuron 1 Contribution, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the eighth neuron Neuron 8 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the tenth neuron Neuron 10 Contribution, the output of the eleventh neuron Neuron 11 Contribution, the output of the thirteenth neuron Neuron 13 Contribution, and the output of the fifteenth neuron Neuron 15 Contribution can be added to generate the programmable segment Model's Output in the fourteenth segment Segment_w14.
[0883] In this regard, the output of the first neuron can be expressed as m1*(n1*x+b1), the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the eighth neuron can be expressed as m8*(n8*x+b8), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the tenth neuron can be expressed as m 10 *(n 10 *x+b 10 ), the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ), the output of the fifteenth neuron can be expressed as m 15 *(n 15 *x+b 15 ).
[0884] Therefore, in the fourteenth segment Segment_w14, the programmable segment can be expressed as m1*(n1*x+b1)+m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m8*(n8*x+b8)+m9*(n9*x+b9)+m 10 *(n 10 *x+b 10 )+m 11*(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 )+m 15 *(n 15 *x+b 15 ).
[0885] In other words, in the fourteenth segment Segment_w14, the programmable segment can be expressed as (m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 +m 15 *n 15 )*x+(m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 +m 15 *b 15 ), which is a first-order function.
[0886] See also Fig.52 In the fifteenth segment Segment_w15, the output of the first neuron Neuron 1 Contribution, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the eighth neuron Neuron 8 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the eleventh neuron Neuron11 Contribution, the output of the thirteenth neuron Neuron 13 Contribution, and the output of the fifteenth neuron Neuron 15 Contribution can be added to generate the programmable segment Model'sOutput in the fifteenth segment Segment_w15.
[0887] In this regard, the output of the first neuron can be expressed as m1*(n1*x+b1), the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the eighth neuron can be expressed as m8*(n8*x+b8), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ), the output of the fifteenth neuron can be expressed as m 15 *(n 15 *x+b 15 ).
[0888] Therefore, in the fifteenth segment Segment_w15, the programmable segment can be expressed as m1*(n1*x+b1)+m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m8*(n8*x+b8)+m9*(n9*x+b9)+m 11 *(n 11 *x+b 11 )+m 13 *(n 13 *x+b 13 )+m 15 *(n 15 *x+b 15 ).
[0889] In other words, in the fifteenth segment Segment_w15, the programmable segment can be expressed as (m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 11 *n 11 +m 13 *n 13 +m 15 *n 15 )*x+(m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 11 *b 11 +m 13 *b 13 +m 15 *b 15 ), which is a first-order function.
[0890] See also Fig.53In the sixteenth segment Segment_w16, the output of the first neuron Neuron 1 Contribution, the output of the second neuron Neuron 2 Contribution, the output of the fifth neuron Neuron 5 Contribution, the output of the sixth neuron Neuron 6 Contribution, the output of the eighth neuron Neuron 8 Contribution, the output of the ninth neuron Neuron 9 Contribution, the output of the eleventh neuron Neuron11 Contribution, the output of the thirteenth neuron Neuron 13 Contribution, the output of the fourteenth neuron Neuron 14 Contribution, and the output of the fifteenth neuron Neuron 15 Contribution can be added to generate the programmable segment Model's Output in the sixteenth segment Segment_w16.
[0891] In this regard, the output of the first neuron can be expressed as m1*(n1*x+b1), the output of the second neuron can be expressed as m2*(n2*x+b2), the output of the fifth neuron can be expressed as m5*(n5*x+b5), the output of the sixth neuron can be expressed as m6*(n6*x+b6), the output of the eighth neuron can be expressed as m8*(n8*x+b8), the output of the ninth neuron can be expressed as m9*(n9*x+b9), and the output of the eleventh neuron can be expressed as m 11 *(n 11 *x+b 11 ), the output of the thirteenth neuron can be expressed as m 13 *(n 13 *x+b 13 ), the output of the fourteenth neuron can be expressed as m 14 *(n 14 *x+b 14 ), the output of the fifteenth neuron can be expressed as m 15 *(n 15 *x+b 15 ).
[0892] Therefore, in the sixteenth segment Segment_w16, the programmable segment can be expressed as m1*(n1*x+b1)+m2*(n2*x+b2)+m5*(n5*x+b5)+m6*(n6*x+b6)+m8*(n8*x+b8)+m9*(n9*x+b9)+m 11 *(n 11 *x+b 11 )+m13 *(n 13 *x+b 13 )+m 14 *(n 14 *x+b 14 )+m 15 *(n 15 *x+b 15 ).
[0893] In other words, in the sixteenth segment Segment_w16, the programmable segment can be expressed as (m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 11 *n 11 +m 13 *n 13 +m 14 *n 14 +m 15 *n 15 )*x+(m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 11 *b 11 +m 13 *b13+m 14 *b 14 +m 15 *b 15 ), which is a first-order function.
[0894] As mentioned above, by concatenating each of the multiple calculated programmable segments, an approximate programmed activation function can be obtained.
[0895] Then, in step S330 , the programmable segment in the form of a first-order function derived from each of the plurality of segments may be converted into a slope and an offset and stored in a lookup table.
[0896] That is, as shown in Table 10, the lookup table may store the slope and offset representing the first-order function in each of the plurality of segments.
[0897] part A B 1 A1 B1 2 A2 B2 3 A3 B3 4 A4 B4 5 A5 B5 6 A6 B6 7 A7 B7 8 A8 B8 9 A9 B9 10 A10 B10 11 A11 B11 12 A12 B12 13 A13 B13 14 A14 B14 15 A15 B15 16 A16 B16
[0898] The programmable segments derived as a plurality of segments as described above may be expressed as slopes and offsets.
[0899] In the first segment Segment_w1, the slope A1 can be m3*n3+m4*n4+m7*n7+m 10 *n 10 +m 12 *n 12, the offset B1 can be m3*b3+m4*b4+m7*b7+m 10 *b 10 +m 12 *b 12 .
[0900] In the second segment Segment_w2, the slope A2 can be m4*n4+m7*n7+m 10 *n 10 +m 12 *n 12 , the offset B2 can be m4*b4+m7*b7+m 10 *b 10 +m 12 *b 12 .
[0901] In the third segment Segment_w3, the slope A3 can be m7*n7+m 10 *n 10 +m 12 *n 12 , offset B3 can be m7*b7+m 10 *b 10 +m 12 *b 12 .
[0902] In the fourth segment Segment_w4, the slope A4 can be m7*n7+m 10 *n 10 +m 11 *n 11 +m 12 *n 12 , the offset B4 can be m7*b7+m 10 *b 10 +m 11 *b 11 +m 12 *b 12 .
[0903] In the fifth segment Segment_w5, the slope A5 can be m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 +m 12 *n 12 , offset B5 can be m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 12 *b 12 .
[0904] In the sixth segment Segment_w6, the slope A6 can be m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 , the offset B6 can be m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 .
[0905] In the seventh segment Segment_w7, the slope A7 can be m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , offset B7 can be m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 .
[0906] In the eighth segment Segment_w8, the slope A8 may be m5*n5+m6*n6+m7*n7+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , offset B8 can be m5*b5+m6*b6+m7*b7+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 .
[0907] In the ninth segment Segment_w9, the slope A9 can be m5*n5+m6*n6+m7*n7+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , offset B9 can be m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13.
[0908] In the tenth segment Segment_w10, the slope A10 can be m2*n2+m5*n5+m6*n6+m7*n7+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , the offset B10 can be m2*b2+m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 .
[0909] In the eleventh segment Segment_w11, the slope A11 can be m2*n2+m5*n5+m6*n6+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , offset B11 can be m2*b2+m5*b5+m6*b6+m7*b7+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 .
[0910] In the twelfth segment Segment_w12, the slope A12 may be m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 , offset B12 may be m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 .
[0911] In the thirteenth segment Segment_w13, the slope A13 may be m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n10 +m 11 *n 11 +m 13 *n 13 +m 15 *n 15 , offset B13 can be m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 +m 15 *b 15 .
[0912] In the fourteenth segment Segment_w14, the slope A14 may be m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 10 *n 10 +m 11 *n 11 +m 13 *n 13 +m 15 *n 15 , offset B14 can be m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 10 *b 10 +m 11 *b 11 +m 13 *b 13 +m 15 *b 15 .
[0913] In the fifteenth segment Segment_w15, the slope A15 may be m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 11 *n 11 +m 13 *n 13 +m 15 *n 15 , offset B15 can be m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 11 *b 11 +m 13 *b 13 +m 15 *b 15 .
[0914] In the sixteenth segment Segment_w16, the slope A16 may be m1*n1+m2*n2+m5*n5+m6*n6+m8*n8+m9*n9+m 11 *n 11 +m 13 *n 13 +m 14 *n 14 +m 15 *n 15 , offset B16 can be m1*b1+m2*b2+m5*b5+m6*b6+m7*b7+m8*b8+m9*b9+m 11 *b 11 +m 13 *b 13 +m 14 *b 14 +m 15 *b 15 .
[0915] As mentioned above, according to the activation function programming method of the present disclosure, the target activation function can be approximated to a programmed activation function with the minimum approximation error through machine learning of an artificial neural network.
[0916] The approximated programmed activation function can then be converted to slopes and offsets and stored in a lookup table.
[0917] The slopes and offsets listed in the lookup table may then be stored in the memory 300 of the NPU 1000 and used for calculations of the programmed activation function execution unit 500 .
[0918] Finally, various nonlinear activation functions can be converted into computationally optimized programmed activation functions having multiple piecewise-linear forms through machine learning of artificial neural networks. Therefore, the computational speed and power consumption of the programmed activation function execution unit 500 of the NPU 1000 can be optimized.
[0919] A method for programming activation functions optimized taking into account the hardware structure of the PAFE unit is described below.
[0920] According to another example of the present disclosure, an activation function programming method may approximate a target activation function as a programmed activation function, so that it may be optimized for a hardware structure of a PAFE unit.
[0921] That is, the number of programmable segments of the programmable activation function may be limited by the number of comparators, which are hardware information of the PAFE unit 500 of the NPU 1000 .
[0922] As described above, the number of the plurality of programmable segments corresponds to the number of the plurality of intervals determined by the breakpoints of the outputs of the plurality of neurons of the artificial neural network.
[0923] Therefore, by controlling the number of breakpoints of the outputs of the plurality of neurons of the artificial neural network corresponding to the hardware information of the PAFE unit, the number of the plurality of programmable segments can be controlled.
[0924] In other words, the number of neurons of the plurality of neural networks may be set to correspond to the hardware information of the PAFE unit.
[0925] For example, when designing an artificial neural network, the number of the plurality of neurons may be set to be smaller than the number of comparators, which are hardware information of the PAFE unit 500 .
[0926] Alternatively, the method of controlling the number of neurons of a plurality of neural networks may include pruning at least one of the neurons of the plurality of neural networks.
[0927] Fig.54 is a graph of an artificial neural network, wherein at least one of a plurality of neurons of the artificial neural network is pruned.
[0928] By pruning at least one of the plurality of neurons of the artificial neural network, the number of breakpoints in the output of the plurality of neurons of the artificial neural network may be reduced.
[0929] Pruning is a method to reduce the parameters of an artificial neural network by removing less important neurons from its weights.
[0930] More specifically, the pruning may be magnitude pruning, which removes neurons based on the magnitude of their weights.
[0931] Referring to Table 9, neurons of the artificial neural network can be deleted if the absolute value of the weight (n) of the first layer or the weight (m) of the second layer is less than or equal to 0.1.
[0932] In Table 9, the absolute value of the weight of the second layer of the third neuron Neuron 3 is 0.04777, the absolute value of the weight of the second layer of the eleventh neuron Neuron 11 is 0.079995, the absolute value of the weight of the first layer of the fourteenth neuron Neuron 14 is 0.019641, and the absolute value of the weight of the second layer is 0.001106.
[0933] Therefore, if Fig.54 As shown, the third neuron Neuron 3, the eleventh neuron Neuron 11 and the fourteenth neuron Neuron 14 can be pruned.
[0934] Therefore, by pruning the third neuron Neuron 3, the first segment and the second segment separated by the third breakpoint bp3 can be merged into a single segment.
[0935] Therefore, by pruning the eleventh neuron Neuron 11, the eleventh breakpoint bp 11 The separate third and fourth paragraphs were merged into a single paragraph.
[0936] Therefore, by pruning the fourteenth neuron Neuron 14, the fourteenth breakpoint bp 14 The separate paragraphs 15 and 16 were merged into a single paragraph.
[0937] That is, at least one of the plurality of neurons of the artificial neural network may be pruned as described above to reduce the number of the plurality of programmable segments of the programmed activation function.
[0938] In other words, through pruning, the number of neurons of the artificial neural network can be set to be smaller than the number of comparators, which is hardware information of the PAFE unit 500 .
[0939] As described above, the activation function programming method can approximate the target activation function as a programmed activation function so that it can be optimized for the hardware structure of the PAFE unit.
[0940] Therefore, a programmed activation function optimized for the PAFE unit can be processed, which not only improves the computational efficiency of the PAFE unit but also optimizes the power consumption efficiency.
[0941] According to an example of the present disclosure, an activation function conversion program unit is provided. The activation function conversion program unit may be configured to approximate a target activation function to a programmed activation function through machine learning of an artificial neural network.
[0942] The artificial neural network may include a first layer including a plurality of neurons and a second layer including a plurality of neurons, a rectified linear unit (ReLU) function may be applied to outputs of the plurality of neurons of the first layer, and values to which the ReLU function may be applied are inputs to the plurality of neurons of the second layer.
[0943] The artificial neural network may include a first layer including a plurality of neurons and a second layer including a plurality of neurons, each of the plurality of neurons in the first layer may include a weight and a bias, and the plurality of neurons in the second layer may include only weights.
[0944] Artificial neural networks can perform machine learning to minimize the error between the target activation function and the programmed activation function.
[0945] The programmed activation function may include a plurality of segments including a programmable segment implemented in the form of a first-order function.
[0946] The artificial neural network may include a plurality of neurons, and the programmed activation function may include a plurality of programmable segments, wherein the plurality of programmable segments are respectively separated by breakpoints of outputs of the plurality of neurons.
[0947] The programmed activation function may include multiple programmable segments, the number of which may correspond to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and the hardware information may correspond to a comparator included in the PAFE Unit.
[0948] The artificial neural network may include a plurality of neurons, and at least one of the outputs of the plurality of neurons may be pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0949] The artificial neural network may include a plurality of neurons, and the number of the plurality of neurons may be less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0950] According to an example of the present disclosure, an activation function programming method may be provided. The activation function programming method may include setting a target activation function, approximating the target activation function to a programmed activation function by allowing an artificial neural network to perform machine learning, and converting the programmed activation function into a slope and an offset and storing them in a lookup table.
[0951] The artificial neural network may include a first layer including multiple neurons and a second layer including multiple neurons, a rectified linear unit (ReLU) function may be applied to outputs of the multiple neurons of the first layer, and values of the ReLU function may be applied as inputs to the multiple neurons of the second layer.
[0952] The artificial neural network may include a first layer including a plurality of neurons and a second layer including a plurality of neurons, each of the plurality of neurons in the first layer may include a weight and a bias, while the plurality of neurons in the second layer may include only a weight.
[0953] Artificial neural networks can be used for machine learning to minimize the error between the target activation function and the programmed activation function.
[0954] The programmed activation function may include a plurality of segments including a programmable segment implemented in the form of a first-order function.
[0955] The artificial neural network may include a plurality of neurons, and the programmed activation function may include a plurality of programmable segments, wherein the plurality of programmable segments are respectively separated by breakpoints of outputs of the plurality of neurons.
[0956] The programmed activation function may include multiple programmable segments, the number of the multiple programmable segments may correspond to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and the hardware information may correspond to a comparator included in the PAFE Unit.
[0957] The artificial neural network may include a plurality of neurons, and at least one of the outputs of the plurality of neurons may be pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0958] The artificial neural network may include a plurality of neurons, and the number of the plurality of neurons may be less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
[0959] The examples of the present disclosure disclosed in this specification and the accompanying drawings are presented only as specific examples to facilitate the explanation of the technical content of the present disclosure and to help understand the present disclosure, and are not intended to limit the scope of the present disclosure. It is obvious to those skilled in the art that in addition to the examples disclosed herein, other modified examples based on the technical spirit of the present invention may also be implemented.
[0960] [National R&D project supporting this invention]
[0961] [Task Number] 1711195792
[0962] [Task Number] 00228938
[0963] [Department Name] Ministry of Science and ICT
[0964] [Project Management (Professional) Organization Name] Korea Information and Communication Planning and Evaluation Institute
[0965] [Research Project Name] Development of AI Semiconductor SW Integration Platform Technology
[0966] [Research Project Name] Commercial Edge AI SoC Semiconductor SW Development Platform Technology Development
[0967] [Contribution rate] 1 / 1
[0968] [Project implementation organization name] DeepX
[0969] [Study period] 2023.04.01~2023.12.31.
Claims
1. An activation function conversion program unit, which is configured to approximate a target activation function into a programmed activation function through machine learning of an artificial neural network.
2. The activation function conversion program unit according to claim 1, in, The artificial neural network comprises a first layer including a plurality of neurons and a second layer including a plurality of neurons, wherein a rectified linear unit (ReLU) function is applied to the outputs of the plurality of neurons of the first layer, and The value of the ReLU function is applied as input to the multiple neurons of the second layer.
3. The activation function conversion program unit according to claim 1, in, The artificial neural network comprises a first layer including a plurality of neurons and a second layer including a plurality of neurons, wherein each of the plurality of neurons in the first layer comprises a weight and a bias, and Wherein, the plurality of neurons in the second layer include only weights.
4. The activation function conversion program unit according to claim 1, in, The artificial neural network is allowed to perform machine learning so as to minimize the error between the target activation function and the programmed activation function.
5. The activation function conversion program unit according to claim 1, in, The programmed activation function includes a plurality of segments including a programmable segment implemented in the form of a first-order function.
6. The activation function conversion program unit according to claim 1, in, The artificial neural network comprises a plurality of neurons, and The programmed activation function includes a plurality of programmable segments, and the plurality of programmable segments are respectively separated by breakpoints of outputs of the plurality of neurons.
7. The activation function conversion program unit according to claim 1, in, The programmed activation function includes a plurality of programmable segments, The number of the plurality of programmable segments corresponds to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and The hardware information corresponds to the comparator included in the PAFE Unit.
8. The activation function conversion program unit according to claim 1, in, The artificial neural network comprises a plurality of neurons, and Wherein, at least one of the outputs of the plurality of neurons is pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
9. The activation function conversion program unit according to claim 1, in, The artificial neural network comprises a plurality of neurons, and Wherein, the number of the plurality of neurons is less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
10. An activation function programming method, comprising: Set the target activation function; Approximating the target activation function to a programmed activation function by subjecting an artificial neural network to machine learning; as well as The programmed activation function is converted to slope and offset and stored in a lookup table.
11. The activation function programming method according to claim 10, The artificial neural network comprises a first layer including a plurality of neurons and a second layer including a plurality of neurons, wherein a rectified linear unit (ReLU) function is applied to the outputs of the plurality of neurons of the first layer, and The values to which the ReLU function is applied are inputs of the multiple neurons of the second layer.
12. The activation function programming method according to claim 10, The artificial neural network comprises a first layer including a plurality of neurons and a second layer including a plurality of neurons, wherein each of the plurality of neurons of the first layer comprises a weight and a bias, and The plurality of neurons in the second layer include only weights.
13. The activation function programming method according to claim 10, in, The artificial neural network is allowed to perform machine learning so as to minimize the error between the target activation function and the programmed activation function.
14. The activation function programming method according to claim 10, in, The programmed activation function includes a plurality of segments including a programmable segment implemented in the form of a first-order function.
15. The activation function programming method according to claim 10, in, The artificial neural network comprises a plurality of neurons, and The programmed activation function includes a plurality of programmable segments, and the plurality of programmable segments are respectively separated by breakpoints of outputs of the plurality of neurons.
16. The activation function programming method according to claim 10, in, The programmed activation function includes a plurality of programmable segments, The number of the plurality of programmable segments corresponds to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function, and The hardware information corresponds to the comparator included in the PAFE Unit.
17. The activation function programming method according to claim 10, in, The artificial neural network comprises a plurality of neurons, and Wherein, the number of the plurality of neurons is less than or equal to the number of comparators included in a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
18. The activation function programming method according to claim 10, in, The artificial neural network comprises a plurality of neurons, and Wherein, at least one of the outputs of the plurality of neurons is pruned according to hardware information of a programmed activation function execution unit (PAFE Unit) that executes the programmed activation function.
Citation Information
Patent Citations
Barium titanate powder and its manufacturing method, and filler for sealing material
KR1020230153485A