Electronic device and control method thereof
Patent Information
- Application Number
- CN202580018646.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-22
- Filing Date
- 2025-02-26
- Publication Date
- 2026-09-29
AI Technical Summary
然而,当使用一个LUT处理多个非线性激活函数时,或者当使用一个LUT处理对整个区段的输入值的运算时,可能存在对准确度的限制
Smart Images

Figure CN122847713A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to an electronic device and a control method for the electronic device, and more specifically, to an electronic device and a control method thereof capable of accelerating the computation of a neural network model. Background Technology
[0002] Recently, research on neural network accelerators for efficiently processing neural network model operations has attracted attention.
[0003] Linear and nonlinear activation functions can be used to perform operations through layers of a neural network model. Therefore, a neural network accelerator can include a matrix operation accelerator (which may be called a multiply-accumulate (MAC) array) for accelerating operations based on linear activation functions, and activation units for accelerating operations based on nonlinear activation functions.
[0004] Here, operations based on non-linear activation functions may be more complex than operations based on linear activation functions. Therefore, when non-linear activation functions are treated as integer operations, the efficiency of neural network accelerators can be significantly improved.
[0005] One approach involves performing a polynomial approximation on the nonlinear activation function. However, this approach can be limited because it significantly reduces accuracy. Furthermore, to address this accuracy reduction, fine-tuning through retraining the neural network model should be performed, or the model should be trained from scratch using the approximation function, potentially requiring multiple iterations.
[0006] Alternatively, a lookup table (LUT) may be used to process non-linear activation functions. However, when using a single LUT to process multiple non-linear activation functions, or when using a single LUT to process input values over an entire segment, there may be limitations on accuracy. Summary of the Invention
[0007] [Technical Solution]
[0008] An electronic device and its control method are provided that can accurately and effectively process operations on nonlinear activation functions associated with neural network models.
[0009] Other aspects will be set forth in part in the description which follows, and in part will be apparent from the description, or may be learned by practice of the embodiments presented.
[0010] According to one aspect of this disclosure, an electronic device includes: at least one processor configured to accelerate at least one operation of a neural network model; and a memory configured to store data associated with the neural network model and instructions, the instructions which, when executed by the at least one processor, cause the electronic device to: identify a range of input values corresponding to each of a plurality of layers associated with a nonlinear activation function of the neural network model; divide the range of input values into a plurality of segments having a first predetermined number for each layer based on the range of input values; quantize the input data corresponding to the neural network model based on the plurality of segments; obtain a first lookup table (LUT) for each layer, the first lookup table including a plurality of integer input values corresponding to the plurality of segments and a plurality of integer output values corresponding to the plurality of integer input values; and process the operation according to the nonlinear activation function into integer operations based on the first LUT.
[0011] The plurality of layers may include a first layer corresponding to a first range divided into a first plurality of segments, and a second layer corresponding to a second range divided into a second plurality of segments, and when an instruction is executed by at least one processor, the instruction may also cause the electronic device to determine that the size of the first plurality of segments is greater than the size of the second plurality of segments based on the fact that the first range is wider than the second range.
[0012] When the instruction is executed by the at least one processor, the electronic device may also: divide at least one of the plurality of segments into a plurality of detailed segments having a second predetermined number of segments based on an error in determining that at least one of the plurality of segments has an output value greater than or equal to a threshold; obtain a second LUT, the second LUT including a plurality of detailed integer input values and a plurality of detailed integer output values corresponding to the plurality of detailed segments; and process the operation into integer operation based on the first LUT and the second LUT.
[0013] When the instruction is executed by at least one processor, the instruction may also cause the electronic device to: identify the segment corresponding to the first integer output value among the plurality of segments as the at least one segment based on the average error between the first integer output value and the real number output value corresponding to the first integer output value being greater than or equal to a threshold.
[0014] The second predetermined quantity of multiple detailed segments can be less than the first predetermined quantity of multiple segments.
[0015] The input data can be quantized to sixteen bits, the multiple integer input values and multiple integer output values included in the first LUT can be quantized to eight bits, and the multiple detailed integer input values and multiple detailed integer output values included in the second LUT can be quantized to four bits.
[0016] When executed by at least one processor, the instructions may also cause the electronic device to: identify a first segment corresponding to the first input value and a second segment following the first segment using a first LUT based on the received first input value; obtain an integer output value corresponding to the real number input value by performing linear interpolation on a first integer output value corresponding to the first segment and a second integer output value corresponding to the second segment among the plurality of integer output values; and obtain a real number output value corresponding to the real number input value by converting the integer output value to a real number based on parameters used for quantization.
[0017] The multiple integer input values and multiple integer output values included in the first LUT can be quantized based on the values corresponding to the high eight bits of the sixteen bits, and when the instruction is executed by at least one processor, the instruction can also cause the electronic device to use the values corresponding to the low eight bits of the sixteen bits as weights to perform linear interpolation on the first integer output value and the second integer output value.
[0018] When the instructions are executed by at least one processor, the instructions may also cause the electronic device to: divide the nonlinear activation function into multiple operations based on the fact that the nonlinear activation function is an activation function in which each input value is affected by different input values, obtain a first LUT for performing at least one nonlinear operation among the multiple operations, and process the at least one nonlinear operation into integer operations based on the first LUT.
[0019] According to one aspect of this disclosure, a method for controlling an electronic device includes: identifying a range of input values corresponding to each of a plurality of layers associated with a nonlinear activation function of a neural network model; dividing the range of input values into a plurality of segments based on the range of input values, the plurality of segments having a first predetermined number for each layer; obtaining a first lookup table (LUT) for each layer by quantizing input data corresponding to the neural network model based on the plurality of segments, the first lookup table including a plurality of integer input values corresponding to the plurality of segments and a plurality of integer output values corresponding to the plurality of integer input values; and processing operations according to the nonlinear activation function into integer operations based on the first LUT.
[0020] The plurality of layers may include a first layer corresponding to a first range divided into a first plurality of segments, and a second layer corresponding to a second range divided into a second plurality of segments, and the method may further include: determining that the size of the first plurality of segments is greater than the size of the second plurality of segments based on the fact that the first range is wider than the second range.
[0021] The method may further include: dividing at least one of the multiple segments into a plurality of detailed segments having a second predetermined number, based on determining that at least one segment has an output value error greater than or equal to a threshold; obtaining a second LUT, the second LUT may include a plurality of detailed integer input values and a plurality of detailed integer output values corresponding to the plurality of detailed segments; and processing the operation into integer operations based on the first LUT and the second LUT.
[0022] Dividing at least one segment into multiple detailed segments may include: identifying the segment corresponding to the first integer output value as at least one segment based on the average error between the first integer output value and the real number output value corresponding to the first integer output value being greater than or equal to a threshold.
[0023] The second predetermined quantity of multiple detailed segments can be less than the first predetermined quantity of multiple segments.
[0024] The input data can be quantized to sixteen bits, the multiple integer input values and multiple integer output values included in the first LUT can be quantized to eight bits, and the multiple detailed integer input values and multiple detailed integer output values included in the second LUT can be quantized to four bits. Attached Figure Description
[0025] The above and other aspects, features, and advantages of certain embodiments of this disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, wherein:
[0026] Figure 1 This is a block diagram schematically illustrating the configuration of an electronic device according to one or more embodiments of the present disclosure;
[0027] Figure 2 This is a flowchart describing an embodiment related to generating a second lookup table (LUT) for multiple detail segments;
[0028] Figure 3 This is a graph used to describe an embodiment related to obtaining a second LUT for multiple detailed segments;
[0029] Figure 4 This is a diagram illustrating one or more embodiments related to dividing a nonlinear activation function into multiple operations;
[0030] Figure 5 It is a block diagram illustrating in detail the configuration of an electronic device according to one or more embodiments of the present disclosure; and
[0031] Figure 6 This is a flowchart illustrating a control method for an electronic device according to one or more embodiments of the present disclosure. Detailed Implementation
[0032] Because this disclosure can be modified in various ways and has several embodiments, some specific embodiments of this disclosure are shown in the accompanying drawings and described in detail in the specific embodiments. However, it should be understood that this disclosure is not limited to these specific embodiments, but includes various modifications, equivalents, and / or substitutions of embodiments according to this disclosure. Throughout the drawings, similar components may be represented by similar reference numerals.
[0033] In describing this disclosure, detailed descriptions of certain functions or configurations related to this disclosure may be omitted if it is determined that such detailed descriptions may unnecessarily obscure the main points of this disclosure.
[0034] Furthermore, the specific embodiments described below can be modified in many different ways, and the scope and spirit of this disclosure are not limited to these specific embodiments. Rather, the following description is provided to convey the technical spirit of this disclosure to those skilled in the art.
[0035] The terminology used in this disclosure is for describing particular embodiments only and is not intended to limit the scope of this disclosure. Unless the context clearly indicates otherwise, the singular form includes the plural form.
[0036] In the specification, expressions such as “have,” “may have,” “include,” and “may include” may be used to indicate the presence of a corresponding feature (e.g., numerical value, function, operation, component such as a part), and do not exclude the presence of additional features.
[0037] In this disclosure, expressions such as “A or B”, “at least one of A and / or B”, and “one or more of A and / or B” can include all possible combinations of the items listed together. For example, expressions such as “A or B”, “at least one of A and B”, and “at least one of A or B” can indicate all of the following: 1) the case that includes at least one A, 2) the case that includes at least one B, and 3) the case that includes both at least one A and at least one B.
[0038] The terms “first” and “second” used in this disclosure may refer to various components regardless of the order and / or importance of the components, and may be used only to distinguish one component from other components, without limiting the corresponding components.
[0039] When a component (e.g., the first component) is described as being (operably or communicatively) coupled to or connected to another component (e.g., the second component), it should be understood that the component may be directly coupled to the other component, or may be coupled to the other component through a different component (e.g., the third component).
[0040] However, when any component (e.g., the first component) is described as being “directly coupled” or “directly connected” to another component (e.g., the second component), it will be understood that there is no different component (e.g., the third component) between that component and the other component.
[0041] Depending on the context, the expression “configured (or set) to” as used in this disclosure may be replaced by the expressions “suitable for,” “capable of,” “designed to,” “suitable for,” “manufactured to,” or “capable of.” The term “configured (or set) to” may not necessarily mean “specifically designed to” in the hardware.
[0042] Conversely, in some cases, the phrase "a device configured to..." can mean that the device can "do" something together with other devices or components. For example, "a processor configured (or set) to perform A, B, and C" can refer to a dedicated processor (e.g., an embedded processor) for performing the corresponding operations or a general-purpose processor (e.g., a central processing unit (CPU) or application processor) that can perform the corresponding operations by executing one or more software programs stored in a memory device.
[0043] In embodiments, modules or elements ending with a suffix such as "er / or" (which may be referred to as "-er / or" elements) can perform at least one function or operation and can be implemented by hardware or software or a combination of hardware and software. Furthermore, aside from modules or "-er / or" elements that require specific hardware implementation, multiple modules or multiple "-er / or" elements can be integrated into at least one module and implemented by at least one processor.
[0044] According to one or more embodiments, various elements and areas in the accompanying drawings may be schematically illustrated. Therefore, the scope of this disclosure is not limited to the relative sizes or spacing shown in the drawings.
[0045] The embodiments of this disclosure will be described in detail below with reference to the accompanying drawings, so that those skilled in the art to which this disclosure pertains can more easily practice this disclosure.
[0046] Figure 1 This is a block diagram schematically illustrating the configuration of an electronic device 100 according to one or more embodiments of the present disclosure.
[0047] like Figure 1 As shown, the electronic device 100 according to this disclosure may include a memory 110 and a processor 120.
[0048] At least one instruction relating to the electronic device 100 may be stored in the memory 110. The memory 110 may store an operating system (O / S) for driving the electronic device 100. Additionally, the memory 110 may store various software programs or applications for operating the electronic device 100 according to various embodiments of the present disclosure. Furthermore, the memory 110 may include at least one of a semiconductor memory such as flash memory and a magnetic storage medium such as a hard disk.
[0049] For example, various software modules for operating the electronic device 100 according to various embodiments of the present disclosure can be stored in the memory 110, and the processor 120 can execute the various software modules stored in the memory 110 to control the operation of the electronic device 100. That is, the memory 110 can be accessed by the processor 120, and the reading, recording, correction, deletion, updating, etc. of data in the memory 110 can be performed by the processor 120.
[0050] In this disclosure, the term "memory 110" may be used to refer to at least one of the memory 110 as described above, the read-only memory (ROM) in the processor 120, the random access memory (RAM), and the memory card (e.g., a micro secure digital (SD) card and a memory stick) installed in the electronic device 100.
[0051] In one or more embodiments, memory 110 may store data for a neural network model. For example, memory 110 may store data including various parameters of the neural network model, such as layers, activation functions, and weights. Furthermore, memory 110 may store various data for accelerating the execution of the neural network model. For example, memory 110 may store data for lookup tables (LUTs) according to this disclosure, as well as data for the quantization process (or calibration process) of the data for the neural network model.
[0052] In addition, various information within the scope of achieving the purposes of this disclosure can be stored in memory 110, and the information stored in memory 110 can be updated when received from an external device or by user input.
[0053] The processor 120 can control the overall operation of the electronic device 100. For example, the processor 120 can be connected to the configuration of the electronic device 100 including the memory 110, and can execute at least one instruction stored in the memory 110 as described above to control the overall operation of the electronic device 100.
[0054] Processor 120 can be implemented in various ways. For example, processor 120 can be implemented by at least one of application-specific integrated circuit (ASIC), embedded processor, microprocessor, hardware control logic, hardware finite state machine (FSM), and digital signal processor (DSP), or may include at least one of them. In this disclosure, the term "processor 120" can be used to refer to at least one of central processing unit (CPU), graphics processing unit (GPU), microprocessor unit (MPU), etc.
[0055] In one or more embodiments, processor 120 can perform integer operations corresponding to the nonlinear activation functions of the neural network model. Various embodiments implemented using processor 120 are described below.
[0056] Processor 120 can accelerate the computation of neural network models. Processor 120 can process at least some of the computations performed by the neural network model into integer operations, thereby improving the computational speed of the neural network model and reducing power consumption. In other words, processor 120 can be or can include hardware designed to efficiently perform computations of neural network models, and can be referred to by terms such as neural network accelerator, neural processing unit (NPU), and artificial intelligence (AI) accelerator.
[0057] The following describes various embodiments implemented by processor 120 to accelerate the operation of neural network models.
[0058] The processor 120 can identify the range of input values for each of the multiple layers associated with the nonlinear activation function of the neural network model.
[0059] According to the embodiments, the operations performed by the neural network model can be divided into linear operations and nonlinear operations. For example, the neural network model can perform linear operations such as transforming input values using a weight matrix, and can also perform nonlinear operations using various types of nonlinear activation functions, examples of which are described below.
[0060] An activation function can be a function that determines whether to activate the output values of previous layers in a neural network model and generate output values. For example, in a feedforward process where input values are sent from the input layer to the output layer while simultaneously acquiring the output values, the activation function can determine whether to send the input values passed from the previous layer to the next layer, and if so, whether to convert the input values into any output values.
[0061] A non-linear activation function is a function that performs non-linear operations as the input values passed from the previous layer are sent to the next layer. When using non-linear activation functions, non-linearity can be added to the neural network model, thus enabling more sophisticated implementations of the neural network model.
[0062] Examples of nonlinear activation functions may include a variety of functions such as the sigmoid function, the rectified linear unit (ReLU) function, the exponential linear unit (ELU) function, the Gaussian error linear unit (GELU) function, and the Swish function, but the nonlinear activation functions according to this disclosure are not limited to a particular type of activation function.
[0063] The processor 120 can acquire information about the structure of the multiple layers included in the neural network model, the operations performed by each of the multiple layers, and the types of activation functions associated with each of the multiple layers, based on data about the neural network model. Specifically, the processor 120 can identify at least one non-linear activation function from the various activation functions included in the neural network model.
[0064] When at least one non-linear activation function is identified, processor 120 can identify the range of input values for each of the multiple layers associated with the non-linear activation function. For example, processor 120 can identify the range of input values for each of the multiple layers by allowing the neural network model to input at least some input values included in the training or validation data of the neural network model, or by allowing the neural network model to input at least some input values included in sample data having a distribution similar to the training or validation data.
[0065] Here, the range of input values can refer to the difference between the maximum and minimum values of the input values fed into each of the multiple layers. The range of input values can be different for each of the multiple layers. For example, the range of input values for the first layer in a neural network model could be from the value "1" to the value "25600", and the range of input values for the second layer in the neural network model could be from the value "1" to the value "256". Since the input value of a particular layer may be the output value of the previous layer, the range of input values for each of the multiple layers can include at least some of the range of output values for each of the multiple layers. Therefore, the term range of input values can be referred to as the range of input / output values.
[0066] The processor 120 can divide the entire segment encompassed within the range of input values identified for each of the multiple layers into a predetermined number (which may be referred to as a preset number) of segments for each of the multiple layers.
[0067] Here, the predetermined number can refer to the number of segments, indicating how many segments are divided into which the entire range of input values is divided. In embodiments, the entire range of input values can refer to, for example, the entire range of input values. The predetermined number can be determined based on the number of bits (e.g., eight bits, sixteen bits, etc.) set in the quantization process described below, and can be changed by the developer or the user.
[0068] The range of input values identified for each of the multiple layers may differ, but the predetermined number of segments used to divide the entire range of input values can be the same. Therefore, the entire range of input values for each of two different layers can be divided into multiple segments of different sizes. For example, processor 120 can identify or determine multiple segments of different sizes for each of the multiple layers based on the range of input values identified for each of the multiple layers.
[0069] For example, when the range of input values corresponding to the first layer is greater than the range of input values corresponding to the second layer, the processor 120 can determine that the size of the multiple segments corresponding to the first layer is greater than the size of the multiple segments corresponding to the second layer.
[0070] For example, when the predetermined number of segments to be distinguished is 256, if the range of the input values of the first layer is from the value "1" to the value "25600", and the range of the input values of the second layer is from the value "1" to the value "256", then the processor 120 can divide the entire segment included in the range of input values of the first layer into 256 segments of size "1", and divide the entire segment included in the range of input values of the second layer into 256 segments of size "100".
[0071] The processor 120 can quantize the input data of the neural network model based on multiple segments divided by multiple layers, thereby obtaining a first LUT for each of the multiple layers, the first LUT including multiple integer input values corresponding to each segment of the multiple segments and multiple integer output values corresponding to each segment of the multiple integer input values.
[0072] Quantization, or quantization, can refer to the process of converting data represented in relatively high-precision units into data with relatively low precision. For example, quantization according to this disclosure can refer to the process of converting data represented as real numbers in the first bit range into data represented as integers in the second bit range, which is smaller than the first bit range. For example, when quantization is performed on data, real weight data represented in 32-bit floating-point (FP32) format can be converted into integer weight data represented in eight bits.
[0073] In this disclosure, the input value before quantization can be referred to as a real number input value, and the output value before quantization can be referred to as a real number output value. Furthermore, the input value after quantization can be referred to as an integer input value, and the output value after quantization can be referred to as an integer output value.
[0074] For example, processor 120 can identify or determine the scale and zero point that may be parameters used for quantization, and perform quantization on the input value using the identified scale and the identified zero point.
[0075] A scale can refer to a parameter that indicates the conversion ratio between actual floating-point values (e.g., real numbers) and integer values. For example, a scale can be calculated by dividing the difference between the maximum and minimum values included in the range of input values by a predetermined number. For example, a scale can indicate the size of each of the multiple segments identified as described above.
[0076] When the range of input values includes negative numbers, the zero point can refer to the parameter that represents an integer value when the input value is "0". The zero point can be calculated by rounding the result of dividing the value obtained by multiplying the minimum value included in the range of input values by "-1" by a proportional value.
[0077] Once the scale and zero point are calculated, the integer input value based on the quantization result can be calculated by rounding the result of dividing the real number input value by the scale and adding the zero point to the rounded value.
[0078] For example, when performing 8-bit quantization, if the input value ranges from "-3.5" to "3.2", the value "0.0263", obtained by dividing the difference by "255", can be calculated as the scale. Then, the value "133" can be calculated as the zero point, obtained by rounding the result of multiplying the minimum input value by -1 by the calculated scale. Then, when the real number input value is "1.5", the result of dividing the real number input value by the scale is rounded, and the value "190", obtained by adding the zero point to the rounded value, can be calculated as the integer input value corresponding to the real number input value "1.5".
[0079] As described above, when acquiring multiple integer input values corresponding to each of the multiple segments, the processor 120 can acquire multiple output values corresponding to each of the multiple integer input values by rounding the result of inputting the multiple integer input values into the nonlinear activation function.
[0080] Additionally, the processor 120 can perform the quantization process described above for each of the multiple layers and obtain a LUT based on the quantization result. For example, the processor 120 can obtain a first LUT for each of the multiple layers, which includes multiple integer input values and multiple integer output values corresponding to each of the multiple integer input values, and store the first LUT in the memory 110.
[0081] Here, LUT can refer to a table that stores pre-computed values, allowing for quick referencing of pre-computed results for a specific input value. Specifically, according to this disclosure, a LUT can refer to a table that stores a plurality of integer input values and a plurality of integer output values corresponding to those integer input values. Any dataset that includes a plurality of integer input values, a plurality of integer output values, and information about the correspondence between the plurality of integer input values and the plurality of integer output values can correspond to a LUT according to this disclosure, regardless of the terminology used to refer to such a dataset. For example, the term LUT can be replaced by terms such as dictionary or dataset.
[0082] A first LUT can refer to a LUT that includes multiple integer input values encompassed within the entire range of input values, and multiple integer output values corresponding to each of the multiple integer input values. The term "first LUT" is used to distinguish it from a second LUT, which includes only information about some segments (e.g., detailed segments) of the entire range of input values. See reference. Figure 2 and Figure 3 Describe the meaning of the second LUT in detail and provide an example of its acquisition and processing.
[0083] Processor 120 can process operations based on at least one nonlinear activation function into integer operations using a first LUT. For example, when acquiring multiple first LUTs corresponding to each of multiple layers associated with a nonlinear activation function, processor 120 can store the multiple first LUTs in memory 110. Subsequently, when input values are input to the neural network model, processor 120 can use the multiple first LUTs stored in memory 110 to process at least some of the operations performed by the neural network model into integer operations.
[0084] For example, when a first input value is input, the processor 120 can identify or select a first segment and a second segment corresponding to the first input value based on a first LUT, wherein the second segment is a segment following the first segment (e.g., a segment after the first segment). When identifying or selecting the first segment and the second segment, the processor 120 can obtain an integer output value corresponding to the real number input value by linearly interpolating a first integer output value corresponding to the first segment and a second integer output value corresponding to the second segment among a plurality of integer output values.
[0085] For example, when the first input value is the real number "0.1", the processor 120 can determine that the segment corresponding to the first input value "0.1" is the first segment among a plurality of segments. Additionally, the processor 120 can determine that "2" is the first integer output value corresponding to the first segment among the plurality of segments, and that "3" is the second integer output value corresponding to the second segment, which is the segment following the first segment. The processor can obtain the integer output value corresponding to the first input value by performing linear interpolation on the value "2" (e.g., the first integer output value) and the value "3" (e.g., the second integer output value).
[0086] Here, information about the lower bits rather than the higher bits used in the first LUT can be used as weights for linear interpolation between the first integer output value and the second integer output value. For example, the multiple integer input values and multiple integer output values included in the first LUT can be quantized based on values corresponding to the higher eight bits of the sixteen-bit LUT. In this case, the processor 120 can perform linear interpolation on the first integer output value and the second integer output value by using the values corresponding to the lower eight bits of the sixteen-bit LUT as weights.
[0087] When acquiring an integer output value, processor 120 can acquire a real number output value corresponding to a real number input value by converting the integer output value to a real number based on the parameters used for quantization. However, when the input of the next layer is an integer value, processor 120 can input the integer output value of that layer to the next layer without converting the integer output value to a real number.
[0088] The above describes an example of identifying or selecting a first segment corresponding to a first input value and a second segment as the next segment after the first segment based on a first LUT. However, the processor 120 may also select a first segment corresponding to the first input value and a third segment as the preceding segment of the first segment, and obtain an integer output value corresponding to the real number input value by performing linear interpolation on a first integer output value corresponding to the first segment and a third integer output value corresponding to the third segment among a plurality of integer output values.
[0089] The above describes an example of obtaining integer output values using linear interpolation. However, in addition to linear interpolation, various interpolation methods such as polynomial interpolation can be used.
[0090] Based on the above example, the electronic device 100 can obtain LUTs with different precisions according to the range of input values of each of the multiple layers, and based on this, accurately and efficiently process integer operations corresponding to the nonlinear activation functions of the neural network model.
[0091] For example, electronic device 100 can use parameters obtained from the quantization (or calibration) process of data from a neural network model to generate a LUT for treating nonlinear activation functions as integer operations without a separate retraining or iteration process.
[0092] Furthermore, the electronic device 100 can generate a LUT suitable for the range of input values of each of the multiple layers, instead of using a single LUT to process multiple nonlinear activation functions, thereby enabling a high-precision integer arithmetic accelerator even under constraints of hardware size and limited resources.
[0093] For example, even if a LUT can accurately handle operations at the first level where the input values range from "1" to "25600", accuracy loss may be unavoidable when using a LUT to handle operations at the second level where the input values range from "1" to "256". However, according to this disclosure, a different LUT can be provided for each layer based on the range of input values identified for each of the multiple layers.
[0094] The above effect can be further maximized when operations are performed on neural network models (e.g., transformers) that include segments with large variability in output values depending on changes in input values (e.g., activation functions with large nonlinearity (e.g., 1 / x, rsqrt(x))).
[0095] Figure 2 This is a flowchart describing an embodiment related to generating a second LUT for multiple detail sections. Figure 3 It is a graph used to describe an embodiment related to obtaining a second LUT for multiple detailed segments.
[0096] refer to Figure 2 In operation S210, processor 120 can obtain a first LUT for each of the multiple layers. As described above, processor 120 can identify or determine the range of input values for each of the multiple layers associated with the nonlinear activation function of the neural network model, and processor 120 can divide the entire segment included in the range of input values into multiple segments for each of the multiple layers based on the range of input values identified for each of the multiple layers, and can obtain a first LUT including multiple integer input values corresponding to each of the multiple segments and multiple integer output values corresponding to the multiple integer input values of each of the multiple layers by quantizing the input data of the neural network model based on the multiple segments divided for each of the multiple layers.
[0097] Based on the acquired first LUT, the processor 120 can identify or determine whether there exists at least one segment among multiple segments whose output value error is greater than or equal to a threshold. For example, based on the fact that the average error between a first integer output value and the corresponding real number output value among multiple integer output values is greater than or equal to a threshold, the processor 120 can identify or select the segment corresponding to the first integer output value among multiple segments as at least one segment. Hereinafter, a segment among multiple segments whose output value error is greater than or equal to a threshold can be referred to as at least one segment.
[0098] Here, the average error can be one of various average errors, or it can be determined using one of various average errors, such as the mean absolute error (MAE) which indicates the average of the absolute differences between integer output values and real output values, the mean square error (MSE) which indicates the average of the squared differences between integer output values and real output values, and the root mean square error (RMSE) which indicates the square root of the mean square error.
[0099] For example, Figure 3 The graph of the function y = 1 / x is shown, which is an example of a non-linear activation function, where x represents the real input value and y represents the real output value. It includes several segments, including regions where the real output value has high variability according to changes in the real input value (such as...). Figure 3 A portion of region 310 can be represented by a wide range of output values using a relatively narrow range of input values, so the average error can be greater than or equal to the threshold.
[0100] However, in several segments, including regions where the variability of the real-number output value is small according to changes in the real-number input value (such as... Figure 3 The portion in region 320 represents a narrow range of output values with a relatively wide range of input values, so the average error may be less than the threshold.
[0101] Therefore, in Figure 3 In the example, processor 120 can identify or select items included in... Figure 3 The segment in region 310 is considered as at least one segment, and may be either unrecognized or selectively included. Figure 3 The segment in region 320 is considered as at least one segment.
[0102] When at least one segment among multiple segments whose output value error is greater than or equal to a threshold is not identified (N at operation S220), processor 120 can process the operation according to at least one nonlinear activation function into integer operations based on the first LUT at operation S230. For example, when the nonlinear activation function does not include a part with large nonlinearity, processor 120 can, as referenced... Figure 1 and Figure 2The first LUT described herein will process the operation based on at least one nonlinear activation function as an integer operation, without performing the additional processing described below.
[0103] When it is determined that the output value error of at least one of the multiple segments is greater than or equal to a threshold (e.g., when identifying or selecting the aforementioned at least one segment) (Y in operation S220), the processor 120 may divide the at least one segment into a predetermined number of detailed segments in operation S240. Then, in operation S250, the processor 120 may obtain a second LUT, which includes a plurality of detailed integer input values and a plurality of detailed integer output values corresponding to each of the plurality of detailed segments.
[0104] For example, processor 120 may divide an entire segment encompassing a range of input values into multiple segments, and then further divide at least one of the multiple divided segments (e.g., at least one segment identified as having an output value error greater than or equal to a threshold) into multiple detailed segments. For example, processor 120 may divide each of the at least one segment (such as those included in...) into multiple detailed segments. Figure 3 The area 310 is divided into multiple detailed segments, each with an average error greater than or equal to a threshold.
[0105] When dividing into multiple detailed segments, the processor 120 can obtain multiple detailed integer input values and multiple detailed integer output values corresponding to each of the multiple detailed segments, and can obtain a second LUT that includes the multiple detailed integer input values and multiple detailed integer output values of each of the multiple detailed segments.
[0106] In one or more embodiments, the number of multiple detail segments may be less than the number of multiple segments. For example, when the input data is quantized to sixteen bits, and the multiple integer input values and multiple integer output values included in the first LUT are quantized to eight bits, the multiple detailed integer input values and multiple detailed integer output values included in the second LUT may be quantized to four bits.
[0107] For example, because a smaller number of quantization bits may be sufficient to generate a LUT by re-dividing a specific segment compared to generating a LUT for the entire segment, the number of multiple detailed segments can be less than the number of multiple segments. However, the embodiments are not limited to this.
[0108] In operation S260, processor 120 can process the operation according to at least one non-linear activation function into integer operations based on a first LUT and a second LUT. For example, processor 120 can identify or determine whether an input value corresponds to at least one segment. Then, processor 120 can use the second LUT for the input value corresponding to at least one segment to process the operation according to at least one non-linear activation function into integer operations. However, processor 120 can use the first LUT to process the operation according to at least one non-linear activation function into integer operations for input values that do not correspond to at least one segment.
[0109] In the above example, one embodiment is described, wherein when at least one segment of the entire segment of the input value is identified or selected where the error of the output value is greater than or equal to a threshold, the operation of the first LUT and the second LUT are processed into integer operations according to the operation of at least one nonlinear activation function. However, it is also possible to replace at least one segment of the first LUT where the output value error is greater than or equal to the threshold with the second LUT to obtain a single LUT, and process the integer operations of the entire segment with a single LUT.
[0110] Based on the above reference Figure 3 and Figure 4 In the described embodiment, the electronic device 100 can obtain a first LUT with different precision based on the input value range of each of the multiple layers, and when the nonlinear activation function includes a portion of nonlinearity greater than or equal to a threshold level, the electronic device 100 can process more accurate integer operations for the portion of nonlinearity greater than or equal to the threshold by using a second LUT that is more detailed than the first LUT.
[0111] For example, in the case of a neural network model (e.g., a transformer) that includes activation functions (e.g., 1 / x, rsqrt(x)) with highly nonlinear components, the following problem may occur: in some parts (e.g., including...) Figure 3 In the section of region 320, high-precision integer operations can be performed using only the first LUT, but in the part where the output value is highly variable depending on the input value (e.g., including...), Figure 3 In a segment of region 310, it may be difficult to perform high-precision integer operations using only the first LUT. According to the above embodiment, this problem can be solved by using a detailed second LUT together with the first LUT.
[0112] Figure 4 This is a diagram illustrating one or more embodiments related to dividing a nonlinear activation function into multiple operations.
[0113] According to one or more embodiments, processor 120 can divide a nonlinear activation function into multiple operations and apply a LUT to the nonlinear operations among the multiple operations.
[0114] Processor 120 can identify or determine whether a nonlinear activation function is one in which each input value is influenced by other input values. For example, processor 120 can identify activation functions defined as producing an output value by using each input value together with other input values, such as the softmax function and layer normalization function in a neural network model.
[0115] For example, Figure 4 The graph illustrates the softmax function, a type of activation function. The softmax function applies an exponential function to each element of a given input vector (e.g., the input value) and normalizes the result to convert the exponential function into a probability distribution. For example, for a given input vector z = [z1, z2, ..., zn], it can be done as follows: Figure 4 The softmax function is defined as shown in equation (a). Here, e zi It can represent the value of an exponential function, and the softmax function can return a normalized probability value by dividing each input value by the sum of all the exponential function values.
[0116] When the nonlinear activation function is one in which each input value is affected by another input value, the processor 120 can divide the nonlinear activation function into multiple operations. For example, the processor 120 can use a polynomial approximation to divide the nonlinear activation function into multiple operations. Here, dividing the nonlinear activation function into multiple operations can mean using a polynomial approximation to divide the operations included in the nonlinear activation function into simpler operations.
[0117] For example, Taylor series expansion can be used to... Figure 4 The softmax function is divided into multiple polynomials. When the exponential function e of the softmax function... zi When expanded into a Taylor series, it can be like Figure 4 It can be expressed as in equation (b). Then, when the exponential function of each input value (zi) is approximated by an nth-order Taylor polynomial and then normalized, it can be expressed as follows: Figure 4 Equation (c) is shown.
[0118] Processor 120 can acquire a first LUT for performing at least one nonlinear operation among a plurality of operations. Processor 120 can process the at least one nonlinear operation into an integer operation based on the first LUT.
[0119] For example, when a nonlinear activation function is divided into multiple operations, at least some of the multiple operations can be linear operations, while others can be nonlinear operations. The processor 120 can acquire a first LUT for each of at least one function to process at least one function performing a nonlinear operation into an integer operation, and use the acquired first LUT to process at least one nonlinear operation into an integer operation.
[0120] For example, according to Figure 4 The polynomial approximation of equation (c) can be divided into z i 2 Nonlinear operations and linear operations such as summation are performed. Processor 120 can acquire data including z... i 2 The first LUT used to perform nonlinear operations.
[0121] After using the first LUT to process the nonlinear operation into an integer operation, the processor 120 can perform linear operations to finally obtain the output value of the softmax function corresponding to the input value.
[0122] Based on the above reference Figure 4 The described example shows that even for neural network models (e.g., transformers) that include complex activation functions (e.g., softmax functions) with both linear and nonlinear characteristics, the electronic device 100 can perform accurate and efficient integer operations.
[0123] Figure 5 This is a block diagram illustrating in detail the configuration of an electronic device 100 according to one or more embodiments of the present disclosure.
[0124] like Figure 5 As shown, in addition to the memory 110 and processor 120, the electronic device 100 according to one or more embodiments of the present disclosure may also include a communication unit 130, an input unit 140, and an output unit 150. However, Figure 1 and Figure 5 The configuration shown is merely exemplary, and in performing this disclosure, in addition to Figure 1 and Figure 5 In addition to the configuration shown, you can add new configurations or omit some configurations.
[0125] The communication unit 130 may include circuitry and may perform communication with external devices. For example, the processor 120 may receive various data or information from external devices connected via the communication unit 130 and may send various data or information to external devices.
[0126] The communication unit 130 may include at least one of a WiFi module, a Bluetooth module, a wireless communication module, an NFC module, and an ultra-wideband (UWB) module. For example, the WiFi module and the Bluetooth module may each perform communication using WiFi and Bluetooth methods, respectively. When using a WiFi module or a Bluetooth module, various connection information such as SSID is first sent and received, communication is established using the connection information, and then various information can be sent and received.
[0127] In addition, the wireless communication module can perform communication according to various communication protocols such as the Institute of Electrical and Electronics Engineers (IEEE), Zigbee, 3G, 3GPP, LTE, and 5G. The NFC module can perform near-field communication (NFC) using the 13.56MHz band from various radio frequency identification (RFID) bands such as 135kHz, 13.56MHz, 433MHz, 860-960MHz, and 2.45GHz. Furthermore, the UWB module can accurately measure the time of arrival (ToA) (which can be the time it takes for a pulse to reach the target) and angle of arrival (AoA) (which is the angle at which the pulse reaches the transmitting device) through communication between UWB antennas, thus enabling accurate distance and location identification within an error range of tens of centimeters indoors.
[0128] In one or more embodiments, the processor 120 may control the communication unit 130 to send data from an external device regarding various parameters (including layers, activation functions, and weights of the neural network model), various data for performing neural network model acceleration, data regarding the LUT according to this disclosure, data regarding the quantization process (or calibration process) of the neural network model data, etc.
[0129] The input unit 140 may include circuitry, and the processor 120 may receive user commands for controlling the operations of the electronic device 100 through the input unit 140. For example, the input unit 140 may be configured to include components such as a microphone, camera, and remote control signal receiving unit. The input unit 140 may be a touchscreen and may be implemented as an integrated display. Specifically, the microphone may receive voice signals and convert the received voice signals into electrical signals.
[0130] In one or more embodiments, the processor 120 may receive user input for acquiring the LUT, user input for setting the number of multiple segments, and user input for providing information about the acquired LUT (e.g., providing information about the LUT via a display, providing information about the LUT to a user terminal) through the input unit 140.
[0131] The output unit 150 includes circuitry, and the processor 120 can output various functions that the electronic device 100 can perform through the output unit 150. Additionally, the output unit 150 may include at least one of a display, a speaker, and an indicator.
[0132] The display can output video data under the control of the processor 120. For example, the display can output video pre-stored in the memory 110 under the control of the processor 120. Specifically, the display according to one or more embodiments of the present disclosure can display a user interface stored in the memory 110. The display can be implemented as a liquid crystal display panel (LCD), an organic light-emitting diode (OLED), etc., and in some cases, the display can also be implemented as a flexible display, a transparent display, etc. However, the display according to the present disclosure is not limited to a particular type.
[0133] The speaker can output audio data under the control of the processor 120. The indicator can be turned on under the control of the processor 120. For example, the indicator can be turned on with various colors under the control of the processor 120. The indicator can be implemented as a light-emitting diode (LED), a liquid crystal display panel (LCD), a vacuum fluorescent display (VFD), etc., but is not limited to these.
[0134] In one or more embodiments, the processor 120 may control the display to show information about the acquired LUT.
[0135] Figure 6 This is a flowchart of a method for controlling an electronic device 100 according to one or more embodiments of the present disclosure.
[0136] refer to Figure 6 In operation S610, electronic device 100 can identify or determine the range of input values for each of a plurality of layers associated with a nonlinear activation function of a neural network model. Electronic device 100 can identify or determine at least one nonlinear activation function among various activation functions included in the neural network model based on data about the neural network model. When identifying or determining at least one nonlinear activation function, electronic device 100 can identify or determine the range of input values for each of the plurality of layers associated with the nonlinear activation function.
[0137] Electronic device 100 can identify or determine the range of input values for each of a plurality of layers by allowing the neural network model to input at least some input values included in the training or validation data of the neural network model, or by allowing the neural network model to input at least some input values included in sample data having a distribution similar to the training or validation data.
[0138] In operation S620, based on the range of input values, the entire segment encompassed within the input value range of each of the multiple layers can be divided into a predetermined number of segments. The range of input values identified for each of the multiple layers may differ, but the predetermined number of segments used to divide the entire segment encompassed within the range of input values may be the same. Therefore, the entire segment of the input values for each of two different layers can be divided into multiple segments of different sizes.
[0139] Electronic device 100 can quantize the input data of a neural network model based on multiple segments divided by multiple layers, thereby obtaining a first LUT for each of the multiple layers in operation S630. The first LUT includes multiple integer input values corresponding to each of the multiple segments and multiple integer output values corresponding to each of the multiple integer input values. For example, electronic device 100 can obtain a first LUT including multiple integer input values and multiple integer output values corresponding to each of the multiple integer input values for each of the multiple layers, and store the first LUT in memory 110.
[0140] In operation S640, electronic device 100 can process operations based on at least one nonlinear activation function into integer operations using a first LUT. For example, when acquiring multiple first LUTs corresponding to each of multiple layers associated with a nonlinear activation function, electronic device 100 can store the multiple first LUTs in electronic device 100. Subsequently, when input values are input to a neural network model, electronic device 100 can use the multiple first LUTs stored in electronic device 100 to process at least some of the operations performed by the neural network model into integer operations.
[0141] The control method of the electronic device 100 according to the above embodiments can be implemented as a program and provided to the electronic device 100. Specifically, the program, which includes the control method of the electronic device 100, can be provided by storing it in a non-transitory computer-readable medium.
[0142] For example, in a non-transitory computer-readable recording medium including a program for performing a control method of electronic device 100, the control method of electronic device 100 may include: identifying or determining a range of input values for each of a plurality of layers associated with a nonlinear activation function of a neural network model; dividing an entire segment of the range of input values for each of the plurality of layers into a predetermined number of segments based on the range of input values; obtaining a first LUT for each of the plurality of layers by quantizing the input data of the neural network model based on the plurality of segments, the first LUT including a plurality of integer input values corresponding to each of the plurality of segments and a plurality of integer output values corresponding to each of the plurality of integer input values; and processing the operation according to at least one nonlinear activation function into integer operations based on the first LUT.
[0143] In the above description, examples of control methods for electronic device 100 and computer-readable recording media including programs for executing control methods of electronic device 100 have been briefly described, but this is merely to omit redundant descriptions, and it goes without saying that various embodiments of electronic device 100 are also applicable to computer-readable recording media including control methods for electronic device 100 and programs for executing control methods of electronic device 100.
[0144] The processor 120 and memory 110 of the electronic device 100 can be used to compute or perform artificial intelligence-related functions according to this disclosure.
[0145] Processor 120 may include one or more processors. In this case, one or more processors 120 may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and a neural processing unit (NPU), but are not limited to the examples of processor 120 described above.
[0146] The CPU can be a general-purpose processor 120, capable of performing not only general arithmetic but also artificial intelligence arithmetic, and can efficiently execute complex programs through a multi-layered cache structure. The CPU may be advantageous for serial processing methods, allowing for organic connections between the results of previous and next operations through sequential computation. The general-purpose processor 120 is not limited to the examples described above.
[0147] GPUs can be used as processors 120 for large-scale computations (such as floating-point operations used in graphics processing), and can perform large-scale computations in parallel by integrating a large number of cores. In particular, GPUs may be more advantageous than CPUs in parallel processing methods such as convolution operations. Furthermore, GPUs can be used as coprocessors 120 to supplement the functionality of CPUs. The processors 120 used for large-scale computations are not limited to the examples described above.
[0148] An NPU is a processor 120 dedicated to artificial intelligence operations using artificial neural networks, and each layer included in the artificial neural network can be implemented in hardware (e.g., silicon). In this case, the NPU can be specifically designed to meet specific requirements, thus it can have fewer degrees of freedom than a CPU or GPU, but can efficiently handle artificial intelligence operations corresponding to the requirements. According to one or more embodiments, as a processor 120 dedicated to artificial intelligence operations, the NPU can be implemented in various forms, such as a tensor processing unit (TPU), an intelligent processing unit (IPU), and a vision processing unit (VPU). The artificial intelligence processor 120 is not limited to the examples described above.
[0149] Alternatively, one or more processors 120 may be implemented as a system-on-a-chip (SoC). In this case, in addition to one or more processors 120, the SoC may also include memory 110 and network interfaces, such as a bus for data communication between the processors 120 and memory 110.
[0150] When the SoC included in the electronic device 100 includes multiple processors 120, the electronic device 100 can use some of the multiple processors 120 to perform artificial intelligence-related operations (e.g., artificial intelligence operations related to model learning or inference). For example, the electronic device 100 can use at least one of the multiple processors 120, such as a GPU, NPU, VPU, TPU, or a hardware accelerator dedicated to artificial intelligence operations (e.g., convolution operations and matrix multiplication operations), to perform artificial intelligence-related operations. However, this is merely an example, and it goes without saying that a general-purpose processor 120, such as a CPU, can be used to handle artificial intelligence-related operations.
[0151] Furthermore, the electronic device 100 can use multiple cores (e.g., dual-core, quad-core, etc.) included in a processor 120 to perform operations related to artificial intelligence. Specifically, the electronic device 100 can use the multiple cores included in the processor 120 to perform artificial intelligence operations in parallel, such as convolution operations and matrix multiplication operations.
[0152] One or more processors 120 perform control to process input data according to predefined operational rules or artificial intelligence models stored in memory 110. The predefined operational rules or AI models are characterized by being created through training.
[0153] Here, "creating through learning" can mean creating an artificial intelligence model of predefined motion rules or desired characteristics by applying a learning algorithm to multiple training datasets. This training can be performed on the device itself that executes the AI according to this disclosure, or it can be performed via a separate server / system.
[0154] AI models may include multiple neural network layers. At least one layer may have at least one weight value and may perform layer computations based on the computation results of previous layers and at least one defined computation. Examples of neural networks may include convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), deep Q-networks, and transformers, and the neural networks in this disclosure are not limited to the examples described above, except as specified.
[0155] A learning algorithm can be a method of training a predetermined target device (e.g., a robot) using a large amount of training data, enabling the predetermined target device to make decisions or predictions on its own. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, and unless explicitly stated otherwise, the learning algorithms in this disclosure are not limited to the examples described above.
[0156] Machine-readable storage media may be provided in the form of non-transitory storage media. Here, non-transitory storage media can mean that the storage media is a tangible device, not just a temporary signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently on the storage media and cases where data is temporarily stored on the storage media. For example, non-transitory storage media may include buffers for temporarily storing data.
[0157] According to one or more embodiments, methods according to various embodiments disclosed in this document may be included and provided in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., an optical disc read-only memory (CD-ROM)) or may be distributed through an app store (e.g., the Play Store). TM The computer program product (e.g., download or upload) can be distributed online or directly between two user devices (e.g., smartphones). In the case of online distribution, at least some of the computer program product (e.g., downloadable application) can be temporarily stored in a machine-readable storage medium, such as the memory 110 of a manufacturer's server, an app store server, or a relay server, or can be temporarily generated.
[0158] Each component (e.g., module or program) according to the various embodiments of this disclosure as described above may include a single entity or multiple entities, and some of the corresponding sub-components described above may be omitted, or other sub-components may be further included in the various embodiments. Alternatively or additionally, some components (e.g., modules or programs) may be integrated into one entity and may perform functions in the same or similar manner that were performed by the respective corresponding components prior to integration.
[0159] Operations performed by modules, programs or other components according to different embodiments may be performed sequentially, in parallel, iteratively or heuristically, and at least some operations may be performed in a different order or omitted, or other operations may be added.
[0160] According to one or more embodiments, the terms "unit" or "module" as used in this disclosure may include units configured by hardware, software, or firmware, and may be used compatiblely with terms such as, for example, logic, logic block, component, circuit, etc. An "device / or" element or "module" may be a component configured as a whole or a minimum unit performing one or more functions, or a portion thereof. For example, a module may be configured by an application-specific integrated circuit (ASIC).
[0161] The embodiments of this disclosure can be implemented by software including instructions stored in a machine-readable storage medium (e.g., a computer-readable storage medium). The machine can be a device that invokes the stored instructions from the storage medium and can perform operations according to the invoked instructions, and may include an electronic device (e.g., electronic device 100) according to the disclosed embodiments.
[0162] When a processor executes a command, the processor can directly perform the function corresponding to the command, or other components can perform the function corresponding to the command under the control of the processor. The command may include code created or executed by a compiler or interpreter.
[0163] In the following description and illustration, exemplary embodiments of the present disclosure are illustrated and described. However, the present disclosure is not limited to the specific exemplary embodiments described above, and various modifications can be made by those skilled in the art to which this disclosure pertains without departing from the spirit of the present disclosure as disclosed in the appended claims. Such modifications should also be understood to fall within the scope and spirit of the present disclosure.
Claims
1. An electronic device, comprising: At least one processor is configured to accelerate at least one operation of the neural network model; and The memory is configured to store data associated with the neural network model and instructions, which, when executed by the at least one processor, cause the electronic device to: Identify the range of input values corresponding to each of the multiple layers in a neural network model, which is associated with the nonlinear activation function. The range of input values is divided into multiple segments with a first predetermined number for each layer, based on the range of input values. By quantizing the input data corresponding to the neural network model based on the multiple segments, a first lookup table (LUT) is obtained for each layer. The first lookup table (LUT) includes multiple integer input values corresponding to the multiple segments and multiple integer output values corresponding to the multiple integer input values. The first LUT will process the operations based on the nonlinear activation function into integer operations.
2. The electronic device according to claim 1, wherein, The plurality of layers includes a first layer corresponding to a first range divided into a first plurality of segments, and a second layer corresponding to a second range divided into a second plurality of segments, and When the instruction is executed by the at least one processor, it also causes the electronic device to determine that the size of the first plurality of segments is greater than the size of the second plurality of segments based on the fact that the first range is wider than the second range.
3. The electronic device according to claim 1, wherein, When executed by the at least one processor, the instruction also causes the electronic device to: Based on the error in determining that at least one segment has an output value greater than or equal to a threshold, the at least one segment among the plurality of segments is divided into a plurality of detailed segments having a second predetermined number. Obtain a second LUT, the second LUT including multiple detailed integer input values and multiple detailed integer output values corresponding to the multiple detailed segments, and The operations are processed as integer operations based on the first LUT and the second LUT.
4. The electronic device according to claim 3, wherein, When executed by the at least one processor, the instruction also causes the electronic device to: Based on the fact that the average error between the first integer output value and the real number output value corresponding to the first integer output value is greater than or equal to a threshold, the segment corresponding to the first integer output value among the plurality of segments is identified as the at least one segment.
5. The electronic device according to claim 3, wherein, The second predetermined number of the plurality of detailed segments is less than the first predetermined number of the plurality of segments.
6. The electronic device according to claim 5, wherein, The input data is quantized to sixteen bits. The plurality of integer input values and the plurality of integer output values included in the first LUT are quantized to eight bits, and The plurality of detailed integer input values and the plurality of detailed integer output values included in the second LUT are quantized to four bits.
7. The electronic device according to claim 6, wherein, When executed by the at least one processor, the instruction also causes the electronic device to: Based on the received first input value, the first LUT is used to identify the first segment corresponding to the first input value and the second segment following the first segment. The integer output value corresponding to the real number input value is obtained by linear interpolation of the first integer output value corresponding to the first segment and the second integer output value corresponding to the second segment among the plurality of integer output values, and The real number output value corresponding to the real number input value is obtained by converting the integer output value to a real number based on the parameters used for quantization.
8. The electronic device according to claim 7, wherein, The plurality of integer input values and the plurality of integer output values included in the first LUT are quantized based on values corresponding to the high eight bits of the hexadecimal bits, and When the instruction is executed by the at least one processor, it also causes the electronic device to use the value corresponding to the lower eight bits of the sixteen bits as a weight to perform linear interpolation on the first integer output value and the second integer output value.
9. The electronic device according to claim 1, wherein, When executed by the at least one processor, the instruction also causes the electronic device to: Since a nonlinear activation function is an activation function where each input value is affected by different input values, the nonlinear activation function is divided into multiple operations. Obtain a first LUT for performing at least one nonlinear operation among the plurality of operations, and The at least one nonlinear operation is processed into an integer operation based on the first LUT.
10. A method for controlling an electronic device, the method comprising: Identify the range of input values corresponding to each of the multiple layers associated with the nonlinear activation function of the neural network model; Based on the range of input values, the range of input values is divided into multiple segments with a first predetermined number for each layer; By quantizing the input data corresponding to the neural network model based on the multiple segments, a first lookup table (LUT) is obtained for each layer. The first lookup table includes multiple integer input values corresponding to the multiple segments and multiple integer output values corresponding to the multiple integer input values. as well as The first LUT will process the operations based on the nonlinear activation function into integer operations.
11. The method according to claim 10, wherein, The plurality of layers includes a first layer corresponding to a first range divided into a first plurality of segments, and a second layer corresponding to a second range divided into a second plurality of segments, and The method further includes: The size of the first plurality of segments is determined to be greater than the size of the second plurality of segments based on the fact that the first range is wider than the second range.
12. The method of claim 10, further comprising: Based on the error in determining that at least one segment has an output value greater than or equal to a threshold, the at least one segment among the plurality of segments is divided into a plurality of detailed segments having a second predetermined number; Obtain a second LUT, which includes multiple detailed integer input values and multiple detailed integer output values corresponding to the multiple detailed segments; as well as The operations are processed as integer operations based on the first LUT and the second LUT.
13. The method according to claim 12, wherein, Dividing the at least one segment into the plurality of detailed segments includes: Based on the fact that the average error between the first integer output value and the real number output value corresponding to the first integer output value is greater than or equal to a threshold, the segment corresponding to the first integer output value among the plurality of segments is identified as the at least one segment.
14. The method according to claim 12, wherein, The second predetermined number of the plurality of detailed segments is less than the first predetermined number of the plurality of segments.
15. The method according to claim 14, wherein, The input data is quantized to sixteen bits. The plurality of integer input values and the plurality of integer output values included in the first LUT are quantized to eight bits, and The plurality of detailed integer input values and the plurality of detailed integer output values included in the second LUT are quantized to four bits.