Approximation method of softmax function calculation, and neural network to which same approximation method of softmax function calculation is applied
The approximation method for softmax function calculation in neural networks uses leaky ReLU and low-degree polynomial functions to reduce computational intensity and energy consumption, effectively addressing the inefficiencies of existing methods.
Patent Information
- Application Number
- JP2024006426
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-01-18
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-01-18
AI Technical Summary
Existing softmax function calculations in neural networks are computationally intensive and energy-consuming due to the high degree of the exponential function and the use of float32 format for input values.
An approximation method for softmax function calculation that converts input values into exponential approximation values using a leaky ReLU function and polynomial calculations of low degrees, followed by addition and division operations to obtain the output values.
This method significantly reduces the computational load and energy consumption while maintaining small output errors, thereby shortening calculation time and conserving energy.
Smart Images

Figure 2025084650000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an approximation method for softmax function calculation and a neural network applying this approximation method for softmax function calculation, and particularly to an approximation method for softmax function calculation used in a classification machine of an artificial intelligence deep learning model and a neural network applying this approximation method for softmax function calculation.
Background Art
[0002] Artificial intelligence (AI) is generally a technology that represents human intellectual ability by ordinary computer programs. The most important part of AI is a mathematical model or computational model that mimics the structure and function of a biological neural network in the fields of machine learning and cognitive science, and further an approximate artificial neural network that infers a function. Common neural network models include, for example, convolutional neural networks and recurrent neural networks (CNN and RNN). In recent years, the Transformer model has been developed, and the Transformer model is gradually tending to replace convolutional neural networks and recurrent neural networks (CNN and RNN) and is becoming the most popular deep learning model.
[0003] As shown in FIG. 1, almost all of the above neural models include at least a feature learning machine 81 and a classification machine 82. In the classification machine 82 of the neural model, a softmax function 821 is usually adopted to output at least one output value between 0 and 1.
[0004] The expansion formula of the conventional softmax function is Equation 1.
[0005]
Number
[0006] Generally, the calculation by the softmax function (Equation 1) converts the input value of an i-dimensional vector into an output value of a j-dimensional vector. Each output value of the j-dimensional vector is usually a numerical value between 0 and 1, and the sum of these numerical values is 1.
[0007] Furthermore, most of the softmax function calculations related to commercially available GPUs (e.g., Nvidia) adopt the softmax function of Equation 1, and the input values are in the float32 format to realize the calculation of the softmax function. However, when actually calculating the exponential function, the degree of the polynomial is large, and since the input values adopt the float32 format, it is necessary to process a large amount of numerical calculations in the calculation process of the classification machine, resulting in problems of time and energy consumption. Therefore, it is a very important issue to provide a simple softmax function calculation and to shorten the calculation time and save energy in the calculation process of the neural network classification machine.
Summary of the Invention
Problems to be Solved by the Invention
[0008] An object of the present invention is to provide an approximation method for softmax function calculation that can reduce the calculation time and also reduce the energy consumption. Another object of the present invention is to provide a neural network applying this approximation method for softmax function calculation that can reduce the calculation time and also reduce the energy consumption.
Means for Solving the Problems
[0009] In order to achieve the above object, based on the approximation method of the softmax function calculation of the present invention, the input value of the k-dimensional vector is converted into the output value of the m-dimensional vector. The approximation method of the softmax function calculation obtains the normalized function calculation value by performing the leaky rectified linear unit calculation (leaky ReLU) on the input value of the k-dimensional vector, and then performs the polynomial function calculation of a certain degree based on this normalized function calculation value to obtain the exponential approximation value. After that, by repeating the order with another input value, an exponential approximation calculation order for obtaining another exponential approximation value, an addition calculation order for obtaining the sum value obtained by adding the exponential approximation value and another exponential approximation value, and a division calculation order for obtaining the output value of the m-dimensional vector by dividing at least one of the exponential approximation values obtained in the exponential approximation calculation order by the sum value are provided.
[0010] In one embodiment, in the exponential approximation calculation order, after performing the clamp function calculation on the input value first, the leaky ReLU calculation is further performed.
[0011] In one embodiment, in the addition calculation order, a protection value is further added to the sum value to ensure that the absolute value of the sum value is greater than zero.
[0012] In one embodiment, the polynomial function calculation of a certain degree is a polynomial calculation from the second degree to the fifth degree.
[0013] In one embodiment, the exponential approximation calculation order is repeated until each input value passes through the leaky ReLU calculation to obtain the normalized function calculation value, and then performs the polynomial function calculation of a certain degree based on the normalized function calculation value to obtain the exponential approximation value corresponding to each input value. The addition calculation order obtains the sum value obtained by adding all the exponential approximation values. The division calculation order divides each exponential approximation value of the exponential approximation values obtained in the exponential approximation calculation order by the sum value to obtain the output value of the m-dimensional vector corresponding to the plurality of k-dimensional vectors.
[0014] In one embodiment, the input value is an integer-type numerical value.
[0015] Also, based on a neural network applying the approximation method of the softmax function calculation of the present invention, a softmax function calculation module is provided in the classification machine of the neural network to convert the input value of the k-dimensional vector into the output value of the m-dimensional vector. The softmax function calculation module obtains a normalization function calculation value by performing a leaky rectified linear unit function (LeakyReLU function) calculation on the input value of the k-dimensional vector, and then performs a polynomial function calculation of a certain degree based on this normalization function calculation value to obtain an exponential approximation value. After that, another input value is repeatedly subjected to the leaky rectified linear unit function calculation and the polynomial function calculation of the certain degree to obtain another exponential approximation value. The softmax function calculation module includes an exponential approximation calculation unit, an addition calculation unit, and a division calculation unit. The exponential approximation calculation unit obtains the normalization function calculation value by performing the leaky rectified linear unit function calculation on the input value of the k-dimensional vector, and then performs a polynomial function calculation of a certain degree based on this normalization function calculation value to obtain an exponential approximation value. After that, another input value is repeatedly subjected to the leaky rectified linear unit function calculation and the polynomial function calculation of the certain degree to obtain another exponential approximation value. The addition calculation unit obtains a sum value by adding the exponential approximation value and another exponential approximation value. The division calculation unit obtains the output value of the m-dimensional vector by dividing at least one of the exponential approximation values obtained by the exponential approximation calculation unit by the sum value.
[0016] In another embodiment, in the exponential approximation calculation unit, the input value is first subjected to a clamp function calculation and then a leaky ReLU calculation.
[0017] In another embodiment, in the addition calculation unit, the sum value further adds a protection value to ensure that the absolute value of the sum value is greater than zero.
[0018] In another embodiment, the polynomial function calculation of a certain degree is a polynomial calculation from the second degree to the fifth degree.
[0019] In another embodiment, in the exponential approximation calculation unit, after each input value passes through the leaky ReLU calculation to obtain a normalization function calculation value, the polynomial function calculation of a certain degree is further performed by this normalization function calculation value until the exponential approximation value corresponding to each input value is obtained. The addition calculation unit obtains a sum value by adding all the exponential approximation values. The division calculation unit obtains the output value of the m-dimensional vector corresponding to a plurality of k-dimensional vectors by dividing each exponential approximation value of the exponential approximation values obtained by the exponential approximation calculation unit by the sum value.
[0020] In one embodiment, the input value is an integer numerical value.
Advantages of the Invention
[0021] In the approximation method of the softmax function calculation of the present invention, the exponential function e (xk) is limited to a low-order polynomial (for example, a quadratic polynomial), and a clamp function and a leaky ReLU calculation are adopted. Therefore, when calculating the exponential function e (xk) the calculation result shown by the dashed curve can approximate the result of the high-order calculation shown by the solid curve. In other words, the output error of the softmax function calculation formula of the present invention is all small.
[0022] In particular, since all elements of the input vector are integers and the exponential function e (xk) is limited to a low-order polynomial (for example, a quadratic polynomial), compared with the conventional float32 format and the calculation of high-order polynomials, the amount of calculation can be significantly reduced, so the calculation time can be shortened and the energy consumption can also be suppressed.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Embodiments for Carrying Out the Invention
[0024] Using each drawing, better embodiments related to the approximation method for calculating the softmax function of the present invention and the neural network applying this approximation method for calculating the softmax function will be described below.
[0025] Before specifically describing the embodiments of the present invention, first, it should be noted that in this embodiment, since the softmax function can convert the input value of a k-dimensional vector into the output value of an m-dimensional vector, the softmax function of this embodiment can be expressed as in Equation 2.
[0026]
Number
[0027] In the softmax function calculation of the above formula (2), the most difficult and time-consuming calculation is to calculate the exponential function exp(x) (i.e., e (xk) ) for the input value of the k-dimensional vector. Actually, when calculating the value of e (xk) , generally, it is calculated using the Taylor expansion formula, that is, calculated using the calculation formula of formula (3).
[0028]
Number
[0029]
Number
[0030] Subsequently, as shown in FIG. 2, the exponential function value exp(x) calculated by the high-order polynomial shown in formula (3) is represented by a solid-line curve, and the exponential function value (2nd-taylor exp(x)) calculated by the quadratic Taylor polynomial shown in formula (4) is represented by a dashed-line curve. As shown in FIG. 2, in order to simplify the exponential calculation, if the entire calculation order of the Taylor exponential expansion formula (formula (3)) is reduced to the calculation of a lower-order polynomial, for example, restricted to the calculation of a quadratic polynomial, the exponential function value (2nd-taylor exp(x)) represented by the dashed-line curve as shown in FIG. 2 can be obtained. However, as can be seen from FIG. 2, the difference between the calculation result of the quadratic polynomial (shown by the dashed line) and the calculation result of the high-order polynomial (shown by the solid line) is large.
[0031] Thus, when restricting the exponential function e (xk) to the calculation of a quadratic polynomial in order to simplify the exponential function calculation of the softmax function, in order to make the simplified calculation result approach the calculation result before simplification, it is necessary to improve the calculation method of formula (4) to make its calculation result approach the solid-line curve shown in FIG. 2.
[0032] As shown in FIGS. 3 and 4, the approximation method for softmax function calculation of the present invention includes an exponential approximation calculation order S1, an addition calculation order S2, and a division calculation order S3. The exponential approximation calculation order S1 performs a leaky rectified linear unit function (LeakyReLU function) calculation step S11 and a quadratic polynomial exponential approximation calculation step S12 shown in FIG. 3. Also, as shown in FIG. 4, the output result of the above formula 2 is obtained by these calculation processes.
[0033] Furthermore, as shown in FIG. 3, in this embodiment, leaky ReLU is used in the calculation of formula 4, and the equation of formula 4 can be expressed as the equation of formula 5.
[0034]
Equation
[0035] By using leaky ReLU, the normalization function calculation value L 1 is obtained in step S11, and when the normalization function calculation value L 1 is substituted into formula 5, formula 5 can be expressed as formula 6. By calculating formula 6, the exponential approximation calculation value exp1(X k ) is obtained from the calculation of step S12. That is, by the calculations of step S11 and step S12, the exponential approximation calculation value exp1(X k ) can be calculated from the exponential approximation calculation order S1 shown in FIG. 4. As shown by the dashed curve in FIG. 5, this calculation result restricts the exponential function e (xk) to a quadratic polynomial and adopts leaky ReLU calculation. When X k is negative, the approximate calculation result (shown by the dashed line) can be closer to the solid curve.
[0036]
Equation
[0037] The calculation result of formula 5 or formula 6 is better than the calculation result of formula 4. However, as shown in FIG. 5, when X kWhen it is a negative number or greater than 1, there are certain differences in the approximate calculation results. To solve the problem of certain differences, in a better embodiment of the present invention, as shown in FIGS. 6 and 7, the exponential approximation calculation order S1' is to first input the value X k After performing the clamp function calculation step S10, perform the leaky normalization linear unit function calculation step S11', and further perform the quadratic polynomial exponential approximation calculation step S12'. That is, it is the calculation order as shown in FIG. 6. At this time, the quadratic polynomial exponential approximation calculation formula may be expressed as in Equation 7.
[0038]
Equation
[0039] Using the clamp function and leaky ReLU, the normalization function calculation value L 2 is obtained by the calculation in step S10. Substituting the normalization function calculation value L 2 into Equation 7, Equation 7 can be expressed as Equation 8.
[0040]
Equation
[0041] In particular, if F(X k ) = LeakyReLU(Clamp(X k , min, max)) and the quadratic polynomial coefficients are taken into account simultaneously, the exponential approximation calculation value of the present invention can be expressed by a general formula as in Equation 9.
[0042]
Equation
[0043] In other words, when substituting the normalization function calculation value L 1 or the normalization function calculation value L 2 into Equation 9, Equation 6 can be expressed as Equation 10, and Equation 8 can be expressed as Equation 11.
[0044]
Number
[0045]
Number
[0046] As can be seen from FIG. 8, when the exponential function e (xk) is restricted to a quadratic polynomial and the clamp function calculation and the leaky ReLU calculation are adopted, that is, when calculating using Equation 8 or Equation 11, the calculation result shown by the dashed curve approaches the high-order calculation result shown by the solid curve in the interval where X k is a negative number or greater than 1. Furthermore, according to the exponential calculation of Equation 8 or Equation 11, in the calculation of step S12’, the exponential approximation calculation value exp2(X k ) can be obtained. In other words, in the exponential approximation calculation order S1’ shown in FIG. 7, the exponential approximation calculation value exp2(X k ) can be calculated.
[0047] The following specifically describes the actual calculation of the approximation method of the softmax function calculation of the present invention with reference to FIG. 9. It should be particularly noted that the softmax function calculation of this embodiment (that is, Equation 2) performs the actual calculation by adopting the softmax function calculation layer of the Transformer model.
[0048] Furthermore, as shown in FIG. 9, when the input vector value is [-2, 0, 8], if the exponential function formula adopts Equation 3 and the softmax function calculation formula adopts Equation 2, when each vector value of the input vector value is substituted into Equation 3 for calculation, e (-2) = 0.13, e (0) = 1, e (8)The result of =2980 is obtained. Then, if the results calculated by Equation 3 are substituted into Equation 2 respectively, the calculated value (output value) of the softmax function calculation formula (Equation 2) shown in FIG. 9 is obtained, and the output vector is [0.000045, 0.000335, 0.999619].
[0049] As shown in FIG. 9, when the input vector value is [-2, 0, 8], if the exponential function formula adopts Equation 10 and the softmax function calculation formula adopts Equation 2, the coefficients of Equation 10 are a = 1, b = 2, c = 1. After substituting each vector of the input vector into Equation 10 for calculation respectively, the output vector of the calculated value (output value) of the softmax function calculation formula (Equation 2) is [0.003040, 0.012158, 0.984802]. As can be seen from FIG. 9, for the input vector amount with a relatively large reaction ratio (for example, inputting "8" in this embodiment), the reaction error is within 1-2%.
[0050] However, in the above description, if the exponential function formula adopts Equation 10, the softmax function calculation formula adopts Equation 2, and the coefficients of Equation 10 are a = 1, b = 2, c = 1, when the input vector value is [-4, -4, -4], as shown in FIG. 10, the calculated value (output value) of the softmax function calculation formula (Equation 2) cannot be calculated. The reason for this phenomenon is that the denominator of the softmax function calculation formula (Equation 2) is close to 0.
[0051] To solve the problem that it cannot be calculated normally because the denominator of the softmax function calculation formula may approach 0, please refer to FIG. 7. In this embodiment, the approximation method of the softmax function calculation of the present invention includes an exponential approximation calculation order S1', an addition calculation order S2, and a division calculation order S3. After the addition calculation order S2 is completed, it also includes a protection value calculation order S21 in which the calculated value and the protection value eps are added to each other. Thus, the softmax function calculation formula of the present invention can be expressed as Equation 12. The function of the protection value eps is to ensure that the denominator of Equation 2 does not approach 0 or is not 0. If the denominator of Equation 2 approaches 0 or is 0, the softmax' (xk)m Calculation cannot obtain a result.
[0052]
Number
[0053] Subsequently, as shown in FIG. 11, if the exponential function formula adopts Formula 11 and the softmax function calculation formula adopts Formula 12, and the coefficients of Formula 11 are set as a = 1, b = 2, c = 1, even if the input vector value is [-4, -4, -4], as shown in FIG. 11, the calculated value (output value) of the softmax function calculation formula (Formula 12) can also be calculated. Further, if the exponential function formula adopts Formula 11, the softmax function calculation formula adopts Formula 12, the coefficients shown in Formula 11 are set as a = 1, b = 2, c = 1, and eps shown in Formula 12 is set as 1, even if the input vector value is [-2, 0, 8], as shown in FIG. 12, the calculated value (output value) of the softmax function calculation formula (Formula 12) is calculated as [0.007007, 0.016016, 0.976977]. As can be seen by comparing with the output value 0.984802 (adopting Formula 2) shown in FIG. 12, for the input vector amount with a relatively large reaction ratio (for example, inputting "8" in this embodiment), the reaction error is kept within 1 to 2%.
[0054] In summary, in the approximation method of the softmax function calculation of the present invention, the exponential function e (xk) is limited to a low-order polynomial (for example, a quadratic polynomial), and since the clamp function and the leaky ReLU calculation are adopted, even if the exponential function e (xk) is calculated by either Formula 8 or Formula 11, the calculation result shown by the dashed curve can approximate the result of the high-order calculation shown by the solid curve. Also, in the approximation method of the softmax function calculation of the present invention, if the softmax function calculation formula adopts Formula 12 and appropriate adjustments are made to the coefficients shown in Formula 11 (for example, a = 1, b = 2, c = 1), for the main reaction values, the output error of the softmax function calculation formula of the present invention is all small.
[0055] In particular, in this embodiment, all elements of the input vector are integers, and the exponential function e (xk) is restricted to a low-order polynomial (for example, a quadratic polynomial). Therefore, compared with the conventional float32 format and the calculation of high-order polynomials, the amount of calculation can be significantly reduced, so the calculation time can be shortened and the energy consumption can also be suppressed.
[0056] Another embodiment of the present invention provides a neural network that applies this approximation method of softmax function calculation. Since the specific description of the neural network applying this approximation method of softmax function calculation of the present invention is substantially the same as the foregoing method, it is omitted here. The only thing to be specifically explained is that in another embodiment of the present invention, the neural network is not limited to the neural network of the Transformer model.
[0057] The above are only examples of the present invention and are not restrictive. Any modifications or changes that do not depart from the spirit and scope of the present invention should all belong to the scope of the claims.
Industrial Applicability
[0058] The present invention relates to an approximation method of softmax function calculation used in a classification machine of an artificial intelligence deep learning model and a neural network applying this approximation method of softmax function calculation.
Explanation of Signs
[0059] 81 Feature learning machine 82 Classification machine S1, S1’ Exponential approximation calculation order S10 Clamp function calculation step S11, S11’ Leaky ReLU function calculation step S12, S12’ Quadratic polynomial exponential approximation calculation step S2 Addition calculation order S21 Protection value calculation order S3 Division calculation order L1 normalization function calculated value L2 normalization function calculated value
Claims
1. An approximation method for softmax function calculation that converts an input value of a k-dimensional vector into an output value of an m-dimensional vector, comprising the steps of: an exponential approximation calculation sequence in which a normalization function calculation value is obtained by performing a leaky normalized linear unit function (LeakyReLU function) calculation on an input value of the k-dimensional vector, a polynomial function calculation of a certain degree is performed based on the normalization function calculation value to obtain an exponential approximation value, and another exponential approximation value is obtained by repeating the above sequence on another input value of the k-dimensional vector; an additive calculation sequence in which a sum of the exponential approximation value and the other exponential approximation value is obtained; and a division calculation sequence in which at least one exponential approximation value among the exponential approximations obtained in the exponential approximation calculation sequence is divided by the sum to obtain an output value of the m-dimensional vector.
2. 2. The method of claim 1, wherein in the exponential approximation calculation sequence, a clamp function calculation is performed on an input value first, and then the leaky normalized linear unit function calculation is performed.
3. 3. The method of claim 2, wherein in the additive calculation sequence, the sum further includes a guard value to ensure that the absolute value of the sum is greater than zero.
4. 2. The method of approximating a softmax function according to claim 1, wherein the constant degree polynomial function calculation is a second to fifth degree polynomial calculation.
5. 2. The method of claim 1, wherein the exponential approximation calculation sequence is repeated until each input value undergoes the leaky normalized linear unit function calculation to obtain each normalized function calculation value, and then a certain degree of polynomial function calculation is performed using each normalized function calculation value to obtain an exponential approximation value corresponding to each input value; the additive calculation sequence is repeated until a total value is obtained by adding up all the exponential approximations; and the division calculation sequence is repeated until an output value of an m-dimensional vector corresponding to a plurality of k-dimensional vectors is obtained by dividing all the exponential approximations of the exponential approximations obtained in the exponential approximation calculation sequence by the total value.
6. 2. The method of claim 1, wherein the input values are integer type numerical values.
7. A neural network comprising: A softmax function calculation module is provided for converting the input value of the k-dimensional vector into the output value of the m-dimensional vector in the neural network classification machine; The softmax function calculation module: an exponential approximation calculation unit that obtains a normalized function calculation value by performing a leaky normalized linear unit function calculation on an input value of the k-dimensional vector, and then performs a polynomial function calculation of a fixed degree based on the normalized function calculation value to obtain an exponential approximation value, and then repeats the leaky normalized linear unit function calculation and the polynomial function calculation of the fixed degree on another input value to obtain another exponential approximation value; an addition calculation unit that obtains a sum value by adding the exponential approximation value and the other exponential approximation value; and a division calculation unit that divides at least one of the exponential approximations obtained by the exponential approximation calculation unit by the sum value to obtain an output value of the m-dimensional vector.
8. 8. The neural network according to claim 7, wherein the exponential approximation calculation sequence includes first performing a clamp function calculation on the input values, and then performing the leaky normalized linear unit function calculation.
9. 9. The neural network of claim 8, wherein said sum further includes a guard value to ensure that the absolute value of said sum is greater than zero.
10. 8. The neural network of claim 7, wherein said constant degree polynomial calculation is a second to fifth degree polynomial calculation.
11. The neural network of claim 7, wherein the exponential approximation calculation unit performs the leaky normalized linear unit function calculation for each input value to obtain each normalized function calculation value, and then performs a polynomial function calculation of a certain degree using each normalized function calculation value to obtain an exponential approximation value corresponding to each input value; the addition calculation unit obtains a sum value by adding up all the exponential approximations; and the division calculation unit divides all the exponential approximations of the exponential approximations obtained by the exponential approximation calculation unit by the sum value to obtain an output value of an m-dimensional vector corresponding to a plurality of k-dimensional vectors.
12. 8. The neural network of claim 7, wherein said input values are integer type numeric values.