Fast and resource efficient approximation of exponential functions

By simplifying the calculation of exponential functions through Taylor expansion and bit shifting operations, the complexity of softmax function calculation is solved, enabling efficient calculation of neural network output in embedded systems and reducing hardware requirements and power consumption.

CN121050686APending Publication Date: 2025-12-02ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510688917.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-29
Filing Date
2025-05-27
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

The computational overhead of the exponential function exponent in the softmax function is large, which makes neural network computation complex and hardware requirements high, making it difficult to execute efficiently in resource-constrained embedded systems.

Method used

The exponential function approximation method using Taylor expansion is adopted, and bit shifting operations are used to replace division, simplifying the calculation of the exponential function. By decomposing the independent variable and approximating with factorial, the hardware circuit requirements are reduced.

Benefits of technology

It significantly reduces computation time and hardware resource requirements, lowers power consumption, and allows the hardware platform to be scaled down while maintaining the accuracy of neural network outputs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121050686A_ABST
    Figure CN121050686A_ABST
Patent Text Reader

Abstract

Fast and resource efficient approximation of exponential functions is provided. A method (100) for calculating an approximation A of an exponential function ex of an argument x, comprising the steps of: approximating (110) ex with a Taylor expansion T around x = 0, said Taylor expansion T comprising a predetermined number n terms, where the i power xi of the argument x is divided by the factorial of the respective i, i = 1,..., n; and in the calculation of each term, approximating (140) the factorial of i to the most recent power of 2, p (i!) ).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to computing exponential functions in a manner that can be performed more efficiently on a computing platform, thereby saving processing time and allowing for a reduction in the size of the computing platform. Background Technology

[0002] When evaluating the output of a neuron in a neural network, the inputs of that particular neuron are aggregated into a weighted sum, and the result is processed into the final output through a non-linear activation function. A very common activation function is the softmax function. In particular, this activation function is placed in layers where probabilities need to be normalized.

[0003] The cost of calling the softmax function very frequently is its high computational overhead. The main reason for this computational complexity is the exponential function e^x. x This requires a large number of floating-point operations or large lookup tables.

[0004] As we all know, e x It is approximately a Taylor series near x = 0, and its form is:

[0005]

[0006] However, this still involves calculating powers of x, factorials, and division. Summary of the Invention

[0007] This invention provides an exponential function e for calculating the independent variable x. x The method for approximating A. This method is based on the known e^(-e ... x Approximately, this approximation includes the i-th power of the independent variable x. i Divide by a predetermined number of n terms of the factorial of the corresponding i, where i = 1, ..., n, such that

[0008]

[0009] In the calculation of each term, the factorial of i is approximately the nearest power of 2, p(i!). That is, 2! = 2 is already a power of 2, 3! = 6 is approximately 4 or 8, 4! = 24 is approximately 16 or 32, and 5! = 120 is approximately 128.

[0010] It has been found that the "next power of 2" approximation brings a surprisingly large saving in complexity because it eliminates the expensive hardware implementation of division. Instead, multiplication or division of powers of 2 can be implemented using a simple bit shift operation, one of the most fundamental computational operations on many hardware architectures, and therefore very fast and requiring a small hardware block. That is, the hardware platform doesn't even need to be equipped with a larger circuit capable of computing division. This elimination of the circuit means that the complete circuitry required to compute the neural network output O fits into a smaller area on the chip. Simultaneously, this reduces power consumption and, more importantly, energy leakage. Of course, the approximation of the nearest power of 2 will introduce some error. However, it has been found, surprisingly, that this introduces only a small error in the expected final result, i.e., the final output O of the neural network. That is, in a neural network that computes many exponents and produces many approximation errors, their effects on the final result at least partially cancel each other out.

[0011] In a further particularly advantageous embodiment, the independent variable x is decomposed into integers X. q Non-integer scaling factor Δ x The product of the product. This scaling factor Δ x Represented as a power of 2, with the exponent being Make In this way, x raised to the power of x i The calculation is simplified to an integer X q integer powers The calculation of the sum and the multiplication by powers of 2. Similarly, this multiplication can be implemented using appropriate bit shifting operations. Due to the exponent... For non-integer terms, discretizing to shifts corresponding to multiplications of integer powers of 2 will introduce some approximation errors. However, this discretization is similar to the "nearest power of 2" approximation of the factorial in the denominator of each term. This means that the final output O of the neural network will tolerate these approximation errors just as well. The only hard work left is... The calculation. However, the hardware implementation for calculating integer powers requires a much smaller structure than the hardware implementation for floating-point calculations.

[0012] In a further particularly advantageous embodiment, The solution is approximately decomposed into The product of the remainder f. Expression Appearing in e x The approximate calculation ends here. Depending on further processing of the results, this expression may also appear in the calculation of e. x The fraction on the other side, and cancels it out.

[0013] For example, the calculated e xAn approximation A can be used for the elements y of an input vector y with m elements. k Calculation of the softmax function:

[0014]

[0015] This softmax function normalizes the m components so that they are all between 0 and 1, and their sum is 1. Here, if we set e... x The calculation is decomposed as described above. f, item Will appear in S(y) k The expression contains both the numerator and denominator. In the denominator, this can be extracted from the sum, therefore in both the numerator and denominator... The two instances cancel each other out. This means that, in a further advantageous embodiment, their calculations can be completely omitted.

[0016] Consider the example where n = 3. Here, the approximate value A is given by the following formula:

[0017]

[0018] Δ x Expressing it as a power of 2, and simultaneously approximating the denominator to the nearest power of 2, we get:

[0019]

[0020] Here, Item It can be dragged to the front to get:

[0021]

[0022] In other words,

[0023]

[0024] Substitute this into S(y) k The expression is obtained as follows:

[0025]

[0026] Here, in the numerator and denominator The two instances cancel each other out, so their calculation can be completely omitted.

[0027] As discussed above, in a further advantageous embodiment, at least one multiplication of a number with a power of 2 and / or division of the number with the power of 2 is performed by bit shifting the number by the number of bits corresponding to the exponent. This eliminates the need for the more complex circuitry originally required to perform the multiplication or division.

[0028] As discussed earlier, when calculating the output O of a neural network, the e given here... x The primary use case for approximation is using e. x The approximate value A is calculated, and / or the value S(y) of the softmax function is calculated. k In particular, when evaluating this output O of a neural network, it is often necessary to e. x The approximation provides a simplification that allows for a more precise understanding of the neural network. Therefore, the time savings resulting from this simplification accumulate to a considerable amount. Furthermore, it eliminates the need for specific structures to perform more complex computational operations. This means that the hardware platform can be scaled down to its required chip area. This is particularly advantageous for applications in embedded systems, such as autonomous driving systems for vehicles or robots, production machines, or quality inspection machines. In these systems, embedded systems running neural networks are often subject to strict size and power constraints.

[0029] In a further particularly advantageous embodiment, e x The approximate value A and / or the calculated value S(y) of the softmax function are calculated. k This refers to the output O of a classifier network used to compute images or other measurement data records, and / or the output O of a multi-head attention module in a transformer network. These network architectures have shown particular resilience to the approximation errors introduced by the approximations presented here. That is, the approximation error cannot affect the final output O of the neural network. In particular, in classification networks, the approximation error cannot switch the class that achieves the highest classification score to another class. Furthermore, these architectures make particularly extensive use of exponential functions, resulting in even more significant overall savings in processing time.

[0030] In many applications of neural networks, not all inputs are equally difficult to process. In particular, in applications involving classification tasks, for some inputs, the final decision becomes clear quickly, while for others, the final decision doesn't become apparent until all layers have processed them. This means that for inputs of varying difficulty, the final output O pairs are determined by e. x The resilience of the approximation error introduced by the approximation can also vary. For less difficult inputs, the power series can be shortened, meaning fewer n terms can be used. The time saved can then be used for more difficult inputs, which may require e terms with higher n. x A more accurate approximation. How difficult a particular input is can be determined, for example, from the confidence level C of the neural network's output O.

[0031] Therefore, in a further particularly advantageous embodiment, the confidence level C of the neural network's output O is determined. In response to this confidence level C that satisfies a predetermined condition, in e xThe number of terms n used in subsequent calculations of the approximation A is modified. In this way, a high number of terms n are used only when truly needed, and less processing power is “wasted” on inputs for which the final result is already clear early in the processing.

[0032] For this purpose, specifically, the number of terms n is controlled to be kept at a minimum value sufficient to achieve the predetermined minimum confidence level C for output O.

[0033] As discussed earlier, the e given here x The lightweight approximation allows for a reduction in the hardware platform required to compute the output O of the neural network. Therefore, in a further particularly advantageous embodiment, the neural network is implemented on a hardware platform that, unlike the non-approximate e, allows for a smaller hardware platform. x Compared to the memory and / or processing resources required to compute output O under the given conditions, this hardware platform has fewer memory and / or fewer processing resources. This applies to both quantitative and qualitative dimensions. Quantitatively, it means that for hardware resources that would require at least one instance regardless of whether the approximation given in this paper is used, fewer instances are required if the approximation is used. Qualitatively, it means that for certain types of circuitry that would be required without using the approximation, no instances need to exist in the hardware platform by using the approximation. This type of qualitative reduction does not slow down the hardware platform, but rather renders it completely unable to perform certain tasks, such as multiplication or division.

[0034] In a further particularly advantageous embodiment, the independent variable x is derived from measurement data acquired using at least one sensor. An actuation signal is calculated based on the output O of the neural network. The actuation signal is used to actuate a vehicle, driver assistance system, robot, quality inspection system, monitoring system, and / or medical imaging system. In this way, the actuation signal can be determined more quickly, and the corresponding actuated technical system requires only a low-power and / or low-capacity embedded system to process the neural network.

[0035] This method can be implemented entirely or partially by a computer and embodied in software. Therefore, the present invention also relates to a computer program having machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method described above. Here, control units for vehicles or robots, and other embedded systems capable of executing machine-readable instructions, are also considered computers. Computing instances include virtual machines, containers, or other execution environments that allow the execution of machine-readable instructions in the cloud.

[0036] Non-transitory storage media and / or downloadable products may include computer programs. A downloadable product is an electronic product that can be sold online and transmitted over a network for immediate fulfillment. One or more computers and / or computing instances may be equipped with the computer programs and / or the non-transitory storage media and / or downloadable products.

[0037] The invention will be described below using accompanying drawings, without any intention to limit the scope of the invention. Attached Figure Description

[0038] Figure 1 : The exponential function e used to calculate the independent variable x x An exemplary embodiment of the method 100 for approximating the value A;

[0039] Figure 2 A simplified illustration based on the approximation of method 100. Detailed Implementation

[0040] Figure 1 It is the exponential function e used to calculate the independent variable x. x A schematic flowchart of an embodiment of the method 100 for approximating the value A.

[0041] According to box 105, the independent variable x can be derived from measurement data acquired using at least one sensor.

[0042] In step 110, e is approximated by a Taylor expansion T around x = 0. x The Taylor expansion T includes a predetermined number of n terms, where the i-th power of the independent variable x is x. i Divide by the factorial of the corresponding i, i = 1, ..., n.

[0043] In step 120, the independent variable x is decomposed into integers X. q Non-integer scaling factor Δ x The product of.

[0044] In step 130, this scaling factor Δ x Represented as a power of 2, with the exponent being Make

[0045] According to box 131, The solution can be approximately decomposed into The product of f and the remaining part.

[0046] In step 140, in the calculation of each term of the Taylor expansion T, the factorial of i is approximately the nearest power of 2, p(i!). The calculation of all terms of the Taylor expansion T yields e. x The approximate value is obtained by solving for it.

[0047] In step 150, the calculated e x The approximation A is used to compute the elements y of an input vector y with m elements. k softmax function

[0048] According to box 151, if A has already been decomposed according to box 131 into The occurrence of S(y) can be omitted. k In the numerator and denominator of ) The calculation of two instances.

[0049] In step 160, the calculated e x The approximate value A and / or the calculated value S(y) of the softmax function. k ) is used to calculate the output O of the neural network.

[0050] Based on box 161, e is calculated. x The approximate value A and / or the calculated value S(y) of the softmax function. k This can be used to calculate the output O of a classifier network for image or other measurement data records, and / or the output O of a multi-head attention module of a transformer network.

[0051] According to box 162, the confidence level C of the neural network's output O can be determined. Then, in box 163, it can be determined whether this confidence level C meets a predetermined condition, such as being higher or lower than a predetermined threshold. If this is the case (truth value 1), then according to box 164, in e x The number of terms n used in subsequent calculations of the approximation A can be modified. In particular, according to box 164a, the number of terms n can be controlled to remain at a minimum value sufficient to achieve a predetermined minimum confidence level C for output O.

[0052] According to box 165, neural networks can be implemented on hardware platforms, similar to those that do not approximate e. x Compared to the memory and / or processing resources required to calculate output O under the same conditions, this hardware platform has less memory and / or fewer processing resources.

[0053] In step 170, the actuation signal 170a is determined from the output O of the neural network.

[0054] In step 180, the vehicle 50, the driver assistance system 51, the robot 60, the quality inspection system 70, the monitoring system 80 and / or the medical imaging system 90 are actuated using the actuation signal 170a.

[0055] Figure 2 The simplification introduced by the approximation according to method 100 is shown.

[0056] Calculate ex The task can be summarized as raising the independent variable x (drawn as a container filled with water) to a given height h. This is difficult because the container filled with water is heavy. This can be solved by factoring the independent variable x into integers X. q (Using a nearly empty container to represent) and non-integer scaling factor Δ x Containers can significantly reduce the workload. Calculate the integer X. q i-th power This corresponds to raising a nearly empty container to a given height h.

[0057] The approximation requires two additional calculations. The scaling factor Δ x The power must be calculated, and a division with the factorial i! is required. The scaling factor Δ x Represented as a power of 2, with the exponent being Make This reduces the impact of the scaling factor Δ x Multiplication or exponentiation simplifies to simple bit shift operations, in Figure 2 The factorial i! is represented as a counterclockwise rotation of a nearly empty container being lifted. The factorial i! is approximately the nearest power of 2, p(i!). This simplifies division by the factorial i! to a bit shift operation in another direction, represented by a clockwise rotation of a nearly empty container being lifted.

[0058] The final result is approximately A, which is close to but not equal to e. x The actual value. This is determined by container A not being in e. x At the full height h and having a higher capacity than container e x Characterized by a low fill level.

Claims

1. An exponential function et for calculating the independent variable x. x The method (100) for approximating the value A includes the following steps: We approximate (110)e using the Taylor expansion T around x = 0. x The Taylor expansion T includes a predetermined number of n terms, where the i-th power of the independent variable x is xi. i Divide by the factorial of the corresponding i, i = 1, ..., n; as well as In the calculation of each term, the factorial of i is approximated (140) as the nearest power of 2, p(i!).

2. The method (100) according to claim 1 further includes: Decompose the independent variable x into integers X (120). q Non-integer scaling factor Δ x The product; And the scaling factor Δ x (130) represents a power of 2, with the exponent being... Make 3. The method (100) according to claim 2, further comprising: Will The solution approximates the decomposition of (131) into The product of the remainder f.

4. The method (100) according to any one of claims 1 to 3, further comprising: The calculated e x The approximation A is used in (150) to compute the elements y of an input vector y with m elements. k softmax function 5. The method (100) according to claims 3 and 4, wherein the omission (151) appears in S(y) k In the numerator and denominator of ) The calculation of two instances.

6. The method (100) according to any one of claims 1 to 5, wherein at least one multiplication of a number with a power of 2 and / or a division of the number with a power of 2 is performed by bit shifting the number by a number of bits corresponding to the exponent.

7. The method (100) according to any one of claims 1 to 6, further comprising: The calculated e x The approximate value A and / or the calculated value S(y) of the softmax function. k ) is used to calculate the output O of the neural network (160).

8. The method (100) according to claim 7, wherein the calculated e x The approximate value A and / or the calculated value S(y) of the softmax function. k (161) is used to calculate the output O of the classifier network for image or other measurement data records, and / or the output O of the multi-head attention module of the transformer network.

9. The method (100) according to any one of claims 7 or 8, further comprising: Determine the confidence level C of the output O of the (162) neural network; as well as In response to the fact that the confidence level C meets the predetermined condition (163), modify (164) in e x The number of terms n used in subsequent calculations of the approximate value A.

10. The method (100) of claim 9, wherein the number of terms n is controlled (164a) to be kept at a minimum value sufficient to achieve a predetermined minimum confidence level C for output O.

11. The method (100) according to any one of claims 7 to 10, wherein the neural network is implemented on a hardware platform (165), and in a non-approximate e x Compared to the memory and / or processing resources required to calculate output O under the condition of value, the hardware platform has less memory and / or fewer processing resources.

12. The method (100) according to any one of claims 7 to 11, wherein the independent variable x is derived (105) from measurement data acquired using at least one sensor, and wherein the method further comprises: The actuation signal (170a) is determined from the output O of the neural network; as well as Actuation signals (170a) are used to actuate (180) a vehicle (50), a driver assistance system (51), a robot (60), a quality inspection system (70), a monitoring system (80), and / or a medical imaging system (90).

13. A computer program comprising machine-readable instructions that, when executed by one or more computers and / or computing instances, cause the one or more computers and / or computing instances to perform the method (100) according to any one of claims 1 to 12.

14. A non-transitory machine-readable storage medium and / or downloadable product having a computer program according to claim 13.

15. One or more computers and / or computing instances having a computer program as claimed in claim 13 and / or a non-transitory machine-readable storage medium and / or downloadable product as claimed in claim 14.