A high-precision softmax hardware implementation method

CN121660011BActive Publication Date: 2026-09-25CLP KESHENTAI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610066286.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-09-25
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

但这些方法存在几种问题,如输入实数有范围限制,否则会导致数据过大溢出或者绝对误差过大导致结果精度不够;消耗了大量面积与时间资源后计算结果精度仍旧不够高;消耗了较多存储资源来存储中间计算结果等

Benefits of technology

(1)通过减去所计算向量元素中的最大值来使指数结果不溢出;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121660011B_ABST
    Figure CN121660011B_ABST
Patent Text Reader

Abstract

The application discloses a high-precision softmax hardware implementation method and belongs to the field of neural network hardware acceleration. The application reduces hardware storage and calculation resource consumption by increasing memory access times and twice look-up table searching. The application prevents overflow of exponential results by subtracting the maximum value in the calculated vector elements. The application separates the exponential and mantissa parts by utilizing the characteristics of the floating-point number format, reduces the numerical range of the input look-up table, uses a small amount of look-up table storage resources, and achieves a high-precision softmax hardware implementation result with an absolute error controlled within 3e-5. The application avoids data conflict and improves processing efficiency by using a pipeline parallel mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of neural network hardware acceleration technology, and in particular to a high-precision softmax hardware implementation method. Background Technology

[0002] Softmax is a function that transforms any real vector into a probability distribution. It belongs to the normalized exponential function family and is commonly used in the output layer of multi-class classification problems, with significant applications in transformer models.

[0003] Softmax requires exponentiation and division operations, which consume significant area and time resources in hardware. Current hardware implementations of softmax computation mainly include polynomial approximation, piecewise linear fitting, direct lookup tables, and approximating the correct result using Newton's iteration method. However, these methods have several problems, such as limitations on the range of input real numbers, which can lead to overflow due to excessively large data or excessive absolute errors resulting in insufficient accuracy; insufficient accuracy despite consuming significant area and time resources; and excessive storage resources for storing intermediate calculation results. Therefore, it is necessary to explore hardware implementations of softmax that offer higher accuracy and consume fewer storage resources. Summary of the Invention

[0004] The purpose of this invention is to provide a high-precision softmax hardware implementation method to solve the problems in the background art.

[0005] To address the aforementioned technical problems, this invention provides a high-precision softmax hardware implementation method. This method utilizes three memory accesses, floating-point processing, and two lookup tables to implement the softmax hardware implementation, and includes the following steps: Step S1: Read the vector elements from memory and obtain the maximum value x. max And store; Step S2, read the vector element x from memory. i and subtract the maximum value x max Obtain the difference; Step S3: Multiply the difference by the constant log2e; Step S4: Process the value obtained in step S3 and separate the integer a and the decimal b; Step S5: Input the decimal b into the small lookup table cell to calculate 2^b; if the decimal b is non-zero, subtract 1 from the integer a and add the floating-point exponent bias value to write it into the first exponent fifo; if the decimal b is 0, add the floating-point exponent bias value to the integer a and write it into the first exponent fifo. Step S6: Concatenate the exponent stored in the first exponent FIFO with the mantissa of the result calculated from the small lookup table to obtain element x. iThe natural index e^(x) i - x max )value; Step S7, repeat steps S2-S6 to calculate the natural exponent e^(x) for each vector element. i – x max The result is then input into the accumulator unit for accumulation, and the accumulation result is processed to separate the exponent and mantissa. Step S8: Input the last digit into the big lookup table cell to calculate 1 / m, where m is e^(x i – x max The sum of the exponents; if the last digit is 1, write the exponent to the second exponent FIFO; if the last digit is not 1, subtract 1 from the exponent and write it to the second exponent FIFO. Step S9: Concatenate the exponents stored in the second exponent FIFO with the mantissas of the results calculated from the large lookup table to obtain the reciprocal of the sum of the natural exponents of all vector elements: 1 / ∑e^(x i - x max ); Step S10, assign the natural exponent e^(x) to each vector element. i - x max The new softmax result vector is obtained by multiplying the inverse of the sum of the natural exponents of all vector elements calculated in step S9.

[0006] In one implementation, the domain of the small lookup table input value in step S5 is [-1, 0], and it includes a uniform sampling point memory for storing the value of the function 2^b. The interval step size step is an integer power of 2, and the sampling point n is written based on the formula 2^(min + step * n), where min is the minimum value of the input range -1.

[0007] In one implementation, the domain of the input value of the large lookup table in step S8 is [1, 2]; it contains more memory than the small lookup table and is used to store the value of the function 1 / x.

[0008] In one embodiment, the small lookup table unit includes 65 sample point memories with a step size of 1 / 64; the large lookup table unit includes 257 sample point memories with a step size of 1 / 256.

[0009] In one implementation, in step S3, log2e is represented using 32-bit binary 00111111101110001010101000111011, and the absolute error between log2e and its exact value is less than 3e-5.

[0010] In one implementation, the domain of integer a in step S4 is (-∞, 0], and the domain of decimal b is (-1, 0).

[0011] In one implementation, the maximum value among the vector elements is obtained by a comparator in step S1.

[0012] In one implementation, the data specifications in steps S1 to S10 are all floating-point numbers.

[0013] The high-precision softmax hardware implementation method provided by this invention has the following beneficial effects: (1) The exponential result is prevented from overflowing by subtracting the maximum value among the elements of the calculated vector; (2) By accessing memory multiple times, the use of large amounts of RAM to store intermediate calculation results is avoided, thus reducing hardware storage resources; (3) By using one lookup table to implement the function 2^x and another lookup table to implement 1 / x, the direct calculation of exponents and division by the hardware is avoided, thus reducing hardware computing resources; (4) By using the format characteristics of floating-point numbers, the domains of the input values ​​of the two lookup tables for the calculation functions 2^x and 1 / x are defined as [-1,0] and [1,2], respectively, so that high-precision calculation can be achieved with a small amount of lookup table storage resources, and the absolute error of the calculation results of the two functions is below 3e-5; (5) The two lookup table units can be computed in parallel, avoiding data conflicts and improving efficiency. Attached Figure Description

[0014] Figure 1 This is a schematic diagram of the hardware architecture of a specific embodiment of the present invention; Figure 2 This is a schematic diagram of the method flow of a specific embodiment of the present invention. Detailed Implementation

[0015] The following detailed description, in conjunction with the accompanying drawings and specific embodiments, provides a further detailed explanation of the high-precision softmax hardware implementation method proposed in this invention. The advantages and features of this invention will become clearer from the following description. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise proportions, and are only used to facilitate and clarify the illustration of the embodiments of this invention.

[0016] This invention provides a high-precision softmax hardware implementation method, based on, for example... Figure 1 The hardware architecture shown includes the following steps S1-S9: Step S1: Read the vector elements from memory and use a comparator to obtain the maximum value x. max After converting the maximum value to fp32 format, the value is stored in a 32-bit wide memory cell. Step S2: Read the vector element from memory again, convert the vector element to fp32 format, and then use the fp32 floating-point subtractor to subtract the maximum value obtained in step S1 from the value. The resulting difference is in the range of (-∞, 0] to prevent the exponent value from overflowing. Step S3: Based on the mathematical transformation e^x = 2^(x * log2e) = 2^(integer a + decimal b) = 2^a * 2^b, the difference obtained in step S2 is multiplied by the constant log2e in fp32 format using an fp32 floating-point multiplier to obtain the product. log2e is represented in 32-bit binary 00111111101110001010101000111011, and the absolute error between log2e and the exact value is less than 3e-5. Step S4: Decompose the product of step S3 into an integer a and a decimal b; Step S5: The domain of integer a is (-∞, 0], and the decimal b is output to a small lookup table unit containing 65 sampling points and a step size of 1 / 64 to calculate 2^b. Since the domain of decimal b is (-1, 0] and the range of 2^b is (0.5, 1], when b is not equal to 0, the integer a-1 is added to the fp32 exponent bias of 127, and this value is written to the FIFO used for data synchronization; when b is equal to 0, the integer a is added to the fp32 exponent bias of 127, and this value is written to the FIFO used for data synchronization. The FIFO depth can be obtained according to the number of processing cycles of the lookup table unit. Step S6, based on step S3, the format 2^a*2^b is consistent with the format of the floating-point number 2^exp (exponent)*1.sig (mantissa), where 'a' is the exponent. According to the description in step S5, the value range of 2^b is (0.5, 1]. By concatenating the exponent of the synchronized FIFO and the mantissa of the floating-point result calculated from the lookup table, element x can be obtained. i The natural index e^(x) i - x max )value; If step S2 finishes reading the vector elements, the second repeated reading can begin, and the process proceeds to step S6; calculate the natural exponent e^(x) of each vector element. i – x max ) and accumulate using an fp32 floating-point adder; Step S7: Process the fp32 format result of the summation in step S6 to separate the exponent and mantissa. Step S8: The mantissa input contains a memory with 257 sampling points and a large lookup table unit with a step size of 1 / 256 to calculate 1 / m, where m is e^(x i – x maxThe sum of ) ; since the domain of the mantissa is in [1, 2), the range of the function 1 / x is (0.5, 1]; if the mantissa is 1, the exponent is inversely inverted and written to the exponent fifo; if the mantissa is not 1, the exponent is inversely inverted and subtracted by 1 and written to the exponent fifo. Step S9: Read the exponent stored in the exponent FIFO and concatenate the mantissas of the results calculated by the large lookup table cell to obtain the reciprocal of the sum of the natural exponents of all vector elements, 1 / ∑e^(x). i - x max ); Step S10, each e^(x) i - x max The reciprocal of the sum of the natural exponents of the vector elements calculated in step S9 is multiplied by the fp32 multiplier to obtain a new softmax result vector.

[0017] The specific embodiments described above further illustrate the purpose and technical solutions of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0018] The above description is merely a description of preferred embodiments of the present invention and is not intended to limit the scope of the present invention in any way. Any changes or modifications made by those skilled in the art based on the above disclosure shall fall within the protection scope of the claims.

Claims

1. A high-precision softmax hardware implementation method, characterized in that, The softmax hardware implementation method uses three memory accesses, floating-point processing, and two lookup tables, and includes the following steps: Step S1: Read the vector elements from memory and obtain the maximum value x. max And store; Step S2, then read the vector element x from memory. i and subtract the maximum value x max Obtain the difference; Step S3: Multiply the difference by the constant log2e; Step S4: Process the value obtained in step S3 and separate the integer a and the decimal b; Step S5: Input the decimal b into the small lookup table cell to calculate 2^b; if the decimal b is non-zero, subtract 1 from the integer a and add the floating-point exponent bias value to write it into the first exponent fifo; if the decimal b is 0, add the floating-point exponent bias value to the integer a and write it into the first exponent fifo. Step S6: Concatenate the exponent stored in the first exponent FIFO with the mantissa of the result calculated from the small lookup table to obtain element x. i The natural index e^(x) i - x max )value; Step S7, repeat steps S2-S6 to calculate the natural exponent e^(x) for each vector element. i – x max The sum is then input into the accumulator unit for accumulation. The fp32 format result of the accumulated sum is processed to separate the exponent and mantissa. Step S8: Input the last digit into the large lookup table cell to calculate 1 / m, where m is e^(x i – x max The sum of the sums of the exponents; the domain of the mantissa m is [1, 2]. If the mantissa is 1, the exponent is written into the second exponent FIFO; if the mantissa is not 1, the exponent is reduced by 1 and then written into the second exponent FIFO. Step S9: Concatenate the exponents stored in the second exponent FIFO with the mantissas of the results calculated from the large lookup table to obtain the reciprocal of the sum of the natural exponents of all vector elements: 1 / ∑e^(x i - x max ); Step S10, assign the natural exponent e^(x) to each vector element. i - x max The new softmax result vector is obtained by multiplying the inverse of the sum of the natural exponents of all vector elements calculated in step S9.

2. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, In step S5, the input value domain of the small lookup table is [-1, 0]. It includes a uniform sampling point memory for storing the value of the function 2^b. The interval step size step is an integer power of 2. The sampling point n is written based on the formula 2^(min + step * n), where min is the minimum value of the input range -1.

3. The high-precision softmax hardware implementation method as described in claim 2, characterized in that, In step S8, the domain of the input value of the large lookup table is [1, 2]; it contains more memory than the small lookup table and is used to store the value of the function 1 / x.

4. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, The small lookup table unit contains 65 sample point memories with a step size of 1 / 64; the large lookup table unit contains 257 sample point memories with a step size of 1 / 256.

5. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, In step S3, log2e is represented using 32-bit binary 00111111101110001010101000111011, and the absolute error between log2e and its exact value is less than 3e-5.

6. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, In step S4, the domain of integer a is (-∞, 0], and the domain of decimal b is (-1, 0).

7. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, In step S1, the maximum value among the vector elements is obtained through a comparator.

8. The high-precision softmax hardware implementation method as described in claim 1, characterized in that, The data specifications in steps S1 to S10 are all floating-point numbers.

Citation Information

Patent Citations

  • Optimization method based on exponential function and softmax function, hardware system and chip

    CN114610267A

  • System and method for accelerating calculation of exponential functions

    CN117897688A