A Method and System for Hardware Implementation of Softmax on a Platform with Limited Logical Resources

Through function fitting, serial accumulation and exponential operation unit multiplexing, equivalent transformation and cardinal replacement of Softmax functions are solved, and the problems of excessive Softmax function resource consumption and high complexity in embedded systems are achieved, and the efficient and low-complexity Softmax hardware implementation is achieved.

CN115062768BActive Publication Date: 2025-06-10SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210790639.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-05
Publication Date
2025-06-10
Estimated Expiration
2042-07-05

AI Technical Summary

Technical Problem

When implementing Softmax functions in embedded systems, the prior art consumes too much resource and is complex, making it difficult to efficiently calculate.

Method used

Through function fitting, serial accumulation and exponential operation unit multiplexing, equivalent transformation and cardinal replacement of Softmax functions are performed to reduce the complexity and resource consumption of hardware implementation.

Benefits of technology

It effectively reduces the complexity and power consumption cost of Softmax hardware implementation, improves computing efficiency and throughput, and solves the problem of Softmax function implementation on resource-constrained platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062768B_ABST
    Figure CN115062768B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for hardware implementation of Softmax on a platform with limited logical resources, which is applicable to any n input values x1, x2,...., x n , and completes the conversion from scalar to probability. Through function equivalent transformation, replacement of the base of exponentiation and logarithm, function fitting, serial accumulation, and reuse of the exponential operation unit, the present invention uses only a limited number of basic operation logic units to implement complex functions, transforms the combination of power function and division in the original function into the combination of power function and logarithmic function, and simultaneously performs precision-controllable function fitting according to the operation characteristics and data range, saving a large amount of computing time and iterative process, and effectively reducing the hardware implementation area and power consumption cost by using serial accumulation and function unit reuse.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method and system for hardware implementation of Softmax on a platform with limited logical resources, and belongs to the technical field of hardware implementation of a neural network activation function. Background Art

[0002] In the field of big data, deep neural networks (DNNs) have achieved great success, and efficient hardware architectures have always been the pursuit goals of academia and industry. Among them, the Softmax layer is widely used in different DNNs. The Softmax function is usually used as the activation function of the output layer in classification tasks. It maps the outputs of multiple neurons to the interval (0, 1), and it has very wide applications in machine learning and deep learning. Especially when dealing with multi-classification (C>2) problems, the output unit at the end of the classifier requires numerical processing by the Softmax function. Its expression is: The multiplication and division calculations in the Softmax function are quite costly, especially in embedded systems and cannot be directly calculated. Many existing implementation schemes based on lookup tables will introduce excessive resource consumption and have high complexity. Therefore, a reasonable method is needed to deploy the function on hardware to ensure low resource consumption and efficient implementation. Summary of the Invention

[0003] Technical Problem: Aiming at the above problems, the present invention proposes a hardware implementation of Softmax on a platform with limited logical resources, and uses function fitting and serial accumulation to solve the problems of high cost and low efficiency in the hardware implementation of the Softmax function. The technical solution is as follows:

[0004] Technical Solution: The complete technical solution of the present invention is: A method for hardware implementation of Softmax on a platform with limited logical resources, including the steps of:

[0005] 1) Transform the original Softmax function expression as follows:

[0006]

[0007] where n is the total number of inputs of Softmax; i is the index of the corresponding input and output, i = 1,..., n;

[0008] 2) Calculate the exponential function of each gated input x 1 , x 2 ,...., x n where the base e is replaced by the base 2: as follows:

[0009]

[0010] 3) Calculate the sum of the exponential powers of each input x in step 2 1 , x 2 ,...., x n :

[0011]

[0012] 4) Calculate the natural logarithm of the cumulative result in step 3

[0013] f_ln = ln(f)

[0014] 5) Calculate the exponential powers of each selected input x 1 , x 2 ,...., x n added to the calculation result f_ln in step 4 respectively, where the base e is replaced by base 2

[0015]

[0016] 6) Store each calculation result r(i) in step 5 in a register to obtain the final overall output R

[0017] Furthermore, in steps 2 and 5, the same exponentiation module is used for calculation, with the base e replaced by base 2. The exponentiation module consists of two adders, a constant multiplier, and a shifter

[0018]

[0019] u is the integer part of |y.log 2 e|, that is, the integer part of the fixed-point number is intercepted; v is the decimal part of |y.log 2 e|, that is, the decimal part of the fixed-point number is intercepted. The absolute value is reflected in the fixed-point hardware by determining the sign bit. The absolute value of a positive number is the original value, and the absolute value of a negative number is obtained by taking the complement and adding 1. For example, in steps 2 and 5, when calculating or The first adder calculates x i -0 or x i -f_ln; the constant multiplier calculates (x i -f_ln).log 2 e or x i .log 2 e; the integer part of the absolute value is intercepted to obtain u; the second adder implements the fitting function: 2 v ≈v + b 1 , where b 1 is a constant; finally, the calculation result is obtained by shifting 2 v left or right by u

[0020] Further, in step 3), a serial accumulation module is used for calculation, and the accumulation enable signal counter is controlled.

[0021] Further, in step 4), a logarithmic operation module is used for calculation. The base e is replaced with base 2. The logarithmic operation module consists of a leading one detector (LOD), a decoder, a right shifter, a constant adder, and an adder:

[0022] ln(f) = ln2 * log 2 f = ln2 * (w + log 2 k)

[0023] t is an intermediate value where the highest bit of f is 1 and other bits are 0, w is the index where the highest bit of f is located, k is the remaining number after f is scaled, and k ∈ [1, 2). For example, for a 16-bit fixed-point number:

[0024] If f = 16'b0000_1011_1111.0011,

[0025] then t = 16'b0000_1000_0000.0000, w = 4, k = 16'b0000_1.011_1111_0011;

[0026] If f = 16'b0000_0011_1111.0011,

[0027] then t = 16'b0000_0010_0000.0000, w = 6, k = 16'b0000_001.1_1111_0011.

[0028] Further, when calculating ln(f) in step 4), the LOD is used to obtain the intermediate value t where only the highest bit of f is 1 and other bits are 0; the decoder inputs t to obtain the index w where the highest bit of f is located; the right shifter shifts f to the right by w to obtain k, and the constant adder implements the fitting function: log 2 k ≈ k + b 2 , where b 2 is a constant; finally, the constant multiplier and adder calculate ln2 * (w + k + b 2 ).

[0029] Beneficial effects: The method and system for hardware implementation of Softmax on a platform with limited logic resources of the present invention effectively reduce the complexity of hardware implementation of Softmax. Through function equivalent transformation, replacement of the base of multiplication and exponentiation and the base of logarithm, function fitting, serial accumulation, and reuse of the exponential operation unit, complex functions are implemented only with a limited number of basic operation logic units, effectively reducing the hardware implementation area and power consumption cost.

[0030] Compared with the prior art, it greatly reduces the circuit complexity, power consumption and area cost. At the same time, there is no need for parameter storage based on the LUT lookup table method and no requirement for processing time based on the iteration of the CORDIC method, increasing the throughput rate and solving the problem of difficult hardware implementation of Softmax. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is the structural diagram of the hardware implementation of Softmax on the logic resource - limited platform according to the present invention;

[0032] Figure 2 is the schematic diagram of the implementation of the exponential function circuit;

[0033] Figure 3 is the schematic diagram of the implementation of the logarithmic function circuit. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0034] The technical solution of the present invention will be further described below with reference to the drawings and embodiments.

[0035] As Figure 1 shown, for the hardware implementation system of Softmax on the logic resource - limited platform according to the present invention, the input is x 1 , x 2 ,...., x n , and the output is r 1 , r 2 ,...., r n . It includes the steps:

[0036] (1) After inputting the operands x 1 , x 2 ,...., x n , set LnF_en to 0, enable the counter Cnt[$clog2(n)-1:0], and use Cnt[$clog2(n)-1:0] as the gating signal to sequentially gate x 1 , x 2 ,...., x n , and calculate the exponential function of each gated input x 1 , x 2 ,...., x n where the base e is replaced by the base 2:

[0037] (2) Calculate each input x in step (1) 1 , x 2 ,...., x nThe cumulative sum of the exponential powers, set Acc_en to 1, and set Acc_en to 0 when Cnt[$clog2(n)-1:0] is equal to the number of inputs n, completing all x 1 ,x 2 ,....,x n The cumulative sum of the exponential powers.

[0038] (3) Calculate the natural logarithm f_ln = ln(f) of the cumulative result in step (2), where the base e is replaced by the base 2.

[0039] (4) Reset and re-enable the counter Cnt[$clog2(n)-1:0], set LnF_en to 1, and calculate each gated input x 1 ,x 2 ,....,x n The exponential powers obtained by adding each of them to the calculation result f_ln in step (3) respectively, where the base e is replaced by the base 2:

[0040] (5) Store each calculation result r(i) in step (4) in the register to obtain the final overall output R.

[0041] Through the design of the present invention, it meets the requirements of high efficiency, low complexity, low area, and low power consumption of the Softmax hardware circuit in practical application scenarios. Theoretically, a single n-input Softmax calculation can be completed in 2n clock cycles, where n is the number of inputs, while ensuring advantages in terms of area and power consumption.

[0042] As Figure 2 shown, the calculation process of the exponential function circuit:

[0043] (1) Multiply the input x i or x i -f_ln by the fixed-point constant log 2 e to obtain the fixed-point output {u i ,v i}, where u i is the integer part and v i is the fractional part, where -1 < v i < 1.

[0044] (2) Use f(v i ) = v i +b to perform function fitting calculation 2 v , although x i ≥0 but x i -f_ln may have both positive and negative values, so it is necessary to use a two-segment function fitting to select different fitting parameters b i according to the sign of u 1or b 2 , the segmented interval is (-1, 0) and [0, 1). For example, if the input is a 16-bit fixed-point number with a 4-bit fractional width, then take b 1 = 6'b01_0100, b 2 = 6'b00_1111, and the coefficient of determination for fitting in the [0, 1) interval is 0.9919. The coefficient of determination for fitting in the (-1, 0) interval is slightly worse. Through piecewise fitting, in the (-1, 0) interval, use f(v i ) = a * v i + b for fitting to improve accuracy, where a = 0.4966, b = 0.9711. For example, if the input is a 16-bit fixed-point number with a 4-bit fractional width, a * vi can be achieved by shifting v i one bit to the right, with no additional cost in hardware. Take b 2 = 6'b00_1111, and the coefficient of determination for fitting is 0.9909.

[0045] (3) Determine the positive or negative by the highest sign bit of u i . If it is negative, take the inverse, add one to get the absolute value, and select. According to the positive or negative of u i , shift the 2 v calculated in step (2) left or right to obtain the final exponential function calculation result:

[0046] As Figure 3 shown, the calculation process of the logarithmic function circuit:

[0047] (1) Use LOD to obtain the intermediate value t where the highest bit of the input f is 1 and other bits are 0.

[0048] (2) Use the decoder to decode the input t to obtain the index w where the highest bit of f is located.

[0049] (3) Use the shifter to shift the input f to the right by w bits to obtain k, where k ∈ (1, 2).

[0050] (4) Use the adder to implement the fitting function to calculate log 2 k: log 2 k ≈ k + b 2 , where b 2 is a constant. For example, in a fixed-point number system, and the total bit width of the output f_ln of the logarithmic function circuit is 6 with a 4-bit fractional width, then take b = -0.9485, which is b = 6'b11.0001 after fixed-point conversion, and the coefficient of determination for fitting is 0.9906.

[0051] (5) Combine w with log 2Sum k, and multiply the accumulated value by ln2 using a constant multiplier to obtain the final logarithmic function calculation result: ln(f) = ln2 * log 2 f = ln2 * (w + log 2 k)

[0052] The present invention adopts function equivalent transformation, replacement of the base of the power and the base of the logarithm, function fitting, serial accumulation, reuse of the exponential operation unit, and realizes complex functions only with a limited number of basic operation logic units. It transforms the combination of the power function and division of the original function into the combination of the power function and the logarithmic function. At the same time, according to the operation characteristics and data range, it performs function fitting with controllable precision, saving a large amount of calculation time and iteration process, and effectively reducing the hardware implementation area and power consumption cost by using serial accumulation and function unit reuse.

Claims

1. A Softmax hardware implementation system for a logic resource-constrained platform, characterized in that, it includes the following units: Exponential operation unit: realizes exponential operation through radix transformation and linear fitting; Serial accumulation unit: calculate the accumulation sum of the exponential powers of each input x 1 , x 2 ,...., x n : Logarithmic operation unit: realizes logarithmic operation through base transformation and linear fitting, calculates the natural logarithm of the accumulation result: f_ln = ln(f); The exponential operation unit includes two adders, a constant multiplier, and a shifter, where the shifter supports left shift and right shift; the first adder calculates x i -0 or x i -f_ln; the constant multiplier calculates (x i -f_ln) × log 2 e or x i × log 2 e; the integer part of the absolute value is intercepted to obtain u; the second adder implements the fitting function: 2 v ≈ v + b 1 , where b 1 is a constant; finally, the calculation result is obtained by shifting 2 v left or right by u; The logarithmic operation unit includes a leading 1 detector, a decoder, a shifter, and a constant adder, where the shifter only supports left shift; It is used to calculate ln(f), uses the leading 1 detector to obtain an intermediate value t with only the bit where the highest bit is 1 and other bits are 0; the decoder inputs t to obtain the index w where the highest bit of f is located; The right shifter shifts f to the right by w to obtain k, and a constant adder is used to implement the fitting function: log 2 k ≈ k + b 2 , where b 2 is a constant; finally, the constant multiplier and adder calculate the result of the logarithmic operation unit: ln(f) = ln2 * (w + k + b 2 ); The serial accumulation unit can accept the accumulation of any n inputs, and the accumulation enable is controlled by a counter. When the counter value is equal to n, the accumulation stops.

Citation Information

Patent Citations

  • Softmax implementation method based on hardware platforms

    CN108021537A

  • Softmax function hardware circuit with variable calculation precision and implementation method thereof

    CN110135086A