Softmax layer int8 quantized hardware CORDIC acceleration system and chip
The CORDIC algorithm converts the exponential operation of the softmax layer into basic addition and subtraction and shift operations, which solves the problem of high computational complexity of the softmax layer and realizes the optimization of calculation accuracy and area overhead.
Patent Information
- Application Number
- CN202411927789.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the software layer has high computational complexity, resulting in high computing resources and time consumption, especially in scenarios with high requirements for computing speed and real-time performance, which affects system performance.
The CORDIC algorithm is used to convert the e-exponential operation into rotation and orientation operations under the hyperbolic coordinate system, gradually approximate the target value, realize the addition and subtraction and shift operations of the exponential operation results, and reduce the multiplication operations.
While ensuring calculation accuracy, it reduces the calculation complexity, reduces the use of multipliers in the circuit, and reduces the system area overhead.
Smart Images

Figure CN119940430A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of hardware acceleration of neural networks, and specifically relates to a hardware CORDIC acceleration system and chip for Softmax layer int8 quantization. Background Art
[0002] With the rapid development of artificial intelligence technology, deep neural networks have achieved remarkable results in many fields such as image recognition, natural language processing, speech recognition, etc. Among them, the softmax layer plays a vital role in the neural network, which converts the output of the neural network into a probability distribution, enabling the model to perform tasks such as classification.
[0003] However, in practical applications, the calculation of the softmax layer often becomes one of the bottlenecks of the entire neural network calculation. In particular, the exponential calculation and division calculation consume a lot of computing resources and time due to their high computational complexity. This will seriously affect the performance of the system in scenarios with high requirements for computing speed and real-time performance, such as autonomous driving and real-time video analysis.
[0004] Current e-exponential acceleration solutions usually use approximate calculations or optimized algorithms and data structures. Approximate calculations reduce computational complexity by approximating exponential functions and divisions. For example, the exponential function is approximated by Taylor series expansion. Optimized algorithms and data structures reduce computational complexity and memory usage by improving the algorithms and data structures of the softmax layer. Use a lookup table. Pre-calculate the exponential function values within a certain range and store them in a lookup table. Use the lookup table to obtain approximate values during actual calculations.
[0005] However, the use of approximate calculation methods will introduce certain errors and affect the calculation accuracy. In addition, methods such as Taylor series expansion introduce a large number of multiplication calculations, which is not very hardware-friendly. The acceleration effect of optimization algorithms and data structures is limited, and it is difficult to meet the scenarios with extremely high requirements for calculation speed. Lookup tables can reduce the amount of calculation to a certain extent, but there are limitations on the size of the lookup table and possible loss of accuracy. When the input value exceeds the range of the lookup table, additional processing is required, which increases the complexity of the calculation.
[0006] Therefore, the current acceleration scheme has low computational accuracy and requires a large hardware area. Summary of the invention
[0007] The embodiment of the present invention provides a hardware CORDIC acceleration system and chip for Softmax layer int8 quantization, which can solve the problems of low calculation accuracy and large required hardware area of current acceleration solutions.
[0008] In a first aspect, an embodiment of the present invention provides a hardware CORDIC acceleration system for Softmax layer int8 quantization, including: multiple e-exponential acceleration modules, a first adder, and a divider;
[0009] The input end of the e-exponential acceleration module receives the exponent to be calculated, the output end of the e-exponential acceleration module is respectively connected to the input end of the first adder and the first input end of the divider, the output end of the first adder is connected to the second input end of the divider, and the output end of the divider outputs a normalized output result;
[0010] The e-exponential acceleration module is used to convert the e-exponential operation of the exponent to be calculated into a rotation and orientation operation in a hyperbolic coordinate system based on a coordinate rotation digital calculation method to obtain an exponential operation result; the first adder is used to add all the exponential operation results to obtain the sum of the exponential operation results, and the divider is used to divide the exponential operation result by the sum of the exponential operation results to obtain a normalized output result.
[0011] In a second aspect, an embodiment of the present invention provides a hardware CORDIC acceleration chip for int8 quantization of the Softmax layer of the system described in the first aspect.
[0012] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: according to the system provided by the present invention, the e-exponential operation is converted into rotation and orientation operations in a hyperbolic coordinate system through the CORDIC algorithm, and the exponential operation can be converted into basic addition, subtraction and shift operations, gradually approaching the target value to obtain the exponential operation result, thereby reducing the calculation complexity while ensuring the calculation accuracy, reducing the use of multipliers in the circuit, and further achieving the purpose of reducing the area overhead of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 A schematic diagram of the structure of a hardware CORDIC acceleration system for Softmax layer int8 quantization provided by an embodiment of the present invention;
[0014] Figure 2 A schematic diagram of a hyperbolic coordinate rotation model provided by an embodiment of the present invention;
[0015] Figure 3 A binary digit structure diagram provided by an embodiment of the present invention;
[0016] Figure 4 A schematic diagram of the structure of an e-exponential acceleration module provided by an embodiment of the present invention;
[0017] Figure 5 A schematic diagram of the structure of a convergence region expansion module provided by an embodiment of the present invention;
[0018] Figure 6 A schematic diagram of the structure of an exponential operation module provided by an embodiment of the present invention;
[0019] Figure 7 A schematic diagram of the structure of a first iterator module provided by an embodiment of the present invention;
[0020] Figure 8 A schematic diagram of the structure of another exponential operation module provided by an embodiment of the present invention;
[0021] Fig. 9 A schematic diagram of the structure of a second iterator module provided in an embodiment of the present invention;
[0022] Fig.10 A schematic diagram of the structure of another exponential operation module provided by an embodiment of the present invention;
[0023] Fig.11 A schematic diagram of the structure of the third iterator module provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present invention. However, it should be clear to those skilled in the art that the present invention may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present invention.
[0025] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0026] It should also be understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0027] As used in the present specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]", depending on the context.
[0028] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0029] References to "one embodiment" or "some embodiments" etc. described in the present specification mean that one or more embodiments of the present invention include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0030] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0031] Example 1
[0032] Figure 1 The structure diagram of a hardware CORDIC acceleration system for Softmax layer int8 quantization provided by an embodiment of the present invention is shown. As an example but not a limitation, the system 100 may include a plurality of e-exponential acceleration modules 1, a first adder A1, and a divider B1.
[0033] For example, see Figure 1 , the input end of the exponential acceleration module 1 can receive the i′th exponent θ to be calculated i′ The output end can be connected to the input end of the first adder A1 and the first input end of the divider B1 respectively. The output end of the first adder A1 can be connected to the second input end of the divider B1. The output end of the divider B1 outputs the normalized output result.
[0034] In one example, the e-exponential acceleration module 1 can be used to convert the e-exponential operation of the exponential to be calculated into a rotation and orientation operation in a hyperbolic coordinate system based on a coordinate rotation digital computer (CORDIC) method to obtain an exponential operation result. The first adder A1 can be used to add all the exponential operation results to obtain the sum of the exponential operation results, and the divider B1 is used to divide the exponential operation result by the sum of the exponential operation results to obtain a normalized output result.
[0035] Exemplarily, the normalized output result can satisfy the following formula:
[0036]
[0037] Among them, y i′ is the i′th normalized output result, represents the sum of the results of exponential operations, is the result of the i′th exponential operation, and N is the total number of exponents to be operated.
[0038] Generally, in a neural network, θ i′ It can be 40 score results of float32 (single-precision 32-bit) type.
[0039] Figure 2 Shown is a schematic diagram of a hyperbolic coordinate rotation model provided by an embodiment of the present invention.
[0040] In some embodiments, see Figure 2 , a point x can be drawn from the hyperbola x 2 -y 2 =C 2 Start the rotation and orientation operation at any point in the first quadrant and move to another point on the hyperbola. Figure 1 point V1 in), the final mapping point (see Figure 1 Vn in the figure), the hyperbola between the initial and final mapping points, the origin (see Figure 1 The area of the closed figure formed by the rays emitted from point O) to these two points is equal to in, θ is the index to be calculated; the area of the closed figure enclosed by the initial mapping point, the intersection of the hyperbola and the x-axis, the origin and the hyperbola is equal to
[0041] Exemplarily, the coordinate values of the final mapping points may satisfy the following formula:
[0042]
[0043] Among them, x n ,y n are the x and y coordinate values of the final mapping point, respectively; x1 and y1 are the x and y coordinate values of the initial mapping point, respectively.
[0044] For example, θ The exponential operation can be decomposed into:
[0045] e θ =coshθ+sinhθ(1.3).
[0046] In one possible implementation, to calculate e θ The value of θDecomposed into a series of θ i (i=1,2,…,n) cumulative sum form; accordingly, the final mapping point can be regarded as the point reached after n rotations (i.e., the mapping point after the nth rotation); approximating e by iteration θ The calculation result of .
[0047] Exemplary, decomposed θ The following formula can be satisfied:
[0048]
[0049] Among them, θ i is the i-th decomposition value of θ, which is equal to the mapping point x after the i-th rotation i d is twice the area of the closed image enclosed by the hyperbola between the initial mapping point and the origin. i ={-1,,1} is the rotation direction factor.
[0050] In one example, if θ i Satisfying the condition tanhθ i =2 -i , then the above formula (1.2) can be expressed as:
[0051]
[0052] Among them, the calibration factor K can satisfy the following formula:
[0053]
[0054] Through the CORDIC algorithm, based on the above formulas (1.1)-(1.6), the exponential operation process can be converted into an iterative process consisting of addition, subtraction and shift operations.
[0055] According to the system provided by the present invention, the e-exponential operation is converted into rotation and orientation operations in a hyperbolic coordinate system through the CORDIC algorithm, and the exponential operation can be converted into basic addition, subtraction and shift operations, gradually approaching the target value to obtain the exponential operation result, thereby reducing the calculation complexity while ensuring the calculation accuracy, reducing the use of multipliers in the circuit and further achieving the purpose of reducing the area overhead of the system.
[0056] Example 2
[0057] As an example, based on Example 1, if the x and y coordinates of the initial mapping point are different, for example, the hyperbola and the initial mapping point satisfy C=1, x1=1, y1=0, and the initial iteration value of the CORDIC algorithm is X1=K, Y1=0, Z1=θ; then the iteration formula satisfies the following formula:
[0058]
[0059] where x i+1 is the x - coordinate of the mapped point after the (i + 1)-th rotation, i is a positive integer less than or equal to n, y i+1 is the y - coordinate of the mapped point after the (i + 1)-th rotation, z i+1 is the remaining angle after the (i + 1)-th rotation, and sgn is the sign function.
[0060] To satisfy the convergence of the iterative sequence, the value of the iterative sequence i starts from the 4th term. When i = k n (k n = 3k n-1 + 1, k1 = 4, n ∈ Z + ), this iteration is repeated, that is, i = 1, 2, 3, 4, 4, …, 13, 13, …, 40, 40, …
[0061] After n rotations, we can get:
[0062]
[0063] Therefore, from the above formulas (1.7) and (1.8), it can be obtained that the first iterative model of the exponential operation result can satisfy the following formula
[0064]
[0065] where θ is the exponent to be operated, e θ is the result of the exponential operation.
[0066] According to the first iterative model (see formula (1.9)), the convergence range of the hyperbolic coordinate CORDIC algorithm can be obtained as: where θ max is the sum of all rotation radians, that is, the maximum radian value for the convergence of this algorithm.
[0067] Since the convergence domain of θ is (-1.1182, 1.1182), its practical application significance is limited. Therefore, it is necessary to expand the convergence domain. Any exponent to be operated can be split into the following form:
[0068] θ = Qln2 + γ(1.10)
[0069] where Q ∈ Z is the quantized integer, and 0 ≤ γ < ln2 = 0.6931 is the exponent after convergence.
[0070] Based on the above formula (1.10), the exponential operation result can be split into: e θ = e Qln2+γ = 2 Q e γTherefore, after determining the fixed-point integer Q, we only need to convert the exponential operation result of the converged exponential into γ Shift left or right |Q| to get e θ ; In this way, e θ The calculation is equivalent to e γ The calculation of the exponential function CORDIC algorithm is realized by the interval compression method.
[0071] In one example, to reduce the use of memory resources, the input of Softmax can be θ Perform int8 quantization, input int8 quantization data, zero point and scaling factor to the hardware circuit, and then restore the exponent to be calculated by calculating the input int8 quantization data, zero point and scaling factor in the circuit.
[0072] For example, the int8 quantization process can be expressed as:
[0073] θ=(qz)*s(1.11)
[0074] Among them, θ is the exponent to be calculated, which can be the score result of float32 type in the Softmax layer, q is int8 quantized data, z is the zero point, and s is the fixed-point scaling factor.
[0075] According to the above formulas (1.10) and (1.11), the exponent to be calculated can be equivalent to (qz)*s=Qln2+γ. By fixing both sides of the equation and dividing by ln2, we can get:
[0076]
[0077] Among them, qz can be represented by a 9-bit signed number, s*2 32 / ln2 is a 32-bit unsigned number.
[0078] For example, see Figure 3 [(qz)*s*2 32 / ln2] corresponding to the binary digit structure diagram, Q is the fixed-point integer part, which can be used Figure 3 H9 in the formula means that H9 = Q; is the fixed-point fractional part, which can be used Figure 3 L 32 Indicates that
[0079] Therefore, the input of the e-exponential acceleration module 1 can be represented by a segment [(qz)*s*2 32 / ln2], considering the signed number, the fixed-point integers Q and
[0080]
[0081] The effect of directly intercepting the integer part of the complement code is to round down the data; for positive numbers, the decimal part is directly intercepted; for negative numbers, the decimal part is +1 to obtain and
[0082] For the L 16 or L1′6 part of the codeword (the following formula (1.14) is uniformly represented by L 16 Instead) the converged exponential multiplication result can be obtained by performing fixed-point extraction using the following formula:
[0083]
[0084] In view of the first iterative model of the exponential operation result proposed by the present invention when the x and y coordinate values of the above-mentioned initial mapping points are different (see the above-mentioned formula (1.9)) and the input method of the exponent to be calculated (see the above-mentioned formulas (1.10)-(1.14)), the present invention proposes an e-exponential acceleration module 1.
[0085] Figure 4 The structure diagram of an e-exponential acceleration module provided by an embodiment of the present invention is shown. As an example but not a limitation, the e-exponential acceleration module 1 may include a convergence region expansion module 11 and an exponential operation module 12.
[0086] Exemplarily, the input end of the convergence region expansion module 11 can be used as the input end of the e-exponential acceleration module 1 , and the output end is connected to the input end of the exponential operation module 12 , and the output end of the exponential operation module 12 is used as the output end of the e-exponential acceleration module 1 .
[0087] Specifically, the convergence domain expansion module 11 can be used to determine the fixed-point integer and the converged exponent that satisfies the convergence domain according to the exponent to be calculated, and the exponential operation module can be used to convert the e-exponential operation of the exponent to be calculated into rotation and orientation operations in a hyperbolic coordinate system to obtain the exponential operation result.
[0088] In a possible implementation, based on the input method of the exponent to be calculated proposed by the present invention (see the above formulas (1.10)-(1.14)), see Figure 5 The convergence region expansion module 11 may include a first subtractor D1, a first multiplier E1, a second adder A2, a data selector U1, and a second multiplier E2.
[0089] In one example, see Figure 5The first input end and the second input end of the first subtractor D1 are respectively used as an input end of the convergence domain expansion module 11 to receive the int8 quantized data q and the zero point z respectively. The output end of the first subtractor D1 is connected to the first input end of the first multiplier E1. The second input end of the first multiplier E1 is used as an input end of the convergence domain expansion module 11 to receive the calculation result s*2 of the fixed-point scaling factor. 32 / ln2, the first output end of the first multiplier E1 outputs a fixed-point integer as an output end of the convergence domain expansion module 11, the second output end and the third output end of the first multiplier E1 are respectively connected to the control input end and the first input end of the data selector U1, the fourth output end of the first multiplier is connected to the first input end of the second adder A2; the output end of the second adder A2 is connected to the second input end of the data selector U1, the output end of the data selector U1 is connected to the input end of the second multiplier E2, and the output end of the second multiplier E2 outputs the converged exponential multiplication result γ*2 as an output end of the convergence domain expansion module 11 32 .
[0090] Specifically, see Figure 5 And the above formulas (1.12)-(1.14), the first subtractor D1 calculates the difference between the int8 quantized data q and the zero point z to obtain (qz), and the first multiplier E1 calculates (qz)*s*2 32 / ln2, its first output terminal outputs (qz)*s*2 32 / ln2 bit [40:32] codeword, that is, the fixed-point integer Q; the second output terminal inputs the sign bit, that is, the 40th codeword, into the control input terminal of the data selector U1, and the third output terminal outputs L 16 The first input terminal of the data selector U1 outputs L1′6. The second adder calculates L1′6 and 2 16 The product of And input it to the second input terminal of the data selector U1. The data selector U1 selects one of the inputs as The second multiplier E2 will Multiply by ln2*2 16 Get the converged exponential multiplication result γ*2 32 .
[0091] In a possible implementation, based on the first iterative model provided by the present invention (see the above formula (1.19)), see Figure 6 The exponential operation module 12 may include a first lookup table TA1, a third adder A3, a first Q shifter Q1 and n cascaded first iterator modules 121.
[0092] For example, see Figure 7 Each first iterative submodule 121 may include a first sign function circuit S1, a first adder / subtractor AS1, a second adder / subtractor AS2, a first shifter T1, a second shifter T2, and a third adder / subtractor AS3.
[0093] In one example, see Figure 6 , the i-th first sign function circuit S1 i The input terminal of the i-1th first adder / subtractor AS1 i-1 The output end is connected to the output end of the i-th first adder / subtractor AS1 i , the i-th second adder / subtractor AS2 i , the i-th third adder / subtractor AS3 i The control input terminal of the i-th first adder / subtractor AS1 is connected; i The first input terminal and the second input terminal of the first lookup table TA1 and the output terminal of the i-1th first adder / subtractor AS1 are respectively connected to the output terminal of the first lookup table TA1 and the output terminal of the i-1th first adder / subtractor AS1. i-1 The output end of the i-th second adder / subtractor AS2 is connected; i The first input terminal of the i-1th second adder / subtractor AS2 i-1 The output end is connected to the first shifter T1, and the second input end is connected to the first shifter T1 i The output end of the i-1 second adder / summer AS2 is connected to i-1 The output of the i-th second shifter T2 i The input end of the i-th third adder / subtractor AS3 is connected; i The first input terminal of the i-th second shifter T2 i The output end is connected to the second input end of the i-1th third adder / subtractor AS3 i-1 The output end of the i-1th third adder / subtractor AS3 is connected to i-1 The output of the i-th first shifter T1 i The first input terminal and the second input terminal of the third adder A3 are respectively connected to the nth second adder / subtractor AS2 n The output terminal of the nth third adder / subtractor AS3 n The output end is connected to the first input end of the first Q shifter Q1; the second input end of the first Q shifter Q1 is connected to the first output end of the first multiplier E1, and the output end serves as the output end of the exponential operation module 12 to output the exponential operation result.
[0094] Exemplarily, the input end of the first first sign function circuit S11 and the second input end of the first first adder / subtractor AS11 are connected to the output end of the second multiplier E2 to receive γ*2 32.
[0095] Exemplarily, the first input terminal of the first second adder / subtractor AS21 and the input terminal of the first second shifter T21 receive the iteration initial value Y1=0 obtained according to the y-coordinate value y1 of the initial mapping point.
[0096] Exemplarily, the second input terminal of the first third adder / subtractor AS31 and the input terminal of the first first shifter T11 receive the x coordinate value x1 of the initial mapping point to obtain the iteration initial value X1=K·2 32 .
[0097] Specifically, see Figure 6 And formula (1.9), in each iteration, the i-th first sign function circuit S1 i According to z i (equal to γ i 2 32 ) to get d i , the i-th first adder / subtractor AS1 i According to d i Select output z i The first lookup table output is 2 32 ·tanh -1 2 -i The sum / difference of i+1 The i-th first shifter T1 i x i Shift right by i bits. The i-th second adder / subtractor AS2 i According to d i Select output x i 2 -i With y i The sum / difference of i+1 The i-th second shifter will shift y i Shift right by i bits and output y i 2 -i The third adder / subtractor AS3 i According to d i Select the output y after right shift i bits i (ie y i 2 -i ) and x i The sum / difference of i+1 Finally, the third adder A3 converts x n+1 and n+1 Add them together to get e γ 2 32 , the first Q shifter performs Q shift on it to obtain e θ 2 32 .
[0098] According to the method provided by the present invention, the exponential operation result is obtained only through addition, subtraction and shift operations based on the coordinate rotation digital calculation method, so that the exponential operation module provided by the present invention can be implemented only by an adder / subtractor, a shifter and a lookup table, without the need for additional multipliers to participate in the circuit, which can save the area overhead of the exponential operation module.
[0099] Example 3
[0100] As an example, based on Example 1 and Example 2, if the x and y coordinates of the initial mapping point are the same, for example, when x1=y1=1, the initial iteration value is X1=Y1=K, Z1=θ and C=0, the hyperbola degenerates into a straight line, then x n ,y n The iteration modes are the same. The two can be combined to omit the iteration process of one item. The iteration formula can satisfy the following formula:
[0101]
[0102] After n rotations, we can get:
[0103]
[0104] Therefore, the second iteration model of the exponential operation result obtained from the above formulas (1.15) and (1.16) can satisfy the following formula:
[0105]
[0106] In the second iteration model, the convergence domain of θ is the same as that of the first iteration model, so the same method can be used to split it. In this case, the same convergence domain expansion module 11 can be used to calculate the multiplication result γ*2 of the fixed-point integer Q and the converged exponential. 32 .
[0107] In view of the second iterative model of the exponential operation result proposed by the present invention when the x and y coordinate values of the initial mapping point are the same (see the above formula (1.17)), the present invention proposes another exponential operation module 12.
[0108] Figure 8 The structure diagram of another exponential operation module provided in an embodiment of the present invention is shown. As an example but not limitation, the exponential operation module 12 may include a second lookup table TA2, a second Q shifter Q2, and n cascaded second iterator modules 122.
[0109] For example, see Fig. 9Each second iterator module 122 may include a second sign function circuit S2, a fourth adder / subtractor AS4, a fifth adder / subtractor AS5, and a third shifter T3.
[0110] In one example, see Figure 8 , the i-th second sign function circuit S2 i The input terminal of the fourth adder / subtractor AS4 of the i-1th i-1 The output end of the ith adder / subtractor AS4 is connected to the output end of the ith adder / subtractor AS4. i , the ith fifth adder / subtractor AS5 i The control input terminal of the i-th fourth adder / subtractor AS4 is connected; i The first input terminal is connected to the output terminal of the second look-up table TA2, and the second input terminal is connected to the i-1th fourth adder / subtractor AS4 i-1 The output end of the i-th fifth adder / subtractor AS5 is connected; i The first input terminal of the i-th third shifter T3 i The second input terminal is connected to the i-1th fifth adder / subtractor AS5 i-1 The output end of the i-1th fifth adder / subtractor AS5 is connected to i-1 The output of the third shifter T3 i The first input terminal and the second input terminal of the second Q shifter Q2 are respectively connected to the nth fifth adder / subtractor AS5 n The output end of is connected to the output end of the second lookup table TA2, and the output end serves as the output end of the exponential operation module 12 to output the exponential operation result.
[0111] Exemplarily, the input end of the first second sign function circuit S21 and the second input end of the first fourth adder / subtractor AS41 are connected to the output end of the second multiplier E2 to receive the converged exponential multiplication result γ*2 32 .
[0112] Exemplarily, the second input terminal of the first fifth adder / subtractor AS51 and the input terminal of the first third shifter T31 receive the iteration initial value X1=Y1=K·2 obtained according to the x coordinate or y coordinate of the initial mapping point. 32 .
[0113] Specifically, see Figure 8 According to the second iteration model (see formula (1.17)), the i-th second sign function circuit S2 i According to z i (also equal to γ i 2 32 ) to get d iThe fourth adder / subtractor AS4 i According to d i Select output z i The first lookup table output is 2 32 ·tanh -1 2 -i The sum / difference of i+1 The i-th third shifter T3 i x i Shift right by i bits and output x i 2 -i , the fifth adder / subtractor AS5 i According to d i The value of output x is selected i 2 -i With x i The sum / difference of i+1 The second Q shifter Q2 converts x n+1 After moving Q, we get e θ 2 32 .
[0114] By starting the iteration from the initial mapping point with the same x and y coordinate values, the 4-dimensional calculation in each iteration can be reduced to 3-dimensional calculation, thereby reducing one iteration circuit in the iterative submodule and further reducing the area overhead of the exponential operation module. At the same time, this method can save about 1 / 3 of hardware resources in hardware implementation.
[0115] Example 4
[0116] As an example, based on Embodiment 1 to Embodiment 3, if the x and y coordinate values of the initial mapping point are the same and the index to be calculated can be recoded by angle polarization, the iterative formula can be further simplified.
[0117] Exemplarily, the index to be calculated after angle polarization recoding may satisfy the following formula:
[0118]
[0119] Where θ is the binary representation of radians or angles (i.e., the exponent to be calculated), θ∈[0,1), b i ={0,,1}.
[0120] Let r i =2b i -1, r i ={-1,1}, we can get:
[0121]
[0122] Combining formula (1.19), formula (1.5) can be converted to:
[0123]
[0124] Among them, the calibration factor
[0125] When C=0, x1=y1=1, the initial value of iteration is X1=Y1=K'(1+tanhφ0), Z=θ, and the iteration formula can satisfy the following formula:
[0126] x i+1 =x i +r i x i tanh2 -(i+1) (1.21)
[0127] Likewise, e θ =sinhθ+coshθ=x n+1 .
[0128] For tanh2 in the above formula (1.21) -(i+1) Taylor expansion can be obtained:
[0129]
[0130] when When , the above formula (1.16) can be equivalent to tanh2 -(i+1) ≈2 -(i+1) , so formula (1.21) can be replaced by:
[0131] x i+1 =x i +r i x i 2 -(i+1) (1.23)
[0132] Among them, r i ={-1,1} According to formula (1.18) b i Knowing this in advance, no calculation of the remaining angles is required.
[0133] Simplifying the above formula (1.23) and performing two-step iteration, we can get:
[0134] x i+2 =x i (1+r i 2 -(i+1) +r i+1 2 -(i+1) +r i r i+1 2 -(i+2) )(1.24)
[0135] Because when When i r i+1 2 -(2i+3) ≈0, so formula (1.24) can be simplified to:
[0136] x i+2 =x i (1+r i 2 -(i+1) +r i+1 2 -(i+2) )(1.25).
[0137] By analogy, when All remaining iterative formulas can be simplified to:
[0138]
[0139] Where k is the number of stride iterations.
[0140] Therefore, if the x and y coordinates of the initial mapping point are the same and the index to be calculated can be recoded by angle polarization, when the number of iterations i is less than When the exponential operation module 12 can iterate based on the above formula (1.15), when the number of iterations exceeds When , the exponential operation module 12 can obtain the exponential operation result based on the following third iterative model:
[0141]
[0142] In the third iteration model, the convergence domain of θ is the same as that of the first iteration model, so the same method can be used to split it. In this case, the same convergence domain expansion module 11 can be used to calculate the multiplication result γ*2 of the fixed-point integer Q and the converged exponential. 32 .
[0143] In view of the third iterative model of the exponential operation result proposed by the present invention when the x and y coordinate values of the initial mapping point are the same and the exponent to be calculated can be angle-dipolarized recoded (see the above formula (1.27)), the present invention proposes another exponential operation module 12.
[0144] Fig.10 The structure diagram of another exponential operation module provided by the embodiment of the present invention is shown. As an example but not limitation, the exponential operation module 12 may include a third lookup table TA3, n / 2 fifth shifters T5, a seventh adder / subtractor AS7, a third Q shifter Q3, The third iterator module 123 of the cascade.
[0145] For example, see Fig.11, each third iterator module 123 may include one sixth adder / subtractor AS6 and one fourth shifter T4.
[0146] For example, see Fig.10 , the mth sixth adder / subtractor AS6 m The first input terminal of the mth fourth shifter T4 m The output end is connected to the m-1th sixth adder / subtractor AS6, and the second input end is connected to the m-1th sixth adder / subtractor AS6 m-1 The output end of the control input end is connected to the output end of the second multiplier E2; the m-1 sixth adder / subtractor AS6 m-1 The output of the mth fourth shifter T4 m The input end of all fifth shifters T5 is connected to the input end of the first Sixth adder / subtractor The output end of the seventh adder / subtractor AS7 is connected to the first input end of the seventh adder / subtractor AS7; the second input end of the seventh adder / subtractor AS7 is connected to the Sixth adder / subtractor The output end of the control input end is connected to the output end of the second multiplier E2; the first input end and the second input end of the third Q shifter Q3 are respectively connected to the output end of the seventh adder / subtractor AS7 and the first output end of the first multiplier E1, and the output end serves as the output end of the exponential operation module 12 to output the exponential operation result.
[0147] For example, m is less than or equal to The input end of the first fourth shifter T41 and the second input end of the first sixth adder / subtractor AS61 are connected to the output end of the third look-up table TA3.
[0148] Specifically, see Fig.10 and the third iteration model (see the above formula (1.27)), the third lookup table TA3 can be pre-obtained according to the above formula (1.15) Iteration results The mth fourth shifter T41 is input to the second input terminal of the first sixth adder / subtractor AS61 and the input terminal of the first fourth shifter T41. m Will Shift Right The mth adder / subtractor AS6 is input to the mth adder / subtractor AS6 m Then the sixth adder / subtractor AS6 m according to Will and The sum / difference of Output. Then Sixth adder / subtractor Will The ath fifth shifter T5 is input to the second input terminal of the seventh adder / subtractor AS7 and the input terminal of each fifth shifter T5. a Will Shift Right The bit is input to the first input terminal of the seventh adder / subtractor AS7, and a is less than or equal to The seventh adder / subtractor AS7 is based on the formula Output x n+1 =e γ 2 32 The third Q shifter Q3 converts x n+1 After moving Q, we get e θ 2 32 .
[0149] For example, when n=32, the initial iteration value can be X1=Y1=K'(1+tanhφ0)·2 according to the above formula (1.15). 32 The calculated x 10 Stored in the third lookup table TA3. That is, when i≤9, the third lookup table TA3 can be used according to γ·2 32 [31:23] Find and output x 10 To the second input terminal of the first sixth adder / subtractor AS61 and the input terminal of the first fourth shifter T41. That is, when 16>i>9, the mth third iterator module 123 receives γ·2 output from the second multiplier E2. 32 The m+9th last digit of the data is r m+9 , for example, the first third iterator module is based on γ·2 32
[22] Get r 10 , the second one is based on γ·2 32
[21] Get r 11 ...; then according to r m+9 And the third iteration model (see the above formula (1.27)) calculates The seventh third iteration submodule 123 outputs x 17 to the second input terminal of the seventh adder / subtractor AS7 and the input terminal of the fifth shifter T5. That is, when i>16, the seventh adder / subtractor TS7 calculates the value of γ·2 32 The remaining bits of data, namely γ·2 32 [15:0], and the third iteration model (see the above formula (1.27)) calculates x 33 to the third Q shifter Q3.
[0150] According to the exponential operation module provided by the present invention, the exponent to be operated is re-encoded by angle two plans, thereby saving the remaining angle z i The iterative process is directly based on γ·2 32 [31:0] performs rotation iteration; this can reduce one iteration circuit in the iterator module, further reducing area overhead. It is only necessary to keep the space of the third lookup table at If the output bit width is larger than that, about 1 / 3 of the hardware resources can be saved in hardware implementation.
[0151] An embodiment of the present invention further provides a hardware CORDIC acceleration chip for Softmax layer int8 quantization. This chip may include the hardware CORDIC acceleration system for Softmax layer int8 quantization provided by any of the above embodiments.
[0152] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0153] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
Claims
1. A hardware CORDIC acceleration system for Softmax layer int8 quantization, characterized in that: include: A plurality of e-exponential acceleration modules (1), a first adder (A1), and a divider (B1); The input end of the e-exponential acceleration module (1) receives the exponent to be calculated, the output end of the e-exponential acceleration module (1) is respectively connected to the input end of the first adder (A1) and the first input end of the divider (B1), the output end of the first adder (A1) is connected to the second input end of the divider (B1), and the output end of the divider (B1) outputs a normalized output result; The e-exponential acceleration module (1) is used to convert the e-exponential operation of the exponential to be calculated into a rotation and orientation operation in a hyperbolic coordinate system based on a coordinate rotation digital calculation method, so as to obtain an exponential operation result; the first adder (A1) is used to add all the exponential operation results to obtain the sum of the exponential operation results, and the divider (B1) is used to divide the exponential operation result by the sum of the exponential operation results to obtain a normalized output result.
2. The system according to claim 1, characterized in that The e-exponential acceleration module (1) comprises a convergence domain expansion module (11) and an exponential operation module (12); The input end of the convergence domain expansion module (11) serves as the input end of the e-exponential acceleration module (1), the output end of the convergence domain expansion module (11) is connected to the input end of the exponential operation module (12), and the output end of the exponential operation module (12) serves as the output end of the e-exponential acceleration module (1); The convergence domain expansion module (11) is used to determine a fixed-point integer and a converged exponent that satisfies the convergence domain when the exponential operation module (12) calculates the exponential operation module (12) according to the exponential to be calculated; the exponential operation module (12) is used to convert the e-exponential operation of the exponential to be calculated into a rotation and orientation operation in a hyperbolic coordinate system according to the fixed-point integer and the converged exponent to obtain the exponential operation result.
3. The system according to claim 2, characterized in that The convergence domain expansion module (11) comprises a first subtractor (D1), a first multiplier (E1), a second adder (A2), a data selector (U1), and a second multiplier (E2); The first input end and the second input end of the first subtractor (D1) are respectively used as an input end of the convergence domain expansion module (11) to receive int8 quantized data and a zero point respectively; the output end of the first subtractor (D1) is connected to the first input end of the first multiplier (E1); the second input end of the first multiplier (E1) is used as an input end of the convergence domain expansion module (11) to receive the calculation result of the fixed-point scaling factor; the first output end of the first multiplier (E1) is used as an output end of the convergence domain expansion module (11) to output the fixed-point integer; the second output end and the third output end of the first multiplier (E1) are respectively connected to the control input end and the first input end of the data selector (U1); the fourth output end of the first multiplier is connected to the first input end of the second adder (A2); The output end of the second adder (A2) is connected to the second input end of the data selector (U1), the output end of the data selector (U1) is connected to the input end of the second multiplier (E2), and the output end of the second multiplier (E2) serves as an output end of the convergence domain expansion module (11) to output the multiplication result of the converged exponent; wherein the exponent to be calculated is composed of the int8 quantized data, the zero point and the fixed-point scaling factor.
4. The system according to claim 3, characterized in that If the x and y coordinate values of the initial mapping point for starting the rotation and orientation operation in the hyperbolic coordinate system are different, the exponential operation result satisfies the following formula: and θ =coshθ+sinhθ=x n+1 +y n+1 θ is the exponent to be calculated, e θ is the exponential operation result; in: x i+1 is the x-coordinate of the mapping point after the i+1th rotation, i is a positive integer less than or equal to n, and the area of the closed image formed by the initial mapping point, the mapping point after the nth rotation, the origin, and the hyperbola passing through the initial mapping point and the mapping point after the nth rotation is equal to y i+1 is the y coordinate of the mapping point after the i+1th rotation, z i+1 is the remaining angle after the i+1th rotation, and sgn is the sign function.
5. The system according to claim 4, characterized in that The exponential operation module (12) comprises a first lookup table (TA1), a third adder (A3), a first Q shifter (Q1), and n cascaded first iterative submodules (121), each of which comprises: a first sign function circuit (S1), a first adder / subtractor (AS1), a second adder / subtractor (AS2), a first shifter (T1), a second shifter (T2), and a third adder / subtractor (AS3); The i-th first sign function circuit (S1 i ) is connected to the input of the first adder / subtractor (AS1 i-1 ) is connected to the output end of the i-th first adder / subtractor (AS1 i ), the i-th second adder / subtractor (AS2 i ), the i-th third adder / subtractor (AS3 i ) is connected to the control input terminal; The i-th first adder / subtractor (AS1 i ) is connected to the output end of the first lookup table (TA1), the output end of the i-1 first adder / subtractor (AS1 i-1 ), wherein the input end of the first first sign function circuit (S11) and the second input end of the first first adder / subtractor (AS11) are connected to the output end of the second multiplier (E2); The i-th second adder / subtractor (AS2 i ) and the first input terminal of the i-1th second adder / subtractor (AS2 i-1 ) is connected to the output end, and the second input end is connected to the i-th first shifter (T1 i ) is connected to the output end of the i-1 second adder / summer (AS2 i-1 ) is also connected to the output of the i-th second shifter (T2 i ), wherein the first input end of the first second adder / subtractor (AS21) and the input end of the first second shifter (T21) receive the iteration initial value Y1 obtained according to the y coordinate value of the initial mapping point; The i-th third adder / subtractor (AS3 i ) and the first input terminal of the i-th second shifter (T2 i ) is connected to the output end thereof, and the second input end thereof is connected to the i-1th third adder / subtractor (AS3 i-1 ) is connected to the output end of the i-1th third adder / subtractor (AS3 i-1 ) is also connected to the output end of the i-th first shifter (T1 i ), wherein the second input end of the first third adder / subtractor (AS31) and the input end of the first first shifter (T11) receive the iteration initial value X1 obtained according to the x coordinate value of the initial mapping point; The first input terminal and the second input terminal of the third adder (A3) are respectively connected to the nth second adder / subtractor (AS2 n ) output terminal, the nth third adder / subtractor (AS3 n ), the output end of which is connected to the first input end of the first Q shifter (Q1); The second input end of the first Q shifter (Q1) is connected to the first output end of the first multiplier (E1), and the output end serves as the output end of the exponential operation module (12) to output the exponential operation result.
6. The system according to claim 3, characterized in that If the x and y coordinate values of the initial mapping point for starting the rotation and orientation operations in the hyperbolic coordinate system are the same, then the exponential operation result satisfies the following formula: and θ =coshθ+sinhθ=x n+1 =and n+1 θ is the exponent to be calculated, e θ is the exponential operation result; in: x i+1 is the x-coordinate of the mapping point after the (i+1)th rotation, i is a positive integer less than or equal to n, and the area of the closed image formed by the initial mapping point, the mapping point after the (n+1)th rotation, the origin, and the hyperbola passing through the initial mapping point and the mapping point after the (n+1)th rotation is equal to z i+1 is the remaining angle after the i+1th rotation, and sgn is the sign function.
7. The system according to claim 6, characterized in that The exponential operation module comprises a second lookup table (TA2), a second Q shifter (Q2) and n cascaded second iterative submodules (122), each of the second iterative submodules (122) comprising: a second sign function circuit (S2), a fourth adder / subtractor (AS4), a fifth adder / subtractor (AS5) and a third shifter (T3); The i-th second sign function circuit (S2 i ) is connected to the input of the i-1th fourth adder / subtractor (AS4 i-1 ) is connected to the output end of the ith fourth adder / subtractor (AS4 i ), the ith fifth adder / subtractor (AS5 i ) is connected to the control input terminal; The i-th fourth adder / subtractor (AS4 i ) is connected to the output of the second lookup table (TA2), and the second input is connected to the i-1th fourth adder / subtractor (AS4 i-1 ), wherein the input end of the first second sign function circuit (S21) and the second input end of the first fourth adder / subtractor (AS41) are connected to the output end of the second multiplier (E2); The i-th fifth adder / subtractor (AS5 i ) and the first input terminal of the i-th third shifter (T3 i ) is connected to the output terminal, and the second input terminal is connected to the i-1th fifth adder / subtractor (AS5 i-1 ) is connected to the output end of the i-1th fifth adder / subtractor (AS5 i-1 ) is also connected to the output terminal of the i-th third shifter (T3 i ), wherein the second input end of the first fifth adder / subtractor (AS51) and the input end of the first third shifter (T31) receive an iteration initial value obtained according to the y coordinate value or the x coordinate value of the initial mapping point; The first input terminal and the second input terminal of the second Q shifter (Q2) are respectively connected to the nth fifth adder / subtractor (AS5 n ) is connected to the output end of the second lookup table (TA2), and the output end serves as the output end of the exponential operation module (12) to output the exponential operation result.
8. The system according to claim 3, characterized in that If the x and y coordinate values of the initial mapping point for starting the rotation and orientation operation in the hyperbolic coordinate system are the same, and the exponent to be calculated can be angle-dipolarized recoded, then the exponential calculation result satisfies the following formula: e θ =sinhθ+coshθ=x n+1 θ is the exponent to be calculated, e θ is the exponential operation result; in: r i =2b i -1 x i+1 is the x-coordinate of the mapping point after the i+1th rotation, i is a positive integer less than or equal to n, and the area of the closed image formed by the initial mapping point, the mapping point after the n+1th rotation, the origin, and the hyperbola passing through the initial mapping point and the mapping point after the n+1th rotation is equal to 9. The system according to claim 8, characterized in that The exponential operation module (12) includes cascaded third iterator modules (123), each of the third iterator modules (123) comprising: a sixth adder / subtractor (AS6), a fourth shifter (T4), the exponential operation module (12) further comprising a third lookup table (TA3), n / 2 fifth shifters (T5), a seventh adder / subtractor (AS7), and a third Q shifter (Q3); The mth sixth adder / subtractor (AS6 m ) and the first input terminal of the mth fourth shifter (T4 m ) is connected to the output end thereof, and the second input end thereof is connected to the m-1th sixth adder / subtractor (AS6 m-1 ) is connected to the output end of the control input end and the output end of the second multiplier (E2); the m-1th sixth adder / subtractor (AS6 m-1 ) is also connected to the output terminal of the mth fourth shifter (T4 m ) is connected to the input terminal; where m is less than or equal to The input end of the first fourth shifter (T41) and the second input end of the first sixth adder / subtractor (AS61) are connected to the output end of the third lookup table (TA3); All the input terminals of the fifth shifter (T5) are connected to the first Sixth adder / subtractor The output ends of the adder / subtractor (AS7) are connected to the first input end of the seventh adder / subtractor (AS7); The second input terminal of the seventh adder / subtractor (AS7) is connected to the first Sixth adder / subtractor The control input end is connected to the output end of the second multiplier (E2); The first input end and the second input end of the third Q shifter (Q3) are respectively connected to the output end of the seventh adder / subtractor (AS7) and the first output end of the first multiplier (E1), and the output end serves as the output end of the exponential operation module (12) to output the exponential operation result.
10. A hardware CORDIC acceleration chip for Softmax layer int8 quantization, characterized in that: The chip comprises the system as claimed in any one of claims 1-9.