Reconfigurable operator structure, computational method, and hardware architecture
By designing a reconfigurable operator structure at the hardware level, storing the segmented approximate function parameters of the activation function, and implementing hardware calculation of the activation function through data comparison and high-order function calculation structure, the problem of slow speed and high power consumption of the calculation of the activation function at the software level is solved, and efficient and low-power activation function calculation is achieved.
Patent Information
- Application Number
- CN202110996875.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2041-08-27
AI Technical Summary
In the prior art, the calculation of activation functions through software level has problems such as slow speed and high power consumption, which cannot meet the pursuit of AI for computing speed and power consumption.
A reconfigurable operator structure is provided, including a parameter storage unit, a data comparison unit and a high-order function calculation structure, and the calculation of the activation function is realized at the hardware level. This structure stores the calculation parameters of the segmented approximation function of the activation function, compares the range of input data, and performs hardware calculations of the higher order function based on the calculation parameters of the target range.
The calculation process of the activation function is realized through hardware, which improves the computing efficiency and reduces power consumption. It has a simple and reliable structure, and is easy to expand and implement. This structure has the reconfigurability to adapt to various activation functions.
Smart Images

Figure CN113780540B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a reconfigurable operator structure, a computing method and a hardware architecture. Background Art
[0002] As an indispensable part of neural networks and deep learning algorithms, the calculation of activation functions directly affects the execution efficiency of the entire algorithm. For example, when implementing CNN (Convolutional Neural Network) on VLSI (Very Large Scale Integration), the activation function layer must be implemented.
[0003] At present, the activation function calculation based on artificial intelligence chips is mostly designed and implemented at the software level. The activation function usually contains a large number of nonlinear operations. For example, the sigmoid function involves exponential operations. Nonlinear operations at the software level will generate a considerable amount of calculations, occupy a large amount of computing resources and computing time, resulting in slow computing speed and high power consumption. The general design method at the software level can no longer meet the pursuit of speed and power consumption in the development of artificial intelligence. Therefore, a new implementation method is urgently needed to improve the computing efficiency of the activation function and reduce its computing power consumption. Summary of the invention
[0004] The present invention provides a reconfigurable operator structure, a calculation method and a hardware architecture, which are used to solve the problems of slow speed and high power consumption in calculating activation functions at the software level in the prior art.
[0005] The present invention provides a reconfigurable operator structure, comprising:
[0006] A parameter storage unit, used to store calculation parameters of each piecewise approximate function of each activation function, wherein the piecewise approximate function is a high-order function;
[0007] A data comparison unit, used for selecting an input range to which the input data belongs as a target range from each input range of each piecewise approximate function of the target activation function in each activation function;
[0008] A control unit, the control unit being connected to the parameter storage unit and the data comparison unit respectively, and being used for selecting, from the parameter storage unit, calculation parameters of the piecewise approximation function corresponding to the target range of the target activation function;
[0009] A high-order function calculation structure, wherein the high-order function calculation structure is connected to the parameter storage unit, and the high-order function calculation structure is a hardware structure for calculating the high-order function, and is used to determine the calculation result of the input data under the target activation function based on the calculation parameters of the piecewise approximate function corresponding to the target range.
[0010] According to a reconfigurable operator structure provided by the present invention, the high-order function is in the form of multiple iterations of the splitting function;
[0011] The high-order function calculation structure includes a calculation unit, and the calculation unit is used to calculate the split function;
[0012] The output end of the calculation unit is connected to the parameter input end of the calculation unit, and the input parameters of the next splitting function calculation input by the parameter input end include the calculation result of the current splitting function output by the output end.
[0013] A reconfigurable operator structure provided according to the present invention further includes:
[0014] A range storage unit, the range storage unit is connected to the data comparison unit, and is used to store each input range of each piecewise approximation function of each activation function.
[0015] According to a reconfigurable operator structure provided by the present invention, the high-order function calculation structure further includes an input register, a parameter register, a data selector and a calculation unit register;
[0016] The input of the input register is the input data, and the output of the input register is connected to the variable input of the calculation unit;
[0017] The input of the parameter register is the calculation parameter, and the output of the parameter register is directly connected to the parameter input of the calculation unit, or is connected to the parameter input of the calculation unit through the data selector;
[0018] The output end of the calculation unit is connected to the parameter input end of the calculation unit through the calculation unit register and the data selector.
[0019] According to a reconfigurable operator structure provided by the present invention, the parameter register includes a first parameter register and a second parameter register;
[0020] The output end of the first parameter register is connected to the first parameter input end of the calculation unit through the data selector, and the output end of the second parameter register is directly connected to the second parameter input end of the calculation unit.
[0021] According to a reconfigurable operator structure provided by the present invention, the input of the first parameter register is the highest-order coefficient in the calculation parameters of the piecewise approximate function corresponding to the target range, and the input of the second parameter register includes coefficients other than the highest-order coefficient in the calculation parameters of the piecewise approximate function corresponding to the target range.
[0022] According to a reconfigurable operator structure provided by the present invention, the input end of the first parameter register is connected to the output end of the second parameter register.
[0023] According to a reconfigurable operator structure provided by the present invention, the data comparison unit includes a plurality of parallel comparators;
[0024] The plurality of parallel comparators are used to compare the input data in parallel with the upper and lower limits of the input range of each piecewise approximate function to obtain a comparison result, and the comparison result is used to determine the input range to which the input data belongs.
[0025] According to a reconfigurable operator structure provided by the present invention, each piecewise approximate function of each activation function is determined based on an approximate calculation method adapted to each activation function.
[0026] The present invention also provides a calculation method based on the above reconfigurable operator structure, comprising:
[0027] Determine the target activation function and input data;
[0028] The input data is input into the reconfigurable operator structure, the reconfigurable operator structure is controlled to perform calculation of the target activation function, and a calculation result output by the reconfigurable operator structure is obtained.
[0029] The present invention also provides a hardware architecture of a reconfigurable operator structure, including the above-mentioned reconfigurable operator structure.
[0030] The reconfigurable operator structure, calculation method and hardware architecture provided by the present invention store the calculation parameters of each piecewise approximate function of each activation function through a parameter storage unit, select the target range to which the input data belongs through a data comparison unit, and determine the calculation result based on the calculation parameters of the piecewise approximate function corresponding to the target range through a high-order function calculation structure. The calculation process of the activation function is all realized by hardware, because its structure and specific calculation tasks are simple, the calculation efficiency is higher, the calculation process is more reliable, and it is easier to expand and more convenient to realize. And for different activation functions, only the parameter storage unit is required to pre-store the calculation parameters of the piecewise approximate function of the activation function, and the calculation parameters in the high-order function calculation structure can be replaced during the specific calculation. The transformation between various activation functions is very simple and has the reconfigurability to adapt to various different activation functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly described below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0032] Figure 1 One of the structural schematic diagrams of the reconfigurable operator structure provided by the present invention;
[0033] Figure 2 One of the structural schematic diagrams of the high-order function calculation structure provided by the present invention;
[0034] Figure 3 The second structural diagram of the high-order function calculation structure provided by the present invention;
[0035] Figure 4 The third structural diagram of the high-order function calculation structure provided by the present invention;
[0036] Figure 5 A schematic diagram of the structure of an 8-bit data comparator provided by the present invention;
[0037] Figure 6 The second structural schematic diagram of the reconfigurable operator structure provided by the present invention;
[0038] Figure 7 The third structural diagram of the reconfigurable operator structure provided by the present invention;
[0039] Figure 8 A schematic diagram of a flow chart of a calculation method for a reconfigurable operator structure provided by the present invention;
[0040] Fig. 9 The hardware architecture of the reconfigurable operator structure provided by the present invention;
[0041] Description of reference numerals:
[0042] a1-input register; a2-first parameter register; a3-second parameter register;
[0043] a4-computation unit register; b1-data selector; b2-computation unit. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Figure 1 One of the structural diagrams of the reconfigurable operator structure provided by the present invention is as follows: Figure 1 As shown, the structure includes:
[0046] A parameter storage unit 130, used to store calculation parameters of each piecewise approximate function of each activation function, wherein the piecewise approximate function is a high-order function;
[0047] A data comparison unit 110, configured to select an input range to which the input data belongs as a target range from each input range of each piecewise approximate function of the target activation function in each activation function;
[0048] A control unit 120, the control unit 120 is connected to the parameter storage unit 130 and the data comparison unit 110, and is used to select the calculation parameters of the piecewise approximation function corresponding to the target range of the target activation function from the parameter storage unit 130;
[0049] The high-order function calculation structure 140 is connected to the parameter storage unit 130. The high-order function calculation structure 140 is a hardware structure for calculating high-order functions, and is used to determine the calculation results of the input data under the target activation function based on the calculation parameters of the piecewise approximate function corresponding to the target range.
[0050] Specifically, the reconfigurable operator structure provided by the embodiment of the present invention is a hardware architecture, in which the data comparison unit 110 can be implemented by a comparator circuit structure, the control unit 120 can be implemented by a hardware logic circuit with data search and mapping functions, the parameter storage unit 130 is a hardware memory, and the high-order function calculation structure 140 is used to calculate high-order functions, which can be implemented by a logic circuit adapted to high-order function calculation. Compared with the activation function calculation method at the software level, the activation function calculation at the hardware level has a simple structure and specific calculation tasks, higher calculation efficiency, and a more reliable calculation process. At the same time, it is easier to expand and more convenient to implement.
[0051] It should be noted that the reconfigurable operator structure provided in the embodiment of the present invention can be used to realize the calculation of any activation function, without limiting the type and parameters of the activation function. Any activation function can be piecewise fitted by an approximate algorithm, thereby converting a complex activation function into a high-order function including multiple pieces, so that each piece of the high-order function can be calculated by the activation calculation unit built into the reconfigurable operator structure.
[0052] Here, the parameter storage unit 130 is used to store the calculation parameters of each piecewise approximate function of each activation function. Among them, each activation function can be one or more arbitrary activation functions, such as the common Relu (Rectified Linear Unit, linear rectifier function), Sigmoid (S-type function), Tanh (Hyperbolic Tangent, hyperbolic tangent), etc. For a single activation function included therein, a piecewise approximate fitting can be performed by an approximate calculation method to obtain each piecewise approximate function of the activation function. The piecewise approximate fitting referred to here is to divide a plurality of input ranges according to the input size of the activation function, and perform approximate fitting on the segments of the activation function under each input range, so as to obtain the piecewise approximate function under each input range. In order to ensure the calculation accuracy of the activation function while ensuring the hardware feasibility and calculation simplification of the subsequent activation calculation unit, each piecewise approximate function of each activation function in the embodiment of the present invention is a high-order function. Compared with a linear function, a high-order function can more accurately fit a complex and changeable activation function. The high-order function referred to here can be a cubic, a quartic, etc. For example, a cubic function can be expressed as p1x 3 +p2x 2 +p3x+p4, where the values of p1, p2, p3 and p4 are the calculation parameters of the piecewise approximate function.
[0053] Specifically, in the calculation process of the activation function, any one or more of the activation functions can be used as the target activation function. The choice of the target activation function can be determined by the user. When facing different problems and training different models, it is necessary to select a suitable activation function as the target activation function. It should be noted that when there are multiple target activation functions, the calculation process of each target activation function is consistent with the calculation process of a single target activation function. The calculation process of multiple target activation functions can be executed in parallel or sequentially. For ease of understanding, the following is an example of a single target activation function.
[0054] The input data applied to the target activation function calculation is input into the data comparison unit 110. The data comparison unit 110 has a data comparison function, and can compare the input data with the input range of each piecewise approximate function of the target activation function. Here, the input range of each piecewise approximate function can be used as the range of the input data of each piecewise approximate function. For example, the target activation function can be divided into 4 pieces of piecewise approximate functions, and the input range of each piecewise approximate function can be [0, A1), [A1, A2), [A2, A3) and [A3, ∞). The data comparison unit 110 can determine which input range the input data specifically belongs to through the data comparison function, and thus select the input range to which the input data belongs as the target range.
[0055] The control unit 120 is connected to the data comparison unit 110 and the parameter storage unit 130. After the target range is selected, the control unit 120 can receive the target interval output by the data comparison unit 110, and by accessing the parameter storage unit 130, from the calculation parameters of each piecewise approximation function of each activation function pre-stored in the parameter storage unit 130, according to the corresponding relationship between each piecewise approximation function under the target activation function and the input range, select the calculation parameters of the piecewise approximation function corresponding to the target range of the target activation function, that is, the piecewise approximation function of the target activation function used for the input data calculation, and control the parameter storage unit 130 to input the calculation parameters of the piecewise approximation function corresponding to the target range of the target activation function into the high-order function calculation structure 140.
[0056] After receiving the calculation parameters of the piecewise approximate function corresponding to the target range transmitted by the parameter storage unit 130, the high-order function calculation structure 140 can realize the calculation of the input data under the target activation function and obtain the calculation result of the input data under the target activation function. Here, the calculation task undertaken by the high-order function calculation structure 140 is to use the input data as the independent variable, the calculation parameters of the piecewise approximate function corresponding to the target range to which the input data belongs as the calculation parameters of the high-order function, and to perform calculations with the fixed calculation logic of the high-order function. Therefore, for different input data or different activation functions, the calculation logic executed by the high-order function calculation structure 140 is the same, and only the independent variables and calculation parameters involved in the calculation change.
[0057] The high-order function calculation structure 140 itself is a logic calculation hardware structure, and its calculation logic is a fixed high-order function calculation logic. Considering that all the piecewise approximate functions obtained by piecewise approximate fitting for different activation functions are high-order functions of the same degree, the high-order function calculation structure 140 can be a logic calculation hardware structure fixed for a high-order function of one degree. When calculating for different input data and different types of activation functions, the high-order function calculation structure 140 itself does not require any adjustment, and only the input of the high-order function calculation structure 140 needs to be changed.
[0058] The reconfigurable operator structure provided by the embodiment of the present invention stores the calculation parameters of each piecewise approximate function of each activation function through a parameter storage unit, selects the target range to which the input data belongs through a data comparison unit, and determines the calculation result based on the calculation parameters of the piecewise approximate function corresponding to the target range through a high-order function calculation structure. The calculation process of the activation function is implemented by hardware. Because of its simple structure and specific calculation tasks, the calculation efficiency is higher, the calculation process is more reliable, and it is easier to expand and more convenient to implement. And for different activation functions, only the parameter storage unit is required to pre-store the calculation parameters of the piecewise approximate function of the activation function, and the calculation parameters in the high-order function calculation structure can be replaced during the specific calculation. The transformation between various activation functions is very simple and has the reconfigurability to adapt to various different activation functions.
[0059] In addition, the application of high-order function computing structure to realize the calculation of piecewise approximate functions in pure hardware can ensure the universality of various activation function calculations while giving full play to the advantages of high efficiency and strong reliability of the hardware structure itself in data calculation.
[0060] Based on the above embodiment, the input of the high-order function calculation structure includes input data and calculation parameters of the piecewise approximate function corresponding to the target range. The calculation parameters include multiple coefficients of different orders, and the order in which the multiple coefficients of different orders are input into the high-order function calculation structure is not specifically limited in the embodiment of the present invention.
[0061] Furthermore, the input of the high-order function calculation structure includes two parts, one part is the independent variable, that is, the input data x, and the other part is the parameters required for calculation, that is, the calculation parameters of the piecewise approximate function corresponding to the target range from the parameter storage unit, such as the values of p1, p2, p3 and p4 under the cubic function. These two parts of input can be input into the high-order function calculation structure according to the pre-set input logic to participate in the high-order function calculation, so as to obtain the calculation result of the input data under the target activation function.
[0062] Regarding the order in which the calculation coefficients of the high-order function are input into the high-order function calculation structure, considering that the higher the order, the more times the multiplication operation with the independent variable is involved, in the high-order function calculation structure that is input sequentially according to the clock, the input order is higher. Therefore, for coefficients of different orders in the calculation parameters, the higher the order, the earlier they are input into the high-order function calculation structure, and the lower the order, the later they are input into the high-order function calculation structure.
[0063] In addition, regarding the order in which the calculation coefficients of the high-order function are input into the high-order function calculation structure, the coefficients of different orders can also be input in order from low to high order, or input in parallel, and the embodiment of the present invention does not make specific limitations on this.
[0064] Based on any of the above embodiments, the high-order function is represented by a form of multiple iterations of the splitting function;
[0065] The high-order function calculation structure includes a calculation unit, and the calculation unit is used to calculate the split function;
[0066] The output end of the calculation unit is connected to the parameter input end of the calculation unit, and the input parameters of the next splitting function calculation input by the parameter input end include the calculation result of the current splitting function output by the output end.
[0067] Specifically, based on the functional characteristics of the high-order function itself, the high-order function can be split to represent it, thereby obtaining a high-order function representation that iterates the split function multiple times. For example, the multiplication and addition operation can be used as a split function, and the corresponding cubic function p1x 3 +p2x 2 +p3x+p4 can be expressed in the form of [(p1x+p2)·x+p3]·x+p4. The expression of the cubic function here can be understood as multiple iterations of the multiplication and addition operation O=W*Y+Z. Corresponding to the cubic function, O in the multiplication and addition operation is the output of a single multiplication and addition operation, W and Y are two multipliers in the multiplication and addition operation, one of the multipliers is the coefficient in the calculation parameter, and the other multiplier is the input data. Z is the addend in the multiplication and addition operation, specifically the coefficient in the calculation parameter.
[0068] By expressing a high-order function as a form of multiple iterations of a split function, the part of the high-order function calculation structure used to implement the high-order function calculation, that is, the calculation unit, can be set as a unit for implementing the split function calculation, and through the connection relationship outside the calculation unit, that is, connecting the output end of the calculation unit with the parameter input end of the calculation unit, the calculation unit itself has the ability of iterative calculation. Thus, the high-order function calculation under a single calculation unit is realized through the ability of the calculation unit itself to calculate the split function and the ability of the calculation unit determined by the connection relationship of the calculation unit to iterate the calculation unit.
[0069] Similarly, taking the cubic function as a form of three iterations of the split function, that is, [(p1x+p2)·x+p3]·x+p4 as an example, when the calculation unit performs the first split function calculation, since there is no calculation result of the previous split function, the calculation parameters p1, p2, and input data x are used as inputs to perform the split function calculation, and the calculation result p1x+p2 of the first split function is obtained;
[0070] The calculation result p1x+p2 of the first splitting function is transmitted to the parameter input end of the computing unit through the output end of the computing unit; when performing the second splitting function calculation, since there is the calculation result p1x+p2 of the previous splitting function, the previous calculation result p1x+p2, the calculation parameter p3, and the input data x are used as inputs to perform the splitting function calculation, and the calculation result (p1x+p2)·x+p3 of the second splitting function is obtained;
[0071] The calculation result (p1x+p2)·x+p3 of the second splitting function is transmitted to the parameter input end of the computing unit through the output end of the computing unit; when performing the third splitting function calculation, since there is the calculation result (p1x+p2)·x+p3 of the previous splitting function, the previous calculation result (p1x+p2)·x+p3, the calculation parameter p4, and the input data x are used as input to execute the splitting function calculation, and the calculation result of the third splitting function [(p1x+p2)·x+p3]·x+p4=p1x is obtained. 3 +p2x 2 +p3x+p4, at this time the three iterations of the splitting function are completed and the higher-order function calculation is completed.
[0072] The reconfigurable operator structure provided by the embodiment of the present invention realizes the calculation of the high-order function by representing the high-order function as a form of multiple iterations of the split function, and setting a calculation unit for performing multiple iterative calculations on the split function. The calculation of the high-order function is realized in the form of iterative calculation. Compared with the traditional solution of setting a dedicated calculation unit for each operation in the high-order function calculation, the reuse capability of the calculation unit is enhanced, the hardware scale required for the high-order function calculation is reduced, the hardware cost is reduced, and it is helpful to improve the integration level of the reconfigurable operator structure.
[0073] And because each piecewise approximate function of different activation functions is a high-order function, each piecewise approximate function of different objective functions can be expressed as multiple iterations of the same splitting function. Therefore, when performing operations on different activation functions, a high-order function calculation structure can also be applied.
[0074] Based on any of the above embodiments, the high-order function calculation structure further includes an input register, a parameter register, a data selector and a calculation unit register;
[0075] The input of the input register is input data, and the output of the input register is connected to the variable input of the calculation unit;
[0076] The input of the parameter register is a calculation parameter, and the output of the parameter register is directly connected to the parameter input of the calculation unit, or is connected to the parameter input of the calculation unit through a data selector;
[0077] The output end of the calculation unit is connected to the parameter input end of the calculation unit through the calculation unit register and the data selector.
[0078] Specifically, the input register, parameter register and calculation unit register are registers for input data, calculation parameters and calculation unit output results, respectively. In the embodiment of the present invention, the input register may be one or more. When there are multiple input registers, the high-order function calculation structure can support the parallel calculation of multiple input data under the target activation function. The parameter register in the embodiment of the present invention can also be one or more single-input registers or multiple-input registers. When there is only one single-input register in the parameter register, the coefficients of multiple different orders in the calculation parameters need to be input one by one into the single output register in a pre-set order. When the parameter register contains a multi-input register or multiple registers, each input port may correspond to a coefficient of one order, or one input port may correspond to coefficients of multiple orders. The embodiment of the present invention does not specifically limit this.
[0079] The calculation unit is a calculation structure used to implement fixed operation logic, which can be used to implement various calculation operations, such as addition, multiplication, subtraction, shift, etc. The result output by the calculation unit during the activation function operation can be understood as the intermediate result of the activation function operation, and the intermediate result can be input to the data selector through the calculation unit register.
[0080] The input data output by the input register can be used as an independent variable during calculation and directly input into the variable input terminal of the calculation unit. The calculation parameters output by the parameter register can be used as parameters during calculation and directly input into the parameter input terminal of the calculation unit, or they can be input into a data selector, which selects whether to apply the calculation parameters or the intermediate results in the operation process as the parameters during calculation and input into the parameter input terminal of the calculation unit. It should be noted that the variable input terminal and the parameter input terminal of the calculation unit are different input terminals of the calculation unit. For example, when the calculation unit includes two or more input terminals, one or more of them can be selected as the variable input terminal, and the remaining one or more can be selected as the parameter input terminal.
[0081] In addition, in the case where there are multiple parameter registers or the parameter register is a multi-input and multi-output register, the output ends of some parameter registers can be directly connected to the parameter input ends of the computing unit, and the output ends of other parameter registers can be connected to the parameter input ends of the computing unit through a data selector.
[0082] Based on any of the above embodiments, Figure 2 One of the structural diagrams of the high-order function calculation structure provided by the present invention is as follows: Figure 2 As shown, the parameter register includes a first parameter register a2 and a second parameter register a3;
[0083] The output end of the first parameter register a2 is connected to the first parameter input end of the calculation unit b2 through the data selector b1, and the output end of the second parameter register a3 is directly connected to the second parameter input end of the calculation unit b2. Here, the first parameter register a2 and the second parameter register a3 can be one or more registers. "First" and "second" are only used to distinguish between the two types of registers. The difference between the two types of registers lies in the different direct connection relationship between the output end and the parameter input end of the calculation unit b2.
[0084] The first parameter register a2 is connected to the first parameter input terminal of the calculation unit b2 through the data selector b1, and the data selector b1 is also connected to the output terminal of the calculation unit register a4. The data selector b1 can select the output of the first parameter register a2 or the output of the calculation unit register a4 as the input of the first parameter input terminal of the calculation unit b2. The output terminal of the second parameter register a3 is directly connected to the second parameter input terminal of the calculation unit b2, and the output of the second parameter register a3 is the input of the second parameter input terminal of the calculation unit b2.
[0085] Based on any of the above embodiments, the calculation unit in the high-order function calculation structure may be a multiplier-adder, wherein the input of the input register is input data, and the output of the input register is connected to the first multiplier terminal of the multiplier-adder, where the first multiplier terminal is the variable input terminal;
[0086] The inputs of the first parameter register and the second parameter register are calculation parameters, the output end of the first parameter register is connected to the first input end of the data selector, and the output end of the second parameter register is connected to the addend end of the multiplier-adder, where the addend end is the second parameter input end;
[0087] The output end of the data selector is connected to the second multiplier end of the multiplier-adder, and the output end of the multiplier-adder is connected to the second input end of the data selector through the multiplier-add register, where the second multiplier end is the first parameter input end.
[0088] Specifically, considering that high-order function calculations can be split into multiple multiplication and addition operations, for example, the cubic function p1x 3 +p2x 2 +p3x+p4 can also be expressed as [(p1x+p2)·x+p3]·x+p4. Therefore, the high-order function calculation structure used to realize high-order function calculation includes a multiplier for performing multiplication and addition operations. The multiplication and addition operation here can be understood as O=W*Y+Z, O is the output of a single multiplication and addition operation, W and Y are the inputs of the first multiplier terminal and the second multiplier terminal of the multiplier, and Z is the input of the addend terminal of the multiplier. The multiplier can be implemented by a DSP (Digital Signal Process) or by a combination of a multiplier and an adder, and the embodiment of the present invention does not specifically limit this.
[0089] Since the calculation of a high-order function needs to be split into multiple multiplication and addition operations, registers need to be set separately for the input data, calculation parameters, and multiplication and addition results, wherein the input register is used to implement the participation of the input data in the multiplication and addition operation at each clock, the first parameter register and the second parameter register are used to implement the participation of the calculation parameters in the multiplication and addition operation at each clock, and the multiplication and addition register is used to implement the participation of the multiplication and addition result at each clock in the next multiplication and addition operation.
[0090] Furthermore, considering that in different multiplication and addition operations, one multiplier in the multiplication and addition operation is always the input data, and the other multiplier is the highest-order coefficient in the calculation parameter in the first multiplication and addition operation, and the result of the previous multiplication and addition operation in each subsequent multiplication and addition operation. Therefore, for the transformation of the other multiplier, a data selector is set in the high-order function calculation structure, which is used to select the calculation parameter or the result of the multiplication and addition operation as the input of the second multiplier terminal in different multiplication and addition operations.
[0091] Based on any of the above embodiments, the input of the first parameter register is the highest order coefficient in the calculation parameters of the piecewise approximate function corresponding to the target range, and the input of the second parameter register includes the coefficients other than the highest order coefficient in the calculation parameters of the piecewise approximate function corresponding to the target range.
[0092] Here, the first parameter register and the second parameter register input the coefficients of different orders in the calculation parameters at the same clock. Taking the cubic function as an example, the orders of p1, p2, p3 and p4 decrease. When the first parameter register inputs p1, the second parameter register inputs p2. Thereafter, the order of the coefficients input to the second parameter register decreases with each clock.
[0093] Figure 3 The second structural diagram of the high-order function calculation structure provided by the present invention is as follows: Figure 3As shown, the calculation unit b2 is specifically a multiplier adder, wherein the variable input terminal is the first multiplier terminal W, the first parameter input terminal is the second multiplier terminal Y, and the second parameter input terminal is the addend terminal Z.
[0094] X is the channel of input data, A is the input channel of the first parameter register a2, and B is the input channel of the second parameter register a3. The input data x is connected to the first multiplier terminal W of the calculation unit b2 through the input register a1, the input of the first parameter register a2 is the high-order coefficient, and the input of the second parameter register a3 is the low-order coefficient. Here, the high-order coefficient is the highest-order coefficient, and the low-order coefficient is the coefficient other than the high-order coefficient.
[0095] The two inputs of the data selector b1 are the output of the first parameter register a2 and the output of the calculation unit register a4. When the calculation unit register a4 has no output, the output of the first parameter register a2 is selected to be output. When the calculation unit register a4 has an output, the output of the calculation unit register a4 is selected to be output. The output of the data selector b1 is connected to the second multiplier terminal Y of the calculation unit b2.
[0096] The output of the second parameter register a3 is connected to the addend terminal a3 of the calculation unit b2, and the output terminal O of the calculation unit b2 is connected to the calculation unit register a4.
[0097] Taking the piecewise approximate function as a cubic function as an example, Figure 3 The operation flow of the high-order function calculation structure shown is as follows:
[0098] In the first clock, channel X inputs x, channel A inputs p1, and channel B inputs p2.
[0099] At the second clock, channel B inputs p3;
[0100] The input register a1 outputs x to the first multiplier terminal W of the calculation unit b2, the first parameter register a2 outputs p1 to the second multiplier terminal Y of the calculation unit b2, the second parameter register a3 outputs p2 to the addend terminal Z of the calculation unit b2, and the output terminal O of the calculation unit b2 outputs output=p1x+p2.
[0101] At the third clock, channel B inputs p4;
[0102] The input register a1 outputs x to the first multiplier terminal W of the calculation unit b2, the data selector b1 is selected so that the input of the second multiplier terminal Y of the calculation unit b2 is the output output of the second clock, the second parameter register a3 outputs p3 to the addend terminal Z of the calculation unit b2, and the output terminal O of the calculation unit b2 outputs output=(p1x+p2)x+p3.
[0103] At the fourth clock, the input register a1 outputs x to the first multiplier terminal W of the computing unit b2, the data selector b1 is selected so that the input of the second multiplier terminal Y of the computing unit b2 is the output output of the third clock, the second parameter register a3 outputs p4 to the addend terminal Z of the computing unit b2, and the output terminal O of the computing unit b2 outputs output = [(p1x + p2)x + p3]x + p4 = p1x 3 +p2x 2 +p3x+p4.
[0104] Based on any of the above embodiments, Figure 4 The third structural diagram of the high-order function calculation structure provided by the present invention is as follows: Figure 4 As shown, the calculation unit b2 is specifically a multiplier adder, wherein the variable input terminal is the first multiplier terminal W, the first parameter input terminal is the second multiplier terminal Y, and the second parameter input terminal is the addend terminal Z.
[0105] Different from Figure 3 Enter the coefficients of two orders in the calculation parameters at once, Figure 4 The input end of the first parameter register a2 is connected to the output end of the second parameter register a3. Figure 4 In the figure, X is the channel of the input data, and A is the input channel of the second parameter register a2.
[0106] Taking the piecewise approximate function as a cubic function as an example, Figure 4 The operation flow of the high-order function calculation structure shown is as follows:
[0107] At the first clock,
[0108] Channel X input is x, channel A input is p1.
[0109] At the second clock, channel A inputs p2;
[0110] The input register a1 outputs x to the first multiplier terminal W of the calculation unit b2, and the second parameter register b3 outputs p1 to the first parameter register b2.
[0111] At the third clock, channel A inputs p3;
[0112] The input register a1 outputs x to the first multiplier terminal W of the calculation unit b2, the first parameter register b2 outputs p1 to the second multiplier terminal Y of the calculation unit b2, the second parameter register a3 outputs p2 to the addend terminal Z of the calculation unit b2, and the output terminal O of the calculation unit b2 outputs output=p1x+p2.
[0113] At the fourth clock, channel A inputs p4;
[0114] The input register a1 outputs x to the first multiplier terminal W of the calculation unit b2, the data selector b1 is selected so that the input of the second multiplier terminal Y of the calculation unit b2 is the output output of the third clock, the second parameter register a3 outputs p3 to the addend terminal Z of the calculation unit b2, and the output terminal O of the calculation unit b2 outputs output=(p1x+p2)x+p3.
[0115] At the fifth clock, the input register a1 outputs x to the first multiplier terminal W of the computing unit b2, the data selector b1 is selected so that the input of the second multiplier terminal Y of the computing unit b2 is the output output of the fourth clock, the second parameter register a3 outputs p4 to the addend terminal Z of the computing unit b2, and the output terminal O of the computing unit b2 outputs output = [(p1x + p2)x + p3]x + p4 = p1x 3 +p2x 2 +p3x+p4.
[0116] Based on any of the above embodiments, the data comparison unit includes a plurality of parallel comparators;
[0117] A plurality of parallel comparators are used to compare the input data with the upper and lower limits of the input range of each piecewise approximate function in parallel to obtain a comparison result, and the comparison result is used to determine the input range to which the input data belongs.
[0118] Specifically, the data comparison unit can be implemented by multiple parallel comparators. For example, when judging whether the input data belongs to the input range of any piecewise approximate function, the input data can be compared with the upper and lower limits of the input respectively. If the input data is greater than or equal to the lower limit and less than or equal to the upper limit, it means that the input data belongs to the input range. If the input data is greater than the upper limit or less than the lower limit, it means that the input data does not belong to the input range. Here, the task of comparing the input data with the upper limit or lower limit is implemented by a comparator. Comparing multiple comparators in parallel can improve the comparison efficiency, especially when there are many piecewise approximate functions divided by a single activation function, which can greatly shorten the time required for comparison.
[0119] Furthermore, when data comparison requires a comparator with a larger number of bits, the comparator can be obtained by cascading existing comparators. For example, an 8-bit data comparator can be obtained by cascading two 4-bit data comparators. Figure 5 The schematic diagram of the structure of the 8-bit data comparator provided by the present invention is as follows: Figure 5 As shown, the comparison result of the 4-bit data comparator 1 is connected to the input end of the 4-bit data comparator 2. Where a represents input data, b represents the upper limit or lower limit of the input range of a piecewise approximate function, and Output represents the final comparison result of the 8-bit data comparator thus constructed.
[0120] Based on any of the above embodiments, Figure 6 The second structural diagram of the reconfigurable operator structure provided by the present invention is as follows: Figure 6 As shown, the reconfigurable operator structure also includes:
[0121] The range storage unit 150 is used to store each input range of each piecewise approximation function of each activation function.
[0122] Specifically, the range storage unit 150 may be used to provide a comparison basis for the data comparison unit 110 to determine the target range.
[0123] The range storage unit 150 stores the input ranges of each piecewise approximate function of each activation function. When determining the target activation function and calculating the target activation function, the range storage unit 150 can transmit the input ranges of each piecewise approximate function under the target activation function stored in itself to the data comparison unit 110. The data comparison unit 110 can compare the input data with the upper and lower limits of each input range transmitted by the range storage unit 150, thereby determining the input range to which the input data belongs as the target range, and transmitting the determined target range to the control unit 120. After receiving the target range, the control unit 120 can extract the calculation parameters of the piecewise approximate function corresponding to the target range under the target activation function from the parameter storage unit 130, and transmit the extracted calculation parameters to the high-order function calculation structure 140. The high-order function calculation structure 140 can determine the calculation result of the input data under the target activation function based on the calculation parameters of the piecewise approximate function corresponding to the target range.
[0124] It should be noted that in the high-order function calculation structure 140, the parameter register can be used as an input interface for calculating parameters, for example Figure 1 and Figure 6 The position indicated by the bold arrow connecting the parameter storage unit 130 to the high-order function calculation structure 140 reflects the input interface of the calculation parameters. Specifically, the input interface of the calculation parameters may include one or more ports for inputting the calculation parameters. The number of ports may be adjusted according to the number of parameter registers and the connection relationship between the parameter registers. For example, Figure 2 and Figure 3 In the example, the number of ports available for inputting calculation parameters can be 2. Figure 4 The number of ports available for calculation parameter input can be 1.
[0125] Based on any of the above embodiments, Figure 7 The third structural diagram of the reconfigurable operator structure provided by the present invention is as follows: Figure 7 As shown, compared to Figure 6The data comparison unit 110 includes a plurality of parallel comparators 111, each of which can compare the input data with the upper and lower limits of the input range, thereby obtaining the target range.
[0126] Based on any of the above embodiments, each piecewise approximate function of each activation function is determined based on an approximate calculation method adapted for each activation function.
[0127] Specifically, considering that different types of activation functions have different functional characteristics, different activation functions are adapted to different approximate calculation methods. Common approximate calculation methods can include Taylor expansion method, coordinate rotation digital calculation method CORDIC, polynomial approximate calculation method, piecewise linear interpolation method, etc.
[0128] Taking the Sigmoid function as an example, the Sigmoid function is a symmetrical S-shaped growth curve function, which can be applied to the piecewise Taylor expansion for piecewise approximate fitting. The results of the piecewise approximate fitting are shown in the following table:
[0129] Input range Taylor series expansion Maximum precision loss [0,4) <![CDATA[3.56*10 -3 *x 3 -5.71*10 -2 *x 2 +2.93*10 -1 *x+4.92*10 -1 ]]> <![CDATA[7.53*10 -3 <!-- 10 -->]]> [4,8) <![CDATA[4.96*10 -4 *x 3 -1.05*10 -2 *x 2 +7.51*10 -2 *x+8.19*10 -1 ]]> <![CDATA[5.23*10 -4 ]]> [8,15) <![CDATA[3.21*10 -6 *x 3 -1.22*10 -4 *x 2 +1.54*10 -3 *x+9.94*10 -1 ]]> <![CDATA[3.71*10 -5 ]]> [15,∞) 1 <![CDATA[3.0*10 -7 ]]>
[0130] This gives the calculation parameters for the input range [0,4):
[0131] p1=3.56*10 -3 , p2=-5.71*10 -2 , p3=2.93*10 -1 , p4=4.92*10 -1
[0132] Calculation parameters for input range [4,8):
[0133] p1=4.96*10 -4 , p2=-1.05*10 -2 , p3=7.51*10 -2 , p4=8.19*10 -1
[0134] Calculation parameters for input range [8,15):
[0135] p1=3.21*10 -6 , p2=-1.22*10 -4 , p3=1.54*10 -3 , p4=9.94*10 -1
[0136] Calculation parameters for input range [15,∞):
[0137] p1=0, p2=0, p3=0, p4=1
[0138] Specifically, when the data comparison unit performs comparison, the input data x may be compared with 4, 8, and 15 respectively.
[0139] Based on any of the above embodiments, Figure 8 A schematic diagram of a calculation method for a reconfigurable operator structure provided by the present invention is shown in FIG. Figure 8 As shown, the method includes:
[0140] Step 810, determine the target activation function and input data.
[0141] Step 820, inputting the input data into the reconfigurable operator structure, controlling the reconfigurable operator structure to perform calculation of the target activation function, and obtaining the calculation result output by the reconfigurable operator structure.
[0142] Specifically, the target activation function can be any activation function, and the input data is the data to be calculated relative to the target activation function. After obtaining the two, the two can be input into the reconfigurable operator structure, and the data comparison unit in the reconfigurable operator structure is called to compare the input data with the input range of each piecewise approximate function of the target activation function, so as to determine the target range, and the high-order function calculation structure is called to use the input data as the independent variable, and the calculation parameters of the piecewise approximate function corresponding to the target range of the input data are used as the calculation parameters of the high-order function, and the calculation is performed with the fixed calculation logic of the high-order function to obtain the calculation result.
[0143] The calculation method based on the reconfigurable operator structure provided by the embodiment of the present invention only needs to determine the target activation function and input data, and the activation function calculation based on hardware can be realized through the reconfigurable operator structure. Since the reconfigurable operator structure itself and the specific calculation task are simple, the calculation efficiency is high, the calculation process is reliable, and the implementation is convenient. And for different activation functions, only the parameter storage unit inside the reconfigurable operator structure needs to pre-store the calculation parameters of the piecewise approximate function of the activation function, and the calculation parameters in the high-order function calculation structure can be replaced during the specific calculation. The transformation between various activation functions is very simple and has the reconfigurability to adapt to various different activation functions.
[0144] Based on any of the above embodiments, Fig. 9 The hardware architecture of the reconfigurable operator structure provided by the present invention is as follows: Fig. 9 As shown, the hardware architecture of the reconfigurable operator structure includes the reconfigurable operator structure 90 as described in any of the above embodiments.
[0145] Here, the hardware architecture of the reconfigurable operator structure may be an ASIC (Application Specific Integrated Circuit).
[0146] Therefore, the hardware architecture of the reconfigurable operator structure including any of the above-mentioned embodiments also has all the advantages of the above-mentioned reconfigurable operator structure 90. Among them, the above-mentioned reconfigurable operator structure can be encapsulated to obtain the hardware architecture of the reconfigurable operator structure. The hardware architecture thus obtained stores the calculation parameters of each piecewise approximate function of each activation function through the parameter storage unit inside the reconfigurable operator structure, determines the target range to which the input data belongs through the data comparison unit, and determines the calculation result based on the calculation parameters of the piecewise approximate function corresponding to the target range through the high-order function calculation structure. The calculation process of the activation function is all implemented by hardware, because its structure and specific calculation tasks are simple, the calculation efficiency is higher, the calculation process is more reliable, and it is easier to expand and more convenient to implement. And for different activation functions, it is only necessary to pre-store the calculation parameters of the piecewise approximate function of the activation function in the parameter storage unit inside the reconfigurable operator structure, and replace the calculation parameters in the high-order function calculation structure during the specific calculation. The transformation between various activation functions is very simple and has the reconfigurability to adapt to various different activation functions.
[0147] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0148] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A hardware architecture of a reconfigurable operator structure, characterized in that: A reconfigurable operator structure is included, wherein the reconfigurable operator structure includes: A parameter storage unit, used to store calculation parameters of each piecewise approximate function of each activation function, wherein the piecewise approximate function is a high-order function; A data comparison unit, configured to select, from the input ranges of the respective piecewise approximation functions of the target activation function in the respective activation functions, an input range to which the input data belongs as a target range; A control unit, the control unit being connected to the parameter storage unit and the data comparison unit respectively, and being used for selecting, from the parameter storage unit, calculation parameters of the piecewise approximation function corresponding to the target range of the target activation function; A high-order function calculation structure, the high-order function calculation structure is connected to the parameter storage unit, the high-order function calculation structure is a hardware structure for calculating the high-order function, and is used to determine the calculation result of the input data under the target activation function based on the calculation parameters of the piecewise approximate function corresponding to the target range; The high-order function is represented by a form of multiple iterations of a split function, and the split function is a multiplication and addition operation; The high-order function calculation structure includes a calculation unit, the calculation unit is used to calculate the split function, and the high-order function calculation structure is applicable to the calculation of different activation functions; The output end of the calculation unit is connected to the parameter input end of the calculation unit, and the input parameters of the next splitting function calculation input by the parameter input end include the calculation result of the current splitting function output by the output end.
2. The hardware architecture of the reconfigurable operator structure according to claim 1, characterized in that: The reconfigurable operator structure also includes: A range storage unit, the range storage unit is connected to the data comparison unit, and is used to store each input range of each piecewise approximation function of each activation function.
3. The hardware architecture of the reconfigurable operator structure according to claim 1, characterized in that: The high-order function calculation structure also includes an input register, a parameter register, a data selector and a calculation unit register; The input of the input register is the input data, and the output of the input register is connected to the variable input of the calculation unit; The input of the parameter register is the calculation parameter, and the output of the parameter register is directly connected to the parameter input of the calculation unit, or is connected to the parameter input of the calculation unit through the data selector; The output end of the calculation unit is connected to the parameter input end of the calculation unit through the calculation unit register and the data selector.
4. The hardware architecture of the reconfigurable operator structure according to claim 3, characterized in that: The parameter registers include a first parameter register and a second parameter register; The output end of the first parameter register is connected to the first parameter input end of the calculation unit through the data selector, and the output end of the second parameter register is directly connected to the second parameter input end of the calculation unit.
5. The hardware architecture of the reconfigurable operator structure according to claim 4, characterized in that: The input of the first parameter register is the highest-order coefficient in the calculation parameters of the piecewise approximation function corresponding to the target range, and the input of the second parameter register includes the coefficients of the calculation parameters of the piecewise approximation function corresponding to the target range except the highest-order coefficient.
6. The hardware architecture of the reconfigurable operator structure according to claim 4, characterized in that: An input terminal of the first parameter register is connected to an output terminal of the second parameter register.
7. The hardware architecture of the reconfigurable operator structure according to any one of claims 1 to 6, characterized in that: The data comparison unit includes a plurality of parallel comparators; The plurality of parallel comparators are used to compare the input data in parallel with the upper and lower limits of the input range of each piecewise approximate function to obtain a comparison result, and the comparison result is used to determine the input range to which the input data belongs.
8. The hardware architecture of the reconfigurable operator structure according to any one of claims 1 to 6, characterized in that: Each piecewise approximate function of each activation function is determined based on an approximate calculation method adapted to each activation function.
9. A calculation method based on the reconfigurable operator structure according to any one of claims 1 to 8, characterized in that: include: Determine the target activation function and input data; The input data is input into the reconfigurable operator structure, the reconfigurable operator structure is controlled to perform calculation of the target activation function, and a calculation result output by the reconfigurable operator structure is obtained.
Citation Information
Patent Citations
Hardware circuit for processing activation function and chip
CN112651496A
Signal converter and signal conversion method
JP2000183753A