A circuit for implementing an activation function and a processor including the circuit
By designing a computing circuit containing multiple computing units, the problem of insufficient calculation performance of activation function and single hardware acceleration solution in the prior art is solved, and efficient and multifunctional activation function hardware acceleration is achieved, suitable for intelligent devices and edge computing.
Patent Information
- Application Number
- CN201911133061.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-19
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2039-11-19
AI Technical Summary
When performing the calculation of activation functions in neural networks, the existing technology has problems with poor computing performance, which affects the training and inference speed of neural networks. The existing hardware acceleration scheme is only designed for certain activation functions, lacking integration and versatility.
An operation circuit including an exponential operation unit, a division operation unit, a logistic regression function operation unit and a selection unit is designed. Through the combination of these units, hardware accelerated operations of a variety of activation functions are realized, including sigmoid, tanh and ReLU functions.
This design significantly improves the efficiency and accuracy of activation function operations, reduces the repeated design of circuits and improves integration, and is suitable for efficient neural network computing in smart devices and edge computing devices.
Smart Images

Figure CN112906876B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of processors, and particularly to circuit designs for neural network operations. Background Art
[0002] In the field of artificial intelligence, various neural network models have been proposed to handle problems in aspects such as image, video, audio, and natural language processing. A neural network model generally includes multiple network layers, with a large number of nodes in each layer. The values at each layer of nodes can be passed to the next layer of nodes and activated by the values of the previous layer of nodes. Therefore, in a neural network model, a large number of activation function calculations are required.
[0003] In the prior art field, activation function calculations are usually performed in software. Since this method only uses conventional processor instructions for processing, when performing a large number of activation function calculations, there is a problem of poor operation performance, which significantly affects the training and inference speeds of neural networks.
[0004] In the prior art field, there are also hardware acceleration solutions. However, existing hardware acceleration solutions only target one or several activation functions for hardware acceleration processing, without seeing the integration of operations of various activation functions and providing a hardware acceleration solution for multiple activation function operations simultaneously.
[0005] Therefore, a new operation circuit design solution is needed that can provide efficient hardware acceleration processing for activation function operations. Summary of the Invention
[0006] For this purpose, the present invention provides a new operation circuit, processor, and system-on-chip to attempt to solve or at least alleviate at least one of the above problems.
[0007] According to one aspect of the present invention, an operation circuit is provided that is suitable for performing operations on data. The operation circuit includes an exponential operation unit, a division operation unit, a logistic regression function operation unit, and a selection unit. The exponential operation unit is suitable for performing exponential operations on data; the division operation unit is suitable for performing division operations on data; the logistic regression (sigmoid) function operation unit is suitable for performing logistic regression function operations on data; and the selection unit is suitable for selecting, according to the operation mode, the operation output of the exponential operation unit, the division operation unit, or the logistic regression function operation unit as the operation result. The logistic regression function operation unit is coupled to the exponential operation unit and the division operation unit, and is suitable for, when performing a logistic regression function operation on data, using the exponential operation unit to perform the exponential operation in the logistic regression function operation and using the division operation unit to perform the division operation in the logistic regression function operation.
[0008] Optionally, in the arithmetic circuit according to the present invention, the exponential operation unit includes: an exponential data input module adapted to receive data to be subjected to exponential operation; an exponential selection module adapted to determine the interval to which the data belongs according to the magnitude of the data to be operated; an exponential parameter module adapted to provide a parameter corresponding to the determined interval, where the parameter is a parameter required for performing Taylor expansion on the exponential operation; an exponential multiplication module adapted to perform a multiplication operation using the provided parameter and the data to perform Taylor expansion processing on the data; and an exponential result output module adapted to obtain the Taylor expansion processing result of the multiplication module to provide an exponential operation result.
[0009] Optionally, in the arithmetic circuit according to the present invention, the exponential operation unit further includes an exponential operation control module adapted to, when receiving a valid input indication signal, instruct the exponential operation unit to start the exponential operation processing on the received data, and generate a valid output indication signal after a predetermined operation period of the exponential operation unit has passed, to indicate that an exponential operation result of the data is provided at the exponential result output module.
[0010] Optionally, in the arithmetic circuit according to the present invention, the division operation unit includes: a division data input module adapted to receive data to be subjected to division operation; a first division parameter selection module adapted to determine a first division parameter according to the magnitude of the data for performing division operation; a first division shift module adapted to shift the data for performing division operation according to the first division parameter; a second division parameter selection module adapted to select a second or third division parameter according to the data; a multiplication module adapted to perform a multiplication operation on the output of the first division shift module and the output of the second division parameter selection module; a second division shift module adapted to perform a shift operation on the output of the multiplication module; and a division result output module adapted to obtain the output of the second division shift module as an operation result of performing division operation on the said data.
[0011] Optionally, in the arithmetic circuit according to the present invention, the division operation unit further includes a division operation control module adapted to, when receiving a valid input indication signal, instruct the division operation unit to start the division operation processing on the received data, and generate a valid output indication signal after a predetermined operation period of the division operation unit has passed, to indicate that a division operation result of the data is provided at the division result output module.
[0012] Optionally, in the arithmetic circuit according to the present invention, the logistic regression function arithmetic unit includes: a logistic regression arithmetic data input module adapted to receive data to be subjected to a logistic regression function arithmetic; a first logistic regression arithmetic module coupled to the exponential arithmetic unit and adapted to send the data to the exponential arithmetic unit for exponential arithmetic and obtain an exponential arithmetic result; a second logistic regression arithmetic module coupled to the division arithmetic unit and adapted to send the exponential arithmetic result obtained by the first arithmetic module to the division arithmetic unit to obtain a division arithmetic result and multiply the exponential arithmetic result and the division arithmetic result to obtain a multiplication result; a logistic regression shift module adapted to perform a shift operation on the multiplication result of the second logistic regression arithmetic module; and a logistic regression arithmetic result output module adapted to obtain the shift output of the logistic regression shift module as the result of performing a logistic regression arithmetic on the data.
[0013] Optionally, in the arithmetic circuit according to the present invention, the logistic regression function arithmetic unit further includes: a logistic regression arithmetic control module adapted to, upon receiving a valid input indication signal, instruct the logistic regression arithmetic unit to start a logistic regression arithmetic process on the received data, and generate a valid output indication signal after a predetermined arithmetic period of the logistic regression arithmetic unit has elapsed to indicate that the logistic regression arithmetic result of the data is provided at the logistic regression arithmetic result output module.
[0014] Optionally, the arithmetic circuit according to the present invention further includes a hyperbolic tangent (tanh) function arithmetic unit adapted to perform a hyperbolic tangent function arithmetic on data; a selection unit adapted to select, according to an arithmetic mode, the arithmetic output of the exponential arithmetic unit, the division arithmetic unit, the logistic regression function arithmetic unit, or the hyperbolic tangent function arithmetic unit as an arithmetic result, and wherein the hyperbolic tangent function arithmetic unit is coupled to the logistic regression function arithmetic unit and is adapted to utilize the logistic regression function arithmetic unit to perform a logistic regression function arithmetic in the hyperbolic tangent function arithmetic when performing a hyperbolic tangent function arithmetic on data.
[0015] Optionally, in the arithmetic circuit according to the present invention, the hyperbolic tangent function arithmetic unit includes: a hyperbolic tangent arithmetic data input module adapted to receive data to be subjected to a hyperbolic tangent function arithmetic; a hyperbolic tangent arithmetic module coupled to the logistic regression function arithmetic unit and adapted to perform an inversion process on the data to be arithmetically operated on, send the data after the inversion process to the logistic regression function arithmetic unit to perform a logistic regression function arithmetic, and perform a shift process on the logistic regression function arithmetic result; and a hyperbolic tangent arithmetic result output module adapted to obtain the shift output of the hyperbolic tangent arithmetic module as the result of performing a hyperbolic tangent function arithmetic on the data.
[0016] Optionally, in the arithmetic circuit according to the present invention, the hyperbolic tangent function arithmetic unit further includes: a hyperbolic tangent function arithmetic control module, adapted to, when receiving a valid input indication signal, instruct the hyperbolic tangent function arithmetic unit to start the hyperbolic tangent function arithmetic processing on the received data, and after a predetermined arithmetic period of the hyperbolic tangent function arithmetic unit has passed, generate a valid output indication signal to indicate that the hyperbolic tangent function arithmetic result of the data is provided at the hyperbolic tangent function arithmetic result output module.
[0017] Optionally, the arithmetic circuit according to the present invention further includes a rectified linear unit (ReLU) arithmetic unit adapted to perform rectified linear function arithmetic on data, and a selection unit adapted to select, according to an arithmetic mode, the arithmetic output of the exponential arithmetic unit, the division arithmetic unit, the logistic regression function arithmetic unit, the hyperbolic tangent function arithmetic unit, or the rectified linear unit (ReLU) arithmetic unit as the arithmetic result.
[0018] Optionally, in the arithmetic circuit according to the present invention, the rectified linear function arithmetic unit includes: a rectified linear selection unit adapted to output a value of 0 when the value of the data to be subjected to rectified linear function arithmetic is less than 0; and output the data as the arithmetic result of the rectified linear function arithmetic when the value of the data is not less than 0.
[0019] Optionally, in the arithmetic circuit according to the present invention, when the most significant bit of the binary representation of the data is 1, it indicates that the value of the data is less than 0.
[0020] According to another aspect of the present invention, there is provided a processor including the arithmetic circuit according to the present invention.
[0021] According to still another aspect of the present invention, there is provided a system on chip including the instruction processing device or the processor according to the present invention.
[0022] According to still another aspect of the present invention, there is further provided an intelligent device including the system on chip according to the present invention.
[0023] According to the solution of the present invention, by specially designing the exponential arithmetic unit and the division arithmetic unit, the design of the activation function (such as the logistic regression (sigmoid) function) arithmetic unit based on the exponential arithmetic and the division arithmetic can be simplified, and the applicable range of the circuit is increased.
[0024] In addition, according to the solution of the present invention, for the exponential arithmetic unit, the data is divided into multiple intervals according to the size of the data to be subjected to exponential arithmetic, and different parameters are used for Taylor expansion in each interval to approximate the exponential arithmetic result, providing a high-precision and efficient exponential arithmetic circuit design solution.
[0025] In addition, according to the solution of the present invention, arithmetic circuits for other activation functions (such as the hyperbolic tangent tanh function and the rectified linear unit ReLU function) can also be provided. These circuits can also utilize the already designed exponential operation, division operation, and logistic regression (sigmoid) function operation circuits, thereby simplifying the circuit design.
[0026] In this way, according to the solution of the present invention, the hardware implementations of multiple activation functions can be integrated into a single circuit unit, reducing the redundant design of the circuit and improving the integration degree of the circuit. Integrating the arithmetic circuit according to the present invention into a processing chip such as a system-on-chip (SoC) can not only significantly improve the speed of the chip for performing neural network calculations, but also significantly reduce the energy consumption, making such a processing chip more suitable for edge and end computing devices with high requirements for energy consumption, such as smart devices and IoT devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and the drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. The above and other purposes, features, and advantages of the present disclosure will become more apparent by reading the following detailed description in conjunction with the drawings. Throughout the present disclosure, the same reference numerals generally refer to the same components or elements.
[0028] Figure 1 FIG. shows a schematic diagram of the circuit design of an arithmetic circuit according to an embodiment of the present invention;
[0029] Figure 2 FIG. shows a schematic diagram of the circuit design of an exponential operation unit according to an embodiment of the present invention;
[0030] Figure 3 FIG. shows a schematic diagram of the circuit design of a division operation unit according to an embodiment of the present invention;
[0031] Figure 4 FIG. shows a schematic diagram of the circuit design of a logistic regression (sigmoid) function operation unit according to an embodiment of the present invention;
[0032] Figure 5 FIG. shows a schematic diagram of the circuit design of a hyperbolic tangent function operation unit according to an embodiment of the present invention;
[0033] Figure 6 FIG. shows a schematic diagram of the circuit design of a rectified linear unit operation unit according to an embodiment of the present invention;
[0034] Figure 7Shows a schematic diagram of an instruction processing apparatus according to an embodiment of the present invention;
[0035] Figure 8 Shows a schematic diagram of a processor according to an embodiment of the present invention; and
[0036] Figure 9 Shows a schematic diagram of a system-on-chip (SoC) according to an embodiment of the present invention. Detailed implementation manners
[0037] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be completely conveyed to those skilled in the art.
[0038] Figure 1 Shows a schematic diagram of the circuit design of an arithmetic circuit 100 according to an embodiment of the present invention. Figure 1 The shown circuit design diagram is the overall design diagram of the arithmetic circuit 100. As Figure 1 shown, the arithmetic circuit 100 includes an exponential operation unit 110, a division operation unit 120, a logistic regression (sigmoid) function operation unit 130, and a selection unit 140. After receiving the data to be operated on at the data input terminal data_in 150, the arithmetic circuit 100 determines the type of operation to be performed according to the received operation mode (mode_sel), and sends the data to the corresponding operation unit (110, 120, or 130). The operation results of each operation unit are sent to the selection unit 140. The selection unit 140 selects the operation result of the corresponding operation unit as the final operation result according to the operation mode, and sends it to the data output terminal data_out 160.
[0039] Optionally, the arithmetic circuit 100 further includes a control unit FSM 170, which controls the execution process of the arithmetic circuit 100. When the FSM 170 receives the valid input indication signal valid_in, it determines that the data to be processed has been sent to the data input terminal 150, and thus instructs the arithmetic circuit 100 to start processing the data at the data input terminal 150. At the same time, the FSM 170 starts counting the operation cycle. Since different arithmetic units have different operation cycles, when the FSM 170 determines that the operation cycle of the arithmetic unit corresponding to the operation mode has been reached (which means that the corresponding operation has been completed and the operation result has been provided at the data output terminal 160), the FSM 170 can generate the valid output indication signal valid_out. In this way, other modules coupled to the arithmetic circuit 100 can obtain the operation result for the input data data_in from the data output terminal data_out 160 when receiving the valid_out signal.
[0040] As Figure 1 shown, the logistic regression (sigmoid) function arithmetic unit 130 is coupled to the exponential arithmetic unit 110 and the division arithmetic unit 120. When performing the sigmoid function operation on data, the exponential operation provided by the exponential arithmetic unit 110 and the division operation provided by the division arithmetic unit 120 can be used to perform some operations related to the exponential operation and the division operation in the sigmoid function operation.
[0041] In the present invention, the logistic regression function refers to the sigmoid function, and the two can be used interchangeably throughout the text. The geometric shape of the sigmoid function is a sigmoid curve (S-shaped curve), which maps a value between 0 and 1. Therefore, in a neural network, the sigmoid function is often used as an activation function. In a neural network, the activation function of a node defines the output of the node given the input or set of inputs.
[0042] For the input z, the sigmoid function g(z) can be expressed as:
[0043]
[0044] Figure 2 shows a schematic diagram of the circuit design of the exponential arithmetic unit 110 in the arithmetic circuit 100 according to an embodiment of the present invention. In Figure 2In the circuit design shown, the Taylor expansion method is used to approximate the result of the exponential operation. The Taylor expansion, also known as the Taylor formula, is a formula that describes the values of a function in the vicinity using the information of the function at a certain point. If the function is smooth enough, given the values of the derivatives of the function at a certain point, the Taylor formula can construct a polynomial with these derivative values as coefficients to approximate the values of the function in the neighborhood of this point. That is, the result of the Taylor expansion with specific parameters can be used as the result of the exponential operation.
[0045] According to an embodiment of the present invention, for an input input, an 8th - order Taylor expansion can be used for the exponential operation, that is:
[0046] m1 = 1 + (input - a);
[0047] m2 = (input - a)^2;
[0048] m3 = (input - a)^3;
[0049] m4 = (input - a)^4;
[0050] m5 = (input - a)^5;
[0051] m6 = (input - a)^6;
[0052] m7 = (input - a)^7;
[0053] m8 = (input - a)^8;
[0054] y = A1 * (m1 + B1 * m2 + C1 * m3 + D1 * m4 + E1 * m5 + F1 * m6 + G1 * m7 + H1 * m8)
[0055] Where a is the starting point of the input interval, and A1, B1, C1, D1, E1, F1, G1, and H1 are the parameters to be used in this interval.
[0056] To provide high accuracy, the input can be divided into multiple intervals, and different parameters are used for the Taylor expansion in each interval. According to an embodiment of the present invention, the input can be divided into 18 intervals, and an 8th - order Taylor expansion is performed in each interval.
[0057] It should be understood that the present invention is not limited to the number of intervals and the order of the Taylor expansion. The number of intervals and the order of expansion can be changed according to the accuracy of the exponential operation without departing from the protection scope of the present invention.
[0058] Such as Figure 2As shown, the exponential operation unit 110 includes a data input module data_in 112, a selection module SEL 114, an exponential parameter module 116, a multiplication operation module 118, and a result output module 119.
[0059] The data input module data_in 112 receives the data to be subjected to the exponential operation. Subsequently, the selection module SEL 114 coupled to the data input module 112 determines the interval in which the data is located according to the magnitude of the data. For example, which of the 18 intervals mentioned above the data belongs to.
[0060] Subsequently, the exponential parameter module 116 selects the parameters corresponding to the determined interval, such as the a value and the parameters A1, B1, C1, D1, E1, F1, G1, and H1 mentioned above.
[0061] The multiplication operation module 118 performs the multiplication operations and corresponding addition operations required for the Taylor expansion. According to an embodiment of the present invention, the multiplication operation module 118 may include 4 multipliers so that 4 multiplication operations can be performed in parallel within one operation cycle, thereby improving the execution efficiency. The present invention is not limited to the number of multipliers, and the number of multipliers can be increased or decreased according to actual needs without departing from the protection scope of the present invention.
[0062] When the multiplication operation module 118 calculates the result of the Taylor expansion, that is, the result of the exponential operation, the operation result data_out can be output at the result output unit 119.
[0063] Optionally, as Figure 2As shown, the exponential operation unit 110 further includes a control circuit. The control circuit includes a valid input indication signal valid_in, an operation control module FSM 111, and a valid output indication signal valid_out. After an external module coupled to the exponential operation unit 110 sends the data to be exponentiated to the data input module data_in 112, it sets valid_in to indicate that the data is in place. After determining that valid_in is set, the operation control module 111 instructs the exponential operation unit 110 to start the exponential operation on the data at data_in and simultaneously starts counting the operation cycle, i.e., the clock cycle of the circuit. As described above, the exponential operation unit 110 has a specific clock cycle for performing the exponential operation. When the operation control module 111 determines that this specific clock cycle has been reached, it means that the exponential operation on this data has been completed and the exponential operation result has started to be provided at data_out 119. Therefore, at this time, the operation control module 111 outputs a valid output indication signal, or sets valid_out. In this way, the external module can determine that the operation result is ready after finding that valid_out is set, and thus can obtain the exponential operation result from data_out 119.
[0064] According to an embodiment of the present invention, when 4 multipliers are used in the multiplication operation module 118, if only one multiplication can be achieved in one clock cycle, the exponential operation unit 110 requires 23 clock cycles to complete one exponential operation. It should be noted that the present invention is not limited thereto, and the number of clock cycles required by the exponential operation unit 110 can be reset according to the Taylor expansion series and the number of multipliers.
[0065] Figure 3 FIG. shows a schematic diagram of the circuit design of the division operation unit 120 in the operation circuit 100 according to an embodiment of the present invention. In Figure 3 the shown circuit design, the Newton-Raphson method is used to perform the division calculation, that is, U = 1 / a in the division can be rewritten as:
[0066] For the function f(U) = 1 / U - a, find the value of U when the function value is 0.
[0067] And then it can be written as U n+1 = Un*(2 - a*Un). By performing iterative processing on the above formula, when Un+1 converges rapidly to be close to Un, it means that the value of Un is very close to 1 / a, thus achieving the result of the division operation.
[0068] According to an embodiment of the present invention, the algorithm for performing the division calculation is as follows:
[0069] Let n = floor(log2(a));
[0070] a = a * 2^-n;
[0071] x = x0;
[0072] Loop 5 times:
[0073] x = x * (2 - a * x);
[0074] Loop end;
[0075] x = x * 2^-n;
[0076] Return x;
[0077] That is, in the circuit design shown in Figure 3 Using the Newton - Raphson method, multiple multiplication and shift operations are used to obtain the division result. It should be noted that the present invention is not limited to the specific parameters n, x0, and the number of iterations used in this method. All methods that can use multiplication and shift operations to obtain the division result are within the protection scope of the present invention.
[0078] As Figure 3 shown, the division operation unit 120 includes a data input module data_in 121, a first parameter selection module SEL 122, a first shift module 123, a second parameter selection module 124, a multiplication operation module 125, a second shift module 126, and a result output module data_out 127.
[0079] The data input module data_in 121 receives the data to be divided. Subsequently, the first parameter selection module SEL 122 coupled to data_in 121 determines the shift parameter n according to the size of the data to be divided. For example, n can be set to be equal to floor(log2(a)), where a is the data to be divided.
[0080] The first shift module 123 shifts a according to the shift parameter n determined by the first parameter selection module 122. For example, in the binary case, a is shifted n bits to the right to obtain the value a = a * 2^-n.
[0081] The second parameter selection module 124 determines the initial parameter according to the value of the data a to be divided. For example, one of coe_0 and coe_1 can be selected as the initial value of the iteration. According to an embodiment of the present invention, coe_0 can be set to 0 and coe_1 can be set to 0.75. When the value of a is greater than 0.5, coe_1 is selected, otherwise coe_0 is selected as the iteration initial value.
[0082] The multiplication operation module 125 performs a multiplication operation on the data after being shifted by the first shift module 123 and the initial value selected by the second parameter selection module 124, such as performing iterative operations in the Newton - Raphson method, to obtain the result after multiple iterations. For example, in the above example, in the multiplication operation module 125, multiple multiplication iterations of x = x * (2 - a * x) (for example, 5 iterations) are performed to obtain the value of x after multiple iterations.
[0083] Subsequently, in the second shift module 126, the result of multiple iterations output by the multiplication operation module 125 is shifted again to obtain the final division result. For example, in the above example, an operation of shifting x = x * 2^-n to the right by n bits is performed to obtain the final result x as the division operation result and output it to the result output module 127 as data_out.
[0084] Optionally, as Figure 3 shown, the division operation unit 120 further includes a control circuit. The control circuit includes a valid input indication signal valid_in, an operation control module FSM 128, and a valid output indication signal valid_out. After an external module coupled to the division operation unit 120 sends the data to be exponentiated to the data input module data_in 121, it sets valid_in to indicate that the data is in place. After determining that valid_in is set, the operation control module 128 instructs the division operation unit 120 to start the exponentiation operation on the data at data_in and simultaneously starts counting the operation cycle, that is, the clock cycle of the circuit. As described above, the division operation unit 120 has a specific clock cycle for performing the division operation. When the operation control module 128 determines that this specific clock cycle has been reached, it means that the division operation for this data has been completed and the division operation result has started to be provided at data_out 127. Therefore, at this time, the operation control module 128 outputs a valid output indication signal, or sets valid_out. In this way, the external module can determine that the operation result is ready after finding that valid_out is set, and thus can obtain the division operation result from data_out 127.
[0085] Figure 4 FIG. shows a schematic diagram of the circuit design of a logistic regression (sigmoid) function operation unit 130 according to an embodiment of the present invention.
[0086] The sigmoid function is an activation function. For an input z, the sigmoid function g(z) can be expressed as:
[0087]
[0088] This includes exponentiation operations on z and corresponding division operations. Therefore, the sigmoid function operation unit 130 utilizes the exponentiation operation unit 110 and the division operation unit 120 to perform the exponentiation and division operations in the sigmoid function.
[0089] As Figure 4 shown, the sigmoid function operation unit 130 includes a data input module data_in 131, a first sigmoid operation module 132, a second sigmoid operation module 133, a shift module 134, and a result output module data_out 135.
[0090] The data input module 131 receives the data to be subjected to sigmoid operations. The first sigmoid operation module 132 is coupled to the data input module 131 and the exponentiation operation unit 110, receives data from the data input module 131, sends it to the exponentiation operation unit 110 for exponentiation operations, and obtains the exponentiation operation result from the exponentiation operation unit 110. The circuit design of the exponentiation operation unit 110 has been described above with reference to Figure 2 According to one embodiment, the first sigmoid operation module 132 sends the data to the data input module 112 of the exponentiation operation unit 110, and at the same time sets the valid input indication signal valid_in of the exponentiation operation unit 110. In this way, the exponentiation operation unit 110 starts the exponentiation operation on this data, and when obtaining the exponentiation operation result, sets the valid output indication signal valid_out. The first sigmoid operation module 132 obtains the exponentiation operation result from the result output unit 119 of the exponentiation operation unit 110 when receiving the valid valid_out signal. After obtaining the exponentiation operation result, the first sigmoid operation module 132 also performs an addition calculation to calculate 1 + e in the sigmoid function z .
[0091] The second sigmoid operation module 133 is coupled to the first sigmoid operation module 132 and the division operation unit 120, and sends the exponentiation operation result 1 + e z calculated by the first sigmoid operation module 132 to the division operation unit 120 to obtain the division operation result 1 / (1 + e z ). The above reference Figure 3The circuit design of the division operation unit 120 has been described. According to one embodiment, the second sigmoid operation module 133 sends data to the data input module 121 of the division operation unit 120 while setting the valid input indication signal valid_in of the division operation unit 120. In this way, the division operation unit 120 starts the division operation on this data, and when the division operation result is obtained, the valid output indication signal valid_out is set. When the second sigmoid operation module 133 receives the valid valid_out signal, it obtains the division operation result from the result output unit 127 of the division operation unit 120.
[0092] The second sigmoid operation module 133 also multiplies the exponential operation result e z by the division operation result to obtain the sigmoid function operation result.
[0093] Optionally, according to one embodiment, the operation results of the exponential operation unit 110 and the division operation unit 120 are both 32-bit numbers. In this way, the multiplication result will produce a 64-bit result. To finally obtain a 32-bit result, the shift module 134 is coupled to the second sigmoid operation module 133 and performs a shift operation on the multiplication result of the second sigmoid operation module 133 to obtain the final 32-bit result of the sigmoid function operation, and outputs it to the result output module 135 as the result of the sigmoid function operation on the data.
[0094] It should be noted that the shift module 134 is optional, and the shift number is also optional. Whether to use the shift module 134 and the shift amount in the shift module 134 can be selected according to the actual requirements of the sigmoid function operation without departing from the protection scope of the present invention.
[0095] Optionally, as Figure 4As shown, the sigmoid function operation unit 130 further includes a control circuit. The control circuit includes a valid input indication signal valid_in, an operation control module FSM 136, and a valid output indication signal valid_out. After an external module coupled to the sigmoid function operation unit 130 sends the data to be sigmoid-operated to the data input module data_in 131, it sets valid_in to indicate that the data is in place. After determining that valid_in is set, the operation control module 136 instructs the sigmoid function operation unit 130 to start the exponential operation on the data at data_in, and at the same time starts to count the operation cycle, that is, the clock cycle of the circuit. As described above, the sigmoid function operation unit 130 has a specific clock cycle for performing the sigmoid function operation. When the operation control module 136 determines that this specific clock cycle has been reached, it means that the sigmoid function operation for this data has been completed, and the sigmoid function operation result has started to be provided at data_out 135. Therefore, at this time, the operation control module 136 will output a valid output indication signal, or set valid_out. In this way, the external module can determine that the operation result is ready after finding that valid_out is set, and thus can obtain the sigmoid function operation result from data_out 135.
[0096] The above refers to Figure 1-4 the circuit design of the operation circuit 100 according to the present invention. The operation circuit 100 provides a hardware implementation for the sigmoid function operation, and additionally designs relatively independent exponential operation and division operation modules, simplifies the implementation logic of the sigmoid function operation, and can be easily extended to other function implementation fields that require exponential and division operations.
[0097] In addition, in the operation circuit 100, by implementing Taylor expansion in hardware and dividing the data into multiple intervals and designing different Taylor expansion parameters for each interval, the exponential operation can be performed efficiently and with high precision, which significantly improves the hardware execution efficiency of the exponential operation.
[0098] In addition, optionally, according to an embodiment of the present invention, the operation circuit 100 further includes other activation function operation circuits. For example, as Figure 1As shown, the arithmetic circuit 100 further includes a hyperbolic tangent (tanh) function arithmetic unit 180. The hyperbolic tangent function is another activation function. Therefore, in the arithmetic circuit 100, the hyperbolic tangent function arithmetic unit 180 is juxtaposed with other arithmetic units, such as the sigmoid function arithmetic unit 130, the exponential arithmetic unit 110, and the division arithmetic unit 120. The arithmetic circuit 100 can select whether to perform an operation by the hyperbolic tangent function arithmetic unit 180 according to the operation mode, and the selection unit 140 can select the operation result output by the hyperbolic tangent function arithmetic unit 180 as the final operation result according to the operation mode (mode_sel).
[0099] Considering the correlation between the hyperbolic tangent function and the sigmoid function, the hyperbolic tangent function arithmetic unit 180 is coupled to the sigmoid function arithmetic unit 130, and uses the sigmoid function operation performed by the sigmoid function arithmetic unit 130 to perform the hyperbolic tangent function operation.
[0100] Figure 5 The circuit design schematic diagram of a hyperbolic tangent (tanh) function arithmetic unit according to an embodiment of the present invention is shown.
[0101] The hyperbolic tangent function g(z) can be described as:
[0102]
[0103] Therefore, the relationship between the hyperbolic tangent (tanh(z)) function and the sigmoid (sigmoid(z)) function is as follows:
[0104] tanh(z) = 2 * sigmoid(2z) - 1
[0105] As Figure 5 As shown, the hyperbolic tangent function arithmetic unit 180 includes a data input module data_iin 181, a hyperbolic tangent (tanh) arithmetic module 182, and a result output module data_out 183.
[0106] The data input module data_in 181 receives the data to be subjected to the tanh function operation. The hyperbolic tangent arithmetic module 182 is coupled to the data input module 181 and the sigmoid function arithmetic unit 130, receives the data to be operated from the data input module 181, preprocesses the data according to the correlation between tanh and sigmoid, then sends it to the sigmoid function arithmetic unit 130 for sigmoid function operation, obtains the operation result for post-processing, and sends the result after post-processing to the result output module data_out 183 for output as the tanh function operation result.
[0107] According to one embodiment, the hyperbolic tangent operation module 182 first performs an inversion process on the data, and sends the data after the inversion process to the sigmoid function operation unit 130 for sigmoid function operation to obtain an operation result, and performs a shift process on the obtained sigmoid function operation result, so as to use the shifted output as the tanh function operation result and send it to the result output module 183.
[0108] It has been referred to above Figure 4 The circuit design of the sigmoid function operation unit 130 has been described. According to one embodiment, the hyperbolic tangent function operation module 182 sends the data to the data input module 131 of the sigmoid function operation unit 130, and at the same time sets the valid input indication signal valid_in of the sigmoid function operation unit 130. In this way, the sigmoid function operation unit 130 starts the sigmoid function operation on the data, and when the sigmoid function operation result is obtained, sets the valid output indication signal valid_out. When the hyperbolic tangent function operation module 182 receives the valid valid_out signal, it obtains the sigmoiid function operation result from the result output unit 135 of the sigmoid function operation unit 130.
[0109] Optionally, as Figure 5As shown, the tanh function operation unit 180 further includes a control circuit. The control circuit includes a valid input indication signal valid_in, an operation control module FSM 184, and a valid output indication signal valid_out. After an external module coupled to the tanh function operation unit 180 sends the data to be operated on by the tanh function to the data input module data_in 181, it sets valid_in to indicate that the data is in place. After determining that valid_in is set, the operation control module 184 instructs the tanh function operation unit 180 to start the tanh function operation on the data at data_in, and at the same time starts counting the operation cycle, that is, the clock cycle of the circuit. As described above, the tanh function operation unit 180 has a specific clock cycle for performing the tanh function operation. When the operation control module 184 determines that this specific clock cycle has been reached, it means that the tanh function operation on this data has been completed, and the tanh function operation result has started to be provided at data_out 183. Therefore, at this time, the operation control module 183 will output a valid output indication signal, or set valid_out. In this way, the external module can determine that the operation result is ready after finding that valid_out is set, and thus can obtain the tanh function operation result from data_out 183.
[0110] In addition, optionally, according to an embodiment of the present invention, the operation circuit 100 further includes other activation function operation circuits. For example, as Figure 1 shown, the operation circuit 100 further includes a rectified linear unit (ReLU) operation unit 190. The ReLU function is another activation function. Therefore, in the operation circuit 100, the ReLU function operation unit 190 is juxtaposed with other operation units, such as the tanh function operation unit 180, the sigmoid function operation unit 130, the exponential operation unit 110, and the division operation unit 120. The operation circuit 100 can select whether to perform the operation by the ReLU function operation unit 190 according to the operation mode, and the selection unit 140 can select the operation result output by the ReLU function operation unit 190 as the final operation result according to the operation mode (mode_sel).
[0111] Figure 6 The figure shows a schematic diagram of the circuit design of the rectified linear unit (ReLU) operation unit 190 according to an embodiment of the present invention. For a given input z, ReLU provides the following function calculation:
[0112]
[0113] For this purpose, as Figure 6As shown, the linear rectifier function operation unit 190 includes a data input module data_in 192, a selection module 194, and a result output module 196.
[0114] The data input module 192 receives the data to be subjected to ReLU operation. The selection module 194 is coupled to the data input module 192 and determines the output of the selection module 194 according to the size of the data. Specifically, when the value of the data received by the data input module 192 is less than 0, the output value is 0; and when the value of the data is not less than 0, the data is output. The result output module 196 is coupled to the selection module 194 and obtains the output of the selection module 194 as the ReLU operation result.
[0115] According to an embodiment of the present invention, the data to be subjected to ReLU operation is signed data. Therefore, the highest bit in its binary representation is the sign bit. When the highest bit is 1, it indicates that the data is less than 0, otherwise the data is not less than 0. Therefore, when the data is a 32-bit binary number, the 32nd bit of the data, that is, data_in
[31] , can be sent to the selection module 194 as the selection judgment value, and the selection module 194 can select to output the data data_in itself or a 32-bit binary number 32’b0 with all bits being 0 according to the value of this bit.
[0116] Using the operation circuit 100 according to the present invention, since various functions share hardware resources, the hardware resources are greatly reduced. The exponential operation unit 110 has extremely high precision. The sigmoid function operation unit 130 and the hyperbolic tangent function operation unit 180 are based on the exponential operation unit 110, and their precision depends on the exponential operation unit 110. Therefore, there is no precision loss, so that the precision of the sigmoid function and the hyperbolic tangent function is the same as that of the exponential function and the same error can be ignored.
[0117] The following table gives a comparison of the execution efficiency of the operation circuit 100 and the implementation method using traditional instructions in a processing chip with the same processing performance and the same precision:
[0118]
[0119] It can be seen that the operation circuit 100 according to the present invention can significantly improve the operation efficiency of executing various activation functions. According to an embodiment of the present invention, the operation circuit 100 can be integrated into the processing chips used in smart devices, AIoT, IoT devices, etc. In this way, when artificial intelligence processing is required in these edge devices or terminal devices, higher execution efficiency can be achieved with lower energy consumption.
[0120] It has been combined above with Figure 1-6The circuit design of the arithmetic circuit 100 is described. The arithmetic circuit 100 can be integrated into an instruction processing device such as a processor core as a hardware implementation of various exponential operations, division operations, and activation function operations.
[0121] Figure 7 FIG. 4 is a schematic diagram of an instruction processing device 700 including the arithmetic circuit 100 according to an embodiment of the present invention. In some embodiments, the instruction processing device 700 may be a processor, a processor core of a multi-core processor, or a processing element in an electronic system.
[0122] As Figure 1 shown, the instruction processing device 700 includes an instruction fetch unit 730. The instruction fetch unit 730 can obtain instructions to be processed from the cache 710, the storage device 720, or other sources, and send them to the decoding unit 740. The instructions fetched by the instruction fetch unit 730 include, but are not limited to, high-level machine instructions or macro instructions, etc. The processing device 700 completes specific functions by executing these instructions.
[0123] The decoding unit 740 receives the instructions passed in from the instruction fetch unit 730, and decodes these instructions to generate low-level micro-operations, micro-code entry points, micro-instructions, or other low-level instructions or control signals. They reflect the received instructions or are derived from the received instructions. The low-level instructions or control signals can implement the operations of the high-level instructions through low-level (e.g., circuit-level or hardware-level) operations. Various different mechanisms can be used to implement the decoding unit 740. Examples of suitable mechanisms include, but are not limited to, micro-code, look-up tables, hardware implementations, programmable logic arrays (PLAs). The present invention is not limited to the various mechanisms for implementing the decoding unit 740, and any mechanism that can implement the decoding unit 740 is within the protection scope of the present invention.
[0124] According to one embodiment, in the instruction processing device 700, a dedicated instruction set is customized, which includes dedicated instructions defined for various functions that can be executed in the arithmetic circuit 100, such as exponential functions and various activation functions. The instruction processing device 700 further includes a computing acceleration enabling unit 780. The computing acceleration enabling unit 780 is used to control whether to use the arithmetic circuit 100 for exponential function and activation function operations.
[0125] When the decoding unit 740 decodes the instructions for processing exponential functions and various activation functions, if the computing acceleration enabling unit 780 indicates not to start the arithmetic circuit 100 for hardware acceleration operations, these instructions are decoded into a conventional instruction set for execution by a conventional instruction execution unit. If the computing acceleration enabling unit 780 indicates to start hardware acceleration operations, the decoding unit 740 decodes these instructions into dedicated instructions for execution by a dedicated instruction execution unit.
[0126] Subsequently, these decoded instructions are sent to the execution unit 750 and executed by the execution unit 750. The execution unit 750 includes circuitry operable to execute instructions. When executing these instructions, the execution unit 750 receives data inputs from the register file 770, the cache 710, and / or the storage device 720 and generates data outputs to them. According to one embodiment, the execution unit 750 is also coupled to the arithmetic circuit 100 to perform operations of exponential functions and other activation functions by the arithmetic circuit 100.
[0127] In one embodiment, the register file 770 includes architectural registers, which are also referred to as registers. Unless otherwise specified or clearly apparent, in this document, the phrases architectural register, register file, and register are used to refer to registers that are visible to software and / or the programmer (e.g., software visible) and / or are specified by macroinstructions to identify operands. These registers are different from other non-architectural registers in a given microarchitecture (e.g., temporary registers, reorder buffers, retirement registers, etc.). According to one embodiment, the register file 770 may include a set of vector registers 175, each of which may be 512 bits, 256 bits, or 128 bits wide, or different vector widths may be used. Optionally, the register file 770 may also include a set of general-purpose registers 776. The general-purpose registers 176 may be used when the execution unit executes instructions, such as storing jump conditions, storing instruction operation results, storing addresses of data to be accessed, storing data read from the cache 110 and / or the storage device 120, etc.
[0128] The execution unit 750 may include multiple specific instruction execution units 750a, 750b... 750c, etc., such as arithmetic units, arithmetic logic units (ALUs), integer units, floating-point units, data access units, etc., and may execute different types of instructions respectively.
[0129] According to one embodiment of the present invention, the instruction execution unit 750a is coupled to the arithmetic circuit 100, and when performing operations of exponential functions and various activation functions, sends the data required for the operations and the corresponding calculation modes to the arithmetic circuit 100 and obtains the execution result of the arithmetic circuit 100 as the corresponding operation result.
[0130] To avoid confusion in the description, a relatively simple instruction processing device 700 has been shown and described. It should be understood that the instruction processing device 700 may have different forms. For example, other embodiments of the instruction processing device or the processor may have multiple cores, logical processors, or execution engines.
[0131] The computing acceleration enabling unit 780 can be set according to the actual application environment of the instruction processing device 700. When the instruction processing device 700 is mainly used in neural networks, the computing acceleration enabling unit 780 can enable the arithmetic circuit 100 to accelerate the operations of exponential functions and activation functions using hardware. When the instruction processing device 700 is not used in the field of neural networks, the computing acceleration enabling unit 780 can disable the arithmetic circuit 100 to reduce the additional power consumption brought by the arithmetic circuit 100.
[0132] Processor cores can be implemented in different processors in different ways. For example, a processor core can be implemented as a general-purpose ordered core for general computing, a high-performance general-purpose out-of-order core for general computing, and a dedicated core for graphics and / or scientific (throughput) computing. And a processor can be implemented as a CPU (central processing unit) and / or a coprocessor, where the CPU can include one or more general-purpose ordered cores and / or one or more general-purpose out-of-order cores, and the coprocessor can include one or more dedicated cores. Such combinations of different processors can result in different computer system architectures. In one computer system architecture, the coprocessor is on a chip separate from the CPU. In another computer system architecture, the coprocessor is in the same package as the CPU but on a separate die. In still another computer system architecture, the coprocessor is on the same die as the CPU (in which case, such a coprocessor is sometimes referred to as dedicated logic such as integrated graphics and / or scientific (throughput) logic, or as a dedicated core). In still another computer system architecture called a system-on-chip, the described CPU (sometimes referred to as an application core or application processor), the coprocessor described above, and additional functions can be included on the same die. Subsequently, reference will be made to Figure 8 and 9 to describe exemplary processors and computer architectures.
[0133] Figure 8 FIG. 1100 shows a schematic diagram of a processor 1100 according to an embodiment of the present invention. As Figure 8 shown by the solid-line box in, according to one embodiment, the processor 1110 includes a single core 1102A, a system agent unit 1110, and a bus controller unit 1116. As Figure 8 shown by the dashed-line box in, according to another embodiment of the present invention, the processor 1100 may further include multiple cores 1102A-N, an integrated memory controller unit 1114 in the system agent unit 1110, and dedicated logic 1108.
[0134] According to one embodiment, the processor 1100 may be implemented as a central processing unit (CPU), where the dedicated logic 1108 is integrated graphics and / or scientific (throughput) logic (which may include one or more cores), and the cores 1102A-N are one or more general-purpose cores (e.g., general-purpose in-order cores, general-purpose out-of-order cores, a combination of both). According to another embodiment, the processor 1100 may be implemented as a coprocessor, where the cores 1102A-N are multiple dedicated cores for graphics and / or scientific (throughput). According to still another embodiment, the processor 1100 may be implemented as a coprocessor, where the cores 1102A-N are multiple general-purpose in-order cores. Thus, the processor 1100 may be a general-purpose processor, a coprocessor, or a dedicated processor, such as, for example, a network or communication processor, a compression engine, a graphics processor, a GPGPU (general-purpose graphics processing unit), a high-throughput integrated many-core (MIC) coprocessor (including 30 or more cores), or an embedded processor, etc. The processor may be implemented on one or more chips. The processor 1100 may be part of one or more substrates, and / or may be implemented on one or more substrates using any of a plurality of processing technologies, such as, for example, BiCMOS, CMOS, or NMOS.
[0135] The memory hierarchy includes one or more levels of cache within the cores, one or more shared cache units 1106, and an external memory (not shown) coupled to the integrated memory controller unit 1114. The shared cache unit 1106 may include one or more intermediate-level caches, such as a second-level (L2), third-level (L3), fourth-level (L4), or other-level caches, a last-level cache (LLC), and / or a combination thereof. Although in one embodiment, the ring-based interconnect unit 1112 interconnects the integrated graphics logic 1108, the shared cache unit 1106, and the system agent unit 1110 / integrated memory controller unit 1114, the present invention is not limited thereto, and any number of well-known techniques may be used to interconnect these units.
[0136] The system agent 1110 includes those components that coordinate and operate the cores 1102A-N. The system agent unit 1110 may include, for example, a power control unit (PCU) and a display unit. The PCU may include the logic and components required to adjust the power states of the cores 1102A-N and the integrated graphics logic 1108. The display unit is used to drive one or more externally connected displays.
[0137] The cores 1102A-N may have various core architectures and may be homogeneous or heterogeneous in terms of the architectural instruction set. That is, two or more of these cores 1102A-N may be capable of executing the same instruction set, while other cores may be capable of executing only a subset or a different instruction set of that instruction set.
[0138] Figure 9 FIG. 3 shows a schematic diagram of a system-on-chip (SoC) 1500 according to an embodiment of the present invention. Figure 9 The illustrated system-on-chip includes Figure 8 the illustrated processor 1100, and thus components similar to those in Figure 8 have the same reference numerals. As Figure 9 illustrated, the interconnect unit 1502 is coupled to an application processor 1510, a system agent unit 1110, a bus controller unit 1116, an integrated memory controller unit 1114, one or more coprocessors 1520, a static random access memory (SRAM) unit 1530, a direct memory access (DMA) unit 1532, and a display unit 1540 for coupling to one or more external displays. The application processor 1510 includes a set of one or more cores 1102A-N and a shared cache unit 110. The coprocessor 1520 includes integrated graphics logic, an image processor, an audio processor, and a video processor. In one embodiment, the coprocessor 1520 includes a dedicated processor, such as, for example, a network or communication processor, a compression engine, a GPGPU, a high throughput MIC processor, or an embedded processor, etc.
[0139] The system-on-chip (SoC) or processor according to the present invention can be used in various intelligent devices to implement corresponding functions in the intelligent devices. Such intelligent devices include, but are not limited to, in-vehicle devices, smart speakers, smart display devices, IoT devices, mobile terminals, and personal digital terminals, etc.
[0140] It should be understood that, to streamline the present disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected by the following claims, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0141] Those skilled in the art should understand that the modules or units or components of the devices in the examples disclosed herein can be arranged in the devices as described in this embodiment, or alternatively can be located in one or more devices different from the devices in this example. The modules in the foregoing examples can be combined into one module or further divided into multiple sub-modules.
[0142] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0143] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of the features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0144] As used herein, unless otherwise specified, the use of ordinal numbers "first", "second", "third", etc. to describe ordinary objects only indicates different instances of similar objects, and does not intend to imply that the objects so described must have a given order in terms of time, space, ranking, or in any other way.
[0145] Although the present invention has been described in terms of a limited number of embodiments, those skilled in the art in this technical field will understand, from the above description, that other embodiments can be conceived within the scope of the present invention thus described. In addition, it should be noted that the language used in this specification is mainly selected for readability and teaching purposes, rather than for the purpose of explaining or limiting the subject matter of the present invention. Therefore, many modifications and changes will be obvious to those of ordinary skill in the art without departing from the scope and spirit of the appended claims. For the scope of the present invention, the disclosure of the present invention is illustrative, not restrictive, and the scope of the present invention is defined by the appended claims.
Claims
1. An arithmetic circuit adapted to perform arithmetic operations on data, comprising: An exponential operation unit, adapted to perform exponential operations on the data, wherein the data is divided into multiple intervals according to the size of the data to be exponentially operated, and different parameters are used for Taylor expansion in each interval to approximate the exponential operation result; A division operation unit, adapted to perform division operations on the data, wherein the division parameters are determined according to the size of the data to be divided, and the Newton-Raphson method is used to obtain the division result with multiple multiplication and shift operations; A logistic regression function operation unit, adapted to perform logistic regression function operations on the data; and A selection unit, adapted to select the operation output of the exponential operation unit, the division operation unit, or the logistic regression function operation unit as the operation result according to the operation mode, wherein the logistic regression function operation unit is coupled to the exponential operation unit and the division operation unit, and is adapted to use the exponential operation unit to perform the exponential operation in the logistic regression function operation and use the division operation unit to perform the division operation in the logistic regression function operation when performing the logistic regression function operation on the data.
2. The arithmetic circuit according to claim 1, wherein the exponentiation operation unit comprises: An exponential data input module, adapted to receive the data to be exponentially operated; An exponential selection module, adapted to determine the interval to which the data belongs according to the size of the data to be operated; An exponential parameter module, adapted to provide the parameters corresponding to the determined interval, and the parameters are the parameters required for Taylor expansion of the exponential operation; An exponential multiplication module, adapted to perform multiplication operations on the provided parameters and the data to perform Taylor expansion processing on the data; And An exponential result output module, adapted to obtain the Taylor expansion processing result of the multiplication module to provide the exponential operation result.
3. The arithmetic circuit according to claim 2, wherein the exponentiation operation unit further comprises: An exponential operation control module, adapted to, when receiving a valid input indication signal, instruct the exponential operation unit to start the exponential operation processing on the received data, and generate a valid output indication signal after a predetermined operation cycle of the exponential operation unit to indicate that the exponential operation result of the data is provided at the exponential result output module.
4. The arithmetic circuit according to claim 1, wherein the division operation unit comprises: A division data input module, adapted to receive the data to be divided; A first division parameter selection module, adapted to determine the first division parameter according to the size of the data to be divided; A first division shift module, adapted to shift the data to be divided according to the first division parameter; A second division parameter selection module, adapted to select the second or third division parameter according to the data; A multiplication module, adapted to perform multiplication operations on the output of the first division shift module and the output of the second division parameter selection module; A second division shift module, adapted to perform shift operations on the output of the multiplication module; And A division result output module, adapted to obtain the output of the second division shift module as the operation result of the division operation on the data.
5. The arithmetic circuit according to claim 4, wherein the division operation unit further comprises: The division operation control module is adapted to, when receiving a valid input indication signal, instruct the division operation unit to start the division operation process on the received data, and after a predetermined operation period of the division operation unit has passed, generate a valid output indication signal to indicate that the division operation result of the data is provided at the division result output module.
6. The arithmetic circuit according to any one of claims 1-5, wherein the logistic regression function operation unit comprises: The logistic regression operation data input module is adapted to receive data to be subjected to a logistic regression function operation; The first logistic regression operation module is coupled to the exponential operation unit and is adapted to send the data to the exponential operation unit for exponential operation and obtain an exponential operation result; The second logistic regression operation module is coupled to the division operation unit and is adapted to send the exponential operation result obtained by the first logistic regression operation module to the division operation unit to obtain a division operation result, and multiply the exponential operation result and the division operation result to obtain a multiplication result; The logistic regression shift module is adapted to perform a shift operation on the multiplication result of the second logistic regression operation module; and The logistic regression operation result output module is adapted to obtain the shift output of the logistic regression shift module as the result of the logistic regression operation on the data.
7. The arithmetic circuit according to claim 6, wherein the logistic regression function operation unit further comprises: The logistic regression operation control module is adapted to, when receiving a valid input indication signal, instruct the logistic regression operation unit to start the logistic regression operation process on the received data, and after a predetermined operation period of the logistic regression operation unit has passed, generate a valid output indication signal to indicate that the logistic regression operation result of the data is provided at the logistic regression operation result output module.
8. The arithmetic circuit according to claim 1 further includes a hyperbolic tangent function arithmetic unit adapted to perform a hyperbolic tangent function operation on the data; The selection unit is adapted to select, according to the operation mode, the operation output of the exponential arithmetic unit, the division arithmetic unit, the logistic regression function arithmetic unit, or the hyperbolic tangent function arithmetic unit as the operation result, and wherein the hyperbolic tangent function arithmetic unit is coupled to the logistic regression function arithmetic unit and is adapted to utilize the logistic regression function arithmetic unit to perform the logistic regression function operation in the hyperbolic tangent function operation when performing the hyperbolic tangent function operation on the data.
9. The arithmetic circuit according to claim 8, wherein the hyperbolic tangent function arithmetic unit includes: The hyperbolic tangent operation data input module is adapted to receive data to be subjected to a hyperbolic tangent function operation; The hyperbolic tangent operation module is coupled to the logistic regression function operation unit and is adapted to perform an inversion process on the data to be operated, send the data after the inversion process to the logistic regression function operation unit for logistic regression function operation, and perform a shift process on the logistic regression function operation result; and The hyperbolic tangent operation result output module is adapted to obtain the shift output of the hyperbolic tangent operation module as the result of the hyperbolic tangent function operation on the data.
10. The arithmetic circuit according to claim 9, wherein the hyperbolic tangent function arithmetic unit further includes: The hyperbolic tangent function operation control module is adapted to, when receiving a valid input indication signal, instruct the hyperbolic tangent function operation unit to start the hyperbolic tangent function operation process on the received data, and after a predetermined operation period of the hyperbolic tangent function operation unit has passed, generate a valid output indication signal to indicate that the hyperbolic tangent function operation result of the data is provided at the hyperbolic tangent operation result output module.
11. The arithmetic circuit according to any one of claims 8 - 10 further includes a rectified linear unit (ReLU) arithmetic unit adapted to perform a rectified linear unit operation on the data, and the selection unit is adapted to select, according to the operation mode, the operation output of the exponential arithmetic unit, the division arithmetic unit, the logistic regression function arithmetic unit, the hyperbolic tangent function arithmetic unit, or the rectified linear unit arithmetic unit as the operation result.
12. The arithmetic circuit according to claim 11, wherein the rectified linear unit arithmetic unit includes: The linear rectification selection unit is adapted to output a value of 0 when the value of the data to be subjected to a linear rectification function operation is less than 0; and output the data as the operation result of the linear rectification function operation when the value of the data is not less than 0.
13. In the arithmetic circuit according to claim 12, when the most significant bit of the binary representation of the data is 1, it indicates that the value of the data is less than 0.
14. A processor, comprising: The operation circuit according to any one of claims 1-13 is adapted to receive data to be operated and an operation mode signal, and perform corresponding operations on the received data according to the operation mode.
15. A system - on - chip (SoC) includes the processor according to claim 14.
16. An intelligent device includes the system - on - chip according to claim 15.
Citation Information
Patent Citations
Implementation methods of neural network accelerator and neural network model
CN106485317A
Neural network computing method, apparatus, processor, and computer readable storage medium
CN109284827A