Hardware acceleration circuit, data processing acceleration method, chip and accelerator
By using hardware acceleration circuitry to process the Softmax function layer in neural networks, the computational load of data processing is reduced, the speed of obtaining nonlinear function values is improved, the job migration overhead between DLA/NPU and CPU/GPU is resolved, and system latency and power consumption are reduced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU XIAOPENG CONNECTIVITY TECH CO LTD
- Filing Date
- 2021-12-18
- Publication Date
- 2026-05-19
AI Technical Summary
In neural network processing, when the Softmax function layer is located in the middle layer of the network, it leads to high overhead for job migration between DLA/NPU and CPU/GPU, as well as efficiency problems such as increased system bandwidth and high power consumption.
A hardware acceleration circuit is provided, including an exponential function module, an adder, a first processing circuit, a second processing circuit, and a third processing circuit. By processing the addition result into low-bit-width data and performing preset processing, the reciprocal of the addition result is obtained, thereby reducing the computational load of data processing.
It improves the speed of obtaining nonlinear function values, reduces the amount of computation in data processing, and reduces system latency and power consumption.
Smart Images

Figure CN116306825B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a hardware acceleration circuit, a data processing acceleration method, a chip, and an accelerator. Background Technology
[0002] Nonlinear functions introduce nonlinear characteristics into artificial neural networks, playing a crucial role in their ability to learn and understand complex scenarios. Nonlinear functions include, but are not limited to, the softmax function and the sigmoid function.
[0003] The Softmax function, widely used in deep learning, is a prime example. In related technologies, its value can be calculated using general-purpose computing units such as CPUs or GPUs. However, when neural network processing is executed by hardware circuits such as Deep Learning Accelerators (DLAs) or Neural Network Processing Units (NPUs), if the Softmax function layer is located in the middle layer of the neural network, it leads to job migration overhead between the DLA / NPU and the CPU / GPU. This makes the scheme of using the CPU / GPU to determine the nonlinear function value inefficient, resulting in increased system bandwidth and higher power consumption. Summary of the Invention
[0004] To address or partially address the problems existing in related technologies, this application provides a hardware acceleration circuit, a data processing acceleration method, a chip, and an accelerator, which can reduce the computational load of data processing and thereby improve the speed of obtaining nonlinear function values.
[0005] This application provides a hardware acceleration circuit, the hardware acceleration circuit comprising:
[0006] The exponential function module is used to obtain multiple exponential function values for multiple data elements in a dataset;
[0007] An adder is used to obtain the result of adding the multiple exponential function values;
[0008] A first processing circuit is used to perform preset processing on the addition result to process the addition result into at least first data and second data, wherein the length of the addition result is N1 bits, the length of the first data is N2 bits, the length of the second data is N3 bits, and both N2 and N3 are less than N1.
[0009] The second processing circuit is used to perform preset processing on at least the first data and the second data to obtain the reciprocal of the addition result;
[0010] The third processing circuit is used to perform preset processing on the exponential function value and the reciprocal of the i-th data element among the plurality of data elements to obtain a specific function value of the i-th data element.
[0011] This application also provides an artificial intelligence chip, including the hardware acceleration circuit described above.
[0012] This application also provides a data processing acceleration method applied to an artificial intelligence accelerator, the method comprising:
[0013] Obtain multiple exponential function values for multiple data elements in a dataset;
[0014] Obtain the result of the addition operation of the multiple exponential function values;
[0015] Obtain the reciprocal of the result of the addition operation;
[0016] Based on the exponential function value of the i-th data element among the plurality of data elements and the reciprocal, the specific function value of the i-th data element is obtained;
[0017] The reciprocal of the addition result includes:
[0018] The result of the addition operation shall be processed into at least a first data and a second data;
[0019] Based at least the first and second data, obtain the reciprocal of the addition result;
[0020] The length of the addition result is N1 bits, the length of the first data is N2 bits, and the length of the second data is N3 bits, where N2 and N3 are both less than N1.
[0021] This application further provides an artificial intelligence accelerator, comprising:
[0022] Processor; and
[0023] A memory that stores executable code, which, when executed by the processor, causes the processor to perform the method described above.
[0024] The technical solution provided in this application may include the following beneficial effects:
[0025] The technical solution of this application embodiment processes the addition result of the exponential function values of each data element into at least first data and second data whose lengths are both shorter than the addition result. By performing preset processing on at least the first data and second data, the reciprocal of the addition result is obtained. By reducing the bit width of the processed data, the amount of data processing can be reduced, thereby improving the speed of obtaining nonlinear function values.
[0026] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description
[0027] The above and other objects, features and advantages of this application will become more apparent from the following description of exemplary embodiments of this application in conjunction with the accompanying drawings, wherein the same reference numerals generally represent the same components.
[0028] Figure 1 This is a schematic diagram of the structure of a neural network shown in one embodiment of this application;
[0029] Figure 2 This is a schematic diagram of the structure of a neural network for classification shown in one embodiment of this application;
[0030] Figure 3 This is a structural block diagram of a hardware acceleration circuit according to an embodiment of this application;
[0031] Figure 4A This is a structural block diagram of a hardware acceleration circuit according to another embodiment of this application;
[0032] Figure 4B This is a schematic diagram of the basic lookup table circuit unit according to an embodiment of this application;
[0033] Figure 5 This is a structural block diagram of a hardware acceleration circuit according to another embodiment of this application;
[0034] Figure 6 This is a structural block diagram of a hardware acceleration circuit shown in another embodiment of this application;
[0035] Figures 7 to 9 This is a schematic flowchart of a data processing acceleration method according to some embodiments of this application;
[0036] Figure 10 This is a structural block diagram of an artificial intelligence accelerator according to an embodiment of this application. Detailed Implementation
[0037] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While embodiments of this application are shown in the drawings, it should be understood that this application may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to make this application more thorough and complete, and to fully convey the scope of this application to those skilled in the art.
[0038] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0039] It should be understood that although the terms "first," "second," "third," etc., may be used in this application to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0040] The calculation process of nonlinear functions may involve the operation of exponential functions and / or reciprocals. For example, the operation of the Softmax function may involve the operation of the exponent (exp) and the reciprocal of the sum of exponents (1 / sum_of_exp).
[0041] This application provides a data processing acceleration scheme, which processes the addition result of the exponential function values of each data element into at least first data and second data whose length (i.e. bit width) is lower than the addition result. By performing preset processing on at least the first data and second data, the reciprocal of the addition result is obtained. By reducing the bit width of the processed data, the amount of data processing can be reduced, thereby improving the speed of obtaining nonlinear function values.
[0042] Figure 1 This is a schematic diagram of the structure of a neural network shown in one embodiment of this application.
[0043] See Figure 1 The diagram illustrates the topology of a neural network 100, including an input layer, hidden layers, and an output layer. This neural network 100 is capable of performing calculations or operations based on data elements I1 and I2 received from the input layer, and generating output data O1 and O2 based on the results of the calculations.
[0044] For example, neural network 100 can be a deep neural network (DNN) that includes one or more hidden layers. Figure 1 The neural network 100 in the diagram includes an input layer L1, two hidden layers L2 and L3, and an output layer L4. DNNs include, but are not limited to, Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs).
[0045] It should be noted that, Figure 1 The four layers shown are for illustrative purposes only and should not be construed as limiting the scope of this application. For example, a neural network may include more or fewer hidden layers.
[0046] Nodes in different layers of the neural network 100 can be connected to each other for data transmission. For example, a node can receive data from other nodes to perform calculations on the received data and output the calculation results to nodes in other layers.
[0047] Each node can determine its output data based on the output data and weights received from nodes in previous layers. For example, Figure 1 middle This represents the weight between the first node of the first layer and the first node of the second layer. This represents the output data of the first node in the first layer. Let represent the bias value of the first node in the second layer. Then, the output data of the first node in the second layer can be represented as: The calculation method for the output data of other nodes is similar and will not be described in detail here.
[0048] In some embodiments, the neural network is configured with activation function layers, such as softmax function layers, which can convert the result values for each class into probability values.
[0049] In some embodiments, the neural network is configured with a loss function layer after the flexible maximum function layer, which is capable of calculating the loss as an objective function for training or learning.
[0050] Understandably, a neural network can respond to the data to be processed, and after processing the data, obtain a recognition result; the data to be processed may include at least one of voice data, text data, and image data.
[0051] A typical type of neural network is a classification neural network. Classification neural networks determine the category to which a data element belongs by calculating the probability of that element corresponding to each class.
[0052] Figure 2 This is a schematic diagram of the structure of a neural network for classification, as shown in one embodiment of this application.
[0053] See Figure 2 The neural network 200 used for classification in this embodiment may include a hidden layer 210, a fully connected layer (FC layer) 220, a flexible maximum function layer 230, and a loss function layer 240.
[0054] like Figure 2 As shown, the neural network 200 responds to the data to be processed by sequentially calculating the hidden layer 210 and the fully connected (FC) layer 220. The FC layer 220 outputs a calculation result s, which corresponds to the classification probability of the data element. The FC layer 220 may include multiple nodes corresponding to multiple classes, and each node outputs a result value corresponding to the probability that the data element is classified into the corresponding class. For example, see also... Figure 1 FC layer 220 corresponds to Figure 1 The output layer L4 has two nodes, corresponding to two categories (Class 1 and Class 2). The output value of one node can be a result value representing the probability that a data element is classified into Class 1, and the output value of the other node can be a result value representing the probability that a data element is classified into Class 2. The FC layer 220 outputs the calculation result s to the flexible maximal function layer 230, which converts the calculation result s into a probability value y and can also normalize the probability value y.
[0055] The flexible maximum function layer 230 outputs the probability value y to the loss function layer 240, which can calculate the cross-entropy loss s based on the probability value y. entropyloss) L.
[0056] During the backpropagation learning process, the flexible maximum function layer 230 calculates the gradient of the cross-entropy loss L. Then, the FC layer 220 performs gradient learning based on the cross-entropy loss L. For example, the weights of the FC layer 220 can be updated according to the gradient descent algorithm. Further, subsequent learning processing can be performed in the hidden layer 210.
[0057] The neural network 200 can be implemented in software, in hardware circuits, or a combination of both. For example, in the case of hardware circuit implementation, the hidden layer 210, the fully connected (FC) layer 220, the flexible maximum function layer 230, and the loss function layer 240 are all implemented in hardware circuits, which can be integrated into a single AI chip or distributed across multiple chips. This configuration avoids the data migration between other layers of the neural network and processors such as CPUs / GPUs when the flexible maximum function layer 230 is implemented in a CPU / GPU, thereby improving the efficiency of neural network data processing, reducing data processing latency and power consumption, and avoiding increased bandwidth usage.
[0058] The technical solutions of the embodiments of this application are described in detail below with reference to the accompanying drawings.
[0059] Figure 3 This is a structural block diagram of a hardware acceleration circuit according to an embodiment of this application. In this application, the hardware acceleration circuit can be used, for example, but not limited to, to implement the flexible maximum function layer 230 in the neural network 200 described above. The hardware acceleration circuit can be, for example, but not limited to, a circuit component in a CPLD (Complex Programming logic device) chip, an FPGA (Field Programmable Gate Array) chip, a dedicated chip, etc.
[0060] To facilitate understanding of this application, the Softmax function, which maximizes flexibility, is explained below. Assume we have an array X, then the i-th element x... i The formula for calculating the Softmax function value is shown in equation (1).
[0061] Equation (1)
[0062] In equation (1), σ(x) i Represents the i-th element x i The value of the Softmax function, where e is the natural constant, x i Let represent the i-th element of array X, and let c represent the maximum element of array X. This represents the result of adding the exponential function values of at least some elements in array X.
[0063] See Figure 3 A hardware acceleration circuit includes an exponential function module 11, an adder 21, a first processing circuit 31, a second processing circuit 32, and a third processing circuit 33.
[0064] The exponential function module 11 is used to obtain multiple exponential function values for multiple data elements in a dataset.
[0065] In one embodiment, the exponential function module 11 can respond to the index value of each data element in the data set, output the exponential function value corresponding to each data element based on the first lookup table, and output the exponential function value of each data element in the data set.
[0066] Adder 21 is used to obtain the result of adding multiple exponential function values.
[0067] In one embodiment, adder 21 can perform addition operations on the exponential function values of all data elements based on the exponential function values of all data elements input by exponential function module 11, and output the addition result of the exponential function values of all data elements in the data set.
[0068] It is understandable that the result of adding exponential function values can be obtained by directly adding the exponential function values, or by performing a specific transformation on the exponential function values before adding them. In the case of a specific transformation, the subsequent data processing results may be subjected to an inverse transformation or no inverse transformation may be performed, depending on the type of transformation. Similarly, the processing of other data should be broadly understood to include both of the above situations, and should not be limited to processing only the data itself. Other embodiments are similar and will not be described further below.
[0069] The first processing circuit 31 is used to perform preset processing on the addition operation result, so as to process the addition operation result into at least the first data and the second data.
[0070] In one embodiment, the length of the addition result output by adder 21 is N1 bits. The first processing circuit 31 can perform data conversion on the addition result of length N1 bits and output first data and second data. The length of the first data is N2 bits and the length of the second data is N3 bits. Both N2 and N3 are less than N1.
[0071] The second processing circuit 32 is used to perform preset processing on at least the first data and the second data to obtain the reciprocal of the addition result.
[0072] In one embodiment, the second processing circuit 32 performs preset processing on the first data and the second data, and in response to the first data and the second data, outputs the lookup result corresponding to the first data and the second data based on the corresponding lookup table, and obtains the reciprocal of the addition operation result by performing data operation on the lookup result corresponding to the first data and the second data.
[0073] The third processing circuit 33 is used to perform preset processing on the exponential function value and reciprocal of the i-th data element among multiple data elements to obtain a specific function value of the i-th data element.
[0074] In one embodiment, the third processing circuit 33 can perform a multiplication operation on the exponential function value of the i-th data element among multiple data elements and the reciprocal of the addition operation result, and output the flexible maximum function value of the i-th data element.
[0075] In this embodiment, the addition result of the exponential function values of each data element is processed into at least the first and second data whose lengths (i.e. bit widths) are lower than the addition result. By performing preset processing on at least the first and second data, the reciprocal of the addition result is obtained. By reducing the bit width of the processed data, the amount of computation in data processing can be reduced, thereby improving the speed of obtaining nonlinear function values.
[0076] Figure 4A This is a structural block diagram of a hardware acceleration circuit shown in another embodiment of this application.
[0077] See Figure 4A A hardware acceleration circuit includes an exponential function module 11, an adder 21, a first processing circuit 31, a second processing circuit 32, and a third processing circuit 33.
[0078] The exponential function module 11 includes a first lookup table circuit 1101; the first lookup table circuit 1101 is used to obtain multiple exponential function values corresponding to multiple data elements in the data set based on the first lookup table.
[0079] The first lookup table circuit 1101 can output the exponential function value corresponding to each data element based on the first lookup table, and output the N4-bit exponential function value of each data element in the data set.
[0080] Adder 21 is used to obtain the result of adding multiple exponential function values.
[0081] In one embodiment, adder 21 can perform addition operations on the N4-bit exponential function values of all data elements input from exponential function module 11, and output the addition result of the exponential function values of all data elements in the data set. The addition result can be an N1-bit fixed-point integer.
[0082] The first processing circuit 31 includes an integer-to-floating-point circuit 311; the integer-to-floating-point circuit 311 is used to convert the result of the addition operation from an integer to a floating-point number represented by a first exponent data and a first mantissa data.
[0083] In a specific implementation, the integer-to-floating-point circuit 311 includes: a leading 0 counting circuit or a leading 1 detection circuit, as well as a shifter and a subtractor.
[0084] The leading zero counter circuit outputs the number of leading zeros in the addition result; the number of leading zeros is the number of zeros that appear from the most significant bit of the binary data up to the first 1. The leading 1 detector circuit outputs the number of leading 1s in the addition result; the leading 1 is the first 1 detected from the most significant bit of the binary data.
[0085] The shifter is used to output the first mantissa data of the addition result based on the number of leading zeros or the number of leading 1s. In a specific implementation, the shifter uses the number of leading zeros as the shift number, shifts the addition result left by the shift number, and outputs the shifted data with a bit width of N3 bits. That is, it extracts N3 consecutive bits from the addition result starting from the bit after the leading 1 and moving towards the lower bit direction as the first mantissa data of the addition result.
[0086] The subtractor is used to subtract a preset value from the number of leading zeros or the number of leading 1s to output the first exponent of the addition result.
[0087] The second processing circuit 32 includes a first conversion circuit 321, a second conversion circuit 322, and a third conversion circuit 323.
[0088] The first conversion circuit 321 is used to convert the first exponential data into a negative number.
[0089] In one embodiment, the first conversion circuit 321 includes a second lookup table circuit 3212; the second lookup table circuit 3212 is used to output the negative number corresponding to the first exponent data based on the second lookup table.
[0090] The second conversion circuit 322 is used to convert the fractional part of a floating-point number represented by the first exponent data and the first mantissa data into another floating-point number represented by the second exponent data and the second mantissa data, based on the first mantissa data.
[0091] In one embodiment, the second conversion circuit 322 includes a third lookup table circuit 3223 and a fourth lookup table circuit 3224; the third lookup table circuit 3223 is used to obtain second exponent data exp2 corresponding to the first mantissa data based on the third lookup table; the fourth lookup table circuit 3224 is used to obtain second mantissa data frac1 corresponding to the first mantissa data based on the fourth lookup table.
[0092] The third conversion circuit 323 includes an exponent adder 3231 and a shifter 3232. The exponent adder 3231 is used to obtain the sum of the negative number of the first exponent data and the second exponent data; the shifter 3232 is used to use the sum as a shift parameter to shift the second mantissa data to obtain the reciprocal of the addition result.
[0093] Understandably, shifting the second mantissa data can be done by performing necessary conversions or processing on the second mantissa data (such as padding with 1 as mentioned later) before shifting.
[0094] The third processing circuit 33 includes a multiplier 331, which is used to multiply the exponential function value of the i-th data element among the multiple data elements output by the exponential function module 11 by the reciprocal of the addition result of the multiple exponential function values output by the second processing circuit 32, and output the flexible maximum function value of the i-th data element.
[0095] It is understood that in other embodiments, some or all of the table lookup methods described above may be replaced by software calculations performed by a processor (such as a CPU or GPU).
[0096] The following section provides a more detailed explanation using the formula.
[0097] In one embodiment, the floating-point expression for the addition result fp0 is:
[0098] ,
[0099] Its reciprocal can then be expressed as:
[0100]
[0101] make ,
[0102] but:
[0103]
[0104] Combining the above formulas, the integer-to-floating-point circuit 311 converts the addition result fp0 in fixed-point integer format into a floating-point number represented by the first exponent data exp0 and the first mantissa data frac0. The second lookup table circuit 3212, based on the second lookup table, outputs the negative number -exp0 corresponding to the first exponent data exp0. The third lookup table circuit 3223 of the second conversion circuit 322, based on the third lookup table, outputs the second exponent data exp1 corresponding to the first mantissa data frac0. The fourth lookup table circuit 3224 of the second conversion circuit 322, based on the fourth lookup table, outputs the second mantissa data frac1 corresponding to the first mantissa data frac0.
[0105] As shown in the formula above, the reciprocal of the addition result fp0 is... By and Obtained by multiplication. In a specific implementation, it can be... After padding with 1, the result of -exp0+exp1 is used as the shift parameter. Obtained by shifting.
[0106] It is understandable that the transformations of fp0 and frac0 in the above formula are approximate transformations, and the errors introduced by the transformations have a negligible impact on the accuracy of the calculation in applications.
[0107] The third processing circuit 33 performs a multiplication operation on the reciprocal of the addition result of the N4-bit exponential function value and the N5-bit multiplication of the i-th data element, to obtain the N6-bit multiplication result of the i-th data element. Further, the N6-bit multiplication result can be converted, for example, to a result with a lower bit width of N7 bits. This converted result can be used as the maximum flexible function value of the i-th data element output by the hardware acceleration circuit. Understandably, converting the multiplication result from N6 bits to N7 bits can be achieved through processes such as saturation and rounding. Rounding processes include, for example, rounding up, rounding down, and rounding towards zero.
[0108] In one embodiment, the length of the first exponent data and the second exponent data can both be N2 bits, and the length of the first mantissa data and the second mantissa data can both be N3 bits; the values of N2 and N3 can be in the range of [1, 32], and in some specific instances the value range can be [8, 12]. N2 and N3 can be equal or unequal.
[0109] In this embodiment, the first lookup table, the second lookup table, the third lookup table, and the fourth lookup table can be stored in a storage module, such as RAM (Random-Access Memory), ROM (Read-Only Memory), FLASH, etc.
[0110] In one embodiment, the hardware acceleration circuit includes at least two of the first to fourth lookup table circuits, that is, it includes two, three or all of them, and each of the at least two lookup table circuits has a basic lookup table circuit unit.
[0111] See Figure 4BAs shown, in one embodiment, the basic lookup table circuit unit 20 includes an input terminal group 22, a control terminal group 23, an output terminal group 24, and a logic circuit 25. The input terminal group 22 is connected to the memory 10 and inputs the lookup table data into the logic circuit 25. The logic circuit 25 selects the value corresponding to the index value (also called the address) input from the control terminal group 23 and outputs it from the output terminal group 24. The logic circuit 25 can be, for example, a logic gate circuit or a logic switch circuit. It is understood that in this application, a terminal group refers to a set of connection terminals, including one or more connection terminals. If the control terminal group 23 has A control terminals and the output terminal group 24 has B output terminals, the basic lookup table circuit unit 20 is referred to as A-input B-output.
[0112] The basic lookup table circuit unit 20 can perform lookup output based on the stored lookup table. Taking the first lookup table as an example, the lookup table also has A inputs and B outputs. The data elements of the lookup table are index values with a bit width of A bits, and the output data is the exponential function value with a bit width of B bits. The first lookup table in the storage area stores the true value of the exponential function value, and the basic lookup table circuit unit is used to implement the mapping relationship between the index value and the true value of the exponential function value.
[0113] Taking a hardware acceleration circuit where each of the first to fourth lookup table circuits has a basic lookup table circuit unit as an example, the storage module includes a first storage area to a fourth storage area, and the first to fourth lookup tables are stored in the first to fourth storage areas respectively; the first lookup table circuit includes a first basic lookup table circuit unit, the second lookup table circuit includes a second basic lookup table circuit unit, the third lookup table circuit includes a third basic lookup table circuit unit, and the fourth lookup table circuit includes a fourth basic lookup table circuit unit. The first basic lookup table circuit unit is connected to the first storage area and is used to output the corresponding exponential function value stored in the first lookup table of the first storage area in response to the index value of the i-th data element; the second basic lookup table circuit unit is connected to the second storage area and is used to output the corresponding negative number stored in the second lookup table of the second storage area in response to the index value of the first exponential data; the third basic lookup table circuit unit is connected to the third storage area and is used to output the corresponding second exponential data stored in the third lookup table of the third storage area in response to the index value of the first mantissa data; the fourth basic lookup table circuit unit is connected to the fourth storage area and is used to output the corresponding second mantissa data stored in the fourth lookup table of the fourth storage area in response to the index value of the first mantissa data. In another embodiment, the hardware acceleration circuit includes at least two lookup table circuits from the first to the fourth lookup table circuits, and some of the lookup table circuits share a basic lookup table circuit unit. By reusing the basic lookup table circuit unit, the required basic lookup table circuit units can be reduced, thereby effectively reducing the area and cost of the lookup hardware acceleration circuit.
[0114] Taking a hardware acceleration circuit where the first lookup table circuit and the second lookup table circuit share a basic lookup table circuit unit (e.g., called the first basic lookup table circuit unit) as an example, the first basic lookup table circuit unit includes a first input terminal group, a first control terminal group, a first output terminal group, and a first logic gate circuit. The first input terminal group is connected to the storage module. The first logic gate circuit is used to: respond to the index value of the i-th data element input from the first control terminal group in a first time period, output the exponential function value corresponding to the i-th data element based on the first lookup table input from the first output terminal group; and in a second time period after the first time period, respond to the index value of the first exponential data input from the first control terminal group, output the negative number corresponding to the first exponential data based on the second lookup table output from the first output terminal group.
[0115] Understandably, in one specific implementation of this embodiment, the storage module includes a first storage area, and the first lookup table and the second lookup table are stored in the first storage area in a time-sharing manner. Since only one storage area needs to be configured for time-sharing storage of either the first or second lookup table, the storage space occupied by the lookup tables is effectively reduced, thus reducing hardware costs. In another specific implementation, the storage module includes a first storage area and a second storage area, with the first lookup table stored in the first storage area and the second lookup table stored in the second storage area.
[0116] Understandably, in this application, the index value of a certain data can be the data itself, or it can be obtained after the data has undergone a specific transformation.
[0117] In this embodiment, the exponential function value of each data element is obtained by a hardware lookup table circuit. An adder is used to obtain the addition result of the exponential function value. The addition result is then converted into multiple data parts with lower bit widths. Further, a table lookup and subsequent addition and multiplication processes are used to perform division on the addition result, obtaining the corresponding reciprocal. By avoiding complex exponential and reciprocal operations, the data processing speed in the nonlinear function calculation process is improved, and the nonlinear function value is obtained more quickly. Furthermore, it avoids the excessive hardware circuit area and high cost associated with implementing exponential and reciprocal operations.
[0118] Furthermore, by using lookup tables with three lower bit widths—that is, after converting the addition result from an integer to a floating-point number represented by a first exponent and a first mantissa with reduced bit widths—the negative of the first exponent is obtained through the first lookup table, and the second exponent and second mantissa are obtained through the second and third lookup tables—the dependence of the lookup table on large storage space can be significantly reduced, the area and cost of the lookup table logic circuit can be reduced, the lookup time can be accelerated, and thus the data processing speed can be improved.
[0119] For example, if the result of an addition operation is a 16-bit integer, a direct lookup table would require 2^16 (65536) entries, necessitating significant storage space and resulting in excessively high costs for the lookup table logic circuitry. Furthermore, completing a single lookup result would require 65536 cycles, leading to excessively long processing times. In this application, for instance, the 16-bit addition result can be converted from an integer to a floating-point number represented by an 8-bit first exponent and an 8-bit first mantissa. By configuring the first to third lookup tables with 8 inputs and 8 outputs, the total number of entries in the three lookup tables becomes 3 × 2^36. 8 =768 records; obviously, the latter greatly saves the storage space required for the lookup table, the area and cost of the lookup table logic circuit, and speeds up the lookup process.
[0120] Figure 5 This is a structural block diagram of a hardware acceleration circuit shown in another embodiment of this application.
[0121] See Figure 5 A hardware acceleration circuit includes an exponential function module 11, an adder 21, a first processing circuit 31, a second processing circuit 32, and a third processing circuit 33.
[0122] The exponential function module 11 is used to obtain multiple exponential function values for multiple data elements in a dataset.
[0123] Adder 21 is used to obtain the result of adding multiple exponential function values.
[0124] In this embodiment, the addition result is a floating-point number. Adder 21 can perform addition operations on the N4-bit exponential function values of all data elements input from exponential function module 11, and output the N1-bit floating-point addition result of the exponential function values of all data elements in the data set.
[0125] The first processing circuit 31 includes a third lookup table circuit 313 and a fourth lookup table circuit 314. The third lookup table circuit 313 is used to obtain the exponent data corresponding to the addition operation result based on the third lookup table. The fourth lookup table circuit 314 is used to obtain the mantissa data corresponding to the addition operation result based on the fourth lookup table.
[0126] The second processing circuit 32 is used to perform preset processing on the exponent data and mantissa data to obtain the reciprocal of the addition result.
[0127] The third processing circuit 33 is used to perform preset processing on the exponential function value and reciprocal of the i-th data element among multiple data elements to obtain a specific function value of the i-th data element.
[0128] Figure 6 This is a structural block diagram of a hardware acceleration circuit shown in another embodiment of this application.
[0129] See Figure 6 A hardware acceleration circuit includes a subtractor 61, an exponential function module 11, an adder 21, a first processing circuit 31, a second processing circuit 32, and a third processing circuit 33.
[0130] Subtractor 61 is used to subtract the maximum value among multiple initial data in the initial dataset to obtain a data set containing multiple data elements.
[0131] The exponential function module 11 includes a first lookup table circuit 1101; the first lookup table circuit 1101 is used to obtain multiple exponential function values corresponding to multiple data elements in the data set based on the first lookup table.
[0132] In one specific embodiment, the initial dataset input to the hardware acceleration circuit is mathematically transformed, letting element x i =x i The subtractor 61 performs subtraction operations on each initial data point in the initial dataset X with the maximum value in the initial dataset, outputting the data element corresponding to each initial data point in the initial dataset X. A data set is formed by these data elements, where each data element takes the value 0 or a negative number. By performing subtraction operations on multiple initial data points in the initial dataset using the subtractor 61, the value range of the data elements can be reduced, thus facilitating the implementation of the scheme of this application with less bit width data and corresponding hardware circuitry. Furthermore, since each data element in the data set takes the value of a negative number or 0, the exponential function value of that data element with base e can be normalized to the range (0, 1].
[0133] To better understand the lookup process of this embodiment, Table 1 below shows a specific example of the first lookup table, which has N0-bit input and N4-bit output, where N0 and N4 are both 8 bits. The data elements of the first lookup table can be index values with a width of N0 bits, and the output data can be exponential function values with a width of N4 bits. For ease of understanding, the data in Table 1 is represented in decimal format. It is understood that the first lookup table in the storage module only stores the true value of the exponential function. The first lookup table circuit is used to implement the mapping relationship between the index value and the true value of the exponential function. The data elements and the normalized exponential function value are included in the table for better understanding of this application.
[0134] Table 1
[0135]
[0136] As shown in Table 1, the data elements output by subtractor 61 are negative numbers or 0, and the value range of the data elements is defined as [-10, 0]. For table lookup, the value range [-10, 0] is discretized into 256 as shown in the column "Data Elements" (i.e., ...). There are 10 data points, and the corresponding exponential function value for each point is shown in the column "Normalized Exponential Function Value". Each data element point corresponds to an integer value in the range [0, 255] shown in the column "Index Value". Each normalized exponential function value corresponds to an integer value in the range [0, 255] shown in the column "Exponential Function Value". The data in the column "Exponential Function Value" is stored as true values in the first lookup table of the storage module. The table can be looked up by the index value.
[0137] The implementation of adder 21, first processing circuit 31, second processing circuit 32, and third processing circuit 33 can be found in the previous implementation examples and will not be repeated here.
[0138] In one specific implementation, the data element can be an 8-bit fixed-point integer. Each exponential function value in the first lookup table is an 8-bit fixed-point integer. The result of adding multiple exponential function values is a 32-bit fixed-point integer. The first exponent data, the first mantissa data, the second exponent data, and the second mantissa data of the addition result are all 8-bit fixed-point integers. That is, the second, third, and fourth lookup tables are all 8-input, 8-output tables. The result of multiplication is a 16-bit fixed-point integer. The specific function value converted from the multiplication result is an 8-bit fixed-point integer; that is, N0, N2, N3, N4, N5, and N7 are 8, N1 is 32, and N6 is 16.
[0139] Understandably, in other implementations, N0, N2, N3, N4, N5, and N7 can be other values; for example, the range of values for N0, N2, N3, N4, N5, and N7 can be [1, 32], and in some specific instances, the range can be [8, 12]. For example, N0, N2, N3, N4, N5, and N7 can also be unequal; for example, the values of N0 and N3 can be 9, 10, 11, and 12, while N2, N4, N5, and N7 are 8. Because the dynamic range of the Softmax function is very wide, related technologies often use software modules to implement this function. The embodiments of this application provide a hardware circuit solution that is essentially based on 8 bits and can effectively balance important indicators such as circuit cost, power consumption, bandwidth, performance, and data accuracy.
[0140] In this embodiment, during the process of obtaining the reciprocal of the addition result of the exponential function value of each data element by looking up a table, the addition result is converted to floating-point form to obtain the exponent data and mantissa data of the addition result in floating-point form. Based on the exponent data and mantissa data of the addition result, multiple lookup tables are searched. The reciprocal of the addition result is obtained by looking up multiple tables, which can obtain a more accurate reciprocal.
[0141] Furthermore, by converting the addition result into floating-point form during the softmax function's calculation process, and obtaining the reciprocal of the addition result based on the data from multiple table lookups, and by configuring the input / output data bit width of the multiple table lookups within a small range, the storage resources occupied by the lookup table and the area of the lookup table circuit can be reduced, as well as the bandwidth occupied. On the other hand, within the allowable accuracy range, the lookup speed and fixed-point operation speed can be improved, thereby further accelerating the circuit's response speed and reducing power consumption.
[0142] This application also provides embodiments of data processing acceleration methods.
[0143] Figure 7 This is a schematic flowchart illustrating a data processing acceleration method according to an embodiment of this application.
[0144] See Figure 7 A data processing acceleration method includes:
[0145] In step S110, multiple exponential function values of multiple data elements in the dataset are obtained.
[0146] In step S120, the result of the addition operation of multiple exponential function values is obtained.
[0147] In step S130, the reciprocal of the addition result is obtained.
[0148] In step S140, the specific function value of the i-th data element is obtained based on the exponential function value of the i-th data element among multiple data elements and the reciprocal of the addition operation result.
[0149] In step S130, obtaining the reciprocal of the addition result includes:
[0150] Step S130A: Convert the addition result into at least the first data and the second data.
[0151] Step S130B: Obtain the reciprocal of the addition result based at least on the first data and the second data.
[0152] The result of the addition operation is a data of length N1 bits. The first data is a data of length N2 bits, and the second data is a data of length N3 bits. Both N2 and N3 are less than N1.
[0153] Understandably, the result of adding exponential function values can be obtained by directly adding the exponential function values, or by performing a specific transformation on the exponential function values before adding them. In the case of transformation, the subsequent data processing results may be subjected to an inverse transformation or no inverse transformation may be performed, depending on the type of transformation. Similarly, the processing of other data should be broadly understood to include both of the above situations, and should not be limited to processing only the data itself.
[0154] Figure 8 This is a schematic flowchart illustrating a data processing acceleration method according to another embodiment of this application.
[0155] See Figure 8 A data processing acceleration method includes:
[0156] In step S801, multiple exponential function values corresponding to multiple data elements in the data set are obtained.
[0157] In one embodiment, a first lookup table module can be used to output the exponential function value corresponding to each data element based on the index value of each data element in the data set, in response to the index value of each data element in the data set, and output the N4-bit exponential function value of each data element in the data set.
[0158] In step S802, the result of the addition operation of multiple exponential function values is obtained.
[0159] In one embodiment, an adder can be used to perform addition operations on the N4-bit exponential function values of each of the data elements to obtain the N1-bit addition result of the exponential function values of all data elements in the data set output by the adder. The addition result can be an N1-bit fixed-point integer.
[0160] In step S803, the result of the addition operation is converted from an integer to a floating-point number represented by the first exponent data and the first mantissa data.
[0161] In one embodiment, the addition result output by the adder is an N1-bit fixed-point integer. The fixed-point integer can be converted into N2 bits of first exponent data exp0 and N3 bits of first mantissa data frac0 by an integer-to-floating-point circuit.
[0162] In step S804, the first exponential data is converted into a negative number.
[0163] In one embodiment, a second lookup table circuit can be used to output a negative number exp1 corresponding to the first exponential data based on the second lookup table, in response to the index value of the first exponential data.
[0164] In step S805, the fractional part of the floating-point number is converted into another floating-point number represented by the second exponent data and the second mantissa data, based on the first mantissa data.
[0165] In one embodiment, a third lookup table circuit can be used to output second exponent data exp2 corresponding to the first mantissa data based on the third lookup table in response to the index value of the first mantissa data; a fourth lookup table circuit can be used to output second mantissa data frac1 corresponding to the first mantissa data based on the fourth lookup table in response to the index value of the first mantissa data.
[0166] In step S806, the reciprocal of the addition result is obtained based on the negative number of the first exponent data, the second exponent data, and the second mantissa data.
[0167] In one embodiment, the sum of the negative of the first exponent data and the second exponent data is obtained by an exponent adder. The sum is then used as a shift parameter by a shifter to shift the second mantissa data, resulting in the N5-bit reciprocal of the addition result.
[0168] In step S807, the exponential function value of the i-th data element and the reciprocal of the addition operation result are subjected to preset processing to obtain the specific function value of the i-th data element.
[0169] In one embodiment, a multiplication circuit can multiply the N4-bit exponential function value of the i-th data element with its N5-bit reciprocal to obtain the N6-bit multiplication result of the i-th data element. Further, the N6-bit multiplication result can be converted, for example, to a lower bit width of N7 bits. This converted result can be used as the maximum flexibility function value of the i-th data element output by the hardware acceleration circuit. It is understood that converting the multiplication result from N6 bits to N7 bits can be achieved through processes such as saturation and rounding. Rounding processes include, for example, rounding up, rounding down, and rounding towards zero.
[0170] In one embodiment, the length of the first exponent data and the second exponent data can both be N2 bits, and the length of the first mantissa data and the second mantissa data can both be N3 bits; the values of N2 and N3 can be in the range of [1, 32], and in some specific instances the value range can be [8, 12].
[0171] In this embodiment, the exponential function value of each data element is obtained by looking up the table using a hardware lookup table circuit. An adder is used to obtain the addition result of the exponential function value. The addition result is then converted to floating-point form. Using the exponent and mantissa parts of the floating-point addition result as input, the lookup table circuit performs division on the addition result to obtain its reciprocal. This avoids complex exponential and reciprocal operations, improving data processing speed during the Softmax function calculation and allowing for faster Softmax function value acquisition. Furthermore, it avoids the excessive hardware circuit area and high cost associated with implementing exponential and reciprocal operations.
[0172] Figure 9 This is a schematic flowchart illustrating a data processing acceleration method according to another embodiment of this application.
[0173] See Figure 9 A data processing acceleration method includes:
[0174] In step S901, the maximum value among the multiple initial data in the initial dataset is subtracted from the multiple initial data to obtain a data set containing multiple data elements.
[0175] In step S902, multiple exponential function values corresponding to multiple data elements in the dataset are obtained.
[0176] In step S903, the result of the addition operation of multiple exponential function values is obtained.
[0177] In step S904, the reciprocal of the addition result is obtained.
[0178] In step S905, the specific function value of the i-th data element is obtained based on the exponential function value of the i-th data element among multiple data elements and the reciprocal of the addition operation result.
[0179] The result of addition is a floating-point number;
[0180] Step S904 obtains the reciprocal of the addition result, including:
[0181] Convert the result of the addition operation into exponent data and mantissa data;
[0182] At least the reciprocal of the addition result can be obtained based on the exponent and mantissa data.
[0183] The length of the addition result is N1 bits, the length of the exponent data is N2 bits, and the length of the mantissa data is N3 bits, where N2 and N3 are both less than N1.
[0184] The relevant features of the data processing acceleration method in this application embodiment can be found in the relevant content of the foregoing hardware acceleration circuit embodiment, and will not be repeated here.
[0185] The data processing acceleration method according to the embodiments of this application can be applied to artificial intelligence accelerators. Figure 10 This is a schematic diagram of the structure of an artificial intelligence accelerator according to an embodiment of this application.
[0186] See Figure 10 The AI accelerator 1000 includes a memory 1010 and a processor 1020.
[0187] The artificial intelligence accelerator 1000 can be a general-purpose processor, such as a CPU (Central Processing Unit), or an artificial intelligence processor (IPU) for performing artificial intelligence operations. Artificial intelligence operations can include machine learning operations, neuromorphic operations, etc. Machine learning operations include neural network operations, k-means operations, support vector machine operations, etc. The artificial intelligence processor can, for example, include one or a combination of GPU (Graphics Processing Unit), DLA (Deep Learning Accelerator), NPU (Neural-Network Processing Unit), DSP (Digital Signal Processing Unit), Field-Programmable Gate Array (FPGA), and Application Specific Integrated Circuit (ASIC). This application does not limit the specific type of processor.
[0188] Memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by processor 1020 or other modules of the computer. Permanent storage devices may be read-write storage devices. Permanent storage devices may be non-volatile storage devices that retain stored instructions and data even when the computer is powered off. In some embodiments, permanent storage devices use mass storage devices (e.g., magnetic or optical disks, flash memory) as permanent storage devices. In other embodiments, permanent storage devices may be removable storage devices (e.g., floppy disks, optical drives). System memory may be a read-write storage device or a volatile read-write storage device, such as dynamic random access memory. System memory may store some or all of the instructions and data required by the processor during operation. Furthermore, memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (e.g., DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a removable storage device that is readable and / or writable, such as a laser disc (CD), a read-only digital multifunction optical disc (e.g., DVD-ROM, dual-layer DVD-ROM), a read-only Blu-ray disc, an ultra-high density optical disc, a flash memory card (e.g., SD card, mini SD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not contain carrier waves or transient electronic signals transmitted wirelessly or via wired connections.
[0189] The memory 1010 stores executable code, which, when processed by the processor 1020, can cause the processor 1020 to execute part or all of the methods described above.
[0190] In one possible implementation, the artificial intelligence accelerator may include multiple processors, each of which can independently run various assigned tasks. This application does not limit the processors or the tasks they run.
[0191] It is understood that, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or in software program modules.
[0192] When the integrated unit / module is implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the artificial intelligence processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage module can be any suitable magnetic or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0193] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0194] In one possible implementation, an artificial intelligence chip is also disclosed, which includes the aforementioned hardware acceleration circuitry.
[0195] In one possible implementation, a board is also disclosed, which includes a storage device, an interface device, a controller, and the aforementioned artificial intelligence chip; wherein the artificial intelligence chip is connected to the storage device, the controller, and the interface device respectively; the storage device is used to store data; the interface device is used to realize data transmission between the artificial intelligence chip and external devices; and the controller is used to monitor the status of the artificial intelligence chip.
[0196] In one possible implementation, an electronic device is disclosed that includes the aforementioned artificial intelligence chip. The electronic device includes data processing devices, robots, computers, printers, scanners, tablets, smart terminals, mobile phones, dashcams, navigators, sensors, cameras, servers, cloud servers, cameras, camcorders, projectors, watches, headphones, mobile storage, wearable devices, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound scanners, and / or electrocardiographs.
[0197] Furthermore, the method according to this application can also be implemented as a computer program or computer program product, which includes computer program code instructions for performing some or all of the steps in the method described above.
[0198] Alternatively, this application may be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium) storing executable code (or computer program or computer instruction code) thereon, which, when executed by a processor of an electronic device (or server, etc.), causes the processor to perform part or all of the steps of the methods described above according to this application.
[0199] The various embodiments of this application have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A hardware acceleration circuit, characterized in that, include: The exponential function module is used to obtain multiple exponential function values for multiple data elements in a dataset; An adder is used to obtain the result of adding the multiple exponential function values; A first processing circuit is configured to perform preset processing on the addition result to process the addition result into at least a first data and a second data, wherein the length of the addition result is N1 bits, the length of the first data is N2 bits, and the length of the second data is N3 bits, where N2 and N3 are both less than N1; specifically, the first processing circuit is configured to convert the addition result from an integer into a floating-point number represented by a first exponent data and a first mantissa data; The second processing circuit is used to perform at least preset processing on the first data and the second data to obtain the reciprocal of the addition result; the second processing circuit includes a first conversion circuit to convert the first exponent data into a negative number, a second conversion circuit to convert the fractional part of the floating-point number into another floating-point number represented by the second exponent data and the second mantissa data according to the first mantissa data, and a third conversion circuit to obtain the reciprocal of the addition result according to the negative number, the second exponent data and the second mantissa data. The third processing circuit is used to perform preset processing on the exponential function value and the reciprocal of the i-th data element among the plurality of data elements to obtain a specific function value of the i-th data element; The hardware acceleration circuit is used to implement the flexible maximum function layer of the neural network.
2. The hardware acceleration circuit as described in claim 1, characterized in that, The third conversion circuit includes: An exponential adder is used to obtain the sum of the negative of the first exponential data and the second exponential data; A shifter is used to shift the second mantissa data by using the sum as a shift parameter.
3. The hardware acceleration circuit as described in claim 1, characterized in that: The exponential function module includes: a first lookup table circuit, used to obtain multiple exponential function values corresponding to multiple data elements in the data set based on the first lookup table; and / or, The first conversion circuit includes: a second lookup table circuit, used to obtain the negative number corresponding to the first exponent data based on the second lookup table; and / or, The second conversion circuit includes: a third lookup table circuit for obtaining second exponent data corresponding to the first mantissa data based on the third lookup table; and a fourth lookup table circuit for obtaining second mantissa data corresponding to the first mantissa data based on the fourth lookup table.
4. The hardware acceleration circuit as described in claim 1, characterized in that: The length of the first exponent data and the second exponent data is N2 bits, and the length of the first mantissa data and the second mantissa data is N3 bits. The values of N2 and N3 are in the range [1, 32].
5. The hardware acceleration circuit as described in claim 3, characterized in that: The hardware acceleration circuit includes at least two of the first to fourth lookup table circuits. Each of the at least two lookup table circuits has a basic lookup table circuit unit; or... The at least two lookup table circuits share a basic lookup table circuit unit.
6. The hardware acceleration circuit as described in claim 1, characterized in that: The result of the addition operation is a floating-point number; The first processing circuit includes: a third lookup table circuit, used to obtain exponent data corresponding to the addition result based on the third lookup table; and a fourth lookup table circuit, used to obtain mantissa data corresponding to the addition result based on the fourth lookup table. The second processing circuit is used to perform preset processing on the exponent data and mantissa data to obtain the reciprocal of the addition result; The third processing circuit is used to perform preset processing on the exponential function value and the reciprocal of the i-th data element among the plurality of data elements to obtain a specific function value of the i-th data element.
7. The hardware acceleration circuit as described in claim 1, characterized in that: Also includes: A subtractor is used to subtract the maximum value among multiple initial data in an initial dataset to obtain the data set containing the multiple data elements; The third processing circuit includes a multiplier, which is used to multiply the exponential function value of the i-th data element among the plurality of data elements with the reciprocal, and output the flexible maximum function value of the i-th data element.
8. An artificial intelligence chip, characterized in that, Includes the hardware acceleration circuit as described in any one of claims 1 to 7.
9. A data processing acceleration method, characterized in that, Applied to artificial intelligence accelerators, the method is used to implement a flexible maximum function layer of a neural network, which is used to classify data to be processed; wherein the data to be processed includes at least one of voice data, text data, and image data; The method includes: Obtain multiple exponential function values for multiple data elements in a dataset; Obtain the result of the addition operation of the multiple exponential function values; Obtain the reciprocal of the result of the addition operation; Based on the exponential function value of the i-th data element among the plurality of data elements and the reciprocal, the specific function value of the i-th data element is obtained; The process of obtaining the reciprocal of the addition result includes: processing the addition result into at least a first data and a second data; and obtaining the reciprocal of the addition result based at least on the first data and the second data. The step of processing the addition result into at least a first data and a second data includes: converting the addition result from an integer into a floating-point number represented by a first exponent data and a first mantissa data; The step of obtaining the reciprocal of the addition result based at least on the first data and the second data includes: converting the first exponent data into a negative number, converting the fractional part of the floating-point number into another floating-point number represented by the second exponent data and the second mantissa data based on the first mantissa data, and obtaining the reciprocal of the addition result based on the negative number, the second exponent data, and the second mantissa data. The length of the addition result is N1 bits, the length of the first data is N2 bits, and the length of the second data is N3 bits, where N2 and N3 are both less than N1.
10. The method according to claim 9, characterized in that: The result of the addition operation is a floating-point number; The step of processing the addition result into at least first data and second data includes converting the addition result into exponent data and mantissa data.
11. The method according to claim 9, characterized in that: The result of the addition operation is an integer; The step of processing the addition result into at least a first data and a second data includes: converting the addition result from an integer into a floating-point number represented by a first exponent data and a first mantissa data, and, based on the first mantissa data, converting the fractional part of the floating-point number into another floating-point number represented by a second exponent data and a second mantissa data; Obtaining the reciprocal of the addition result based at least on the first data and the second data includes: obtaining the reciprocal of the addition result based at least on the first exponent data, the second exponent data, and the second mantissa data.
12. The method according to claim 11, characterized in that, Based at least the first exponent data, the second exponent data, and the second mantissa data, the reciprocal corresponding to the addition operation result is obtained, including: Obtain the negative value of the first index data; Obtain the sum of the negative value of the first exponential data and the second exponential data; The sum is used as a shift parameter to shift the second mantissa data.
13. The method according to claim 12, characterized in that, Obtaining multiple exponential function values for multiple data elements in the dataset includes: obtaining multiple exponential function values corresponding to multiple data elements in the dataset based on a first lookup table; and / or, Obtaining the negative number of the first index data includes: obtaining the negative number corresponding to the first index data based on the second lookup table.
14. The method according to claim 11, characterized in that, The step of obtaining the second exponent data and the second tail data based on the first tail data includes: Based on the third lookup table, the second index data corresponding to the first tail number data is obtained; Based on the fourth lookup table, the second tail number data corresponding to the first tail number data is obtained.
15. The method according to claim 11, characterized in that: The length of the first exponent data and the second exponent data is N2 bits, and the length of the first mantissa data and the second mantissa data is N3 bits. The values of N2 and N3 are in the range [1, 32].
16. The method according to claim 9, characterized in that, Before obtaining the multiple exponential function values of multiple data elements in the dataset, the method further includes: subtracting the maximum value among the multiple initial data in the initial dataset from the multiple initial data to obtain the dataset containing the multiple data elements.
17. The method according to claim 9, characterized in that, Obtaining the specific function value of the i-th data element based on the exponential function value of the i-th data element and its reciprocal includes: The exponential function value of the i-th data element among the plurality of data elements is multiplied by the reciprocal to obtain the maximum flexible function value of the i-th data element.
18. An artificial intelligence accelerator, characterized in that, include: processor; as well as A memory having executable code stored thereon, which, when executed by the processor, causes the processor to perform the method as described in any one of claims 9-17.