Data processing circuit, floating point operation method and related device
By introducing arithmetic circuits that process floating-point data in parallel for the non-exponential domain, exponential domain, and shift parameters, the problem of long computation time in arithmetic units of pulsating array circuits is solved, and the computational performance of the data processing circuit is improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2026-05-28
- Publication Date
- 2026-07-21
AI Technical Summary
The pulsating array circuit's arithmetic units take a long time to perform operations on floating-point data, resulting in poor computational performance of the data processing circuit.
By introducing a first arithmetic circuit, a second arithmetic circuit, and a third arithmetic circuit into the data processing circuit, the non-exponential domain, exponential domain, and shift parameters of floating-point data are processed respectively, thereby achieving parallel operation and reducing computation time.
The computing performance of the data processing circuit has been improved. Parallel computing has reduced the time consumption of floating-point data operations and enhanced the computing performance of the processing unit.
Smart Images

Figure CN122433818A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data processing technology, specifically relating to a data processing circuit, a floating-point arithmetic method, and related equipment. Background Technology
[0002] With the continuous development of artificial intelligence technology, neural network models have been widely used in fields such as image recognition, language processing, and autonomous driving. Currently, in the process of processing tasks, neural network models can use multiple arithmetic units of data processing circuits (such as pulsating array circuits) to perform operations on large amounts of floating-point data and output the processing results of the task based on the operation results.
[0003] However, because each arithmetic unit in the above-mentioned pulsating array circuit takes a long time to perform operations on floating-point data, the operation performance of each arithmetic unit is poor, which in turn leads to poor operation performance of the data processing circuit. Summary of the Invention
[0004] The purpose of this application is to provide a data processing circuit, a floating-point arithmetic method, and related equipment, which can improve the computing performance of the data processing circuit.
[0005] In a first aspect, embodiments of this application provide a data processing circuit, which includes a processing unit comprising: a first arithmetic circuit for performing arithmetic on the non-exponential field of a first floating-point data to be processed input at an input terminal of the first arithmetic circuit to obtain a first intermediate result; a second arithmetic circuit for performing arithmetic on the exponential field of the first floating-point data to be processed input at a first input terminal of the second arithmetic circuit and the first exponential field input at a second input terminal of the second arithmetic circuit to obtain a first exponential field result and a first shift parameter, wherein the first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential field to the first exponential field result; and a third arithmetic circuit connected to the first and second arithmetic circuits for performing shift and arithmetic on the first intermediate result, the first shift parameter, and the first non-exponential field input at a first input terminal of the third arithmetic circuit to obtain a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field.
[0006] Secondly, embodiments of this application provide an electronic device, which includes: the data processing circuit as described in the first aspect.
[0007] Thirdly, embodiments of this application provide a floating-point arithmetic method, which is applied to an electronic device as described in the second aspect. The method includes: the electronic device performing arithmetic on the non-exponential field of a first floating-point data to be arithmetic by a first arithmetic circuit of a data processing unit of the electronic device to obtain a first intermediate result; and performing arithmetic on the exponential field and a first exponential field of the first floating-point data to be arithmetic by a second arithmetic circuit of the processing unit to obtain a first exponential field result and a first shift parameter, wherein the first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential field to the first exponential field result; and performing shift and arithmetic on the first intermediate result, the first shift parameter, and the first non-exponential field by a third arithmetic circuit of the processing unit to obtain a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field.
[0008] Fourthly, embodiments of this application provide a floating-point arithmetic device, which includes a processing module. The processing module is configured to perform calculations based on the non-exponential field of a first floating-point data to be calculated, obtaining a first intermediate result; and to perform calculations based on the exponent field and a first exponent field of the first floating-point data to be calculated, obtaining a first exponent field result and a first shift parameter, the first shift parameter indicating the shift amount required to align the first intermediate result and the first non-exponential field to the first exponent field result; and to perform shift and calculations based on the first intermediate result, the first shift parameter, and the first non-exponential field, obtaining a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field.
[0009] Fifthly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the method as described in the third aspect.
[0010] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described in the third aspect.
[0011] In a seventh aspect, embodiments of this application provide a chip including a processor and a communication interface coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method as described in the third aspect.
[0012] Eighthly, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method as described in the third aspect.
[0013] In this embodiment, the data processing circuit includes a first arithmetic circuit, a second arithmetic circuit, and a third arithmetic circuit. The first arithmetic circuit can perform operations based on the non-exponential field of the first floating-point data to be processed, and the second arithmetic circuit can perform operations based on the exponential field and the first exponential field of the first floating-point data to be processed. In other words, the first arithmetic circuit and the second arithmetic circuit can perform operations based on different components of the same floating-point data. That is, the first arithmetic circuit and the second arithmetic circuit can perform operations relatively independently. This reduces the time consumed by the first arithmetic circuit and the second arithmetic circuit in performing operations in parallel, so as to quickly obtain the first intermediate result, the first exponential field result, and the first shift parameter. Thus, the third arithmetic circuit can quickly perform shift and operation based on the first intermediate result, the first shift parameter, and the first non-exponential field to obtain the first non-exponential field result. Therefore, the time consumed by the processing unit in performing operations on floating-point data can be reduced, thereby improving the computing performance of the processing unit. In this way, the computing performance of the data processing circuit can be improved. Attached Figure Description
[0014] Figure 1 This is a schematic diagram of the arithmetic unit in a pulsating array circuit provided by related technologies;
[0015] Figure 2 This is one of the structural schematic diagrams of the data processing circuit provided in the embodiments of this application;
[0016] Figure 3 This is a second schematic diagram of the data processing circuit provided in the embodiments of this application;
[0017] Figure 4 This is the third schematic diagram of the data processing circuit provided in the embodiments of this application;
[0018] Figure 5 This is the fourth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0019] Figure 6 This is the fifth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0020] Figure 7 This is the sixth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0021] Figure 8 This is the seventh schematic diagram of the data processing circuit provided in the embodiments of this application;
[0022] Figure 9 This is the eighth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0023] Figure 10 This is the ninth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0024] Figure 11 This is the tenth schematic diagram of the data processing circuit provided in the embodiments of this application;
[0025] Figure 12 This is eleventh of the structural schematic diagrams of the data processing circuit provided in the embodiments of this application;
[0026] Figure 13 This is a schematic diagram of the chip structure provided in the embodiments of this application;
[0027] Figure 14 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;
[0028] Figure 15 This is a flowchart illustrating the floating-point arithmetic method provided in an embodiment of this application;
[0029] Figure 16 This is a schematic diagram of the structure of the floating-point arithmetic device provided in the embodiments of this application;
[0030] Figure 17 This is one of the hardware structure diagrams of the electronic device provided in the embodiments of this application;
[0031] Figure 18 This is the second schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0032] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0033] The following will explain the technical terms used in the embodiments of this application.
[0034] 1. Pulsating array circuit
[0035] A systolic array circuit is a highly modular and regularized computing architecture comprising multiple arithmetic units. These arithmetic units are arranged in a grid to form an array structure. Floating-point data can be input into the corresponding rows or columns in the systolic array circuit, and then sequentially transferred from one arithmetic unit to the other in the corresponding row or column to perform floating-point multiplication and accumulation operations.
[0036] 2. Floating-point numbers
[0037] Floating-point numbers are a data type in computers that approximates real numbers. Their characteristic is that the position of the decimal point can float in order to balance the range and precision of the numerical value with a fixed number of bits.
[0038] In computers, the formula for representing floating-point numbers is: ;
[0039] Where: S is the sign bit, used to indicate the sign of the floating-point number, taking the value 0 or 1. When S is 0, the floating-point number is positive; when S is 1, the floating-point number is negative. M is the mantissa, used to indicate the significant digits in binary, reflecting the precision of the floating-point number. 2 indicates the base of the true exponent, representing that the floating-point number uses binary notation. E is the true exponent, which can be positive or negative, directly determining the scaling factor of the value.
[0040] To simplify the computer's storage and operation of exponents (avoiding the handling of negative numbers), the true exponent E is not stored directly. Instead, it is converted into an unsigned storage exponent exp using a fixed offset. The conversion method is as follows:
[0041]
[0042] The value of the fixed offset is determined by the precision of the floating-point number. For example, the fixed offset for a single-precision floating-point number is 127, and the fixed offset for a double-precision floating-point number is 1023.
[0043] For example, if a single-precision floating-point number is represented as: Then we can conclude that the sign bit S is 0, meaning the floating-point number is positive, the mantissa M is 1.011, the true exponent E is 2, and the storage exponent exp is 129.
[0044] It should be noted that internally, the computer stores and performs calculations on floating-point numbers based on the stored exponent exp. When the results need to be displayed to the user, the stored exponent exp is restored to the actual exponent E, and then converted back to a regular data format for the user to view.
[0045] 3. Other terms
[0046] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0047] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0048] With the continuous development of artificial intelligence technology, neural network models have been widely used in fields such as image recognition, language processing, and autonomous driving. Currently, in the process of processing tasks, neural network models can use multiple arithmetic units of data processing circuits (such as pulsating array circuits) to perform operations on large amounts of floating-point data and output the processing results of the task based on the operation results.
[0049] However, the applicant found in practical application that the data processing circuit's performance was poor because each arithmetic unit in the pulsating array circuit took a long time to perform operations on floating-point data.
[0050] For example, Figure 1 A schematic diagram of the arithmetic unit in a pulsating array circuit provided for related technologies. (e.g.) Figure 1 As shown, the arithmetic unit includes: two floating-point multipliers (e.g., floating-point multiplier 0 and floating-point multiplier 1), floating-point adder 0, register 0, floating-point adder 1, and register 1.
[0051] The arithmetic unit can perform multiply-accumulate operations on two floating-point data pairs (i.e., floating-point data pair 1 (including data1_0 and data0_0) and floating-point data pair 2 (including data0_1 and data1_1)) and floating-point data output from the previous arithmetic unit (e.g., acc_n).
[0052] The following will provide an illustrative example of the operations performed by the arithmetic unit.
[0053] First, each floating-point multiplier performs a multiplication operation on its corresponding floating-point data pair. For example, floating-point multiplier 0 multiplies data0_0 and data1_0 (i.e., floating-point data pair 1), resulting in a multiplication result of 0. As another example, floating-point multiplier 1 multiplies data0_1 and data1_1 (i.e., floating-point data pair 2), resulting in a multiplication result of 1. It should be noted that the multiplication operations of different floating-point multipliers are performed in parallel. Next, floating-point adder 0 performs an addition operation on the multiplication result 0 and the multiplication result 1, resulting in an addition result of 0, which is then stored in register 0.
[0054] Then, floating-point adder 1 performs floating-point addition on the addition result 0 and the output of the previous arithmetic unit to obtain addition result 1, stores addition result 1 in register 1, and outputs addition result 1 to the next arithmetic unit.
[0055] Specifically, since floating-point data includes a sign field, a mantissa field, and an exponent field, for each floating-point multiplier, during multiplication, it first performs a multiplication operation of the sign field and the mantissa field, and then performs an addition operation of the exponent field to obtain the multiplication result. When floating-point adder 0 performs an addition operation on the multiplication results calculated by floating-point multiplier 0 and floating-point multiplier 1, it needs to perform order alignment processing on all multiplication results. That is, it compares the exponent fields in all multiplication results, unifies the exponent fields in all multiplication results to the one with the largest value among all exponent fields, and then right-shifts the mantissa field of each multiplication result according to the difference between the exponent field of each multiplication result and the unified exponent field, thus completing the order alignment. After completing the order alignment, the addition operation is performed to obtain the addition result 0.
[0056] Similarly, when performing floating-point addition, floating-point adder 1 also needs to perform alignment. Only after alignment can the addition operation be performed to obtain the addition result 1.
[0057] Since the operating logic of each arithmetic unit is the same, the latency of the arithmetic unit's operation is determined by the longer of the first-level combinational logic and the second-level combinational logic. The first-level combinational logic can be understood as the logic of floating-point multiplier 0, floating-point multiplier 1 and floating-point adder 0, and the second-level combinational logic can be understood as the logic of floating-point adder 1.
[0058] As shown above, the second-level combinational logic performs alignment before addition, while the first-level combinational logic performs multiplication of the sign field and mantissa field simultaneously, adds the exponents, aligns the exponents, and then performs addition. Clearly, the first-level combinational logic is significantly longer than the second-level combinational logic. Furthermore, the delay of multiplication is longer than that of addition; therefore, the delay of the first level is longer than that of the second level. Thus, the length of the clock cycle of the arithmetic unit is determined by the delay of the first-level combinational logic.
[0059] However, since the combinational logic of the first stage is mainly serial, meaning that subsequent operations must depend on the results of the preceding operations, the delay of the first stage is relatively long, which in turn leads to a longer delay of the entire arithmetic unit. This means that the operation takes a long time, and the arithmetic unit cannot operate at a higher clock frequency. As a result, the operation performance of the systolic array circuit is poor, which in turn leads to poor operation performance of the data processing circuit.
[0060] To address the aforementioned technical problems, embodiments of this application provide a data processing circuit, a floating-point arithmetic method, and related equipment. The data processing circuit, floating-point arithmetic method, and related equipment provided in this application will be described in detail below with reference to the accompanying drawings and specific embodiments and application scenarios.
[0061] Figure 2 A schematic diagram of the data processing circuit provided in an embodiment of this application is shown. Figure 2 As shown, the data processing circuit provided in this application embodiment includes a processing unit 10, which includes: a first arithmetic circuit 11, used to perform calculations on the non-exponential domain of the first floating-point data to be calculated input at the input terminal of the first arithmetic circuit 11 to obtain a first intermediate result; a second arithmetic circuit 12, used to perform calculations on the exponential domain of the first floating-point data to be calculated input at the first input terminal of the second arithmetic circuit 12 and the first exponential domain input at the second input terminal of the second arithmetic circuit 12 to obtain a first exponential domain result and a first shift parameter, wherein the first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential domain to the first exponential domain result; and a third arithmetic circuit 13, connected to the first arithmetic circuit 11 and the second arithmetic circuit 12, used to perform shift and calculations on the first intermediate result, the first shift parameter, and the first non-exponential domain input at the first input terminal of the third arithmetic circuit 13 to obtain a first non-exponential domain result.
[0062] In this embodiment of the application, the non-exponential field includes the sign field and the mantissa field.
[0063] In some embodiments of this application, the data processing circuit described above may include, but is not limited to, any of the following: a pulsation array circuit, a pulsation vector processing circuit, or a wavefront array circuit.
[0064] The aforementioned pulsating array circuit can be configured in any of the following: Application Specific Integrated Circuit (ASIC), Field-Programmable Gate Array (FPGA), Graphics Processing Unit (GPU), System-on-a-Chip (SOC), Neural Processing Unit (NPU).
[0065] In some embodiments of this application, the processing unit 10 described above can be a basic computing unit of a data processing circuit. For example, the processing unit 10 can be an arithmetic unit.
[0066] In some embodiments of this application, the number of processing units 10 can be at least one. When the number of processing units 10 is at least two, the at least two processing units 10 can be connected end to end or arranged in parallel.
[0067] It should be noted that the above "connected end to end" can be understood as: the output end of one processing unit 10 is connected to the input end of another processing unit 10.
[0068] It is understood that when at least two processing units 10 are connected end-to-end, the first processing unit 10 can perform calculations based on the input floating-point data and default floating-point data, and output the calculated floating-point data to the second processing unit 10. Any processing unit 10 other than the first processing unit 10 can perform calculations based on the input floating-point data and the floating-point data output by the preceding processing unit 10, and output the calculated floating-point data to the following processing unit 10. The aforementioned floating-point data may include, but is not limited to, the aforementioned floating-point numbers, and the default floating-point data may include, but is not limited to, floating-point zero data, that is, the exponent field and non-exponent field of the default floating-point data can both be 0.
[0069] In some embodiments of this application, the number of the first floating-point data to be calculated can be at least one, and each first floating-point data to be calculated can include two floating-point data.
[0070] In some embodiments of this application, the number of input terminals of the first arithmetic circuit 11 can be at least one. Each input terminal may include two sub-input terminals, and the sub-input terminals of each input terminal are used to obtain the non-exponential field of a floating-point data in a first floating-point data to be operated on from the data source, thereby obtaining at least one non-exponential field of the first floating-point data to be operated on.
[0071] For example, combined with Figure 2 ,like Figure 3As shown, the first arithmetic circuit 11 has two input terminals. One input terminal includes a sub-input terminal 111 and a sub-input terminal 112. The sub-input terminal 111 is used to obtain the non-exponential field of one floating-point data in a first floating-point data to be calculated from the data source (e.g., data0_0(sign+mts)). The sub-input terminal 112 is used to obtain the non-exponential field of another floating-point data in the same first floating-point data to be calculated from the data source (e.g., data1_0(sign+mts)). The other input terminal includes a sub-input terminal 113 and a sub-input terminal 114. The sub-input terminal 113 is used to obtain the non-exponential field of one floating-point data in another first floating-point data to be calculated from the data source (e.g., data0_1(sign+mts)). The sub-input terminal 114 is used to obtain the non-exponential field of another floating-point data in the same first floating-point data to be calculated from the data source (e.g., data1_1(sign+mts)). Thus, the first arithmetic circuit 11 can obtain the non-exponential fields of two first floating-point data to be calculated.
[0072] In some embodiments of this application, the first arithmetic circuit 11 can perform multiplication, division, addition, or subtraction operations based on the first floating-point data to be processed. Those skilled in the art can make specific settings according to their needs.
[0073] The following will take the multiplication of the first arithmetic circuit 11 based on the above-mentioned first floating-point data to be operated as an example to illustrate the specific structure of the first arithmetic circuit 11.
[0074] In some embodiments of this application, combined with Figure 2 ,like Figure 4 As shown, the first arithmetic circuit 11 may include a multiplier 115, the input of which is used to receive the non-exponential field of the first floating-point data to be operated on, and the output of which is connected to the third arithmetic circuit 13; wherein, the multiplier 115 is used to perform multiplication operations based on the non-exponential field of the first floating-point data to be operated on.
[0075] It can be understood that the input terminal of multiplier 115 can be regarded as the input terminal of the first operational circuit 11.
[0076] In some examples, the number of the aforementioned multipliers 115 is at least one. Figure 4 The diagram illustrates two multipliers 115. The input of each multiplier 115 is used to receive the non-exponential field of a first floating-point data to be operated on, and the output of each multiplier 115 is connected to the third arithmetic circuit 13. Each multiplier 115 is used to perform multiplication based on the non-exponential field of a first floating-point data to be operated on.
[0077] In some examples, the input of each multiplier 115 may include two sub-inputs, each sub-input of each multiplier 115 may receive a non-exponential field of a floating-point data in a first floating-point data to be computed.
[0078] For example, combined with Figure 4 The input terminals of a multiplier 115 may include sub-input terminals 1151 and 1152. Sub-input terminal 1151 is used to receive a first floating-point data to be operated on (e.g., data0_0(sign+mts)), and sub-input terminal 1152 is used to receive another first floating-point data to be operated on (e.g., data1_0(sign+mts)). Another multiplier 115 may include sub-input terminals 1153 and 1154. Sub-input terminal 1153 is used to receive yet another first floating-point data to be operated on (e.g., data0_1(sign+mts)), and sub-input terminal 1154 is used to receive yet another first floating-point data to be operated on (e.g., data1_1(sign+mts)).
[0079] In some examples, each multiplier 115 may have only one output terminal. The output terminal of each multiplier 115 may be directly connected to the third arithmetic circuit 13, or it may be indirectly connected to the third arithmetic circuit 13 through other components.
[0080] Optionally, combined Figure 4 ,like Figure 5 As shown, the first arithmetic circuit 11 further includes a first register 116, which is located on the path connecting the output of the multiplier 115 and the third arithmetic circuit 13.
[0081] The number of first registers 116 can be at least one, and each first register 116 is respectively set on the path connecting the output of a multiplier 115 and the third arithmetic circuit 13.
[0082] It is understandable that the output of each multiplier 115 can be indirectly connected to the third arithmetic circuit 13 through a first register 116.
[0083] Thus, since the output of the multiplier can also be connected to the third arithmetic circuit through the first register, the result of the multiplication operation can be stored in the connected first register after the multiplier operation is completed. Therefore, the loss of the operation result can be avoided, thus avoiding the situation where the third arithmetic circuit cannot obtain the operation result, and further avoiding the situation where the third arithmetic circuit cannot perform the operation. In this way, the reliability of the processing unit operation can be improved.
[0084] In some examples, at least some of the multipliers 115 in at least one multiplier 115 can operate in parallel.
[0085] In some examples, each multiplier 115 can perform multiplication based on the non-exponential fields of two floating-point data in a first set of floating-point data to be operated on, to obtain a first intermediate result.
[0086] Thus, since the first arithmetic circuit may include a multiplier, and the input terminal of the multiplier is connected to the first input terminal of the first arithmetic circuit, the multiplier can perform multiplication based on the non-exponential domain of the first floating-point data to be operated on, so as to obtain an accurate first intermediate result. Therefore, the operation accuracy of the first arithmetic circuit can be improved.
[0087] In some embodiments of this application, the number of first input terminals of the second arithmetic circuit 12 described above can be at least one. Each first input terminal may include two sub-input terminals, and each sub-input terminal is used to obtain the exponent field of a floating-point data in a first floating-point data to be operated on from the data source, thereby obtaining at least one exponent field of the first floating-point data to be operated on.
[0088] For example, combined with Figure 2 ,like Figure 6 As shown, the second arithmetic circuit 12 has two first input terminals. One input terminal includes sub-input terminal 121 and sub-input terminal 122. Sub-input terminal 121 is used to obtain the exponent field of one floating-point data in the first floating-point data to be calculated from the data source (e.g., data0_1(exp)). Sub-input terminal 122 is used to obtain the exponent field of another floating-point data in the first floating-point data to be calculated from the data source (e.g., data0_0(exp)). The other input terminal includes sub-input terminal 123 and sub-input terminal 124. Sub-input terminal 123 is used to obtain the exponent field of one floating-point data in another first floating-point data to be calculated from the data source (e.g., data1_0(exp)). Sub-input terminal 124 is used to obtain the exponent field of another floating-point data in the other first floating-point data to be calculated from the data source (e.g., data1_1(exp)). Thus, the second arithmetic circuit 12 can obtain the exponent fields of two first floating-point data to be calculated.
[0089] In some examples, such as Figure 6 As shown, the second input terminal of the second operational circuit 12 can be the input terminal 127, so that the input terminal 127 can obtain the first exponent field (e.g., acc_n(exp)).
[0090] In some examples, the first input terminal of the second arithmetic circuit 12 corresponds one-to-one with the first input terminal of the first arithmetic circuit 11, and the first input terminal of the second arithmetic circuit 12 and the first input terminal of the first arithmetic circuit 11 input different data portions of the same first floating-point data to be operated.
[0091] For example, if the first input terminal of the second arithmetic circuit 12 is the exponent field of a certain first floating-point data to be operated on, then the first input terminal of the first arithmetic circuit 11 corresponding to that first input terminal is the non-exponential field of that certain first floating-point data to be operated on.
[0092] In some embodiments of this application, the second input terminal of the second arithmetic circuit 12 can be left floating or connected to the output terminal of the second arithmetic circuit 12 of the previous processing unit 10. It can be understood that the first exponent field can be the exponent field of the default floating-point data (e.g., floating-point zero data) or the exponent field result (e.g., the first exponent field result) output by the output terminal of the second arithmetic circuit 12 of the previous processing unit 10.
[0093] For example, when there is only one processing unit 10, the second input terminal of the second arithmetic circuit 12 can be left floating. When there are at least two processing units 10, the second input terminal of the second arithmetic circuit 12 of the first processing unit 10 can be left floating, and the second input terminals of the second arithmetic circuits 12 of other processing units 10 besides the first processing unit 10 can be connected to the output terminal of the second arithmetic circuit of the previous processing unit 10, so that the output terminal of the second arithmetic circuit 12 of the previous processing unit 10 can output the first exponential domain result, and the second input terminal of the second arithmetic circuit 12 of the other processing unit 10 can use the first exponential domain result as the first exponential domain and perform the operation again.
[0094] In some embodiments of this application, the second arithmetic circuit 12 may have two output terminals: one output terminal is used to output the first exponential domain result, and the other output terminal is connected to the third arithmetic circuit 13, which is used to output the first shift parameter.
[0095] In some embodiments of this application, combined with Figures 2 to 6 ,like Figure 7 As shown, the second arithmetic circuit 12 includes: an exponentiation arithmetic unit 121, the first input of which is used to receive the exponent field of the first floating-point data to be operated on, and the second input of which is used to receive the first exponent field; wherein, the exponentiation arithmetic unit 121 is used to perform exponentiation addition based on the exponent field of the first floating-point data to be operated on, to obtain an intermediate result of the exponent field, and to determine the exponent field result as the first exponent field result by combining the intermediate result of the exponent field and the exponent field with the largest value in the first exponent field, and to calculate the first shift parameter based on the first exponent field result, the intermediate result of the exponent field and the first exponent field.
[0096] It can be understood that the first input terminal of the above-mentioned exponent arithmetic unit 121 can be regarded as the first input terminal of the second arithmetic circuit 12, and the second input terminal of the above-mentioned exponent arithmetic unit 121 can be regarded as the second input terminal of the second arithmetic circuit 12.
[0097] In some examples, combined Figure 6 and Figure 7 The first input terminal of the exponentiation operator 121 includes input terminals 1211, 1212, 1213, and 1214. Input terminal 1211 is used to receive data0_1 (exp), input terminal 1212 is used to receive data0_0 (exp), input terminal 1213 is used to receive data1_0 (exp), and input terminal 1214 is used to receive data1_1 (exp). The second input terminal of the exponentiation operator 121 includes input terminal 1215, which is connected to the aforementioned input terminal 127. Input terminal 1215 is used to receive acc_n (exp).
[0098] In this embodiment, since the processing unit 10 may perform multiplication on the non-exponential field of the first floating-point data to be computed, correspondingly, the processing unit 10 may also need to perform multiplication on the exponential field of the first floating-point data to be computed. Therefore, the exponentiation unit 121 can perform exponential addition based on the exponential field of the first floating-point data to be computed, obtaining an intermediate result in the exponential field. Furthermore, since the processing unit 10 may perform addition on the first intermediate result and the first non-exponential field in subsequent steps, the exponentiation unit 121 can determine the first exponential field result as the exponential field intermediate result and the exponential field with the largest value in the first exponential field, and calculate the first shift parameter based on the first exponential field result, the exponential field intermediate result, and the first exponential field.
[0099] In some examples, one output of the exponentiation unit 121 can be directly connected to the third arithmetic circuit 13, or it can be indirectly connected to the third arithmetic circuit 13 through other components.
[0100] Optionally, combined Figure 7 ,like Figure 8 As shown, the second arithmetic circuit 12 further includes a second register 125, which is located on the path connecting the output of the exponent arithmetic unit 121 and the third arithmetic circuit 13.
[0101] It is understandable that the output of the exponentiation unit 121 can be indirectly connected to the third arithmetic circuit 13 through the second register 125.
[0102] Thus, since the output of the exponent arithmetic unit can also be connected to the third arithmetic circuit through the second register, the result of the operation can be stored in the second register after the exponent arithmetic unit has completed the operation. Therefore, the loss of the operation result can be avoided, thus avoiding the situation where the third arithmetic circuit cannot obtain the operation result, and further avoiding the situation where the third arithmetic circuit cannot perform the operation. In this way, the reliability of the processing unit operation can be improved.
[0103] Furthermore, since the exponentiation unit only needs to perform operations based on floating-point data (i.e., the exponent field and the first exponent field of the first floating-point data to be operated on), rather than based on the exponent field and non-exponent field of the floating-point data, the exponentiation unit can be designed to be smaller and consume less power compared to adders in related technologies. This reduces the size and power consumption of the processing unit, thereby improving the power consumption, performance, area (PPA) of the processing unit.
[0104] In some examples, another output of the aforementioned exponentiation operator 121 can output the first exponent field result.
[0105] In some examples, the number of the above-mentioned exponentiation operators 121 can also be at least two.
[0106] For example, combined with Figure 7 ,like Figure 11As shown, there are two exponentiation operators 121, namely a first exponentiation operator 12111 and a second exponentiation operator 12112. The first exponentiation operator 12111 has its first input terminal used to receive first floating-point data 1 to be processed (including data0_1(exp) and data0_0(exp)) and first floating-point data 2 to be processed (including data1_0(exp) and data1_1(exp)). The first exponentiation operator 12111 can be based on data0_1(exp) and... Data0_0(exp) is subjected to exponential addition to obtain intermediate result 1. Then, data1_0(exp) and data1_1(exp) are subjected to exponential addition to obtain intermediate result 2. The exponential field with the largest value between intermediate result 1 and intermediate result 2 is determined as the first exponential field result 1. Based on the first exponential field result 1, intermediate result 1 and intermediate result 2, a first shift parameter 1 is determined. The first shift parameter 1 indicates the shift amount required to align intermediate result 1 and intermediate result 2 to the first exponential field result 1. The first shift parameter 1 is sent to the third arithmetic circuit 13. The first input terminal of the second exponent arithmetic unit 12112 is connected to the output terminal of the first exponent arithmetic unit 12111. The second input terminal of the second exponent arithmetic unit 12112 is used to obtain the first exponent field. Thus, the second exponent arithmetic unit 12112 can obtain the first exponent field result 1 from the output terminal of the first exponent arithmetic unit 12111. The first exponent field result 1 and the exponent field with the largest value in the first exponent field are determined as the first exponent field result 2. Based on the first exponent field result 2 and the first exponent field, the first shift parameter 2 is calculated. The first shift parameter 2 indicates the shift amount required to align the first exponent field result 2 and the first exponent field to the first exponent field result 2. Therefore, the third arithmetic circuit can first perform a shift and sum operation on the first intermediate result 1 (obtained by the first arithmetic circuit 11) of the first floating-point data 1 to be operated on (including data0_1(exp) and data0_0(exp)), the first intermediate result 2 (obtained by the first arithmetic circuit 11) of the first floating-point data 2 to be operated on (including data1_0(exp) and data1_1(exp)), and the first shift parameter 1, to obtain the first non-exponential domain result 1. Then, it performs a shift and sum operation on the first intermediate result 1, the first intermediate result 2, and the first shift parameter 1 to obtain the first non-exponential domain result 1. Finally, it performs a shift and sum operation on the first non-exponential domain result 1, the first non-exponential domain, and the first shift parameter 2 to obtain the first non-exponential domain result 2. It can be understood that in this example, the output of the processing unit 10 is the first exponential domain result 2 and the first non-exponential domain result 2.
[0107] Furthermore, a second register is provided on the path connecting the first exponent arithmetic unit 12111 and the third arithmetic circuit 13, and a second register is provided on the path connecting the second exponent arithmetic unit 12112 and the third arithmetic circuit 13.
[0108] It should be noted that the second arithmetic circuit 12 may also be equipped with more exponents, so that the second arithmetic circuit 12 can perform operations on a larger number of floating-point data. The embodiments of this application will not be exhaustive here, and those skilled in the art can choose the number of exponents according to their needs.
[0109] In some embodiments of this application, the second input terminal of the third arithmetic circuit 13 can be directly connected to the first arithmetic circuit 11 and the second arithmetic circuit 12, or the second input terminal of the third arithmetic circuit can be indirectly connected to the first arithmetic circuit 11 and the second arithmetic circuit 12 through other components. The third arithmetic circuit 13 may have two second input terminals: one connected to the first arithmetic circuit 11 and the other connected to the second arithmetic circuit 12.
[0110] For example, one second input terminal of the third arithmetic circuit 13 can be connected to the output terminal of at least one multiplier 115 of the first arithmetic circuit 11 through at least one first register 116, and the other second input terminal of the third arithmetic circuit 13 can be connected to the output terminal of the exponentiation unit 121 through the second register 125.
[0111] In some embodiments of this application, the first non-exponential field may include a sign field and a mantissa field.
[0112] In some embodiments of this application, the first input terminal of the third arithmetic circuit 13 may be left floating or connected to the output terminal of the third arithmetic circuit 13 of the previous processing unit 10. It can be understood that the first non-exponential field may be the non-exponential field of the aforementioned default floating-point data (e.g., floating-point zero data) or the non-exponential field result output by the output terminal of the third arithmetic circuit 13 of the previous processing unit 10.
[0113] In some embodiments of this application, when the first non-exponential domain is the non-exponential domain result output by the output terminal of the third arithmetic circuit 13 of the previous processing unit 10, the first exponential domain can also be the exponential domain result output by the output terminal of the second arithmetic circuit 12 of the previous processing unit 10.
[0114] In some examples, the first input terminal of the third arithmetic circuit 13 of the first processing unit 10 can be left floating. In this case, the first input terminal of the third arithmetic circuit 13 of the first processing unit 10 is the non-exponential field of the default floating-point data (e.g., floating-point zero data). The first input terminal of the third arithmetic circuit 13 of other processing units 10 besides the first processing unit 10 can be connected to the output terminal of the third arithmetic circuit 13 of the previous processing unit 10. Thus, the first input terminal of the third arithmetic circuit 13 of other processing units 10 besides the first processing unit 10 is the non-exponential field result (e.g., the first non-exponential field result) output by the output terminal of the third arithmetic circuit 13 of the previous processing unit 10.
[0115] In some embodiments of this application, combined with Figure 2 ,like Figure 9 As shown, the third arithmetic circuit 13 includes an adder 131, the first input terminal of which is connected to the first arithmetic circuit 11, the second input terminal of which is connected to the second arithmetic circuit 12, and the third input terminal of which is used to receive a first non-exponential field; wherein, the adder 131 is used to perform shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field.
[0116] It can be understood that the third input terminal of the above adder 131 can be regarded as the first input terminal of the third operational circuit 13.
[0117] In some examples, combined Figure 9 The first input terminal of adder 131 includes input terminal 1311, which is connected to the first arithmetic circuit 11, for example... Figure 9 The diagram illustrates the connection between input terminal 1311 and the output terminal of the first register 116 of the first arithmetic circuit 11.
[0118] In some examples, combined Figure 9 The second input terminal of adder 131 includes input terminal 1312, which is connected to the second operational circuit 12, for example... Figure 9 The diagram illustrates the connection between input terminal 1312 and the output terminal of the second register 125 of the second operational circuit 12.
[0119] In some examples, combined Figure 9 The third input of adder 131 includes input 1313, which can be regarded as the first input of the third operation circuit 13. Input 1313 is used to receive the first non-exponential domain (e.g., acc_n(sign+mts)).
[0120] In some examples, combined Figure 9The output terminal of adder 131 may include output terminal 1314, which can be connected to other components. The output terminal 1314 of adder 131 is used to output the first non-exponential domain result.
[0121] In some examples, combined Figure 9 ,like Figure 10 As shown, the third arithmetic circuit 13 further includes a third register 132, which is connected to the output terminal 1314 of the adder 131.
[0122] Therefore, since the output of the adder can be connected to a third register, the result of the operation can be stored in the third register after the adder has completed the operation. This can prevent the result from being lost and improve the reliability of the data processed by the processing unit.
[0123] In some examples, adder 131 may first shift the first intermediate result and the first non-exponential field based on the first shift parameter, and then perform addition operation based on the shifted first intermediate result and the first non-exponential field to obtain the result of the first non-exponential field.
[0124] Thus, since the third arithmetic circuit may include an adder, the adder can accurately perform shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field. Therefore, the accurate result of the first non-exponential field can be obtained, thereby ensuring the accuracy of the operation of the third arithmetic circuit.
[0125] In some examples, the number of adders 131 mentioned above can also be at least two.
[0126] For example, combined with Figure 10 ,like Figure 11As shown, there are two adders 131: a first adder 13111 and a second adder 13112. The first input of the first adder 13111 is connected to the first arithmetic circuit 11, and its second input is connected to the second arithmetic circuit 12, for example, to the second register 1251. The first adder 131 performs shift and addition operations based on the first intermediate result 1, the first intermediate result 2, and the first shift parameter 1, outputting the first non-exponential field result 1. The second adder 13112... The first input terminal of the second adder 112 is connected to the third register 1321, and the second input terminal of the second adder 13112 is connected to the second register 1252. The third input terminal of the second adder 13112 is used to receive the first non-exponential domain (e.g., the first non-exponential domain result 1). The first shift parameter 2 can be obtained from the second register 1252, and the first non-exponential domain is obtained from the first input terminal of the third arithmetic circuit 13. Based on the first non-exponential domain result 1, the first non-exponential domain, and the first shift parameter, shift and addition operations are performed to obtain the first non-exponential domain result 2. It can be understood that in this example, the output of the processing unit 10 includes the first non-exponential domain result 2 and the first exponential domain result 2.
[0127] Furthermore, a third register 1321 is provided on the path connecting the first adder 13111 and the second adder 13112, and a third register 1322 is also connected to the output of the second adder 13112.
[0128] It should be noted that the third arithmetic circuit 13 may also be equipped with more adders so that the third arithmetic circuit 13 can perform operations on a larger number of floating-point data. The embodiments of this application will not be exhaustive here, and those skilled in the art can choose the number of adders according to their needs.
[0129] In some embodiments of this application, the output of the processing unit 10 can be a first exponential domain result and a first non-exponential domain result, that is, the first exponential domain result is directly output through one of the output terminals of the second arithmetic circuit 12, and the first non-exponential domain result is output through the output terminal of the third arithmetic circuit 13.
[0130] It should be noted that the above embodiments use any one of the processing units in at least one of the processing units included in the data processing circuit as an example to illustrate the specific structure of the processing unit.
[0131] In some embodiments of this application, the second arithmetic circuit 12 can be considered as a pipeline for operations on the exponential field of floating-point data. The pipeline of the second arithmetic circuit 12 has one stage, meaning the pipeline for operations on the exponential field of floating-point data has one stage. The first arithmetic circuit 11 and the third arithmetic circuit 13 can be considered as pipelines for operations on the non-exponential field of floating-point data. Both the first arithmetic circuit 11 and the third arithmetic circuit 13 have one stage, meaning the pipeline for operations on the non-exponential field of floating-point data has two stages.
[0132] This application provides a data processing circuit, which includes a processing unit comprising: a first arithmetic circuit for performing calculations on the non-exponential field of a first floating-point data to be calculated input from the input terminal of the first arithmetic circuit to obtain a first intermediate result; a second arithmetic circuit for performing calculations on the exponential field of the first floating-point data to be calculated input from the first input terminal of the second arithmetic circuit and the first exponential field input from the second input terminal of the second arithmetic circuit to obtain a first exponential field result and a first shift parameter, wherein the first shift parameter indicates the shift amount required to align the first intermediate result and the first non-exponential field to the first exponential field result; and a third arithmetic circuit connected to the first and second arithmetic circuits for performing shift and calculations on the first intermediate result, the first shift parameter, and the first non-exponential field input from the first input terminal of the third arithmetic circuit to obtain a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field. Since the processing unit of the data processing circuit includes a first arithmetic circuit, a second arithmetic circuit, and a third arithmetic circuit, and the first arithmetic circuit can perform operations based on the non-exponential field of the first floating-point data to be processed, and the second arithmetic circuit can perform operations based on the exponential field and the first exponential field of the first floating-point data to be processed, that is, the first arithmetic circuit and the second arithmetic circuit can perform operations based on different components of the same floating-point data, that is, the first arithmetic circuit and the second arithmetic circuit can perform operations relatively independently. In this way, by making the first arithmetic circuit and the second arithmetic circuit perform operations in parallel, the time consumed by the first arithmetic circuit and the second arithmetic circuit can be reduced, so as to quickly obtain the first intermediate result, the first exponential field result, and the first shift parameter. Thus, the third arithmetic circuit can quickly perform shift and operation based on the first intermediate result, the first shift parameter, and the first non-exponential field to obtain the first non-exponential field result. Therefore, the time consumed by the processing unit to perform operations on floating-point data can be reduced, thereby improving the computing performance of the processing unit. In this way, the computing performance of the data processing circuit can be improved.
[0133] Furthermore, compared to related technologies, the number of adders in the processing unit of the data processing circuit provided in this application embodiment can be significantly higher than that in related technologies (e.g., ...). Figure 1In the related technologies shown, each arithmetic unit has one less adder, and since the algorithm logic of the second arithmetic circuit is relatively simple, the size of the second arithmetic circuit can be smaller than the size of one adder. Therefore, the size of the processing unit of the data processing circuit provided in this application embodiment can be smaller than the size of each arithmetic unit in the related technologies, thereby improving the PPA of the processing unit.
[0134] Of course, since there may be situations where more floating-point data is processed within a single processing unit, multiple arithmetic circuits can also be provided in the processing unit. Some of these arithmetic circuits are used to perform operations on the non-exponential domain of the floating-point data, while others are used to perform operations on the exponential domain of the floating-point data. Still others are used to perform operations based on the results of the operations of these two different arithmetic circuits. An example will be given below.
[0135] In some embodiments of this application, the number of the above-mentioned processing units 10 is M, where M is a positive integer greater than 1; wherein, combined with Figure 2 ,like Figure 12 As shown, the output terminal of the second arithmetic circuit of the j-th processing unit 10 in the M processing units 10 is connected to the second input terminal of the second arithmetic circuit of the (j+1)-th processing unit 10 in the M processing units 10; the output terminal of the third arithmetic circuit of the j-th processing unit 10 is connected to the third arithmetic circuit of the (j+1)-th processing unit 10; j is a positive integer less than M.
[0136] It should be noted that, in Figure 12 The second and third operational circuits are not shown.
[0137] In some examples, the output of the third arithmetic circuit of the j-th processing unit 10 can be connected to the input of the third arithmetic circuit of the (j+1)-th processing unit 10.
[0138] Optionally, if the third arithmetic circuit includes an adder, the output terminal of the adder of the third arithmetic circuit of the j-th processing unit 10 can be connected to the input terminal of the adder of the third arithmetic circuit of the (j+1)-th processing unit 10.
[0139] It can be understood that the result output by the second arithmetic circuit of the j-th processing unit 10 (e.g., the first exponential domain result acc_(n+1)(exp)) can be the first exponential domain input to the second input terminal of the second arithmetic circuit of the (j+1)-th processing unit 10. The result output by the third arithmetic circuit of the j-th processing unit 10 (e.g., the first non-exponential domain result acc_(n+1)(sign+mts)) can be the first non-exponential domain input to the third arithmetic circuit of the (j+1)-th processing unit 10.
[0140] Similarly, the result output by the second arithmetic circuit of the (j+1)th processing unit 10 (e.g., the first exponential domain result acc_(n+2)(exp)) can be the first exponential domain input to the second input terminal of the second arithmetic circuit of the (j+2)th processing unit 10. The result output by the third arithmetic circuit of the (j+1)th processing unit 10 (e.g., the first non-exponential domain result acc_(n+2)(sign+mts)) can be the first non-exponential domain input to the third arithmetic circuit of the (j+2)th processing unit 10, and so on.
[0141] In some examples, the specific structures of the M processing units 10 may be the same or different.
[0142] It should be noted that the specific structure of each of the M processing units 10 can be described by referring to the specific description of the specific structure of the processing unit 10 in the above embodiments, and will not be repeated here in the embodiments of this application.
[0143] In some examples, the last processing unit 10 may also be connected to a normalization unit, which performs normalization processing on the first exponential field result output by the second arithmetic circuit 12 of the last processing unit 10 and the first non-exponential field result output by the third arithmetic circuit 13 of the last processing unit 10 to generate standard floating-point data.
[0144] Thus, since the number of processing units can be M, and the output of the second arithmetic circuit of the j-th processing unit is connected to the second input of the second arithmetic circuit of the (j+1)-th processing unit among the M processing units; the output of the third arithmetic circuit of the j-th processing unit is connected to the third arithmetic circuit of the (j+1)-th processing unit; in this way, the second arithmetic circuit of the j-th processing unit can output a first exponential domain result, so that the second arithmetic circuit of the (j+1)-th processing unit can use this first exponential domain result as the first exponential domain for further calculation; the output of the third arithmetic circuit of the j-th processing unit can output a first non-exponential domain result, so that the third arithmetic circuit of the (j+1)-th processing unit can use this first non-exponential domain result as the first non-exponential domain for further calculation; therefore, the data processing circuit can perform multiple accumulation operations on floating-point data through the M processing units.
[0145] Figure 13 A schematic diagram of the chip structure provided in an embodiment of this application is shown. Figure 13 As shown, the chip 20 provided in this embodiment includes the data processing circuit 21 described in the above embodiment.
[0146] In some embodiments of this application, the chip 20 described above may include, but is not limited to, at least one of the following: SOC, NPU.
[0147] This application provides a chip that includes the data processing circuit described in the above embodiments. Since the processing unit of the data processing circuit includes a first arithmetic circuit, a second arithmetic circuit, and a third arithmetic circuit, and the first arithmetic circuit can perform operations based on the non-exponential domain of the first floating-point data to be processed, and the second arithmetic circuit can perform operations based on the exponential domain and the first exponential domain of the first floating-point data to be processed, that is, the first and second arithmetic circuits can perform operations based on different components of the same floating-point data, i.e., the first and second arithmetic circuits can perform operations relatively independently. This allows the first and second arithmetic circuits to perform operations in parallel, reducing the time consumed by their operations and quickly obtaining the first intermediate result, the first exponential domain result, and the first shift parameter. Consequently, the third arithmetic circuit can quickly perform shifting and operations based on the first intermediate result, the first shift parameter, and the first non-exponential domain to obtain the first non-exponential domain result. Therefore, the time consumed by the processing unit in processing floating-point data can be reduced, thereby improving the processing unit's operational performance. This, in turn, improves the operational performance of the data processing circuit.
[0148] Figure 14 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Figure 14 As shown, the electronic device 22 provided in this application embodiment includes the data processing circuit 23 in the above embodiment.
[0149] This application provides an electronic device including the data processing circuit described in the above embodiments. Since the processing unit of the data processing circuit includes a first arithmetic circuit, a second arithmetic circuit, and a third arithmetic circuit, and the first arithmetic circuit can perform operations based on the non-exponential domain of the first floating-point data to be processed, and the second arithmetic circuit can perform operations based on the exponential domain and the first exponential domain of the first floating-point data to be processed, that is, the first and second arithmetic circuits can perform operations based on different components of the same floating-point data, i.e., the first and second arithmetic circuits can perform operations relatively independently. This allows the first and second arithmetic circuits to perform operations in parallel, reducing the time consumed by their operations and quickly obtaining the first intermediate result, the first exponential domain result, and the first shift parameter. Consequently, the third arithmetic circuit can quickly perform shifting and operations based on the first intermediate result, the first shift parameter, and the first non-exponential domain to obtain the first non-exponential domain result. Therefore, the time consumed by the processing unit in processing floating-point data can be reduced, thereby improving the processing unit's operational performance. This, in turn, improves the operational performance of the data processing circuit.
[0150] It should be noted that the floating-point arithmetic method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as examples to illustrate the floating-point arithmetic method provided in this application.
[0151] Figure 15 A flowchart illustrating the floating-point arithmetic method provided in an embodiment of this application is shown. Figure 15 As shown, the floating-point arithmetic method provided in this application embodiment may include the following steps 101 to 103.
[0152] Step 101: The electronic device performs calculations based on the non-exponential domain of the first floating-point data to be calculated through the first arithmetic circuit of the data processing unit of the electronic device's data processing circuit to obtain a first intermediate result.
[0153] It should be noted that this application embodiment does not limit the execution order of steps 101 and 102. The execution time of step 101 and step 102 at least partially overlaps, meaning that steps 101 and 102 are executed in parallel. In one example, the electronic device may execute step 101 first, followed by step 102. Figure 14 This is an illustration of the execution order; in another example, the electronic device may execute step 102 first, and then execute step 101; in yet another example, the electronic device may execute steps 101 and 102 simultaneously.
[0154] In this embodiment of the application, the non-exponential field includes the sign field and the mantissa field.
[0155] It is understandable that the first intermediate result mentioned above is an intermediate result in the non-exponential field.
[0156] In some examples, step 101 above can be specifically implemented through step 101a below.
[0157] Step 101a: The electronic device performs a multiplication operation based on the non-exponential field of the first floating-point data to be operated on through the multiplier of the first arithmetic circuit.
[0158] Thus, it can be seen that since the electronic device can perform multiplication based on the non-exponential field of the first floating-point data to be operated on through the multiplier of the first arithmetic circuit, the electronic device can obtain an accurate first intermediate result, thereby improving the operation accuracy of the electronic device.
[0159] Step 102: The electronic device performs calculations based on the exponent field and the first exponent field of the first floating-point data to be calculated through the second arithmetic circuit of the processing unit to obtain the first exponent field result and the first shift parameter.
[0160] In this embodiment of the application, the first shift parameter is used to indicate the amount of shift required to align the first intermediate result and the first non-exponential field to the first exponential field result.
[0161] It is understood that the first and second arithmetic circuits perform operations based on different components of the same floating-point data. In other words, the first arithmetic circuit does not need the result of the second arithmetic circuit to perform the operation, and the second arithmetic circuit does not need the result of the first arithmetic circuit to perform the operation. Therefore, the first and second arithmetic circuits can perform operations relatively independently. This way, the time consumed by the first and second arithmetic circuits to perform operations in parallel can be reduced.
[0162] In some embodiments of this application, step 102 can be specifically implemented by steps 102a to 102c as described below.
[0163] Step 102a: The electronic device performs exponential addition based on the exponential field of the first floating-point data to be calculated through the exponential arithmetic unit of the second arithmetic circuit to obtain an intermediate result in the exponential field.
[0164] In some examples, the aforementioned first floating-point data to be computed corresponds one-to-one with the intermediate results in the exponent field. It can be understood that when there is at least one first floating-point data to be computed, there is also at least one intermediate result in the exponent field, and each intermediate result in the exponent field is obtained by performing exponential addition on the exponent fields of the corresponding first floating-point data to be computed, which includes the exponent fields of two floating-point data.
[0165] For example, suppose there are two first floating-point data to be computed. One first floating-point data has an exponent field including data0_0 (exp) and data1_0 (exp), and the other first floating-point data has an exponent field including data0_1 (exp) and data1_1 (exp). Then, the electronic device can use the exponent arithmetic unit to calculate the intermediate result of the exponent field using the following formula:
[0166] exp_sum_0 = data0_0(exp) + data1_0(exp);
[0167] exp_sum_1 = data0_1(exp) + data1_1(exp);
[0168] Wherein, exp_sum_0 is the intermediate result of the exponent field corresponding to one of the first floating-point data to be calculated, and exp_sum_1 is the intermediate result of the exponent field corresponding to the other first floating-point data to be calculated.
[0169] Step 102b: The electronic device uses an exponent calculator to determine the result of the first exponent field as the result of the intermediate result of the exponent field and the exponent field with the largest value in the first exponent field.
[0170] For example, suppose there are two first floating-point data to be computed. One first floating-point data has an exponent field including data0_0 (exp) and data1_0 (exp), and the other first floating-point data has an exponent field including data0_1 (exp) and data1_1 (exp). Then, the electronic device can determine the result of the first exponent field using the following formula through the exponent calculator:
[0171] exp_max = max(exp_sum_0, exp_sum_1, acc_n(exp));
[0172] Where exp_max is the result of the first exponent field, and acc_n(exp) is the result of the first exponent field.
[0173] Step 102c: The electronic device calculates the first shift parameter based on the first exponent field result, the intermediate result of the exponent field, and the first exponent field through the exponent arithmetic unit.
[0174] Thus, it can be seen that since the electronic device can perform exponential addition based on the exponential field of the first floating-point data to be calculated by the exponential arithmetic unit, and determine the intermediate result of the exponential field obtained by the exponential addition operation and the exponential field with the largest value in the first exponential field as the first exponential field result, the electronic device can accurately calculate the first shift parameter based on the first exponential field result, the intermediate result of the exponential field and the first exponential field by the exponential arithmetic unit. Therefore, the calculation accuracy of the electronic device can be improved.
[0175] In some examples, the first shift parameter mentioned above includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the shift amount required to align the first non-exponential domain to the first exponential domain result. Step 102c can be specifically implemented through steps 102c1 and 102c2 described below.
[0176] Step 102c1: The electronic device uses an exponentiation arithmetic unit to determine the difference between the result of the first exponent field and the intermediate result of the exponent field as the first sub-parameter.
[0177] Optionally, the number of first sub-parameters can be at least one, and the first sub-parameters correspond one-to-one with intermediate results in the exponent field.
[0178] For example, suppose there are two first floating-point data to be computed. One first floating-point data has an exponent field including data0_0 (exp) and data1_0 (exp), and the other first floating-point data has an exponent field including data0_1 (exp) and data1_1 (exp). exp_sum_0 is the intermediate result of the exponent field corresponding to the first first floating-point data, and exp_sum_0 is the intermediate result of the exponent field corresponding to the other first floating-point data. Then, the electronic device can calculate the first sub-parameter using the following formula through the exponent calculator:
[0179] exp_sft_0 = exp_max - exp_sum_0;
[0180] exp_sft_1 = exp_max - exp_sum_1;
[0181] Where exp_sft_0 is the first sub-parameter corresponding to exp_sum_0, exp_max is the result of the first exponent field, and exp_sft_1 is the first sub-parameter corresponding to exp_sum_1.
[0182] Step 102c2: The electronic device uses an exponentiation calculator to determine the difference between the result of the first exponent field and the result of the first exponent field as the second sub-parameter.
[0183] For example, an electronic device can calculate the second sub-parameter using an exponential calculator with the following formula:
[0184] exp_sft_2 = exp_max - acc_n(exp);
[0185] Where exp_sft_2 is the second sub-parameter, exp_max is the result of the first exponent field, and acc_n(exp) is the first exponent field.
[0186] Therefore, since the result of the first exponent field can be the exponent field result of the floating-point data processed by the processing unit, the electronic device can determine the difference between the first exponent field result and the intermediate result of the exponent field as the first sub-parameter to obtain an accurate first sub-parameter, and can determine the difference between the first exponent field result and the first exponent field as the second sub-parameter to obtain an accurate second sub-parameter. Thus, in subsequent steps, the electronic device can accurately calculate the first non-exponent field result, thereby improving the computational accuracy of the electronic device.
[0187] Step 103: The electronic device performs shift and operation based on the first intermediate result, the first shift parameter, and the first non-exponential domain through the third arithmetic circuit of the processing unit to obtain the first non-exponential domain result.
[0188] In some embodiments of this application, step 103 described above can be implemented by step 103a as follows.
[0189] Step 103a: The electronic device performs shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field through the adder of the third arithmetic circuit to obtain the result of the first non-exponential field.
[0190] Thus, it can be seen that since the electronic device can accurately perform shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field through the adder, the accuracy of the first non-exponential field result calculated by the electronic device can be improved, thereby improving the computational accuracy of the electronic device.
[0191] In some examples, the first shift parameter mentioned above includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the shift amount required to align the first non-exponential domain to the first exponential domain result. Step 103a can be specifically implemented through steps 103a1 and 103a2 described below.
[0192] Step 103a1: The electronic device shifts the first intermediate result according to the first sub-parameter and the first non-exponential field according to the second sub-parameter using an adder.
[0193] For example, assuming the first sub-parameter includes exp_sft_0 and exp_sft_1, the electronic device can use an adder to shift the first intermediate result according to the first sub-parameter using the following formula:
[0194] mts0_sft = mts0 >> exp_sft_0;
[0195] mts1_sft = mts1 >> exp_sft_1;
[0196] Where mts0_sft is a first intermediate result after shifting, mts0 is the first intermediate result before shifting, and mts0 >> exp_sft_0 means shifting the decimal point of mts0 to the right by exp_sft_0 sign bits. mts1_sft is another first intermediate result after shifting, and mts1 is the other first intermediate result before shifting. mts1 >> exp_sft_1 means shifting the decimal point of mts1 to the right by exp_sft_1 sign bits.
[0197] For example, assuming the second sub-parameter is exp_sft_2, the electronic device can use an adder to shift the first non-exponential field according to the second sub-parameter using the following formula:
[0198] acc_n(sign+mts)_sft = acc_n(sign+mts) >> exp_sft_2;
[0199] Where acc_n(sign+mts)_sft is the first non-exponential field after shifting, acc_n(sign+mts) is the first non-exponential field before shifting, exp_sft_2 is the second sub-parameter, and acc_n(sign+mts) >> exp_sft_2 means shifting the decimal point of acc_n(sign+mts) to the right by exp_sft_2 sign bits.
[0200] Step 103a2: The electronic device performs an addition operation on the shifted first intermediate result and the first non-exponential field through an adder to obtain the result of the first non-exponential field.
[0201] For example, an electronic device can use an adder to perform an addition operation between the shifted first intermediate result and the first non-exponential field using the following formula to obtain the result of the first non-exponential field:
[0202] mts_rslt = mts0_sft + mts1_sft + acc_n(sign+mts)_sft;
[0203] Wherein, mts_rslt is the first non-exponential field result, mts0_sft is a first intermediate result after shifting, mts1_sft is another first intermediate result after shifting, and acc_n(sign+mts)_sft is the first non-exponential field after shifting.
[0204] As can be seen, since the floating-point data corresponding to the first intermediate result and the first exponent field may be different, direct calculation may lead to incorrect results. Therefore, electronic devices can use an adder to first shift the first intermediate result according to the first sub-parameter and then shift the first non-exponent field according to the second sub-parameter, so that the shifted first intermediate result and the first non-exponent field are aligned in the exponent field. Then, the shifted first intermediate result and the first non-exponent field are added to obtain an accurate result for the first non-exponent field. This can improve the calculation accuracy of electronic devices.
[0205] This application provides a floating-point arithmetic method. An electronic device performs arithmetic on the first arithmetic circuit of its data processing unit based on the non-exponential field of the first floating-point data to be processed, obtaining a first intermediate result. Then, through a second arithmetic circuit of the processing unit, it performs arithmetic on the exponential field and a first exponential field of the first floating-point data to be processed, obtaining a first exponential field result and a first shift parameter. The first shift parameter indicates the shift amount required to align the first intermediate result and the first non-exponential field to the first exponential field result. Finally, through a third arithmetic circuit of the processing unit, it performs shift and arithmetic on the first intermediate result, the first shift parameter, and the first non-exponential field, obtaining a first non-exponential field result. The non-exponential field includes a sign field and a mantissa field. Since the electronic device can perform operations based on the non-exponential field of the first floating-point data to be operated on by the first arithmetic circuit, and perform operations based on the exponential field and the first exponential field of the first floating-point data to be operated on by the second arithmetic circuit, that is, the electronic device can perform operations based on different components of the same floating-point data by the first arithmetic circuit and the second arithmetic circuit respectively, that is, the first arithmetic circuit and the second arithmetic circuit can perform operations relatively independently, the electronic device can reduce the time consumed by the first arithmetic circuit and the second arithmetic circuit to perform operations in parallel, so as to quickly obtain the first intermediate result, the first exponential field result and the first shift parameter. Thus, the electronic device can quickly perform shift and operation based on the first intermediate result, the first shift parameter and the first non-exponential field by the third arithmetic circuit to obtain the first non-exponential field result. Therefore, the time consumed by the processing unit of the electronic device to operate on floating-point data can be reduced, thereby improving the operation performance of the processing unit. In this way, the operation performance of the data processing circuit can be improved, thereby improving the operation performance of the electronic device.
[0206] It should be noted that the above embodiments describe the specific operation method of the electronic device when the number of exponents is one and the number of adders is one. The following will illustrate the specific operation method of the electronic device when the number of exponents is two and the number of adders is two.
[0207] When there are two exponentiation operators and two adders, the electronic device can perform exponential addition based on data0_1(exp) and data0_0(exp) using the first exponentiation operator to obtain intermediate result 1, and perform exponential addition based on data1_0(exp) and data1_1(exp) to obtain intermediate result 2. The exponent field with the largest value between intermediate result 1 and intermediate result 2 is determined as the first exponent field result 1. Based on the first exponent field result 1, intermediate result 1, and intermediate result 2, a first shift parameter 1 is determined, which indicates the shift amount required to align intermediate result 1 and intermediate result 2 to the first exponent field result 1, and is sent to the third arithmetic circuit. Furthermore, the second exponentiation operator determines the first exponent field result 1 and the exponent field with the largest value in the first exponent field as the first exponent field result 2, and calculates the first shift parameter 2 based on the first exponent field result 2 and the first exponent field. This first shift parameter 2 indicates the shift amount required to align the first exponent field result 2 and the first exponent field to the first exponent field result 2. Thus, the electronic device can perform a shift and sum operation on the first intermediate result 1 (obtained by the first arithmetic circuit 11) of the first floating-point data 1 to be operated on (including data0_1(exp) and data0_0(exp)), the first intermediate result 2 (obtained by the first arithmetic circuit 11) of the first floating-point data 2 to be operated on (including data1_0(exp) and data1_1(exp)), and the first shift parameter 1, to obtain the first non-exponential domain result 1. Then, based on the first intermediate result 1, the first intermediate result 2, and the first shift parameter 1, a shift and sum operation is performed to obtain the first non-exponential domain result 1. Finally, the second adder of the third arithmetic circuit performs a shift and sum operation on the first non-exponential domain result 1, the first non-exponential domain, and the first shift parameter 2 to obtain the first non-exponential domain result 2. It can be understood that in this example, the output of the processing unit 10 is the first exponential domain result 2 and the first non-exponential domain result 2.
[0208] As can be seen from the above, the improvements in this application may include, but are not limited to:
[0209] (1) The operation logic of the exponential arithmetic unit in the processing unit is relatively simple. Therefore, it is not necessary to keep it on the same pipeline as the operation of the non-exponential domain. If the processing of the non-exponential domain in the processing unit requires two pipelines, such as including the first operation circuit and the third operation circuit, the operation of the exponential domain only needs one pipeline to be processed and output, such as including only the second operation circuit.
[0210] (2) Since many processing units are often connected in series in the data processing circuit, as can be seen from (1), the result of the exponent field (e.g., the result of the first exponent field) can be output in advance. Since the exponent field and non-exponent field of the current processing unit both come from the previous level processing unit, the processing arithmetic unit can obtain the exponent field result of the calculation result of the previous level processing unit in advance (e.g., the result of the first exponent field output by the second operation circuit of the previous level processing unit). The exponent field result of the previous level processing unit obtained in advance can be compared with the exponent field of the current processing unit (i.e., the exponent field of the first floating-point data to be operated on and the first exponent field input by the current processing unit) to directly obtain the maximum exponent field exp_max of the current processing unit, which is also the exponent field result output by the current processing unit (i.e., the result of the first exponent field output by the second operation circuit of the current processing unit).
[0211] (3) Since the final exponent field result (e.g., the first exponent field result) has been calculated in advance in (2), the result of the two multiplication operations of the current processing unit can be directly shifted by the difference between exp_max and its own exponent field (e.g., the first sub-parameter in the above embodiment);
[0212] (4) Since the exponential domain result of the upper-level processing unit is output first, when the non-exponential domain result of the upper-level unit (e.g., the first non-exponential domain result) is output, the shift parameters required for the non-exponential domain result output by the upper-level unit can be calculated in advance, which can reduce the delay of combinational logic.
[0213] (5) The non-exponential domain output of the current arithmetic unit (e.g., the first non-exponential domain result) can be obtained by adding the two shifted results of the current processing unit in (3) and the shifted result of the upper unit in (4).
[0214] Therefore, the advantages of this application over related technologies are: 1. The maximum combinational logic delay in the processing unit of this application is the multiplication delay of the non-exponential field of floating-point data, which is much smaller than the maximum combinational logic delay in related technologies. 2. This application has one less adder than related technologies, reducing area and power consumption.
[0215] In summary, this application reduces the number of pipeline stages for exponent field operations by separating the pipeline for exponent field operations from the pipeline for non-exponential field operations of floating-point data. Within a processing unit, the exponent field result output by the upper-level arithmetic unit arrives before the non-exponential field result, allowing the current processing unit to perform all exponent field operations of floating-point data in advance. This simplifies non-exponential field operations, reducing the number of adders required for non-exponential field operations. Furthermore, since the shift parameters required for non-exponential field operations are pre-generated and stored in registers, the combinational logic latency of the non-exponential field is reduced. Additionally, the combinational logic latency of both pipeline stages for non-exponential field operations in this application is low, allowing the processing unit to operate at higher frequencies.
[0216] It should be noted that each of the above method embodiments, or various possible implementations of each method embodiment, can be executed individually or in combination of any two or more. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.
[0217] The floating-point arithmetic method provided in this application can be executed by a floating-point arithmetic device. This application uses the example of a floating-point arithmetic device executing the floating-point arithmetic method to illustrate the floating-point arithmetic device provided in this application.
[0218] Figure 16 This is a schematic diagram of the structure of the floating-point arithmetic device provided in an embodiment of this application. Figure 16 As shown, the floating-point arithmetic device 300 may include a processing module 301. The processing module 301 is configured to perform calculations based on the non-exponential field of the first floating-point data to be calculated, obtaining a first intermediate result; and to perform calculations based on the exponent field and the first exponent field of the first floating-point data to be calculated, obtaining a first exponent field result and a first shift parameter, the first shift parameter indicating the shift amount required to align the first intermediate result and the first non-exponential field to the first exponent field result; and to perform shift and calculations based on the first intermediate result, the first shift parameter, and the first non-exponential field, obtaining a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field.
[0219] This application provides a floating-point arithmetic device. Since the floating-point arithmetic device can perform operations based on the non-exponential field of the first floating-point data to be operated on by a first arithmetic circuit, and perform operations based on the exponential field and the first exponential field of the first floating-point data by a second arithmetic circuit, meaning the floating-point arithmetic device can perform operations based on different components of the same floating-point data by the first and second arithmetic circuits respectively, i.e., the first and second arithmetic circuits can perform operations relatively independently, the floating-point arithmetic device can reduce the time consumed by the first and second arithmetic circuits to perform operations in parallel, thereby quickly obtaining the first intermediate result, the first exponential field result, and the first shift parameter. Therefore, the floating-point arithmetic device can quickly perform shifting and operations based on the first intermediate result, the first shift parameter, and the first non-exponential field by a third arithmetic circuit to obtain the first non-exponential field result. Therefore, the time consumed by the processing unit of the floating-point arithmetic device to operate on floating-point data can be reduced, thereby improving the operation performance of the processing unit. This, in turn, improves the operation performance of the data processing circuit, thus improving the operation performance of the floating-point arithmetic device.
[0220] In some embodiments of this application, the processing module 301 described above is specifically used to perform multiplication operations based on the non-exponential field of the first floating-point data to be computed.
[0221] In some embodiments of this application, the processing module 301 is specifically used to perform exponential addition based on the exponential field of the first floating-point data to be computed, to obtain an intermediate result in the exponential field; and to determine the intermediate result in the exponential field and the exponential field with the largest value in the first exponential field as the first exponential field result; and to calculate the first shift parameter based on the first exponential field result, the intermediate result in the exponential field and the first exponential field.
[0222] In some embodiments of this application, the first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the shift amount required to align the first non-exponential domain to the first exponential domain result. Specifically, the processing module 301 is used to determine the difference between the first exponential domain result and the exponential domain intermediate result as the first sub-parameter, and to determine the difference between the first exponential domain result and the first exponential domain as the second sub-parameter.
[0223] In some embodiments of this application, the processing module 301 described above is specifically used to perform shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field.
[0224] In some embodiments of this application, the first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the shift amount required to align the first non-exponential domain to the first exponential domain result. Specifically, the processing module 301 is used to shift the first intermediate result according to the first sub-parameter and shift the first non-exponential domain according to the second sub-parameter; and then perform an addition operation between the shifted first intermediate result and the first non-exponential domain.
[0225] The floating-point arithmetic device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0226] The floating-point arithmetic device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0227] The floating-point arithmetic device provided in this application embodiment can achieve... Figures 2 to 12 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.
[0228] In some embodiments of this application, such as Figure 17As shown, this application embodiment also provides an electronic device 400, including a processor 401 and a memory 402. The memory 402 stores a program or instructions that can run on the processor 401. When the program or instructions are executed by the processor 401, they implement the various steps of the above-described floating-point arithmetic method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0229] It should be noted that the electronic devices in the embodiments of this application include the aforementioned mobile electronic devices and non-mobile electronic devices.
[0230] Figure 18 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.
[0231] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 501, network module 502, audio output unit 503, input unit 504, sensor 505, display unit 506, user input unit 507, interface unit 508, memory 509, and processor 510.
[0232] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 18 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0233] The processor 510 is configured to perform calculations based on the non-exponential field of the first floating-point data to be calculated, to obtain a first intermediate result; and to perform calculations based on the exponential field and the first exponential field of the first floating-point data to be calculated, to obtain a first exponential field result and a first shift parameter, wherein the first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential field to the first exponential field result; and to perform shift and calculations based on the first intermediate result, the first shift parameter and the first non-exponential field, to obtain a first non-exponential field result; wherein the non-exponential field includes a sign field and a mantissa field.
[0234] This application provides an electronic device. Because the electronic device can perform calculations based on the non-exponential domain of the first floating-point data to be calculated using a first arithmetic circuit, and perform calculations based on the exponential domain and the first exponential domain of the first floating-point data using a second arithmetic circuit—that is, the electronic device can perform calculations based on different components of the same floating-point data using the first and second arithmetic circuits respectively—the first and second arithmetic circuits can perform calculations relatively independently. This allows the electronic device to reduce the time consumed by the first and second arithmetic circuits in performing calculations in parallel, quickly obtaining the first intermediate result, the first exponential domain result, and the first shift parameter. Subsequently, the electronic device can quickly perform shifting and calculations based on the first intermediate result, the first shift parameter, and the first non-exponential domain using a third arithmetic circuit to obtain the first non-exponential domain result. Therefore, the time consumed by the processing unit of the electronic device in performing calculations on floating-point data can be reduced, thereby improving the computing performance of the processing unit. This, in turn, improves the computing performance of the data processing circuit, thus enhancing the overall computing performance of the electronic device.
[0235] In some embodiments of this application, the processor 510 is specifically used to perform multiplication operations based on the non-exponential field of the first floating-point data to be computed.
[0236] In some embodiments of this application, the processor 510 is specifically used to perform exponential addition based on the exponential field of the first floating-point data to be computed, to obtain an intermediate result in the exponential field; and to determine the intermediate result in the exponential field and the exponential field with the largest value in the first exponential field as the first exponential field result; and to calculate the first shift parameter based on the first exponential field result, the intermediate result in the exponential field and the first exponential field.
[0237] In some embodiments of this application, the first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the amount of shift required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the amount of shift required to align the first non-exponential domain to the first exponential domain result.
[0238] The processor 510 is specifically used to determine the difference between the first exponent field result and the intermediate result of the exponent field as the first sub-parameter; and to determine the difference between the first exponent field result and the first exponent field as the second sub-parameter.
[0239] In some embodiments of this application, the processor 510 is specifically used to perform shift and addition operations based on a first intermediate result, a first shift parameter, and a first non-exponential field.
[0240] In some embodiments of this application, the first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter indicates the amount of shift required to align the first intermediate result to the first exponential domain result, and the second sub-parameter indicates the amount of shift required to align the first non-exponential domain to the first exponential domain result.
[0241] The processor 510 is specifically used to shift the first intermediate result according to the first sub-parameter, and shift the first non-exponential field according to the second sub-parameter; and to perform an addition operation on the shifted first intermediate result and the first non-exponential field.
[0242] It should be understood that, in this embodiment, the input unit 504 may include a graphics processing unit (GPU) 5041 and a microphone 5042. The GPU 5041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 506 may include a display panel 5061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 507 includes at least one of a touch panel 5071 and other input devices 5072. The touch panel 5071 is also called a touch screen. The touch panel 5071 may include a touch detection device and a touch controller. Other input devices 5072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0243] The memory 509 can be used to store software programs and various data. The memory 509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0244] Processor 510 may include one or more processing units; optionally, processor 510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 510.
[0245] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described floating-point arithmetic method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0246] The processor mentioned above is the processor in the electronic device described in the above embodiments. The readable storage medium mentioned above includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0247] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described floating-point arithmetic method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0248] It should be understood that the chip mentioned in the embodiments of this application can be an ASIC chip, an NPU chip, or a system-on-a-chip, or a system chip, chip system, or system-on-a-chip, etc.; this application does not specifically limit the type of chip.
[0249] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described floating-point arithmetic method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0250] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0251] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0252] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A data processing circuit, characterized in that, The data processing circuit includes a processing unit, which includes: The first arithmetic circuit is used to perform calculations on the non-exponential domain of the first floating-point data to be calculated input from the input terminal of the first arithmetic circuit to obtain a first intermediate result; The second arithmetic circuit is used to perform calculations based on the exponent field of the first floating-point data to be calculated input at the first input terminal of the second arithmetic circuit and the first exponent field input at the second input terminal of the second arithmetic circuit, to obtain a first exponent field result and a first shift parameter. The first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential field to the first exponent field result. The third arithmetic circuit, connected to the first arithmetic circuit and the second arithmetic circuit, is used to perform shift and operation based on the first intermediate result, the first shift parameter and the first non-exponential domain input at the first input terminal of the third arithmetic circuit, to obtain the first non-exponential domain result; The non-exponential field includes a sign field and a mantissa field.
2. The circuit according to claim 1, characterized in that, The first operational circuit includes: A multiplier, wherein the input of the multiplier is used to receive the non-exponential field of the first floating-point data to be operated on, and the output of the multiplier is connected to the third operation circuit; The multiplier is used to perform multiplication operations based on the non-exponential field of the first floating-point data to be operated on.
3. The circuit according to claim 1, characterized in that, The second operational circuit includes: An exponentiation arithmetic unit, wherein the first input terminal of the exponentiation arithmetic unit is used to receive the exponent field of the first floating-point data to be calculated, the second input terminal of the exponentiation arithmetic unit is used to receive the first exponent field, and the output terminal of the exponentiation arithmetic unit is connected to the third arithmetic circuit; The exponent arithmetic unit is used to perform exponent addition on the exponent field of the first floating-point data to be calculated, to obtain an intermediate result in the exponent field, and to determine the intermediate result in the exponent field and the exponent field with the largest value in the first exponent field as the first exponent field result. The first shift parameter is calculated based on the first exponent field result, the intermediate result in the exponent field and the first exponent field.
4. The circuit according to claim 1, characterized in that, The third operational circuit includes: An adder, wherein the first input terminal of the adder is connected to the first arithmetic circuit, the second input terminal of the adder is connected to the second arithmetic circuit, and the third input terminal of the adder is used to receive the first non-exponential domain; The adder is used to perform shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field.
5. The circuit according to any one of claims 1 to 4, characterized in that, The number of processing units is M, where M is a positive integer greater than 1; Wherein, the output terminal of the second arithmetic circuit of the j-th processing unit among the M processing units is connected to the second input terminal of the second arithmetic circuit of the (j+1)-th processing unit among the M processing units; The output terminal of the third arithmetic circuit of the j-th processing unit is connected to the third arithmetic circuit of the (j+1)-th processing unit; j is a positive integer less than M.
6. An electronic device, characterized in that, Includes the data processing circuit as described in any one of claims 1 to 5.
7. A floating-point arithmetic method, characterized in that, Applied to the electronic device of claim 6, the method includes: The first arithmetic circuit of the processing unit of the data processing circuit of the electronic device performs calculations based on the non-exponential domain of the first floating-point data to be calculated, and obtains a first intermediate result. The second arithmetic circuit of the processing unit performs calculations based on the exponent field and the first exponent field of the first floating-point data to be calculated, and obtains the first exponent field result and the first shift parameter. The first shift parameter is used to indicate the amount of shift required to align the first intermediate result and the first non-exponential field to the first exponent field result. The third arithmetic circuit of the processing unit performs shift and operation based on the first intermediate result, the first shift parameter, and the first non-exponential field to obtain the first non-exponential field result. The non-exponential field includes a sign field and a mantissa field.
8. The method according to claim 7, characterized in that, The second arithmetic circuit of the processing unit performs calculations based on the exponent field and the first exponent field of the first floating-point data to be calculated, to obtain the first exponent field result and the first shift parameter, including: The exponent arithmetic unit of the second arithmetic circuit performs exponent addition based on the exponent field of the first floating-point data to be processed, and obtains the intermediate result of the exponent field. The exponent arithmetic unit is used to determine the intermediate result of the exponent field and the exponent field with the largest value in the first exponent field as the result of the first exponent field. The first shift parameter is calculated using the exponent arithmetic unit based on the first exponent field result, the intermediate exponent field result, and the first exponent field.
9. The method according to claim 8, characterized in that, The first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter is used to indicate the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter is used to indicate the shift amount required to align the first non-exponential domain to the first exponential domain result. The calculation of the first shift parameter based on the first exponent domain result, the intermediate result of the exponent domain, and the first exponent domain includes: The difference between the first exponent domain result and the intermediate result of the exponent domain is determined as the first sub-parameter; The difference between the first exponent field result and the first exponent field result is determined as the second sub-parameter.
10. The method according to claim 7, characterized in that, The third arithmetic circuit of the processing unit performs shift and operation based on the first intermediate result, the first shift parameter, and the first non-exponential field, including: The adder of the third arithmetic circuit performs shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field.
11. The method according to claim 10, characterized in that, The first shift parameter includes a first sub-parameter and a second sub-parameter. The first sub-parameter is used to indicate the shift amount required to align the first intermediate result to the first exponential domain result, and the second sub-parameter is used to indicate the shift amount required to align the first non-exponential domain to the first exponential domain result. The shift and addition operations based on the first intermediate result, the first shift parameter, and the first non-exponential field include: The first intermediate result is shifted according to the first sub-parameter, and the first non-exponential field is shifted according to the second sub-parameter; The shifted first intermediate result and the first non-exponential field are added together.
12. A floating-point arithmetic device, characterized in that, include: The processing module is used to perform calculations based on the non-exponential field of the first floating-point data to be calculated, and obtain a first intermediate result; The calculation is performed based on the exponent field and the first exponent field of the first floating-point data to be calculated, to obtain the first exponent field result and the first shift parameter. The first shift parameter is used to indicate the shift amount required to align the first intermediate result and the first non-exponential field to the first exponent field result. The shift and sum operation is performed based on the first intermediate result, the first shift parameter and the first non-exponential field to obtain the first non-exponential field result. The non-exponential field includes a sign field and a mantissa field.
13. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 7 to 11.
14. A computer-readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 7 to 11.
15. A chip, characterized in that, The chip includes a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method as described in any one of claims 7 to 11.