Division calculation method and device based on hardware implementation
Through iterative calculation and lookup table processing methods, the problem of low calculation efficiency of hardware dividers is solved, and efficient division calculation is realized without the need for fixed-point to floating point conversion.
Patent Information
- Application Number
- CN202510488200.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
The existing hardware dividers have insufficient computational efficiency, especially when dealing with division between fixed-point numbers and fixed-point numbers, the accuracy cannot be effectively controlled, resulting in low computational efficiency.
By obtaining the dividend and divisor, as well as the preset output data bit width, iterative calculation and shift operations, and finally substituting the result into the lookup table for processing, obtaining the final division result.
This method improves the calculation efficiency of the hardware divider, can obtain division results in accordance with the expected data representation format within the expected error range, and avoids the conversion process from fixed-point to floating point.
Smart Images

Figure CN120010813A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of division calculation based on hardware implementation, and in particular to a division calculation method and device based on hardware implementation. Background Art
[0002] With the development of the fifth generation mobile communication technology (5G), the amount of data that user equipment (UE) needs to process is increasing. However, UE cannot use digital signal processing (DSP) or central processing unit (CPU) to achieve most of the data processing like base stations to ensure higher throughput and faster processing speed for user equipment, so UE needs to rely on hardware to achieve it.
[0003] In the related art, the hardware divider used in the UE is for division between fixed-point numbers. Although floating-point numbers can be represented by shifting fixed-point numbers, the precision cannot be controlled, and only fixed-point division of numbers of the same type is implemented. For example, when calculating the binary number b'1001010 divided by the binary number b'1000, the calculation is subtracted from left to right. The left four bits of the binary number b'1001010 are subtracted from the binary number b'1000 to obtain 1. At this time, 1 is set on the result bit. Since 1 is not enough for four bits, the calculation cannot continue. At this time, two bits are taken to the right in turn, and two bits of 0 are set on the result bit to obtain a four-bit binary number b'1010. Then, a subtraction operation is performed on the binary number b'1000. Finally, the division result is a binary number b'1001 with a remainder of 10. The calculation process is as follows:
[0004] In this example, the expression of converting a binary number to a decimal number is: 74 divided by 8 is 9 with a remainder of 2. It can be seen that the existing divider can only be used for fixed-point integers, and it is also necessary to re-determine a fixed-point to floating-point conversion algorithm based on the data representation range after the division is completed, so as to map and express the current division result (quotient and remainder) in a reasonable way, resulting in low calculation efficiency. Summary of the invention
[0005] The present application provides a division calculation method and device based on hardware implementation to solve the problem of low calculation efficiency of hardware dividers in related technologies.
[0006] In a first aspect, the present application provides a division calculation method based on hardware implementation, comprising: obtaining the dividend 1 and the divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N; setting the initial value of nom to 1, the initial value of denom to x, and performing the following iterative calculation: when performing the nth iterative calculation, judging the high N / 2 of denom obtained by the previous iterative calculation n Check whether there is a valid value; if there is no valid value, shift denom and nom calculated in the previous iteration to the left by N / 2 respectively. n bits, get new denom and nom; until the iteration calculation reaches N / 2 n =1, get the final denom and nom; shift the final denom right by (NQ) bits to get denom>>(NQ), substitute denom>>(NQ) into the preset lookup table to get the lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to get the multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits to get the final calculation result and output it.
[0007] Optionally, when performing the nth iterative calculation, it is determined whether the high N / 2n bits of denom calculated in the previous iterative calculation still have valid values; if valid values exist, no update is performed, and denom and nom calculated in the previous iterative calculation are kept unchanged, and the next iterative calculation is performed.
[0008] Optionally, determine the high N / 2 of denom calculated in the previous iteration n Whether the bit still has a valid value, including: in the n-1th iteration operation, the target data is bitwise ANDed with the denom calculated in the previous iteration; wherein the bit width of the target data is N, and the high N / 2 of the target data n bit is 1, and the remaining bits are 0; when the result of the bitwise AND operation is 0, determine the high N / 2 of denom calculated in the previous iteration n There is no valid value for the bit; if the result of the bitwise AND operation is not 0, determine the high N / 2 of denom calculated in the previous iteration n The bit has a valid value.
[0009] Optionally, the lookup table stores the following corresponding relationship: ,in, The value of is 0, 1, ... , N is the bit width of the divisor x, express The value of is rounded down.
[0010] Optionally, the method includes: when the divisor x is 0, no iterative calculation is performed and the calculation result 2 is directly output Q+1 -1.
[0011] In a second aspect, the present application provides a division calculation device based on hardware implementation, including: an acquisition module, used to obtain the dividend 1 and the divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N; a first processing module, used to set the initial value of nom to 1, the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, determine the high N / 2 of denom obtained by the previous iterative calculation n The second processing module is used to shift denom and nom calculated in the previous iteration to the left by N / 2 respectively if there is no valid value. n bits, and obtain new denom and nom; the third processing module is used until the iteration calculation reaches N / 2 n =1, to obtain the final denom and nom; the fourth processing module is used to shift the final denom right by (NQ) bits to obtain denom>>(NQ), substitute denom>>(NQ) into the preset lookup table to obtain the lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to obtain the multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits to obtain the final calculation result and output it.
[0012] Optionally, the second processing module includes: a first processing unit, configured to perform a bitwise AND operation on the target data and the denom calculated in the previous iteration in the nth iteration operation; wherein the bit width of the target data is N, and the high N / 2 of the target data n The bit is 1, and the remaining bits are 0; the second processing unit is used to determine the high N / 2 of the denom calculated in the previous iteration when the result of the bitwise AND operation is 0 n The bit does not have a valid value; the third processing unit is used to determine the high N / 2 of the denom calculated in the previous iteration when the result of the bitwise AND operation is not 0 n The bit has a valid value.
[0013] In a third aspect, the present application provides a hardware device, including the hardware-based division calculation device described in the second aspect.
[0014] In a fourth aspect, the present application also provides an electronic device comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to execute the hardware-based division calculation method described in the first aspect of the present application.
[0015] In a fifth aspect, the present application further provides a computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute the hardware-based division calculation method described in the first aspect of the present application.
[0016] The above-mentioned technical scheme provided by the embodiment of the present application has the following advantages over the prior art: the method provided by the embodiment of the present application, when performing a division operation between the divisor 1 and the dividend x, first determines the bit width N of the divisor x and the preset output data bit width Q, and then performs iterative operations on the divisor and the dividend based on the N and Q to shift them, and then substitutes the result denom>>(NQ) obtained by right shifting the divisor by (NQ) bits into the preset lookup table LUT (denom>>(NQ)), and multiplies the lookup result LUT (denom>>(NQ)) by the final shifted nom to obtain the multiplication result LUT_div; finally, the multiplication result LUT_div is right shifted by 2×(NQ) bits to obtain the final calculation result. The error between the calculation result and the actual division operation directly performed by the divisor and the dividend is within the expected range, that is, the above method can obtain the division result after the floating point conversion that finally conforms to the expected data representation format, and the process of obtaining the division calculation result by the above method in the embodiment of the present application does not require the conversion from fixed point to floating point, and all data is processed according to the hardware format, thereby improving the calculation efficiency of the hardware divider. BRIEF DESCRIPTION OF THE DRAWINGS The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0017] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0018] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0019] Figure 1 A method flow chart of a division calculation method based on hardware implementation provided in an embodiment of the present application; Figure 2 A schematic diagram of hardware relationship expression for implementing pseudo code provided in an embodiment of the present application; Figure 3 A schematic diagram of the structure of a division calculation device based on hardware implementation provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0021] The disclosure below provides many different embodiments or examples to implement different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.
[0022] Considering a division , which can be converted to: , and the multiplication operation can be simply implemented by the AND gate, so the division operation can be converted to The value of Therefore, in this solution, the dividend is set to 1, and through the division calculation of this solution, we get , and then and The final result is obtained by performing AND gate operation. =0, the calculation result is directly output as 0; =0, the maximum value supported by the current output bit width is output. For other cases, division calculation is performed , the acceptable division error is , then b can be transformed based on p, and we can get , where p is the most significant p bits corresponding to p in b, and l is the number of bits remaining after removing the most significant p bits from b itself. For example, if N=16 is the data bit width, and p=12 is the left (most significant) 12 bits of the corresponding 16-bit data, then l=16-12=4 bits. Further, the above relationship is demonstrated: Transforming the division formula yields:
[0023] Observe the , and Taylor expansion is performed to obtain:
[0024] Therefore, the algorithm based on the shifting of l and p of b can obtain an error of .
[0025] In order to solve the problem of low computational efficiency of hardware dividers in related technologies, the present application provides a division calculation method based on hardware implementation, such as Figure 1 As shown, the steps of the method include: Step 101, obtaining a dividend 1 and a divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, and the output data bit width Q is determined based on an error range of the output data, and Q is less than or equal to N; In this solution, if the current division calculation to be performed is 2÷5, the dividend actually performed in the hardware is 1, and the divisor x is 5; then the obtained divisor 1 and divisor x=5 are divided, and the result of the division calculation is ANDed with 2, then the calculation result of the division calculation 2÷5 can be obtained. Furthermore, in the embodiment of the present application, the bit width of the divisor x is equal to the bit width supported by the hardware currently performing the division calculation, that is, the bit width supported by the hardware currently performing the division calculation is also N, but the bit width corresponding to the divider in different hardware is different, that is, the value of N is determined based on the hardware divider, such as N=16, or N=32, etc. In addition, in the embodiment of the present application, the output data bit width Q is determined based on the error range of the output data. The larger the output data bit width Q, the smaller the error range of the output data, that is, the higher the accuracy, the better the algorithm performance, but the greater the hardware area and power consumption loss; the smaller the output data bit width Q, the larger the error range of the output data, that is, the lower the accuracy, the worse the algorithm performance, but the smaller the hardware area and power consumption loss. Of course, the output data bit width Q needs to be less than or equal to the hardware bit width supported by the hardware. For example, when the hardware bit width supported by the hardware is 22 bits, the output data bit width Q cannot exceed 22 bits, such as the value of Q can be 12, 18, etc.
[0026] It should be noted that when the divisor x is 0, no iterative calculation is performed and the calculation result 2 is directly output. Q+1 -1. In addition, the letter x in the embodiments of the present application represents the data to be calculated in the present application. Of course, other letters can also be used to represent the data to be calculated, which does not constitute a limitation on the present application. The letter Q here represents the output data bit width in the present application. Of course, other letters can also be used to represent the output data bit width, which does not constitute a limitation on the present application.
[0027] Step 102, let the initial value of nom be 1, the initial value of denom be x, and perform the following iterative calculation: when performing the nth iterative calculation, determine the high N / 2 of denom obtained by the previous iterative calculation n Whether the bit has a valid value; Step 103: If there is no valid value, denom and nom calculated in the previous iteration are shifted left by N / 2. n bits, get new denom and nom; In addition, when judging the high N / 2 of denom calculated in the previous iteration n If there is a valid value, no update is done, and the denom and nom calculated in the previous iteration remain unchanged, and the next iteration is performed. For example, the current iteration is the third iteration, and the denom calculated in the second iteration is determined to be N / 2 higher.3 If there are valid values, no update is done, denom and nom obtained from the second iteration remain unchanged, and the fourth iteration is performed.
[0028] Step 104, until the iteration calculation reaches N / 2 n =1, and get the final denom and nom.
[0029] In this case, the initial value of denom is x=5 and N=16. n =1, the final denom and nom will be obtained, and the iterative calculation ends. Therefore, when n=4, 16 / 2 is satisfied 4 =1, that is, when N=16, the number of iterative calculations n=4. It should be noted that in the process of multiple iterative calculations, if the last few iterative calculations determine that the denom obtained in the previous iterative calculation is higher than N / 2 n If all the bits have valid values, the denom and nom calculated in the last iteration will not be shifted in the last few iterations. Instead, the last shifted denom and nom need to be found for subsequent calculations. For example, in the last three iterations, the high N / 2 of the denom calculated in the last iteration is determined. n If all the bits have valid values, you need to find the shifted denom and nom after the fourth-to-last iteration as the final denom and nom.
[0030] Therefore, for the above steps 102 to 104, if the current iteration is n=2, it is necessary to determine the high 16 / 2 of denom obtained by the first iteration. 2 =Whether there is a valid value in the 4 bits. If there is no valid value in the upper 4 bits of denom calculated in the first iteration, denom and nom calculated in the first iteration need to be shifted left by 16 / 2 respectively. 2 =4 bits, get new denom and nom; if the high 4 bits of denom calculated in the first iteration have valid values, no update is done, keep the denom and nom calculated in the previous iteration unchanged, and perform the next iteration. It should be noted that the above is an example of n=2, that is, the second iteration calculation is taken as an example. The iterative calculation process after the second iteration is similar to the second iteration calculation process, but for the first iteration calculation, since there is no denom and nom calculated in the previous iteration, the high N / 2 of denom calculated in the previous iteration is determined when performing the first iteration calculation. n Whether there is a valid value in the 8 bits means: judging whether there is a valid value in the upper 8 bits of x=5.
[0031] Furthermore, in this solution, for the high N / 2 of denom calculated in the previous iteration involved in the above step 102, n The way to determine whether the bit still has a valid value can be further: In the nth iteration, the target data is bitwise ANDed with the denom calculated in the previous iteration; the bit width of the target data is N, and the height of the target data is N / 2. n bit is 1, and the remaining bits are 0; When the result of the bitwise AND operation is 0, determine the high N / 2 of denom calculated in the previous iteration n There is no valid value for the bit; If the result of the bitwise AND operation is not 0, determine the high N / 2 of denom calculated in the previous iteration n The bit has a valid value.
[0032] It can be seen that in this scheme, the high N / 2 of denom calculated in the previous iteration is further judged by bitwise AND. n For this, let's take the initial value of denom x=5 and N=16 as an example. If the current iteration is n=2, the high 16 / 2 of the target data is 2 = 4 bits are 1, and the remaining bits are 0, then the corresponding binary representation of the target data is 11110000000000000, that is, 1111000000000000 is bitwise ANDed with the denom calculated in the previous iteration, and the high 16 / 2 of the denom calculated in the previous iteration is determined based on whether the bitwise AND result is 0. 2 =4 digits whether there is a valid value.
[0033] It should be noted that in this scheme, the way to perform bitwise AND operation on the target data and the denom calculated in the previous iteration is: AND the binary corresponding to the target data and the denom calculated in the previous iteration, that is, when the corresponding binary digits are both 1, the binary digit after AND is 1, and in other cases, the binary digits after AND are all 0.
[0034] The following will fully explain the iterative calculation process in this scheme, taking the initial value of denom, x=5, and N=16 as an example, assuming that the initial value of nom is 1.
[0035] Based on this, the entire iterative calculation process is: First iteration: Since the value of N is 16, the bit width of the target data is 16 bits, and this is the first iteration, the value of n is 1. Based on this, the high 16 / 2 of the binary representation of the target data 1=8 bits are 1, and the remaining bits are 0, that is, 11111111000000000. The binary representation of the initial value of denom x=5 is 101. The two are bitwise ANDed, that is, 1111111100000000&000000000000101==0. It can be seen that the result of the bitwise AND is 0, from which it can be judged that there is no valid value in the upper 8 bits of the binary representation of x=5. Therefore, denom=5 and nom=1 need to be shifted, that is, left shifted 8 bits. Among them, the binary of 1 is still 1, and the binary representation of 5 is 101. The binary representation of nom after left shifting 8 bits is 100000000, and the corresponding decimal representation is 256; the binary representation of denom after left shifting 8 bits is 10100000000, and the corresponding decimal representation is 1280. In this regard, the specific calculation process in hardware can be: if (0xFF00&denom) == 0 is established, then shift left 8 bits, nom after shifting = 256, and denom after shifting = 1280. Among them, 0xFF00 is the hexadecimal representation of the target data 1111111100000000.
[0036] Second iteration calculation: In the second iteration calculation, the value of n is 2, so the high 16 / 2 of the target data 2 =4 bits are 1, and the remaining bits are 0, that is, 1111000000000000. Based on this, the target data is bitwise ANDed with denom=1280 calculated in the first iteration: 1111000000000000&0000010100000000==0. It can be seen that the result of bitwise ANDing is 0, so we can judge the high 16 / 2 of the binary representation of denom=1280. 2 = There is no valid value in 4 bits, so denom=1280 and nom=256 need to be shifted, that is, shifted left by 4 bits. The binary representation of nom after 4 bits is 1000000000000, and the corresponding decimal representation is 4096; the binary representation of denom after 4 bits is 101000000000000, and the corresponding decimal representation is 20480. For this, the specific calculation process in hardware can be: if (0xF000&denom) == 0 is established, then shift left by 4 bits, nom after shifting = 4096, and denom after shifting = 20480. Among them, 0xF000 is the hexadecimal representation of the target data 11110000000000000.
[0037] The third iteration calculation: In the third iteration calculation, the value of n is 3, so the high 16 / 2 of the target data 3=2 bits are 1, and the remaining bits are 0, which is 1100000000000000. Based on this, the target data is bitwise ANDed with denom=20480 calculated by the second iteration: 1100000000000000&010100000000000==1. It can be seen that the result of the bitwise AND is 1, so we can judge the high 16 / 2 of the binary representation of denom=20480 3 = There is a valid value in the 2nd bit, so denom=20480 and nom=4096 are not shifted, and then the denom and nom calculated in this iteration are kept unchanged and the next iteration is performed. For this, the specific calculation process in hardware can be: if (0xC000&denom) == 0 is not true, then denom and nom after the second shift are not shifted. Among them, 0xC000 is the hexadecimal representation of the target data 1100000000000000.
[0038] 4th iteration calculation: In the 4th iteration calculation, the value of n is 4, so the high 16 / 2 of the target data 4 =1 bit is 1, and the remaining bits are 0, that is, 10000000000000000. Based on this, the target data is bitwise ANDed with denom=20480 calculated by the second iteration: 1000000000000000&0101000000000000==0. It can be seen that the result of the bitwise AND is 0, so we can judge the high 16 / 2 of the binary representation of denom=20480. 4 =1 bit does not contain a valid value, so denom=20480 and nom=4096 need to be shifted, that is, left shifted by 1 bit. The binary representation of nom after left shifting by 1 bit is 10000000000000, and the corresponding decimal representation is 8192; the binary representation of denom after left shifting by 1 bit is 1010000000000000, and the corresponding decimal representation is 40960. For this, the specific calculation process in hardware can be: if (0x8000&denom) == 0 is established, then left shift by 1 bit, nom after shifting = 8192, and denom after shifting = 40960. Among them, 0x8000 is the hexadecimal representation of the target data 1000000000000000.
[0039] It can be seen that after the above four iterative calculations, the final dividend nom=8192 and the divisor denom=40960.
[0040] Step 105, shift the final denom right by (NQ) bits to obtain denom>>(NQ), substitute denom>>(NQ) into the preset lookup table to obtain the lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to obtain the multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits to obtain the final calculation result and output it.
[0041] In this regard, in this solution, taking N as 16 and Q as 12 as an example, the final denom is shifted right by 16-12=4 bits, and then the result of the final denom shifted right by 4 bits is substituted into the preset lookup table. For example, the decimal representation of the current final denom is 40960, and the corresponding binary representation is 1010000000000000. After right shifting by 4 bits, the corresponding binary is 1010000000000, and the decimal representation after right shifting is 2560. That is, 2560 is substituted into the lookup table to obtain the search result LUT(denom>>(NQ))= LUT(2560), and the search result LUT(denom>>(NQ)) is multiplied by the shifted nom to obtain the multiplication result LUT_div.
[0042] Furthermore, the lookup table in this solution stores the following corresponding relationship: ,in, The value of is 0, 1, ... , N is the bit width of the divisor x, express The value of is rounded down.
[0043] For this, let's take N as 16, Q as 12, the initial value of the dividend nom as 1, and the initial value of the divisor denom as x=5 as an example. Since N is 16 and Q is 12, the data in the lookup table LUT is:
[0044] It can be seen that the size of the lookup table in this solution is affected by Q. When Q is 12, it can be seen that the maximum number of the lookup table is 4095.
[0045] From the above four iterations of N=16, the initial value of the dividend nom is 1, and the initial value of the divisor denom is x=5, we can know that we finally get: nom=8192, denom=40960. Further, the binary 1010000000000000 corresponding to denom=40960 is shifted right by 16-12=4 bits to get 1010000000000, and the corresponding decimal is 2560, that is, denom>>(NQ)=2560. Based on this, let denom>>(NQ) be m and substitute , after rounding down 25.6, we get LUT(denom>>(NQ))=25. Then LUT_div=LUT(denom>>(NQ)) ×8192=25×8192=204800. After shifting the multiplication result LUT_div to the right by 2×4=8 bits, the final calculation result is 800. The floating point representation of 800 according to Q=12 is: 800 / (2^(12))= 0.1953125. Compared with the calculation result of 1 / 5=0.2 of the division of the initial value of the dividend of 1 and the initial value of the divisor x=5, the error is within the expected range.
[0046] It can be seen that through this scheme, the division calculation , which translates to: After that, first determine through this scheme After calculating the division result of The calculated results are directly calculated with the existing technology Compared with the calculation results of the two, the error is within the expected range. And through this solution, the final division result can be directly obtained based on hardware, avoiding the existing division calculation to complete the data representation range of the division, and it is necessary to re-determine a fixed-point to floating-point conversion algorithm to map the current division result (quotient and remainder) in a reasonable way, thereby improving the efficiency of the division calculation.
[0047] It can be seen from this scheme that when performing a division operation between the divisor 1 and the dividend x, the bit width N of the divisor x and the preset output data bit width Q are first determined, and then the divisor and the dividend are iteratively operated based on the N and Q to shift, and then the result denom>>(NQ) obtained by right shifting the divisor by (NQ) bits is substituted into the preset lookup table LUT (denom>>(NQ)), and the lookup result LUT (denom>>(NQ)) is multiplied by the final shifted nom to obtain the multiplication result LUT_div; finally, the multiplication result LUT_div is right shifted by 2×(NQ) bits to obtain the final calculation result. The error between the calculation result and the actual division operation directly performed by the divisor and the dividend is within the expected range, that is, the method of this scheme can obtain the final division result after the floating point conversion to the fixed point that conforms to the expected data representation format, and the process of obtaining the division calculation result by the above method of this scheme does not require the conversion from fixed point to floating point, and all data is processed according to the hardware format, thereby improving the calculation efficiency of the hardware divider.
[0048] In this solution, the above division calculation process is completed in hardware, and the above division calculation process can be implemented in hardware through pseudo code. Specifically, the hardware relationship of the pseudo code in the embodiment of the present application is expressed as follows: Figure 2 As shown, it can be seen that the division in the embodiment of the present application can be performed within the known required error range through iterative calculation in the embodiment of the present application, and multiplication and shifting based on the prepared lookup table to obtain the final division output result. Based on this, taking N=16 as an example, the division operation process in this solution can be implemented in hardware through the following pseudo code: nom = dividend; denom = divisor; Because (N == 16) { if (0xFF00&denom) == 0) bitwise AND to determine if there are valid values besides the lower 8 bits { denom<<= 8; shift left 8 bits nom<<= 8; } if (0xF000&denom) == 0) bitwise AND to determine if there are valid values besides the lower 12 bits { denom<<= 4; shift left 4 bits nom<<= 4; } if (0xC000&denom) == 0) bitwise AND to determine if there are valid values besides the lower 14 bits { denom<<= 2; shift left 2 bits nom<<= 2; } if (0x8000&denom) == 0) bitwise AND to determine if there are any valid values besides the lower 15 bits { denom<<= 1; shift left 1 bit nom<<= 1; } } LUT_div= nom ×LUT(denom>>(NQ)) Find the corresponding number of denom in the lookup table and multiply it by nom Output = LUT_div>>2×(NQ) Finally, the result is right-shifted 2×(NQ) bits, and then the output is the final division result.
[0049] It should be noted that the values of N and Q can be set accordingly based on the needs of different scenarios. In the current example, N=16, the data shift is set with an exponential relationship of 2, so if each iterative calculation in the above 4 iterative calculations determines that a shift is required, the number of bits shifted is 8, 4, 2, and 1 respectively. Correspondingly, when N=8, if each iterative calculation in the 3 iterative calculations determines that a shift is required, the number of bits shifted is 4, 2, and 1 respectively; and when N=32, the number of shifts is 5. If a shift is required for all 5 iterative calculations, the number of bits shifted for the 5 shifts is 16, 8, 4, 2, and 1 respectively. When the number of bits of the data itself does not satisfy the exponential relationship of 2, it can be achieved by padding the high bits with zeros.
[0050] The following is an example of actual division calculation: Assume that the dividend nom=1, the divisor denom=5 entered in the pseudocode, and N continues to use N=16 and Q=12 in the above example.
[0051] Then the division calculation process can be expressed as: The dividend nom=1, the divisor denom=5, N continues to use the previous example to be 16, and Q is 12.
[0052] The first iteration calculation: (0xFF00&denom) == 0 is true, so shift by 8. After shifting, nom= 256, denom= 1280; Second iteration calculation: (0xF000&denom) == 0 is true, so shift by 4. After shifting, nom= 4096, denom= 20480; The third iteration calculation: (0xC000&denom) == 0 is not true, no shift; The fourth iteration calculation: (0x8000&denom) == 0 is true, shift by 1, after shifting, nom = 8192, denom = 40960; After 4 iterations, the final nom = 8192, denom = 40960 is obtained. Then the result of the final denom shifted right by 4 bits is substituted into the preset LUT, and finally multiplied by the final nom, that is, 8192 × LUT (40960 shifted right by 4 bits) = 8192 × 25 = 204800. Finally, the obtained 204800 is shifted right by 2×(NQ) and the settlement result is output, that is, 204800 is shifted right by 2×(16-12) bits = 800; It can be seen that the result is expressed in floating point form according to Q=12: 800 / (2^(12)) = 0.1953125. It can be seen that compared with the calculation of 1 / 5=0.2, the error is within the expected range.
[0053] It can be seen that the hardware solution for implementing division in the embodiment of the present application is designed based on an acceptable error range, and supports division operations with decimal results. In this solution, there is no need to convert fixed-point to floating-point, all data is processed according to the hardware format, and this solution does not have complex conditions and hardware modules. It only needs to plan a lookup table in advance based on the acceptable error range, and then complete the division operation by shifting and querying and mapping data in the lookup table.
[0054] Corresponding to the above Figure 1 , the embodiment of the present application provides a division calculation device based on hardware implementation, such as Figure 3 As shown, the device comprises: An acquisition module 302 is used to acquire a dividend 1 and a divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, and the output data bit width Q is determined based on an error range of the output data, and Q is less than or equal to N; The first processing module 304 is used to set the initial value of nom to 1, the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, determine the high N / 2 of denom obtained by the previous iterative calculation n Whether the bit has a valid value; The second processing module 306 is used to shift denom and nom calculated in the previous iteration to the left by N / 2 respectively if there is no valid value. n bits, get new denom and nom; The third processing module 308 is used to iterate until N / 2 n =1, get the final denom and nom; The fourth processing module 310 is used to shift the final denom right by (NQ) bits to obtain denom>>(NQ), substitute denom>>(NQ) into a preset lookup table to obtain a lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to obtain a multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits, obtain a final calculation result and output it.
[0055] Through the device of the embodiment of the present application, when performing a division operation between the divisor 1 and the dividend x, the bit width N of the divisor x and the preset output data bit width Q are first calculated, and then the divisor and the dividend are iteratively operated based on the N and Q to shift them, and then the result denom>>(NQ) obtained by right shifting the divisor by (NQ) bits is substituted into the preset lookup table LUT (denom>>(NQ)), and the lookup result LUT (denom>>(NQ)) is multiplied by the final shifted dividend to obtain the multiplication result LUT_div; finally, the multiplication result LUT_div is right shifted by 2 times (NQ) bits to obtain the final calculation result. The error between the calculation result and the actual division operation directly performed by the divisor and the dividend is within the expected range, that is, the above method can obtain the division result after the floating point conversion that finally conforms to the expected data representation format, and the process of obtaining the division calculation result by the above method in the embodiment of the present application does not require the conversion from fixed point to floating point, and all data is processed according to the hardware format, thereby improving the calculation efficiency of the hardware divider.
[0056] In an optional implementation of the embodiment of the present application, the second processing module in the embodiment of the present application includes: a first processing unit, which is used to perform a bitwise AND operation on the target data and the denom calculated in the previous iteration in the nth iteration operation; wherein the bit width of the target data is N, and the high N / 2 of the target data n-1 The bit is 1, and the remaining bits are 0; the second processing unit is used to determine the high N / 2 of the denom calculated in the previous iteration when the result of the bitwise AND operation is 0 n The bit does not have a valid value; the third processing unit is used to determine the high N / 2 of the denom calculated in the previous iteration when the result of the bitwise AND operation is not 0n The bit has a valid value.
[0057] In an optional implementation manner of the embodiment of the present application, the lookup table in the embodiment of the present application stores the following corresponding relationship: ,in, The value of is 0, 1, ... , N is the bit width of the divisor x, express The value of is rounded down.
[0058] In an optional implementation of the embodiment of the present application, when the divisor x is 0, no iterative calculation is performed and the calculation result 2 is directly output. Q+1 -1.
[0059] like Figure 4 As shown, an embodiment of the present application provides an electronic device, including a processor 411, a communication interface 412, a memory 413 and a communication bus 414, wherein the processor 411, the communication interface 412, and the memory 413 communicate with each other through the communication bus 414. Memory 413, used for storing computer programs; In one embodiment of the present application, the processor 411 is used to execute the program stored in the memory 413 to implement the hardware-based division calculation method provided by any of the aforementioned method embodiments, and its role is similar and will not be repeated here.
[0060] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the hardware-based division calculation method provided in any of the aforementioned method embodiments are implemented.
[0061] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0062] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution, in essence, or the part that contributes to the relevant technology, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.
[0063] It should be understood that the terms used herein are only for the purpose of describing specific example embodiments and are not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms "one", "an" and "said" as used herein may also be meant to include plural forms. The terms "include", "comprise", "contain", and "have" are inclusive, and therefore specify the existence of stated features, steps, operations, elements and / or parts, but do not exclude the existence or addition of one or more other features, steps, operations, elements, parts, and / or combinations thereof. The method steps, processes, and operations described herein are not interpreted as necessarily requiring them to be performed in the specific order described or illustrated, unless the execution order is clearly indicated. It should also be understood that additional or alternative steps may be used.
[0064] The foregoing is merely a specific embodiment of the present invention, which enables those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A division calculation method based on hardware implementation, characterized in that: include: Obtaining a dividend 1 and a divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on an error range of the output data, and Q is less than or equal to N; Let the initial value of nom be 1, the initial value of denom be x, and perform the following iterative calculation: When performing the nth iterative calculation, determine the high N / 2 of denom obtained in the previous iterative calculation n Whether the bit still has a valid value; If there is no valid value, denom and nom calculated in the previous iteration are shifted left by N / 2 respectively. n bits, get new denom and nom; Until the iteration reaches N / 2 n =1, get the final denom and nom; Shift the final denom right by (NQ) bits to obtain denom>>(NQ), substitute denom>>(NQ) into the preset lookup table to obtain the lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to obtain the multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits to obtain the final calculation result and output it.
2. The method according to claim 1, characterized in that When performing the nth iteration, determine the N / 2 higher value of denom obtained in the previous iteration. n Whether the bit still has a valid value; If there is a valid value, no update is made, denom and nom obtained from the previous iteration are kept unchanged, and the next iteration is performed.
3. The method according to claim 1, characterized in that The judgment is the high N / 2 of denom calculated in the previous iteration n Whether the bit still has a valid value, including: In the nth iteration operation, the target data is bitwise ANDed with the denom calculated in the previous iteration; wherein the bit width of the target data is N, and the high N / 2 of the target data n bit is 1, and the remaining bits are 0; When the result of the bitwise AND operation is 0, determine the high N / 2 of denom calculated in the previous iteration n There is no valid value for the bit; If the result of the bitwise AND operation is not 0, determine the high N / 2 of denom calculated in the previous iteration n The bit has a valid value.
4. The method according to claim 1, characterized in that The lookup table stores the following corresponding relationships: ,in, The value of is 0, 1, ... , N is the bit width of the divisor x, express The value of is rounded down.
5. The method according to claim 1, characterized in that include: When the divisor x is 0, no iterative calculation is performed and the calculation result 2 is directly output. Q+1 -1.
6. A division calculation device based on hardware implementation, characterized in that: include: An acquisition module, used to acquire a dividend 1 and a divisor x, and a preset output data bit width Q, wherein the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on an error range of the output data, and Q is less than or equal to N; The first processing module is used to set the initial value of nom to 1 and the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, determine the high N / 2 of denom obtained by the previous iterative calculation n Whether the bit has a valid value; The second processing module is used to shift denom and nom calculated in the previous iteration to the left by N / 2 respectively if there is no valid value. n bits, get new denom and nom; The third processing module is used to iterate until N / 2 n =1, get the final denom and nom; The fourth processing module is used to shift the final denom right by (NQ) bits to obtain denom>>(NQ), substitute denom>>(NQ) into a preset lookup table to obtain a lookup result LUT(denom>>(NQ)), multiply the lookup result LUT(denom>>(NQ)) by the final nom to obtain a multiplication result LUT_div, shift the multiplication result LUT_div to the right by 2×(NQ) bits, obtain the final calculation result and output it.
7. The device according to claim 6, characterized in that The second processing module comprises: The first processing unit is used to perform a bitwise AND operation on the target data and the denom calculated in the previous iteration in the nth iteration operation; wherein the bit width of the target data is N, and the high N / 2 of the target data n bit is 1, and the remaining bits are 0; The second processing unit is used to determine the high N / 2 of denom calculated in the previous iteration when the result of the bitwise AND operation is 0 n There is no valid value for the bit; The third processing unit is used to determine the high N / 2 of denom calculated in the previous iteration when the result of the bitwise AND operation is not 0. n The bit has a valid value.
8. A hardware device, characterized in that: The hardware device includes the hardware-based division calculation device described in claim 6 or 7.
9. An electronic device, comprising: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor coupled to the at least one bus; At least one memory connected to the at least one bus, wherein the processor is configured to execute the hardware-implemented division calculation method as described in any one of claims 1 to 5 above.
10. A computer storage medium storing computer executable instructions, wherein the computer executable instructions are used to execute the hardware-based division calculation method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Divider and quotient and remainder solving method
CN106354473A
Lookup table LUT optimization method and device, electronic equipment and storage medium
CN117806589A
Root mean square calculation method and device based on hardware implementation and electronic equipment
CN119357543A
Divider
CN205899527U
Arithmetic unit and operation method
JP2003067182A