Hardware-based Division Calculation Method and Device

Through iterative calculation and lookup table multiplication operation, the problem of low calculation efficiency of hardware dividers is solved, and efficient processing of fixed-point division result is realized, which avoids additional fixed-point to floating-point conversion, and improves the calculation speed and accuracy of hardware dividers.

CN120010813BActive Publication Date: 2025-07-25XIAN WEIHE ZHILIAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510488200.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-25
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Existing hardware dividers are low in computational efficiency and cannot effectively handle division between fixed-point numbers and floating-point numbers, and require additional fixed-point to floating-point conversion algorithms to map division results.

Method used

Through iterative calculation, determine the bit width of the divisor and the dividend, perform shift operations, and substitute the result into the preset lookup table for multiplication, and finally obtain the division result that conforms to the expected data format, avoiding the conversion from fixed-point to floating point.

Benefits of technology

It improves the computing efficiency of hardware dividers, reduces the need for fixed-point to floating-point conversion, and all data is processed in hardware format, improving calculation speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010813B_ABST
    Figure CN120010813B_ABST
Patent Text Reader

Abstract

The present application relates to a hardware-implemented division calculation method and apparatus. The method includes: setting the initial value of nom to 1, the initial value of denom to x, and performing the following iterative calculation: when performing the nth iterative calculation, determining whether there is a valid value in the high N / 2<supgt;n< / supgt> bits of denom obtained from the previous iterative calculation; if there is no valid value, shifting denom and nom obtained from the previous iterative calculation to the left by N / 2<supgt;n< / supgt> bits respectively to obtain new denom and nom; until the iterative calculation reaches N / 2<supgt;n< / supgt>=1 to obtain the final denom and nom; shifting the finally obtained denom to the right by (N-Q) bits to obtain denom>>(N-Q), substituting denom>>(N-Q) into a preset lookup table to obtain a lookup result, and shifting the result obtained by multiplying the lookup result by the finally obtained nom to the right by 2×(N-Q) bits to obtain the final calculation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of hardware-implemented division calculations, and particularly to a hardware-implemented division calculation method and device. Background Art

[0002] With the development of the fifth-generation mobile communication technology (5th Generation Mobile Communication Technology, abbreviated as 5G), the amount of data that user equipment (User Equipment, abbreviated as UE) needs to process is also increasing. However, unlike base stations, in order to ensure that user equipment has higher throughput and faster processing speed, most data processing is implemented through digital signal processing (Digital Signal Process, abbreviated as DSP) or central processing unit (Central Processing Unit, abbreviated as CPU). Therefore, UE needs to rely on hardware to achieve this.

[0003] In related technologies, the hardware dividers used in UE are for division between fixed-point numbers. Although floating-point numbers can be represented by shifting fixed-point numbers, the precision cannot be controlled, and only fixed-point division of the same type of numbers can be achieved. For example, when calculating the binary number b'1001010 divided by the binary number b'1000, the subtraction is performed from left to right. The left four bits of the binary number b'1001010 are subtracted from the binary number b'1000 to get 1, and at this time, 1 is placed in the result bit. Since 1 is less than four bits, the calculation cannot continue. Then, two bits are taken from the right in turn, and two 0s are placed in the result bit to get the four-bit binary number b'1010. Then, the subtraction operation is performed on the binary number b'1000. Finally, the division calculation result is the binary number b'1001, and the remainder is 10. The calculation process is as follows:

[0004]

[0005] In this example, converting the binary number to a decimal number shows that: 74 divided by 8 gives a quotient of 9 and a remainder of 2. It can be seen that the existing dividers can only handle fixed-point integers, and a fixed-point to floating-point conversion algorithm needs to be re-determined for the data representation range after division to map the current division results (quotient and remainder) in a reasonable way, resulting in low calculation efficiency. Summary of the Invention

[0006] This application provides a hardware-implemented division calculation method and device to solve the problem of low calculation efficiency of hardware dividers in related technologies.

[0007] In a first aspect, the present application provides a hardware-based division calculation method, including: obtaining the dividend 1 and the divisor x, as well as a preset output data bit width Q, where the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N; setting the initial value of nom to 1, the initial value of denom to x, and performing the following iterative calculations: when performing the nth iterative calculation, determine whether there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation; if there is no valid value, shift denom and nom obtained from the previous iterative calculation to the left by N / 2 n bits respectively to obtain new denom and nom; until the iterative calculation reaches N / 2 n = 1 to obtain the final denom and nom; shift the finally obtained denom to the right by (N - Q) bits to get denom>>(N - Q), substitute denom>>(N - Q) into a preset look-up table to obtain a look-up result LUT(denom>>(N - Q)), multiply the look-up result LUT(denom>>(N - Q)) by the finally obtained nom to get a multiplication result LUT_div, and after shifting the multiplication result LUT_div to the right by 2×(N - Q) bits, obtain the final calculation result and output it.

[0008] Optionally, when performing the nth iterative calculation, determine whether there is still a valid value in the high N / 2n bits of denom obtained from the previous iterative calculation; if there is a valid value, do not update, keep denom and nom obtained from the previous iterative calculation unchanged, and perform the next iterative calculation.

[0009] Optionally, determining whether there is still a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation includes: in the (n - 1)th iterative operation, performing a bitwise AND operation on the target data and denom obtained from the previous iterative calculation; where the bit width of the target data is N, and the high N / 2 n bits of the target data are 1 and the remaining bits are 0; in the case where the result of the bitwise AND operation is 0, determine that there is no valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation; in the case where the result of the bitwise AND operation is not 0, determine that there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation.

[0010] Optionally, the look-up table stores the following corresponding relationship: , where takes values of 0, 1,... , where N is the bit width of the divisor x, denotes the value rounded down.

[0011] Optionally, the method includes: when the divisor x is 0, no iterative calculation is performed, and the calculation result 2 Q+1 -1 is directly output.

[0012] In a second aspect, the present application provides a hardware-implemented division calculation device, including: an acquisition module for acquiring the dividend 1 and the divisor x, and a preset output data bit width Q, where the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N; a first processing module for setting the initial value of nom to 1, the initial value of denom to x, and performing the following iterative calculation: when performing the nth iterative calculation, determining whether there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation; a second processing module for, if there is no valid value, shifting denom and nom obtained from the previous iterative calculation to the left by N / 2 n bits to obtain new denom and nom; a third processing module for iterating until N / 2 n = 1 to obtain the final denom and nom; a fourth processing module for shifting the finally obtained denom to the right by (N - Q) bits to get denom>>(N - Q), substituting denom>>(N - Q) into a preset look-up table to obtain a look-up result LUT(denom>>(N - Q)), multiplying the look-up result LUT(denom>>(N - Q)) by the finally obtained nom to get a multiplication result LUT_div, and shifting the multiplication result LUT_div to the right by 2×(N - Q) bits to obtain the final calculation result and output it.

[0013] Optionally, the second processing module includes: a first processing unit for performing a bitwise AND operation on the target data and denom obtained from the previous iterative calculation in the nth iterative operation; where the bit width of the target data is N, and the high N / 2 n bits of the target data are 1 and the remaining bits are 0; a second processing unit for determining that there is no valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation when the result of the bitwise AND operation is 0; a third processing unit for determining that there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation when the result of the bitwise AND operation is not 0.

[0014] In a third aspect, the present application provides a hardware device, including the hardware-implemented division calculation device described in the second aspect.

[0015] In a fourth aspect, the present application further provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus, wherein the processor is configured to execute the hardware-implemented division calculation method described in the first aspect of the present application above.

[0016] In a fifth aspect, the present application further provides a computer storage medium storing computer-executable instructions for executing the hardware-implemented division calculation method described in the first aspect of the present application above.

[0017] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: In the method provided by the embodiments of the present application, when performing the division operation of the divisor 1 and the dividend x, first determine the bit width N of the divisor x and the preset output data bit width Q, and then perform iterative operations on the divisor and the dividend based on N and Q for shifting. Then, substitute the result denom>>(N-Q) after shifting the finally shifted divisor to the right by (N-Q) bits into the preset look-up table LUT(denom>>(N-Q)), and multiply the look-up result LUT(denom>>(N-Q)) by the finally shifted nom to obtain the multiplication result LUT_div; finally, shift the multiplication result LUT_div to the right by 2×(N-Q) bits to obtain the final calculation result. The error between this calculation result and the actual direct division operation by the divisor and the dividend is within the expected range, that is, the division result after floating-point to fixed-point conversion that finally conforms to the expected data representation format can be obtained through the above method. Moreover, in the process of obtaining the division calculation result through the above method in the embodiments of the present application, there is no need to perform fixed-point to floating-point conversion, and all data is processed in the format of the hardware, which improves the calculation efficiency of the hardware divider. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0020] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a scale limitation.

[0021] Figure 1 It is a method flowchart of a division calculation method based on hardware implementation provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic diagram of the hardware relationship expression for implementing pseudocode provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic structural diagram of a division calculation device based on hardware implementation provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0026] The following disclosure provides many different embodiments or examples for implementing different structures of the present invention. To simplify the disclosure of the present invention, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present invention. In addition, the present invention may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0027] Considering a division , it can be converted to: , and the multiplication calculation can be simply implemented by an AND gate. Therefore, this division calculation can be converted to a calculation process of multiplying the value of by . Therefore, in this solution, the dividend is set to 1. Through the division calculation of this solution, is obtained, and then is subjected to an AND gate operation with to obtain the final result. It should be noted that when = 0, the calculation result is directly output as 0; when In the case of = 0, the maximum value supported by the current output bit width is output. For other cases, division calculations are performed , and the acceptable division error is , then b can be converted based on p, and we can get , where p is the highest significant p bits corresponding to p in b, and l is the number of bits remaining after removing the highest significant p bits from b itself. For example, when the current N = 16 is the bit width of the data itself, p = 12 is the left (highest significant) 12 bits of the corresponding 16-bit data, then l = 16 - 12 = 4 bits. Further, the above relationship is demonstrated as follows:

[0028] Converting the division formula gives:

[0029] Observing the , performing Taylor expansion on it gives:

[0030]

[0031] Therefore, based on the algorithm of the shift method of l and p of b, the error compared to the traditional division algorithm without expansion is .

[0032] To solve the problem of low calculation efficiency of the hardware divider in the related art, the present application provides a division calculation method implemented based on hardware, as Figure 1 shown. The steps of this method include:

[0033] Step 101, obtain the dividend 1 and the divisor x, and the preset output data bit width Q. Among them, the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N;

[0034] In this solution, if the current division calculation to be performed is 2÷5, the dividend actually used for the division calculation in the hardware is 1, and the divisor x is 5; then the obtained divisor 1 and divisor x = 5 are used for division calculation, and after the result of the division calculation is ANDed with 2, the calculation result of the division calculation 2÷5 can be obtained. Further, in the embodiment of the present application, the bit width of the divisor x is equal to the bit width supported by the hardware for the current division calculation, that is, the bit width supported by the hardware for the current division calculation is also N. However, the bit widths corresponding to the dividers in different hardware are different, that is, the value of N is determined based on the divider of the hardware. For example, N = 16, or N = 32, etc. In addition, in the embodiment of the present application, the output data bit width Q is the bit width determined based on the error range of the output data. The larger the output data bit width Q, the smaller the error range of the output data, that is, the higher the precision and the better the algorithm performance, but the greater the hardware area and power consumption loss; the smaller the output data bit width Q, the larger the error range of the output data, that is, the lower the precision and the worse the algorithm performance, but the smaller the hardware area and power consumption loss. Of course, the output data bit width Q needs to be less than or equal to the hardware bit width supported by the hardware. For example, when the hardware bit width supported by the hardware is 22 bits, the output data bit width Q cannot exceed 22 bits. For example, the value of Q can be 12, 18, etc.

[0035] It should be noted that when the divisor x is 0, no iterative calculation is performed, and the calculation result 2 is directly output Q+1 -1. In addition, the letter x in the embodiment of the present application represents the data to be calculated in the present application. Of course, other letters can also be used to represent the data to be calculated, which does not constitute a limitation to the present application. The letter Q here represents the output data bit width in the present application. Of course, other letters can also be used to represent the output data bit width, which does not constitute a limitation to the present application.

[0036] Step 102, set the initial value of nom to 1, the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, determine whether there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation;

[0037] Step 103, if there is no valid value, shift the denom and nom obtained from the previous iterative calculation to the left by N / 2 n bits to obtain new denom and nom;

[0038] In addition, when determining the high N / 2 of denom obtained from the previous iterative calculation nWhen determining whether there are still valid values in a bit position, if there are valid values, no update is performed, and the denom and nom obtained from the previous iterative calculation are kept unchanged, and the next iterative calculation is carried out. For example, currently it is the 3rd iterative calculation, and in this iterative calculation, it is determined that the high N / 2 of the denom obtained from the 2nd iterative calculation 3 There are valid values, so no update is performed, and the denom and nom obtained from the 2nd iterative calculation are kept unchanged, and the 4th iterative calculation is carried out.

[0039] Step 104, until the iterative calculation reaches N / 2 n = 1, to obtain the final denom and nom.

[0040] Regarding this, in this solution, take the initial value of denom as x = 5 and N = 16 as an example. Since the iterative calculation is at N / 2 n = 1, the final denom and nom will be obtained, and at this point the iterative calculation ends. Therefore, when n = 4, it satisfies 16 / 2 4 = 1, that is, when N = 16, the number of iterative calculations n = 4. It should be noted that during multiple iterative calculations, if during the last few iterative calculations, it is determined that the high N / 2 of the denom obtained from the previous iterative calculation n bits all have valid values, then during the last few iterative calculations, neither the denom nor the nom obtained from the previous iterative calculation is shifted. Instead, the last shifted denom and the shifted nom need to be found for subsequent calculations. For example, if during the last 3 iterative calculations, it is determined that the high N / 2 of the denom obtained from the previous iterative calculation n bits all have valid values, then the shifted denom and the shifted nom after the 4th-to-last iterative calculation need to be found as the final denom and nom.

[0041] Therefore, for the above steps 102 to 104, if currently it is the n = 2nd iterative calculation, it is necessary to determine whether there are valid values in the high 16 / 2 2 = 4 bits of the denom obtained from the 1st iterative calculation. If the high 4 bits of the denom obtained from the 1st iterative calculation do not have valid values, then the denom and nom obtained from the 1st iterative calculation need to be shifted left by 16 / 2 2= 4 bits to obtain new denom and nom; if the upper 4 bits of the denom obtained from the first iterative calculation have valid values, no update is performed, and the denom and nom obtained from the previous iterative calculation are kept unchanged for the next iterative calculation. It should be noted that the above is an example with n = 2, that is, the second iterative calculation. For iterative calculations after the second time, the process is similar to that of the second iterative calculation. However, for the first iterative calculation, since there are no denom and nom obtained from the previous iterative calculation, when performing the first iterative calculation, it is necessary to judge whether the upper N / 2 n bits have valid values, which means judging whether the upper 8 bits of x = 5 have valid values.

[0042] Furthermore, in this solution, for the method of judging whether the upper N / 2 n bits of the denom obtained from the previous iterative calculation still have valid values involved in step 102 above, it can further be:

[0043] In the nth iterative operation, perform a bitwise AND operation on the target data and the denom obtained from the previous iterative calculation; where the bit width of the target data is N, and the upper N / 2 n bits of the target data are 1, and the remaining bits are 0;

[0044] In the case where the result of the bitwise AND operation is 0, it is determined that the upper N / 2 n bits of the denom obtained from the previous iterative calculation do not have valid values;

[0045] In the case where the result of the bitwise AND operation is not 0, it is determined that the upper N / 2 n bits of the denom obtained from the previous iterative calculation have valid values.

[0046] It can be seen that in this solution, it is further judged whether the upper N / 2 n bits of the denom obtained from the previous iterative calculation still have valid values by means of bitwise AND. In this regard, still taking the initial value of denom x = 5 and N = 16 as an example, if the current is the n = 2nd iterative calculation, at this time the upper 16 / 2 2 = 4 bits of the target data are 1, and the remaining bits are 0, then the corresponding binary representation of the target data is 1111000000000000, that is, perform a bitwise AND operation on 1111000000000000 and the denom obtained from the previous iterative calculation, and determine whether the upper 16 / 2 2 = 4 bits of the denom obtained from the previous iterative calculation have valid values according to whether the result of the bitwise AND is 0.

[0047] It should be noted that in this solution, the method of performing a bitwise AND operation on the target data and the denom obtained from the previous iterative calculation is as follows: The binary digits of the target data are ANDed with the corresponding binary digits of the denom obtained from the previous iterative calculation. That is, when the corresponding binary digits are both 1, the binary digit after the AND operation is 1; in other cases, the binary digit after the AND operation is 0.

[0048] Next, with the initial value of nom being 1, taking the initial value of denom as x = 5 and N = 16 as an example, the iterative calculation process in this solution will be fully explained.

[0049] Based on this, the entire iterative calculation process is as follows:

[0050] The first iterative calculation: Since the value of N is 16, the bit width of the target data is 16 bits, and it is the first iterative calculation, the value of n is 1. Based on this, the high 16 / 2 1 = 8 bits of the binary representation of the target data are 1, and the remaining bits are 0, that is, 1111111100000000. The binary representation corresponding to the initial value x = 5 of denom is 101. Performing a bitwise AND operation on the two, that is, 1111111100000000 & 0000000000000101 == 0. It can be seen that the bitwise AND result is 0. From this, it can be judged that there is no valid value in the high 8 bits of the binary representation of x = 5. Therefore, denom = 5 and nom = 1 need to be shifted, that is, shifted left by 8 bits. Among them, the binary of 1 is still 1, the binary representation of 5 is 101, the binary representation of nom after shifting left by 8 bits is 100000000, and the corresponding decimal representation is 256; the binary representation of denom after shifting left by 8 bits is 10100000000, and the corresponding decimal representation is 1280. For this, the specific calculation process in hardware can be: If (0xFF00 & denom) == 0 holds, then shift left by 8 bits. After shifting, nom = 256, and after shifting, denom = 1280. Among them, 0xFF00 is the hexadecimal representation of the target data 1111111100000000.

[0051] The second iterative calculation: In the second iterative calculation, the value of n is 2. Then the high 16 / 2 2 = 4 bits of the target data are 1, and the remaining bits are 0, that is, 1111000000000000. Based on this, performing a bitwise AND operation on the target data and the denom = 1280 obtained from the first iterative calculation is: 1111000000000000 & 0000010100000000 == 0. It can be seen that the bitwise AND result is 0. From this, it can be judged that the high 16 / 2 2There is no valid value in the 4 bits. Therefore, it is necessary to shift denom = 1280 and nom = 256, that is, shift left by 4 bits. The binary representation of nom after shifting left by 4 bits is 1000000000000, and the corresponding decimal representation is 4096; the binary representation of denom after shifting left by 4 bits is 101000000000000, and the corresponding decimal representation is 20480. For this, the specific calculation process in hardware can be: if (0xF000 & denom) == 0 holds, then shift left by 4 bits. After shifting, nom = 4096, and after shifting, denom = 20480. Among them, 0xF000 is the hexadecimal representation of the target data 1111000000000000.

[0052] The 3rd iteration calculation: In the 3rd iteration calculation, the value of n is 3, then the upper 16 / 2 3 = 2 bits of the target data are 1, and the remaining bits are 0, that is, 1100000000000000. Based on this, perform a bitwise AND operation on the target data and the denom = 20480 obtained in the 2nd iteration calculation: 1100000000000000 & 0101000000000000 == 1. It can be seen that the bitwise AND result is 1. From this, it can be judged that the upper 16 / 2 3 = 2 bits of the binary representation of denom = 20480 have valid values. Therefore, do not shift denom = 20480 and nom = 4096, and then keep the denom and nom obtained in this iteration calculation unchanged and perform the next iteration calculation. For this, the specific calculation process in hardware can be: if (0xC000 & denom) == 0 does not hold, then do not shift the denom and nom after the second shift at this time. Among them, 0xC000 is the hexadecimal representation of the target data 1100000000000000.

[0053] The 4th iteration calculation: In the 4th iteration calculation, the value of n is 4, then the upper 16 / 2 4 = 1 bit of the target data is 1, and the remaining bits are 0, that is, 1000000000000000. Based on this, perform a bitwise AND operation on the target data and the denom = 20480 obtained in the 2nd iteration calculation: 1000000000000000 & 0101000000000000 == 0. It can be seen that the bitwise AND result is 0. From this, it can be judged that the upper 16 / 2 4There is no valid value in the 1-bit, so it is necessary to shift denom = 20480 and nom = 4096, that is, shift left by 1 bit. The binary representation of nom after shifting left by 1 bit is 10000000000000, and the corresponding decimal representation is 8192; the binary representation of denom after shifting left by 1 bit is 1010000000000000, and the corresponding decimal representation is 40960. In this regard, the specific calculation process in hardware can be: if (0x8000 & denom) == 0 holds, then shift left by 1 bit, the shifted nom = 8192, and the shifted denom = 40960. Among them, 0x8000 is the hexadecimal representation of the target data 1000000000000000.

[0054] It can be seen that after the above 4 iterations of calculation, the finally obtained dividend nom = 8192 and divisor denom = 40960.

[0055] Step 105: Shift the finally obtained denom to the right by (N - Q) bits to get denom>>(N - Q), substitute denom>>(N - Q) into the preset lookup table to obtain the lookup result LUT(denom>>(N - Q)), multiply the lookup result LUT(denom>>(N - Q)) by the finally obtained nom to get the multiplication result LUT_div, and after shifting the multiplication result LUT_div to the right by 2×(N - Q) bits, obtain the final calculation result and output it.

[0056] In this regard, in this solution, taking N as 16 and Q as 12 as an example, then shift the finally obtained denom to the right by 16 - 12 = 4 bits, and then substitute the result of shifting the finally obtained denom to the right by 4 bits into the preset lookup table. For example, if the decimal representation of the currently finally obtained denom is 40960, the corresponding binary representation is 1010000000000000, and the corresponding binary after shifting to the right by 4 bits is 101000000000, and the decimal representation after shifting to the right is 2560, that is, substitute 2560 into the lookup table to obtain the lookup result LUT(denom>>(N - Q)) = LUT(2560), and multiply the lookup result LUT(denom>>(N - Q)) by the shifted nom to obtain the multiplication result LUT_div.

[0057] Furthermore, the lookup table in this solution stores the following corresponding relationships:

[0058] , where takes values of 0, 1,... , N is the bit width of the divisor x, represents the value of is rounded down.

[0059] For this, take the case where the value of N is 16, the value of Q is 12, the initial value of the dividend nom is 1, and the initial value of the divisor denom is x = 5. Since the value of N is 16 and the value of Q is 12, some data in the lookup table LUT are as follows:

[0060]

[0061] It can be seen that the size of the lookup table in this solution is affected by Q. When Q is 12, it can be known that the maximum number in the lookup table is 4095.

[0062] From the 4 - iteration calculations with N = 16, the initial value of the dividend nom being 1, and the initial value of the divisor denom being x = 5, it can be finally obtained that: nom = 8192 and denom = 40960. Further, for the binary 1010000000000000 corresponding to denom = 40960, after shifting it to the right by 16 - 12 = 4 bits, the resulting binary is 101000000000, and the corresponding decimal is 2560, that is, denom>>(N - Q)=2560. Based on this, let denom>>(N - Q) be set as m and substitute it into , after rounding down 25.6, it is obtained that LUT(denom>>(N - Q)) = 25. Further, LUT_div = LUT(denom>>(N - Q)) × 8192 = 25×8192 = 204800. After shifting the multiplication result LUT_div to the right by 2×4 = 8 bits, the final calculation result is 800. In the floating - point representation with Q = 12, this 800 is: 800 / (2^(12)) = 0.1953125. Compared with the calculation result of 1 / 5 = 0.2 for the division of the initial value of the dividend being 1 and the initial value of the divisor being x = 5, the error is within the expected range.

[0063] It can be seen that through this solution, the division calculation is converted to: After that, first determine the division calculation result of in the way of this solution. Then, the calculation result obtained by ANDing this division calculation result with is compared with the calculation result of directly calculating in the prior art. The error is within the expected range. And through this solution, the final division result can be directly obtained based on hardware, avoiding the need in the existing division calculation to re - determine a fixed - point to floating - point conversion algorithm for the data representation range after the division is completed, so as to map and represent the current division result (quotient and remainder) in a reasonable way, improving the efficiency of the division calculation.

[0064] According to this solution, when performing the division operation of divisor 1 and dividend x, first determine the bit width N of the divisor x and the preset output data bit width Q. Then, based on N and Q, perform iterative operations on the divisor and dividend for shifting. Next, substitute the result of shifting the finally shifted divisor to the right by (N - Q) bits, i.e., denom>>(N - Q), into the preset lookup table LUT(denom>>(N - Q)), and multiply the lookup result LUT(denom>>(N - Q)) by the finally shifted nom to obtain the multiplication result LUT_div. Finally, shift the multiplication result LUT_div to the right by 2×(N - Q) bits to obtain the final calculation result. The error between this calculation result and the actual direct division operation using the divisor and dividend is within the expected range. That is, through the method of this solution, the division result after floating-point to fixed-point conversion that finally conforms to the expected data representation format can be obtained. Moreover, in the process of obtaining the division calculation result through the above method of this solution, there is no need to perform the conversion from fixed-point to floating-point, and all data is processed in the hardware format, which improves the calculation efficiency of the hardware divider.

[0065] In this solution, the above division calculation process is completed in hardware, and the above division calculation process can be implemented in hardware through pseudocode. Specifically, the hardware relationship expression of the pseudocode in the embodiments of this application is as Figure 2 shown. It can be seen that for the division in the embodiments of this application, within the known required error range, through the iterative calculation in the embodiments of this application and based on the prepared lookup table, perform multiplication and then perform shifting to obtain the final division output result. Based on this, taking N = 16 as an example, the division operation process in this solution can be implemented in hardware through the following pseudocode:

[0066] nom = dividend;

[0067] denom = divisor;

[0068] because (N == 16)

[0069] {

[0070] if ((0xFF00 & denom) == 0) bitwise AND to determine if there are valid values except for the lower 8 bits

[0071] {

[0072] denom <<= 8; shift left by 8 bits

[0073] nom <<= 8;

[0074] }

[0075] if ((0xF000 & denom) == 0) bitwise AND to determine if there are valid values except for the lower 12 bits

[0076] {

[0077] denom <<= 4; Shift left by 4 bits

[0078] nom <<= 4;

[0079] }

[0080] if ((0xC000 & denom) == 0) Bitwise AND to determine if there are valid bits other than the lower 14 bits

[0081] {

[0082] denom <<= 2; Shift left by 2 bits

[0083] nom <<= 2;

[0084] }

[0085] if ((0x8000 & denom) == 0) Bitwise AND to determine if there are valid bits other than the lower 15 bits

[0086] {

[0087] denom <<= 1; Shift left by 1 bit

[0088] nom <<= 1;

[0089] }

[0090] }

[0091] LUT_div = nom × LUT(denom >> (N - Q)) Look up the number corresponding to denom in the lookup table and then multiply by nom

[0092] Output = LUT_div >> 2×(N - Q) Finally, shift the result to the right by 2×(N - Q) bits, and the output is the final division result.

[0093] It should be noted that the values of N and Q can be set accordingly based on different scenario requirements. Currently, for example, when N = 16, the data shifting is set in an exponential relationship of 2. Therefore, in each of the above 4 iterative calculation processes, if it is determined to shift, the number of bits shifted is 8, 4, 2, 1 respectively. Correspondingly, when N = 8, in the 3 iterative calculation processes, if it is determined to shift in each iterative calculation, the number of bits shifted is 4, 2, 1 respectively; when N = 32, the number of shifts is 5. If all 5 iterative calculations require shifting, the number of bits shifted in the 5 shifts is 16, 8, 4, 2, 1 in sequence. When the number of bits of the data itself does not satisfy the exponential relationship of 2, it can be achieved by padding zeros at the high bits.

[0094] The following is an example of actual division calculation: Assume that at this time, the dividend nom = 1 and the divisor denom = 5 in the pseudocode, and N continues to use N = 16 and Q = 12 in the above example.

[0095] Then the division calculation process can be expressed as:

[0096] The dividend nom = 1, the divisor denom = 5, N continues to use 16 in the previous example, and Q is 12.

[0097] The first iteration calculation:

[0098] (0xFF00 & denom) == 0 holds, so shift by 8. After shifting, nom = 256 and denom = 1280;

[0099] The second iteration calculation:

[0100] (0xF000 & denom) == 0 holds, so shift by 4. After shifting, nom = 4096 and denom = 20480;

[0101] The third iteration calculation:

[0102] (0xC000 & denom) == 0 does not hold, so no shift;

[0103] The fourth iteration calculation:

[0104] (0x8000 & denom) == 0 holds, shift by 1. After shifting, nom = 8192 and denom = 40960;

[0105] After 4 iteration calculations, the final nom = 8192 and denom = 40960 are obtained. Then, the result after shifting the final denom to the right by 4 bits is substituted into the preset LUT, and finally multiplied by the final nom, that is, 8192 × LUT(40960 shifted to the right by 4 bits) = 8192×25 = 204800;

[0106] Finally, the obtained 204800 is shifted to the right by 2×(N - Q) and the calculation result is output, that is, 204800 shifted to the right by 2×(16 - 12) bits = 800;

[0107] It can be known that the result in the floating-point representation with Q = 12 is: 800 / (2^(12)) = 0.1953125. It can be known that compared with the calculation of 1 / 5 = 0.2, the error is within the expected range.

[0108] It can be seen that the hardware solution for implementing division in the embodiments of the present application is designed based on an acceptable error range and supports division operations where the division result is within the decimal range. In this solution, there is no need to perform fixed-point to floating-point conversion, and all data is processed in the hardware format. Moreover, this solution has no complex conditions and hardware modules. Only a lookup table needs to be pre-planned based on the acceptable error range, and then the division operation can be completed by shifting and querying and mapping data in the lookup table.

[0109] Corresponding to the above Figure 1 , the embodiments of the present application provide a hardware-based division calculation device, as Figure 3 shown, the device includes:

[0110] An acquisition module 302, configured to acquire the dividend 1 and the divisor x, and a preset output data bit width Q, where the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and Q is less than or equal to N;

[0111] A first processing module 304, configured to set the initial value of nom to 1, the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, determine whether there is a valid value in the high N / 2 n bits of denom obtained from the previous iterative calculation;

[0112] A second processing module 306, configured to, if there is no valid value, shift the denom and nom obtained from the previous iterative calculation to the left by N / 2 n bits to obtain a new denom and nom;

[0113] A third processing module 308, configured to until the iterative calculation reaches N / 2 n = 1 to obtain the final denom and nom;

[0114] A fourth processing module 310, configured to shift the finally obtained denom to the right by (N - Q) bits to obtain denom>>(N - Q), substitute denom>>(N - Q) into a preset lookup table to obtain a lookup result LUT(denom>>(N - Q)), multiply the lookup result LUT(denom>>(N - Q)) by the finally obtained nom to obtain a multiplication result LUT_div, and after shifting the multiplication result LUT_div to the right by 2×(N - Q) bits, obtain the final calculation result and output it.

[0115] When the device according to the embodiment of the present application performs the division operation of the divisor 1 and the dividend x, first, the bit width N of the divisor x and the preset output data bit width Q are obtained. Then, based on the N and Q, iterative operations are performed on the divisor and the dividend for shifting. Next, the result of shifting the finally shifted divisor to the right by (N - Q) bits, denom>>(N - Q), is substituted into the preset look-up table LUT(denom>>(N - Q)), and the look-up result LUT(denom>>(N - Q)) is multiplied by the finally shifted dividend to obtain the multiplication result LUT_div. Finally, the multiplication result LUT_div is shifted to the right by 2*(N - Q) bits to obtain the final calculation result. The error between this calculation result and the actual division operation directly performed by the divisor and the dividend is within the expected range. That is, through the above method, the division result after floating-point to fixed-point conversion that finally conforms to the expected data representation format can be obtained. Moreover, in the process of obtaining the division calculation result through the above method in the embodiment of the present application, there is no need to perform the conversion from fixed-point to floating-point, and all data is processed in the format of the hardware, which improves the calculation efficiency of the hardware divider.

[0116] In an alternative embodiment of the embodiment of the present application, the second processing module in the embodiment of the present application includes: a first processing unit, configured to perform a bitwise AND operation on the target data and the denom obtained from the previous iterative calculation in the nth iterative operation; wherein, the bit width of the target data is N, and the high N / 2 n-1 bits of the target data are 1, and the remaining bits are 0; a second processing unit, configured to determine that there is no valid value in the high N / 2 n bits of the denom obtained from the previous iterative calculation when the result of the bitwise AND operation is 0; a third processing unit, configured to determine that there is a valid value in the high N / 2 n bits of the denom obtained from the previous iterative calculation when the result of the bitwise AND operation is not 0.

[0117] In an alternative embodiment of the embodiment of the present application, the look-up table in the embodiment of the present application stores the following corresponding relationship: , where takes values of 0, 1,... , N is the bit width of the divisor x, represents the value of

[0118] In an alternative embodiment of the embodiment of the present application, when the divisor x is 0, no iterative calculation is performed, and the calculation result 2 Q+1 - 1 is directly output.

[0119] Such as Figure 4As shown in the figure, an embodiment of the present application provides an electronic device, including a processor 411, a communication interface 412, a memory 413, and a communication bus 414. Among them, the processor 411, the communication interface 412, and the memory 413 complete mutual communication through the communication bus 414.

[0120] The memory 413 is used to store a computer program.

[0121] In an embodiment of the present application, when the processor 411 is used to execute the program stored on the memory 413, it implements the division calculation method based on hardware provided by any of the foregoing method embodiments, and its function is similar, so it will not be elaborated here.

[0122] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the division calculation method based on hardware provided by any of the foregoing method embodiments.

[0123] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0124] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0125] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0126] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A division calculation method implemented based on hardware, characterized in that, Including: Obtain a dividend 1, a divisor x, and a preset output data bit width Q, where the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and the Q is less than or equal to N; Let the initial value of nom be 1 and the initial value of denom be x, and perform the following iterative calculations: When performing the nth iterative calculation, determine whether there is still a valid value in the upper N / 2 n bits of denom obtained from the previous iterative calculation; if there is a valid value, do not update, keep the denom and nom obtained from the previous iterative calculation unchanged, and perform the next iterative calculation; if there is no valid value, shift the denom and nom obtained from the previous iterative calculation to the left by N / 2 n bits to obtain the new denom and nom; until the iterative calculation reaches N / 2 n = 1 to obtain the final denom and nom; Right-shift the finally obtained denom by (N - Q) bits to get denom>>(N - Q), substitute denom>>(N - Q) into a preset look-up table to obtain a look-up result LUT(denom>>(N - Q)), multiply the look-up result LUT(denom>>(N - Q)) by the finally obtained nom to get a multiplication result LUT_div, and after right-shifting the multiplication result LUT_div by 2×(N - Q) bits, obtain and output the final calculation result; Among them, judging whether there are still valid values in the high N / 2 n bits of the denom obtained in the previous iteration calculation includes: in the nth iteration operation, performing a bitwise AND operation on the target data and the denom obtained in the previous iteration calculation; wherein, the bit width of the target data is N, and the high N / 2 n bits of the target data are 1 and the remaining bits are 0; when the result of the bitwise AND operation is 0, it is determined that there are no valid values in the high N / 2 n bits of the denom obtained in the previous iteration calculation; when the result of the bitwise AND operation is not 0, it is determined that there are valid values in the high N / 2 n bits of the denom obtained in the previous iteration calculation.

2. The method according to claim 1, characterized in that, The look-up table stores the following corresponding relationship: , where takes values of 0, 1, … , N is the bit width of the divisor x represents the value rounded down.

3. The method according to claim 1, characterized in that, Including: When the divisor x is 0, no iterative calculation is performed, and the calculation result 2 is directly output Q+1 -1.

4. A division calculation device implemented based on hardware, characterized in that, Including: An obtaining module, configured to obtain a dividend 1, a divisor x, and a preset output data bit width Q, where the divisor x is a fixed-point number, the bit width of the divisor x is N, the output data bit width Q is determined based on the error range of the output data, and the Q is less than or equal to N; The first processing module is used to set the initial value of nom to 1, the initial value of denom to x, and perform the following iterative calculation: when performing the nth iterative calculation, it is judged whether there is a valid value in the upper N / 2 n bits of the denom obtained from the previous iterative calculation; if there is a valid value, no update is made, and the denom and nom obtained from the previous iterative calculation are kept unchanged, and the next iterative calculation is performed; if there is no valid value, the denom and nom obtained from the previous iterative calculation are shifted left by N / 2 n bits to obtain the new denom and nom; until the iterative calculation reaches N / 2 n = 1 to obtain the final denom and nom; A second processing module, configured to right-shift the finally obtained denom by (N - Q) bits to get denom>>(N - Q), substitute denom>>(N - Q) into a preset look-up table to obtain a look-up result LUT(denom>>(N - Q)), multiply the look-up result LUT(denom>>(N - Q)) by the finally obtained nom to get a multiplication result LUT_div, and after right-shifting the multiplication result LUT_div by 2×(N - Q) bits, obtain and output the final calculation result; Among them, the first processing module includes: a first processing unit, configured to perform a bitwise AND operation on the target data and the denom obtained from the previous iterative calculation during the nth iterative operation; wherein, the bit width of the target data is N, and the upper N / 2 n bits of the target data are 1, and the remaining bits are 0; a second processing unit, configured to determine that there is no valid value in the upper N / 2 n bits of the denom obtained from the previous iterative calculation when the result of the bitwise AND operation is 0; a third processing unit, configured to determine that there is a valid value in the upper N / 2 n bits of the denom obtained from the previous iterative calculation when the result of the bitwise AND operation is not 0.

5. A hardware device, characterized in that, The hardware device includes the hardware-implemented division calculation device described in claim 4.

6. An electronic device, comprising: At least one communication interface; At least one bus connected to the at least one communication interface; At least one processor connected to the at least one bus; At least one memory connected to the at least one bus, where the processor is configured to execute the hardware-implemented division calculation method described in any one of claims 1 to 3 above.

7. A computer storage medium storing computer-executable instructions for executing the hardware-implemented division calculation method described in any one of claims 1 to 3 above.

Citation Information

Patent Citations

  • Divider and quotient and remainder solving method

    CN106354473A

  • Arithmetic unit and operation method

    JP2003067182A