Method and device for realizing integer division operation
Through floating-point operation, the target magic number is generated and the efficiency and performance problems of integer division on GPUs in the prior art are solved, and efficient and low-cost integer division operation is realized.
Patent Information
- Application Number
- CN202510178455.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-20
AI Technical Summary
In the prior art, integer division implementation algorithms are not suitable for modern GPUs, resulting in low computing efficiency, high computing cost and poor performance.
The integer divisor is processed through floating point operation, and the target magic number is generated. Based on this, an underestimated approximate quotient and approximate remainder is generated, and the quotient and remainder are corrected according to the magnitude relationship between the remainder and the divisor, and the integer division operation is realized.
It greatly improves the computing efficiency of the divider, reduces the calculation cost, and improves the performance of integer division operations. It is suitable for modern GPUs.
Smart Images

Figure CN120179207A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a method and device for implementing integer division operations. Background Art
[0002] When large-scale parallel computing tasks need to be processed, multiple Graphics Processing Units (GPUs) are often used to jointly complete the tasks to improve the computing efficiency and speed. Multiple Graphics Processing Units are GPUs. In a GPU, when a divider implements a numerical division operation, it often iteratively subtracts the divisor from the dividend step by step and records the required number of iterations to obtain the quotient and the remainder. Compared with integer addition and integer multiplication, the computational cost of integer division is relatively high. Therefore, integer division is rarely directly implemented in hardware such as GPUs and generally must be simulated and implemented in software.
[0003] In the prior art, many different dividers have tried to find optimized methods for implementing integer division based on available hardware. However, the vast majority of these dividers are not applicable to modern GPUs and mainly have the following problems: (1) The efficiency of floating-point operations is not high; (2) Multiple format conversions are required, resulting in a relatively high computational cost; (3) Due to the need to execute different numbers of instructions or different instruction sets, the performance of implementing integer division operations is poor.
[0004] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and device for implementing integer division operations. The purpose is to implement integer division with a remainder based on floating-point reciprocals and minimized format conversions, greatly improving the computational efficiency of the divider, significantly reducing the computational cost, and improving the performance of implementing integer division operations, and solving the problem that the integer division implementation algorithms of the prior art are not applicable to GPUs.
[0006] The present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides a method for implementing integer division operations, including:
[0008] Processing an integer divisor of a division operation based on floating-point operations to obtain a target magic number; wherein, the target magic number is an integer;
[0009] Generating an underestimated approximate quotient of the division operation according to the target magic number; and obtaining an approximate remainder according to the underestimated approximate quotient;
[0010] According to the magnitude relationship between the approximate remainder and the integer divisor, correct the underestimated approximate quotient and the approximate remainder to obtain the target quotient and the target remainder of the division operation.
[0011] Further, the processing of the integer divisor of the division operation based on floating-point operations to obtain the target magic number includes:
[0012] Determine the overestimated floating-point representation of the integer divisor, and determine the underestimated floating-point representation of the reciprocal of the overestimated floating-point representation;
[0013] According to the underestimated floating-point representation and the maximum positive calculation error of the reciprocal operation of the computer hardware, obtain the floating-point magic number;
[0014] Convert the floating-point magic number to an integer according to the first preset rounding mode to obtain the target magic number.
[0015] Further, the determining the overestimated floating-point representation of the integer divisor and determining the underestimated floating-point representation of the reciprocal of the overestimated floating-point representation includes:
[0016] Convert the integer divisor to the overestimated floating-point representation according to the second preset rounding mode;
[0017] Determine the reciprocal of the overestimated floating-point representation;
[0018] Perform rounding processing on the reciprocal to obtain the corresponding underestimated floating-point representation.
[0019] Further, the converting the integer divisor to the overestimated floating-point representation according to the second preset rounding mode includes:
[0020] When the second preset rounding mode is rounding towards positive infinity, convert the integer divisor from integer type to floating-point type to obtain the overestimated floating-point representation;
[0021] When the second preset rounding mode is rounding to the nearest integer or rounding towards 0, convert the integer divisor from integer type to floating-point type to obtain an intermediate floating-point representation; determine the sum of the intermediate floating-point representation and the minimum precision error unit as the overestimated floating-point representation to ensure that the overestimated floating-point representation is greater than or equal to the integer divisor.
[0022] Further, the obtaining the floating-point magic number according to the underestimated floating-point representation and the maximum positive calculation error of the reciprocal operation includes:
[0023] Before processing the integer divisor of the division operation based on floating-point operations, pre-determine the difference between the bitwise operation value and the maximum positive calculation error as the constant multiplication substitution value;
[0024] Regarding both the underestimated floating-point representation and the constant multiplication substitution value as corresponding integer representations, perform bitwise integer addition on the underestimated floating-point representation and the constant multiplication substitution value to obtain a floating-point magic number.
[0025] Further, generating the underestimated approximate quotient of the division operation according to the target magic number; obtaining the approximate remainder according to the underestimated approximate quotient includes:
[0026] Determine the high-order product of the integer dividend of the division operation and the target magic number as the first iteration result;
[0027] Determine the product of the integer divisor and the first iteration result as the first intermediate value, and determine the difference between the integer dividend and the first intermediate value as the first remainder corresponding to the first iteration result;
[0028] Determine the high-order product of the first remainder and the target magic number as the second iteration result;
[0029] Determine the sum of the first iteration result and the second iteration result as the underestimated approximate quotient;
[0030] Determine the product of the integer divisor and the underestimated approximate quotient as the second intermediate value, and determine the difference between the integer dividend and the second intermediate value as the approximate remainder corresponding to the underestimated approximate quotient.
[0031] Further, when the integer dividend and the integer divisor of the division operation are unsigned integers, correcting the underestimated approximate quotient and the approximate remainder according to the size relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder of the division operation includes:
[0032] When the approximate remainder is greater than or equal to the integer divisor, add one to the underestimated approximate quotient to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the remainder to be checked; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the target quotient; use the approximate remainder as the target remainder;
[0033] When the remainder to be checked is greater than or equal to the integer divisor, add one to the quotient to be checked to obtain the target quotient, and determine the difference between the remainder to be checked and the integer divisor as the target remainder; when the remainder to be checked is less than the integer divisor, use the quotient to be checked as the target quotient, and use the remainder to be checked as the target remainder.
[0034] Further, when the integer dividend and the integer divisor are signed integers, correcting the underestimated approximate quotient and the approximate remainder according to the size relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder of the division operation includes:
[0035] Perform an exclusive OR logical operation on the integer divisor and the integer dividend to obtain an exclusive OR result;
[0036] When the approximate remainder is greater than or equal to the integer divisor, increment the underestimated approximate quotient by one to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the target remainder; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the quotient to be checked; use the approximate remainder as the target remainder;
[0037] When the exclusive OR result is less than 0, determine the negative of the quotient to be checked as the target quotient; when the exclusive OR result is greater than or equal to 0, determine the quotient to be checked as the target quotient.
[0038] In a second aspect, the present invention further provides a device for implementing integer division operations, which is used to implement the method for implementing integer division operations described in the first aspect;
[0039] In a possible implementation, the device for implementing integer division operations includes:
[0040] The device for implementing integer division operations is used to perform a division operation of an integer dividend by an integer divisor to determine a target quotient and a target remainder as the result of the division operation;
[0041] The device for implementing integer division operations includes a magic number generator, a division iterator, and an approximation corrector, where:
[0042] The magic number generator is used to process the integer divisor based on floating-point operations to obtain an integer target magic number;
[0043] The division iterator is used to generate an underestimated approximate quotient of the division operation according to the target magic number; and is also used to obtain an approximate remainder according to the underestimated approximate quotient;
[0044] The approximation corrector is used to correct the underestimated approximate quotient and the approximate remainder according to the size relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder.
[0045] In another possible implementation, the device for implementing integer division operations includes:
[0046] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to perform the method for implementing integer division operations described in the first aspect.
[0047] In a third aspect, the present invention also provides a non-volatile computer storage medium storing computer-executable instructions that, when executed by one or more processors, are used to implement the method for integer division operation described in the first aspect.
[0048] In a fourth aspect, there is provided a chip including a processor and an interface for calling and running a computer program stored in a memory to execute the method for integer division operation as described in the first aspect.
[0049] In a fifth aspect, there is provided a computer program product containing instructions that, when run on a computer or a processor, cause the computer or the processor to execute the method for integer division operation as described in any one of the first to fourth aspects.
[0050] In a sixth aspect, there is provided a system for implementing integer division operation, including the device for implementing integer division operation as described in the second aspect, and using the method for implementing integer division operation as described in the first aspect to complete the interaction of the device for implementing integer division operation as described in the second aspect.
[0051] Different from the prior art, the present invention has at least the following beneficial effects:
[0052] The present invention processes the integer divisor of the division operation through floating-point operations to obtain an integer target magic number, so as to greatly improve the calculation efficiency and hardware friendliness. And since only the integer divisor is processed, there is no need to perform multiple format conversions between integer and floating-point types, greatly reducing the calculation cost. Then, based on the target magic number, an underestimated approximate quotient of the division operation is generated, and an approximate remainder of the underestimated approximate quotient is obtained. Finally, according to the size relationship between the approximate remainder and the integer divisor, the underestimated approximate quotient and the approximate remainder are corrected to obtain the corresponding target quotient and target remainder. Since a predicate expression can be used to correct the underestimated approximate quotient and the approximate remainder based on the size relationship, there will be no branches in the control flow, greatly improving the performance of integer division operation. The method for implementing integer division operation provided by the present invention can be completed by a fixed number of instructions and is suitable for parallel processing on computer graphics processing hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0054] Figure 1 It is a flowchart showing a method for implementing integer division operation provided by an embodiment of the present invention;
[0055] Figure 2 It is a schematic flowchart of step 10 provided by an embodiment of the present invention;
[0056] Figure 3 It is a schematic diagram of rounding towards 0 provided by an embodiment of the present invention;
[0057] Figure 4 It is a schematic diagram of additional processing for other rounding modes provided by an embodiment of the present invention;
[0058] Figure 5 It is a schematic flowchart of step 101 provided by an embodiment of the present invention;
[0059] Figure 6 It is a schematic flowchart of step 1011 provided by an embodiment of the present invention;
[0060] Figure 7 It is a schematic diagram for explaining the ULP concept provided by an embodiment of the present invention;
[0061] Figure 8 It is a schematic flowchart of step 102 provided by an embodiment of the present invention;
[0062] Figure 9 It is a schematic diagram of a floating - point format provided by an embodiment of the present invention;
[0063] Figure 10 It is a schematic diagram of addition provided by an embodiment of the present invention;
[0064] Figure 11 It is a schematic flowchart of step 20 provided by an embodiment of the present invention;
[0065] Figure 12 It is a schematic flowchart of step 30 provided by an embodiment of the present invention;
[0066] Figure 13 It is another schematic flowchart of step 30 provided by an embodiment of the present invention;
[0067] Figure 14 It is a schematic diagram of the architecture of a device for implementing integer division operations provided by an embodiment of the present invention. Detailed implementation manners
[0068] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0069] In the description of the present invention, the orientation or positional relationship indicated by terms such as "inner", "outer", "longitudinal", "transverse", "upper", "lower", "top", "bottom", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention rather than requiring the present invention to be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0070] In the present invention, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0071] In the present application, unless otherwise clearly specified and defined, the term "connection" should be understood in a broad sense. For example, "connection" can be a fixed connection, a detachable connection, or integrated as a whole; it can be directly connected or indirectly connected through an intermediate medium. In addition, the term "coupling" can be a way of realizing electrical connection for signal transmission.
[0072] The division of units in the present invention is only a logical division. In actual implementation, there may be other division methods. For example, multiple units can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection between modules can be electrical or other similar forms, which are not limited herein. And the units or subunits described as separate components may or may not be physically separated, may or may not be physical modules, or may be divided into multiple modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present invention.
[0073] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0074] There are many dividers based on different integer division implementation algorithms in the prior art, but they are not applicable to modern GPUs. The problems existing in the prior art are analyzed below based on the integer division implementation algorithms of the following three dividers:
[0075] (1) Division with a known divisor at compile time
[0076] If the divisor D is known at compile time, integer division can be converted into two steps: (1) multiplying the divisor by its corresponding integer magic number; (2) shifting the result of the multiplication to the right by a specific number of bits.
[0077] Among them, the magic number refers to the specific value directly written in the code in programming (such as the values directly written as numbers like "10", "123", etc.).
[0078] Over time, various algorithms for optimizing constant division have emerged. One popular algorithm, libdivide, is a C / C++ implementation of optimized constant division; in most Central Processing Unit (CPU) architectures, only 2 or 3 instructions are required to implement the division operation. Its main idea is to calculate using the following expression:
[0079]
[0080] Among them, using 2 C Performing an integer division on other dividends is equivalent to performing a right shift of C bits on the dividend, and can be pre-calculated. In the case where the divisor D is a constant, the expression for the magic number M is as follows:
[0081]
[0082] Among them, the symbol ":=" means "defined as"; P represents the precision of the magic number M, and N represents the number of bits parameter of the register for performing the corresponding calculation; when it is determined that M is applicable to an N-bit register, the value of P usually needs to be as large as possible. The value of M can be rounded up or down, and the specific rounding method depends on the value of the divisor D.
[0083] (2) Division algorithms based on non-constant running time
[0084] Many division algorithms are based on the numerical difference between X and Y and require logarithmic running time. Although these division algorithms work well when calculating arbitrary large integer values, they are not suitable for GPUs because GPUs use the Single Instruction Multiple Threads (SIMT) execution mode, and in this parallel execution mode, ideally, the execution time of the same instruction for different data needs to be as the same as possible.
[0085] GPUs typically use "single instruction, multiple threads" to perform calculations. For algorithms with variable numbers of executed instructions and instruction execution branches, depending on the input of the algorithm, different numbers of instructions or different instruction sets are often required. When using "single instruction, multiple threads" to execute such division algorithms, different threads may use different program branches, which can cause other threads to stall and no calculations can be performed until another alternative branch has completed execution. This branch divergence problem caused by such algorithms is a major performance issue in GPU programming.
[0086] To solve the above problems, as Figure 1 shown, an embodiment of the present invention provides a method for implementing integer division operations, including:
[0087] Step 10: Process the integer divisor of the division operation based on floating-point operations to obtain a target magic number; wherein, the target magic number is an integer.
[0088] The following explains relevant concepts in division operations:
[0089] When calculating "X / Y", that is, "X divided by Y", Y is called the divisor, X is called the dividend, and the result of the division is called the quotient.
[0090] When both X and Y are unsigned integers of N bits (bit), that is, the values of X and Y are stored in N bits and can represent any integer within {0, 1, 2,..., 2 N -1}, then the unsigned integer is usually defined as:
[0091] divide_uint:=floor(X / Y)
[0092] wherein, the unsigned integer can be a non-negative integer. floor(Z) represents rounding down the integer Z. For example, floor(3.1415) = 3, floor(128) = 128, floor(-7.5) = -8.
[0093] Regarding unsigned integer division, the corresponding remainder refers to the number remaining at the end of the division. If Q is the result of "X divided by Y", that is, the quotient, and F is the remainder of "X divided by Y", then this relationship can be recorded as:
[0094] Q:=divide_uint(X,Y), F:=X - Q×Y
[0095] In the embodiment of the present invention, when the divisor is not known during compilation, first, a target magic number suitable for 32-bit integers is determined based on the floating-point representation of the integer divisor Y. Then, based on this target magic number, the integer division of the integer X and the integer Y is implemented according to the following expression:
[0096]
[0097] where M represents the target magic number,
[0098] It can be implemented in the GPU as "(X * M) >> 32".
[0099] Step 20: Generate an underestimated approximate quotient for the division operation according to the target magic number; obtain an approximate remainder according to the underestimated approximate quotient.
[0100] In the embodiment of the present invention, it is not only applicable to 32-bit unsigned integers, but also applicable to 32-bit signed integers after certain adjustments. The specific adjustment method will be described below.
[0101] Step 30: Correct the underestimated approximate quotient and the approximate remainder according to the magnitude relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder of the division operation.
[0102] Many GPUs support predicate registers, that is, processing instructions conditionally using bit masks; among them, a predicate register is a special type of processor register used to store boolean values (usually True or False) to control conditional execution in the program flow. Predicates can usually be set by a module to perform comparisons between two or more values; for example, the predicate P0: "Add 5 to H, and then multiply the corresponding sum by 2 when the corresponding sum is greater than 10" can be implemented by the following two instructions (i.e., predicate expressions):
[0103] H := H + 5; P0 := H > 10
[0104] IF (P0) THEN H := H × 2
[0105] Since the process of correcting the underestimated approximate quotient and the approximate remainder can be repeatedly executed based on the predicate expression, avoiding the occurrence of branches in the control flow, it can solve the problem of low performance of integer division operations caused by the need to execute different numbers of instructions or different instruction sets.
[0106] The present invention processes the integer divisor of a division operation through floating-point operations to obtain an integer target magic number, thereby greatly improving the calculation efficiency and hardware friendliness. Moreover, since only the integer divisor is processed and there is no need to perform multiple format conversions between integers and floating-point numbers, the calculation cost is greatly reduced. Furthermore, based on the target magic number, an underestimated approximate quotient of the division operation is generated, and an approximate remainder of the underestimated approximate quotient is obtained. Finally, according to the size relationship between the approximate remainder and the integer divisor, the underestimated approximate quotient and the approximate remainder are corrected to obtain the corresponding target quotient and target remainder. Since the correction of the underestimated approximate quotient and the approximate remainder based on the size relationship can be implemented using a predicate expression, no branches will appear in the control flow, greatly improving the performance of integer division operations. The method for implementing integer division operations provided by the present invention can be completed by a fixed number of instructions and is suitable for parallel processing on computer graphics processing hardware.
[0107] To illustrate the process of determining the target magic number, as Figure 2 shown, step 10 includes:
[0108] Step 101: Determine the overestimated floating-point representation of the integer divisor, and determine the underestimated floating-point representation of the reciprocal of the overestimated floating-point representation.
[0109] Because the magic number needs to adapt to 32-bit integers, and the integer divisor Y in the embodiments of the present invention is unknown at compile time, so determining a target magic number that is strictly underestimated i.e., a magic number lower than its actual value, will ensure that the obtained approximate quotient is less than or equal to the exact quotient Since correcting "one-sided error" is easier than correcting an error that may be positive or negative, the one-sided error also allows the algorithm to more effectively calculate the remainder of the division and set the necessary predicates in the algorithm. Because unsigned integer division is usually defined by rounding the exact result of the division, it is more preferable to use an underestimated target magic number.
[0110] The goal of the embodiments of the present invention is to obtain a floating-point number slightly larger than the integer divisor Y, and the floating-point number of the integer divisor Y is: ideally, the smallest representable floating-point number greater than the integer divisor Y.
[0111] Step 102: Obtain a floating-point magic number according to the underestimated floating-point representation and the maximum positive calculation error of the reciprocal operation of the computer hardware.
[0112] The maximum positive calculation error is the maximum positive calculation error when a specific operation is implemented by GPU hardware and is independent of the integer dividend X and the integer divisor Y; in a specific example, the maximum positive calculation error is obtained from the specification of the hardware implementation of the recip function used.
[0113] Step 103: Convert the floating-point magic number into an integer according to the first preset rounding mode to obtain a target magic number.
[0114] Among them, the first preset rounding mode is selected by those skilled in the art according to the usage scenario (for example, the rounding mode supported by the hardware).
[0115] In a preferred embodiment, the first preset mode is rounding towards zero, that is, using the truncation method, and the target magic number M is:
[0116] M = uint(m)
[0117] Among them, the uint function is used to perform the conversion operation from a floating-point number to an integer according to the rounding mode of rounding towards zero. As Figure 3 shown, since the calculated m is a positive floating-point number, M is less than or equal to m.
[0118] In an alternative embodiment, when the first preset rounding mode is not rounding towards zero, after converting the floating-point magic number into an integer according to the corresponding first preset rounding mode, subtract one from the obtained integer as the target magic number to ensure that the target magic number is not the approximation of (that is, the target magic number can be directly equal to ). As Figure 4 shown, subtract one from the integer G obtained by converting the floating-point magic number according to the corresponding first preset rounding mode, so that M is less than or equal to m.
[0119] To illustrate the process of obtaining the underestimated floating-point representation and the corresponding maximum positive error, as Figure 5 shown, the step 101 includes:
[0120] Step 1011: Convert the integer divisor into the overestimated floating-point representation according to the second preset rounding mode.
[0121] Among them, the second preset rounding mode is selected by those skilled in the art according to the usage scenario (for example, the rounding mode supported by the hardware); in a preferred embodiment, the second preset mode is rounding towards positive infinity.
[0122] Step 1012: Determine the reciprocal of the overestimated floating-point representation.
[0123] The reciprocal of the overestimated floating-point representation is
[0124] Step 1013: Perform rounding processing on the reciprocal to obtain the corresponding underestimated floating-point representation.
[0125] Embodiments of the present invention use the recip function to obtain a rounded approximation of the reciprocal of an overestimated floating-point representation. Denote the reciprocal of the floating-point representation of Y as R, and its expression is:
[0126] R = recip(float(Y))
[0127] where R is also a floating-point number, and recip() represents the recip function. The rounding mode used by the reciprocal function (i.e., the recip function) is not important because different rounding modes do not increase any computational cost; however, the maximum positive computational error of the reciprocal function must be determined.
[0128] The following will specifically illustrate this:
[0129] The recip function is a floating-point function (rather than an integer-based reciprocal function); for a floating-point number F, typically, the value cannot be accurately represented as a floating-point number. When this occurs, the output result of "recip(F)" must necessarily give an answer, thereby introducing some form of error from the true value. If the error of a "recip(F)" is less than 0.5 ULP (i.e., higher or lower than 0.5 ULP), the output result is always rounded to the nearest representable floating-point number to the exact result.
[0130] The maximum positive computational error refers to how much larger the output result is than the exact result. Taking the recip function as an example, the maximum positive computational error when calculating the reciprocal of a floating-point number F is equal to where the maximum positive computational error can be an absolute value, but more commonly, it is represented using ULP, that is, "the number of ULPs exceeding is recip(F)".
[0131] It should be noted that in hardware, discussing the ULPs of fractions has no meaning, and it can only handle the ULPs of integers. However, in floating-point error calculations, dealing with the ULPs of fractions is very useful; the maximum positive computational error is the maximum positive error that can be obtained among all possible inputs of the recip function (i.e., the worst error case). For any operation implemented in hardware, the maximum positive computational error is usually known, and thus it is a known fixed value for those writing programs using this hardware function. Similarly, if the hardware does not directly support the recip function but is constructed from more primitive floating-point functions to achieve the corresponding function, the maximum positive computational error can be calculated analytically or approximated as accurately as possible.
[0132] To illustrate the process of obtaining an overestimated floating-point representation, as Figure 6As shown, step 1011 includes:
[0133] Step 10111: When the second preset rounding mode is rounding towards positive infinity, convert the integer divisor from integer type to floating-point type to obtain the overestimated floating-point representation.
[0134] Step 10112: When the second preset rounding mode is rounding to the nearest integer or rounding towards zero, convert the integer divisor from integer type to floating-point type to obtain an intermediate floating-point representation; determine the sum of the intermediate floating-point representation and the minimum precision error unit as the overestimated floating-point representation to ensure that the overestimated floating-point representation is greater than or equal to the integer divisor.
[0135] Wherein, when the rounding mode is rounding to the nearest integer or rounding towards zero, the value obtained by converting the integer type to floating-point type (i.e., the intermediate floating-point representation) is added with 1 ULP (Unit in the Last Place, the minimum precision error unit) as the result of converting the integer type to floating-point type (i.e., the overestimated floating-point representation).
[0136] Embodiments of the present invention can use rounding towards positive infinity, or can use other different rounding modes (for example, rounding to the nearest integer or rounding towards zero); however, when using other different rounding modes, it is necessary to increase the minimum precision error unit, that is, increase 1 ULP, to ensure that the overestimated floating-point representation of the integer divisor Y is greater than or equal to the integer divisor Y, that is, to achieve the following expression:
[0137] float(Y)≥Y
[0138] The following describes ULP. The definition of ULP for a single real value is: for a given floating-point format, the ULP of a specific real value is the distance between the two closest floating-point numbers to this real value.
[0139] Taking the 32-bit floating-point number in the IEEE754 standard as an example, this format cannot accurately represent the real value 0.1, and can only use the closest floating-point number that can be accurately represented to represent 0.1; there are two numbers closest to 0.1, denoted as A and B. The hexadecimal representation of A is: 0x3dcccccc; the decimal representation of A: 0.099999994039536; the hexadecimal representation of B is: 0x3dcccccd; the decimal representation of B is: 0.10000000149012. Then the ULP of 0.1 is: |A - B| = 0.00000000745076.
[0140] If the downward approximation method is used and A is used to represent 0.1, then the error is:
[0141] 0.1 - 0.099999994039536 = 0.000000005960464 =
[0142] (0.000000005960464 / 0.00000000745076) * ULP ≈ 0.8 ULPs.
[0143] If the upward approximation method is used and B is used to represent 0.1, then the error is:
[0144] 0.10000000149012 - 0.1 = 0.00000000149012 = (0.00000000149012 / 0.00000000745076) * ULP ≈ 0.2 ULPs.
[0145] The calculation of the ULP corresponding to the function will be explained below:
[0146] All functions discussed in the following explanations are functions that return floating-point numbers, because integer calculations do not have ULP errors.
[0147] Definition: When discussing the overall accuracy of a function rather than a specific real value, the number of ulps cited is the worst-case error for any parameter. If the error of a function is always less than 0.5 ulps, then the function always returns a floating-point number closest to the exact result, and such a function is correctly rounded.
[0148] When evaluating the accuracy of a function, the ULP of a single value cannot be used to represent the overall accuracy of the entire function, so it is necessary to define the overall accuracy of a function. ULPs is the plural form of ULP. 0.8 ULPs can be regarded as 0.8 ULPs, and here 0.8 is the number of ULPs.
[0149] For example, when the output error generated by a certain function for any input is not greater than 0.5 ULPs, then the worst-case error for any parameter of this function, that is, the number of ULPs cited by this function is 0.5 ULPs. If the error of a function is always less than 0.5 ULPs, then the function always returns a floating-point number closest to the exact result, and the function is correctly rounded. When the error of a certain function is sometimes greater than 0.5 ULPs, then the function is a wrongly rounded function. The reason for using 0.5 ULPs as the standard for judging the accuracy of a function is as follows: As Figure 7 shown, let the exact solution of the function be set as: res, the closest floating-point number to the left of the exact solution be set as: res 左 , the closest floating-point number to the right of the exact solution be set as: res 右 , then the error is minimized as much as possible, specifically as follows:
[0150] error min= min(|res - res 左 |, |res - res 右 |)
[0151] In an embodiment of the present invention, since an integer is represented as a fixed-point number in hardware, that is, the difference between an integer Z and the smallest integer greater than Z is always a fixed fixed-point number (this fixed fixed-point number is 1). If the prefix "0x" is used to represent the value in a hexadecimal register, then the integer "42" will be represented as a 32-bit integer 0x2A; it should be noted that the leading zeros of the representation "0x2A" are omitted, and it should actually be "0x0000002A".
[0152] Adding "1 ULP" to an integer requires increasing the value of the register by 0x1. For example:
[0153] 42 + '1 ULP' = 0x2A + 0x1 = 0x2B = 43 = 42 + 1
[0154] Among them, for an integer, this example is easy to understand because the absolute value of "1 ULP" is always 1. For a floating-point number, the difference between a floating-point number F and the smallest integer greater than the floating-point number F is no longer a fixed number, and this difference depends on the magnitude of the floating-point number F and the floating-point representation used for the floating-point number F. For example, in an embodiment of the present invention, the floating-point representation adopts the IEEE-754 standard. Representing 42.0 as a 32-bit floating-point number is 42.0 = 0x42280000, representing 84.0 as a 32-bit floating-point number is 84.0 = 0x42A80000. Adding "1 ULP" to these two numbers, the results are as follows:
[0155] 42.0 + '1 ULP' = 0x42280000 + 0x1 = 0x42280001
[0156] = 42.000003814697265625
[0157] 84.0 + '1 ULP' = 0x42A80000 + 0x1 = 0x42A80001
[0158] = 84.000007629394531250
[0159] Among them, the increase of "1 ULP" increases the floating-point number 42.0 by "3.81e-6", but the increase of "1 ULP" increases the floating-point number 84.0 by "7.63e-6". Since "ULPs" can describe the smallest difference in the floating-point representation (that is, the difference between a floating-point number F and the smallest integer greater than the floating-point number F), it is therefore used to describe the relative error that may occur in floating-point operations.
[0160] Based on the above considerations, in the embodiments of the present invention, the result that an accurate recip function may produce is a result that is higher or lower than the actual value by 0.5 ULP. Therefore, the parameter E in the following text is equal to 0.5.
[0161] To illustrate the process of calculating the floating-point magic number, as Figure 8 shown, step 102 includes:
[0162] Step 1021: Before processing the integer divisor of the division operation based on floating-point arithmetic, the difference between the bitwise operation value and the maximum positive calculation error is determined in advance as the constant multiplication substitution value.
[0163] It should be noted that the constant multiplication substitution value does not depend on a specific divisor. It only depends on the maximum positive calculation error of the recip function. Usually, when implemented in hardware, the corresponding maximum positive calculation error is known (usually given by ULPs).
[0164] Step 1022: Regarding both the underestimated floating-point representation and the constant multiplication substitution value as the corresponding integer representations, perform a bitwise integer addition on the underestimated floating-point representation and the constant multiplication substitution value to obtain the floating-point magic number.
[0165] Among them, the bitwise operation value is 0xFFFFFFF; the constant multiplication substitution value is A, and its expression is:
[0166] A = 0xFFFFFFF - ceil(E)
[0167] Among them, ceil(Z) represents rounding up the integer Z. For example, ceil(3.1415) = 4, ceil(128) = 128, ceil(-7.5) = -7. ceil(E) is the maximum positive calculation error of the computer hardware reciprocal operation. Therefore, A does not need to be calculated when performing integer division, and A is also a known constant.
[0168] B + 0xFFFFFFF = B × (2 32 -1)
[0169] In the IEEE-754 floating-point representation, multiplying a floating-point number by 2 32 is equivalent to adding 32 to the exponent. R multiplied by 2 32 can be expressed as R + 0x10000000. Since 0x10000000 - 1 ULP = 0x0FFFFFFF, so R + 0x0FFFFFFF = R × 2 32 -1 ULP.
[0170] The following is an explanation of the floating-point numbers under the IEEE754 standard:
[0171] As shown Figure 9 in the figure, a 32-bit single-precision floating-point number consists of 1 sign bit, 8 exponent bits, and 23 mantissa bits in sequence. Sign is the sign bit and is arranged at the highest bit; among them, Sign = 0 represents a positive number, and Sign = 1 represents a negative number. Mantissa is the mantissa and is arranged at the lowest bit. The position of the decimal point of the floating-point number is to the right of the highest (i.e., the leftmost) significant bit in the mantissa field. Exponent is the exponent. The sign of the exponent is represented in an implicit way, that is, the offset method is used to represent positive and negative exponents; when using this method, when converting the true value e of the exponent of the floating-point number into the exponent Exponent, a fixed offset value 127 (i.e., 01111111 in binary) should be assumed for the exponent e, and the expression is Exponent = e + 127. The true value representation of a 32-bit floating-point number f is as follows:
[0172] f = (-1) sign × (1.Mantissa) × 2 Exponent-127
[0173] where the value represented by the mantissa field is "1.Mantissa"; since the highest significant bit of the mantissa field is always 1, this bit is often not stored and is considered to be hidden in the coordinate of the decimal point. Therefore, 23-bit fields can store 24-bit significant digits.
[0174] For single-precision IEEE-754 floating-point numbers, it is equivalent to performing an integer addition of 0x10000000 on the register containing the floating-point number bits. "Increasing or decreasing the value of the floating-point number by 1 ULP" corresponds to "adding or subtracting 0x1 from the bits of the floating-point number". So there is "0x10000000 - 0x1 = 0xFFFFFFF", that is, "adding 0xFFFFFFF to the bits of the floating-point number", which is equal to "multiplying the floating-point number by 2 32 and then subtracting 1 ULP". It should be noted that the premise of this method is: assuming that there is no risk of overflow or underflow errors for the floating-point number, and it will not enter the NaN range of the floating-point number; however, this will never happen because in the algorithm of the present invention, the value of the floating-point number is always within the range between 0 and 2 32 .
[0175] According to the bit arrangement method of floating-point numbers under the IEEE-754 standard, by performing only one integer addition operation, R can be multiplied by 2 32 , and subtracting ULPs to offset the positive error of the recip function, achieving a significant improvement in calculation efficiency. Specifically, adding 32 to the exponent part of the floating-point number is equivalent to multiplying the value of the floating-point number by 2 32 . Regarding the floating-point number R in the 32-bit register as an integer and adding A to this integer can achieve the effect of "multiplying R by 2 32 " (that is, achieving "multiplying R by 2 32”, and obtain its equivalent alternative). Among them, 1 ULP needs to be subtracted to ensure that the obtained result is a kind of underestimated approximation; then subtract “ceil(E)” ULPs. Subtracting “ceil(E)” here can ensure that any overestimation caused by the error in the recip function is deleted; the above two subtraction operations are implemented based on the expression of A.
[0176] In an alternative embodiment, the expression of the floating-point magic number m is as follows:
[0177] m = R + A
[0178] Among them, R is the underestimated floating-point representation, and A is the constant multiplication substitution value. As Figure 10 shown, the “+” here means performing an integer addition of bits on the two, that is, performing an integer addition on the corresponding bits in the underestimated floating-point representation (floating-point number) and the constant multiplication substitution value (floating-point number). The finally obtained floating-point magic number m is also a floating-point number.
[0179] To illustrate the method of performing division based on the target magic number in the embodiments of the present invention, as Figure 11 shown, the step 20 includes:
[0180] Step 201: Determine the high-order product of the integer dividend of the division operation and the target magic number as the first iteration result.
[0181] Among them, the high-order product is the result obtained by performing high-order multiplication on the integer dividend and the target magic number; the high-order multiplication of integers means: for two N-bit integers X and Y, their non-overflow product X×Y is a 2N-bit integer. For example, the maximum value that an unsigned 32-bit integer can represent is 2 32 -1, so the maximum value of the non-overflow product of two unsigned 32-bit integers is (2 32 -1) 2 = 2 64 -2 33 +1, which is a 64-bit integer.
[0182] In an alternative embodiment, the high-order multiplication is equivalent to calculating a non-overflow product and then performing an N-bit right shift on it, as shown in the following formula:
[0183] hmul(X, Y) = (X×Y) >> N
[0184] It should be noted that since many bits are immediately discarded in this process, for a product, performing an N-bit right shift on it is not the most efficient method. CPUs and GPUs usually support the ability to directly perform high-order multiplication of 32-bit integers in hardware. Therefore, it is usually not necessary to perform high-order multiplication and right shift separately in the program.
[0185] The embodiments of the present invention use a target magic number to perform a downward approximation on the result of unsigned integer division. The specific expression is as follows:
[0186] divide_uint32(X,Y)≈hmul(X,M)
[0187] Among them, for 32-bit division, "hmul(X,Y)" calculates 32 important preset bits of the product of X and M, that is, "(X*M)>>32", where ">>32" represents a right shift of 32 bits.
[0188] The downward approximation is obtained by performing two division iterations. The goal of the first iteration is to conservatively remove as many Y factors as possible from X using M; because So where the data processing unit quantity is 2 32 , the first division intermediate value is The first iteration result is the corresponding value. Denote the first iteration result as Q1, and use the hmul function to implement this process. The input of the hmul function is the integer dividend X and the target magic number M. Then the specific expression of this process is as follows:
[0189] Q1 = hmul(X,M)
[0190] Step 202: Determine the product of the integer divisor and the first iteration result as the first intermediate value, and determine the difference between the integer dividend and the first intermediate value as the first remainder corresponding to the first iteration result.
[0191] Define the first remainder as L1, and the specific expression is as follows:
[0192] L1 = X - Y×Q1
[0193] Among them, Y×Q1 is the first intermediate value.
[0194] Step 203: Determine the high-order product of the first remainder and the target magic number as the second iteration result.
[0195] The second iteration of the embodiments of the present invention is to use the target magic number to divide the first remainder L1 of the first iteration, aiming to ensure that the finally obtained quotient is at most only 1 less than the actual result.
[0196] Denote the second iteration result as Q2, and use the hmul function to implement this process. The input of the hmul function is the first remainder L1 and the target magic number M. Then the specific expression of this process is as follows:
[0197] Q2 = hmul(L1,M)
[0198] Step 204: Determine the sum of the first iteration result and the second iteration result as the underestimated approximate quotient.
[0199] The underestimated approximate quotient corresponding to the final obtained division operation is denoted as Q, and its expression is:
[0200] Q = Q1 + Q2
[0201] The approximate division in the embodiment of the present invention ensures that when the difference in the numerical magnitudes of X and Y is less than 20 bits (i.e., conforming to the following expression), the range lower than the actual result of the division is within 1.
[0202]
[0203] Step 205: Determine the product of the integer divisor and the underestimated approximate quotient as the second intermediate value, and determine the difference between the integer dividend and the second intermediate value as the approximate remainder corresponding to the underestimated approximate quotient.
[0204] The underestimated approximate quotient obtained according to the foregoing steps is at most 1 or 2 less than the actual result of divide_uint32(X, Y). To correct the underestimated approximate quotient Q, calculate the approximate remainder L, and its expression is:
[0205] L = X - Y × Q
[0206] Since the divisor is not known during compilation, the foregoing process all uses the underestimated approximation method. To correct the obtained underestimated approximate quotient, as Figure 12 shown, when the integer dividend and the integer divisor of the division operation are unsigned integers, the step 30 includes:
[0207] Step 301a: When the approximate remainder is greater than or equal to the integer divisor, add 1 to the underestimated approximate quotient to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the remainder to be checked; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the target quotient; use the approximate remainder as the target remainder.
[0208] Step 302a: When the remainder to be checked is greater than or equal to the integer divisor, add 1 to the quotient to be checked to obtain the target quotient, and determine the difference between the remainder to be checked and the integer divisor as the target remainder; when the remainder to be checked is less than the integer divisor, use the quotient to be checked as the target quotient, and use the remainder to be checked as the target remainder.
[0209] Finally, to correct the situation where the "underestimated approximate quotient is lower than the exact quotient of the integer division by 2" may occur, it is also necessary to perform repeated checks according to the method of the above step 301a.
[0210] It should be noted that when the quotient of the division operation only requires a signed or unsigned integer type, there is no need to calculate the target remainder.
[0211] After completing the above steps, at this time, divide_uint32(X, Y) = Q, that is, the obtained target quotient is the accurate result of the corresponding division operation in the 32-bit register.
[0212] Based on the foregoing method for implementing unsigned integer division, the embodiment of the present invention further provides a method for implementing signed division. Specifically, as Figure 13 shown, when the integer dividend and the integer divisor are signed integers, the step 30 includes:
[0213] Step 301b: Perform an exclusive OR logical operation on the integer divisor and the integer dividend to obtain an exclusive OR result.
[0214] Since signed numbers may be negative, when both X and Y are signed two's complements, in order to ensure the correct sign of the final quotient, calculate the exclusive OR result S, and its expression is as follows:
[0215] S := X XOR Y
[0216] Where XOR represents the exclusive OR logical operation, and its specific execution logic is: if the two values a and b are different, the exclusive OR result is 1; if the two values a and b are the same, the exclusive OR result is 0.
[0217] Step 302b: When the approximate remainder is greater than or equal to the integer divisor, add one to the underestimated approximate quotient to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the target remainder; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the quotient to be checked; use the approximate remainder as the target remainder.
[0218] Theoretically, only the exclusive OR result of the sign bit of X and the sign bit of Y needs to be calculated, and the target quotient is obtained according to the underestimated approximate quotient, where the underestimated approximate quotient here is obtained according to the method of correcting unsigned numbers described above.
[0219] However, according to the hardware conditions for implementing this division operation, a most efficient implementation method is as follows:
[0220] Since the underestimated approximate quotient obtained according to the foregoing steps is at most 1 or 2 less than the actual result of divide_uint32(X, Y). Among them, the condition for the case where it is 2 less than the exact quotient of integer division is as follows: the target magic number M is rounded to 1, and the integer dividend X > 2 31 ; this means that when using the integer division implementation algorithm of the embodiment of the present invention to perform division operations on signed integers, since signed integers are strictly less than 231 , so when there is a situation where the underestimated approximate quotient is 2 less than the exact quotient of integer division, the correction of Q can be omitted.
[0221] For explanation, first calculate the quotient Q of the division operation according to the following expression:
[0222] Q := divide_uint32(abs(X), abs(Y))
[0223] Among them, the function abs(Z) represents the absolute value of the integer Z; that is, when the signed integer Z is negative, Z is changed to -Z, and when the signed integer Z is non - negative, Z remains unchanged.
[0224] Since both X and Y are signed integers, abs(X) and abs(Y) are both less than or equal to 2 31 . According to the occurrence condition mentioned before that the underestimated approximate quotient is 2 less than the exact quotient of integer division, it can be known that there cannot be a situation where the underestimated approximate quotient of the division step of the embodiment of the present invention is 2 less than the exact quotient of integer division. Therefore, when correcting signed numbers, the second check can be omitted, that is, only check once whether the approximate remainder L is greater than the integer divisor Y. When there is no need to obtain the remainder of the unsigned division operation, when the approximate remainder L is greater than or equal to the integer divisor Y, there is no need to process the approximate remainder L.
[0225] Step 303b: When the XOR result is less than 0, determine the negative of the quotient to be checked as the target quotient; when the XOR result is greater than or equal to 0, determine the quotient to be checked as the target quotient.
[0226] When the sign bit of the XOR result S is set, that is, S < 0, the finally obtained target quotient Q is negative; vice versa, the corresponding expression is as follows:
[0227] IF S < 0 THEN Q := 0 - Q
[0228] At this time, divide_int32(X, Y) = Q, that is, the obtained target quotient is the accurate result of the corresponding division operation in the 32 - bit register.
[0229] To specifically illustrate the beneficial effects, the embodiment of the present invention also provides a specific embodiment of a method for implementing integer division operation, as follows:
[0230] The hardware functions required to implement the method for implementing integer division operation of the embodiment of the present invention are as follows:
[0231] (1) Support floating - point numbers of the IEEE - 754 standard.
[0232] (2) Support integer to floating - point conversion: Ideally, support rounding operations towards positive infinity; or, support other rounding modes.
[0233] (3) Support calculation of floating - point reciprocals: Be able to implement the recip function of the embodiments of the present invention, that is, obtain the result of 1 / x through recip(x); where x is a floating - point number. When the error bound is known and not overly large, the type of rounding mode has no effect on the result.
[0234] (4) Support floating - point to integer conversion: Ideally, support rounding operations towards 0; or, support other rounding modes.
[0235] (5) Support standard integer arithmetic operations: Support addition, subtraction, and multiplication, as well as comparison operations of greater than or equal to (GreaterThan or Equal, abbreviated as GTE).
[0236] (6) Support high - order multiplication: For two N - bit integers X and Y, their product X×Y is a 2N - bit integer without overflow. Among them, a high - order multiplication instruction hmul(X,Y) only calculates N significant bits of the product. In an alternative embodiment, it can also be simulated using standard integer multiplication or standard integer addition and bit operations (such as, shift, bit - wise OR, and bit - wise AND), although the hardware efficiency of using this method is lower than that of high - order multiplication.
[0237] The input of the integer division implementation algorithm of the embodiments of the present invention is: a 32 - bit unsigned integer dividend X and a divisor Y; the output is: the quotient Q of X divided by Y: = divide_uint32(X,Y); when the remainder corresponding to the unsigned division is required, it also includes the remainder L of X divided by Y: = X - Y×Q.
[0238] The specific execution steps of the unsigned integer division implementation algorithm of the embodiments of the present invention are as follows:
[0239] Step S1: D: = float(Y)
[0240] Step S2: D: = recip(D)
[0241] Step S3: D: = D+(0xFFFFFFF - ceil(E))
[0242] Step S4: M: = uint(D)
[0243] Step S5: Q1 = hmul(X,M)
[0244] Step S6: L1 = X - Y×Q1
[0245] Step S7: Q2 = hmul(L1,M)
[0246] Step S8: Q = Q1 + Q2
[0247] Step S9: L = X - Y × Q
[0248] Step S10: IF L ≥ Y THEN Q := Q + 1
[0249] Step S11: IF L ≥ Y THEN L := L - Y
[0250] Step S12: IF L ≥ Y THEN Q := Q + 1
[0251] Step S13: IF L ≥ Y THEN Q := L - Y
[0252] Among them, Steps S1 to S4 are used to calculate the magic number, Steps S5 to S9 are used for division iteration based on the magic number, and Steps S10 to S13 are used for approximate correction of the result of the division iteration.
[0253] In Step S1, the ideal rounding mode is rounding towards positive infinity. When the hardware does not support rounding towards positive infinity but supports rounding to the nearest integer or rounding towards 0, 1 ULP can be added to D between Step S1 and Step S2. This process is based on floating-point operations, which is generally much more efficient than integer division and can usually be directly implemented in hardware.
[0254] In Step S3, ceil(E) is known at compile time, so the operation of Step S3 can be implemented by only one integer addition instruction.
[0255] In Step S4, the rounding mode is rounding towards 0. When rounding towards 0 is not supported, M can be decremented by 1 between Step S4 and Step S5 to ensure that the resulting result has no rounding, improving the generality of the integer division implementation algorithm in the embodiments of the present invention.
[0256] Many modern GPUs support multiply-add operations (madd operations). Multiply-add operations are common mathematical operations in graphics and compute-intensive tasks, and are usually used to perform various calculations such as matrix multiplication and convolution, etc.; it involves multiplying two numbers and adding the result to an accumulator. Specifically, the multiply-add operation is performed according to the following expression:
[0257] A := B + (C × D)
[0258] Therefore, generally, the operation of any one of Steps S6 to S9 can be implemented by only one integer instruction.
[0259] In arithmetic integer instructions, many GPUs can set an output predicate based on comparing a specific value with the computed output value. Thus, in step S9, a predicate expression can be set when L≥Y, and this predicate expression can be repeatedly used to execute step S10 and step S11, while avoiding any control flow branches and avoiding repeatedly checking whether L is greater than or equal to Y. Since GPUs typically use "single instruction, multiple threads" to perform calculations, if different algorithms will execute different numbers of instructions or different instruction sets for different input data, it may cause differences in execution time due to the GPU executing different branches of instructions for different threads, thereby affecting the performance of the GPU; if different branches of instructions are executed, due to certain input data, it may cause the threads on the GPU to stop until the alternative branch is completed and then start again, which greatly reduces the performance of the GPU. Since the number of execution instructions required by the integer division implementation algorithm of the embodiments of the present invention is fixed for different input data, it can solve the problem of reduced execution performance caused by the need to execute different numbers of instructions or different instruction sets for different input data.
[0260] Similarly, in step S11, the value of the predicate expression can be updated and then repeatedly used in step S12 and step S13.
[0261] When the remainder of the unsigned division operation does not need to be obtained, step S13 can be omitted. For signed numbers, steps S12 and S13 do not need to be executed; when the remainder of the signed division operation does not need to be obtained, step S11 can also be omitted.
[0262] The main advantages of the method for implementing integer division operations of the present invention are as follows: (1) It can perform floating-point operations in a more efficient manner; (2) Only one integer-to-floating-point format conversion and one floating-point-to-integer format conversion are required; (3) It has sufficient tolerance for errors and can still give correct results when using different rounding modes; (4) It can determine an approximation of the magic number for the division operation when the divisor is known only at runtime (i.e., the divisor is not known at compile time).
[0263] It should be noted that the integer division implementation algorithm of the embodiments of the present invention is used to run on a GPU. Therefore, some of its steps can effectively update the conditional predicate mask of the predicate expression while executing the required instructions. The running time of this integer division implementation algorithm is constant. On the premise that the running time of the hardware instructions used for implementation is constant, when the register size is fixed, the running time is independent of the size of the input. The constant running time ensures that when using Single Instruction Multiple Threads (SIMT) to run, there will be no thread divergence (that is, there will be no branch divergence problem), thus making more effective use of the GPU pipeline.
[0264] The application scenarios of the method for implementing integer division operations in the embodiments of the present invention are described below:
[0265] Integer division is widely used in encryption algorithms and decryption algorithms. Such algorithms usually rely on performing modulo operations (Mod), which involve calculating the remainder (i.e., modulus) of certain arithmetic operations. Such encryption algorithms and decryption algorithms are applied in hash algorithms, cryptocurrencies, and network security.
[0266] For example, in the field of network security, modulo operations can be used to implement various security protocols and algorithms. One method is that in the process of digital signature and authentication, modulo operations can be used to verify the validity of the signature. For example, in the process of digital signature, the sender uses its private key to sign the hash value of the message. This signature process may include: the sender performs modular exponentiation on the hash value of the message using its private key. This modular exponentiation process includes a modulo operation, that is, after performing the exponentiation operation, the result of the exponentiation operation is taken modulo with the user modulus to prevent the result from overflowing and keep the signature of a fixed length; among them, the private key is a concept in blockchain technology. It is a set of passwords that users must use in cryptocurrency transactions, digital signatures, and other operations; in a blockchain network, each user has a pair of public keys and private keys; the user modulus is the modulus in the process of generating public keys and private keys, specifically a certain large prime number.
[0267] The specific expression of the above modulo operation process is as follows:
[0268] Digital signature result = Result of exponentiation operation Mod User modulus
[0269] Using the method for implementing integer division operations in the embodiments of the present invention, the specific steps for implementing the calculation of the above formula in a GPU are as follows:
[0270] First, process the user modulus (i.e., the integer divisor of the division operation) based on floating-point operations to obtain the integer replacement value corresponding to the user modulus (i.e., the target magic number).
[0271] Step A1: When the hardware for operation supports rounding towards positive infinity, convert the data format of the user modulus from integer type to floating-point type according to the rounding mode of rounding towards positive infinity to obtain its overestimated floating-point representation. When the hardware for operation does not support rounding towards positive infinity and supports rounding to the nearest integer or rounding towards zero, convert the data format of the user modulus from integer type to floating-point type according to the selected rounding mode to obtain an intermediate floating-point representation; add 1 ULP to the intermediate floating-point representation to obtain the overestimated floating-point representation of the user modulus.
[0272] Step A2: Use the recip function to determine the underestimated floating-point representation of the reciprocal of the overestimated floating-point representation and the corresponding maximum positive calculation error; any rounding mode supported by the hardware for operation can be used here to perform rounding on the reciprocal of the overestimated floating-point representation.
[0273] Step A3: Determine the floating-point substitute value as the sum of the underestimated floating-point representation and the constant multiplication substitute value.
[0274] Step A4: When the hardware for operation supports rounding towards zero, convert the floating-point substitute value in floating-point type to integer type according to the rounding mode of rounding towards zero to obtain an integer substitute value. When the hardware for operation does not support rounding towards zero, convert the floating-point substitute value in floating-point type to integer type according to the available selected rounding mode to obtain an intermediate integer representation; subtract one from the intermediate integer representation to obtain an integer substitute value.
[0275] After obtaining the integer substitute value, it is necessary to generate an underestimated approximation value (i.e., underestimated approximate quotient) corresponding to the digital signature result based on this, and determine the underestimated error value (i.e., approximate remainder) corresponding to the underestimated approximation value.
[0276] Step A5: Calculate the ratio of the power operation result (i.e., integer dividend) to 2 32 to obtain a first division intermediate value, and calculate the product of the obtained first division intermediate value and the function value B substitute value to obtain a first iteration result.
[0277] Calculate the product of the user modulus and the obtained first iteration result to obtain a first intermediate value; calculate the difference between the power operation result and the obtained first intermediate value to obtain the first remainder corresponding to the first iteration result.
[0278] Step A6: Calculate the ratio of the obtained first remainder to 2 32 to obtain a second division intermediate value; calculate the product of the obtained second division intermediate value and the function value B substitute value to obtain a second iteration result.
[0279] Step A7: Calculate the sum of the obtained first iteration result and the obtained second iteration result to obtain an underestimated approximation value.
[0280] Calculate the product of the user modulus and the obtained underestimated approximation to get a second intermediate value; calculate the difference between the power operation result and the obtained second intermediate value to get the underestimated error value corresponding to the underestimated approximation.
[0281] Finally, according to the magnitude relationship between the calculated underestimated error value and the user modulus, correct the underestimated approximation to obtain the digital signature result as a signed integer (i.e., the target remainder).
[0282] Step A8: Perform an exclusive OR logical operation on the user modulus and the power operation result to obtain an exclusive OR result.
[0283] Step A9: When the obtained underestimated approximation is greater than or equal to the user modulus, add one to the obtained underestimated approximation as the value to be checked, and use the difference between the underestimated error value and the user modulus as the remainder to be checked; when the obtained underestimated approximation is less than the user modulus, use the underestimated approximation as the value to be checked and the underestimated error value as the remainder to be checked.
[0284] Step A10: Repeat the check in Step A9. When the remainder to be checked is greater than or equal to the user modulus, determine the difference between the remainder to be checked and the user modulus as the digital signature result; when the remainder to be checked is less than the user modulus, use the remainder to be checked as the digital signature result. Among them, due to the modulo operation, there is no need to determine the final target quotient.
[0285] The embodiment of the present invention also provides a device for implementing integer division operations. The device for implementing integer division operations is used to perform the division operation of an integer dividend by an integer divisor to determine the target quotient and target remainder as the result of the division operation. Based on the hardware pipeline of the device for implementing integer division operations in the embodiment of the present invention and the software operation logic for controlling the hardware pipeline, the method for implementing integer division operations in the embodiment of the present invention can be realized.
[0286] In a possible implementation manner, the device for implementing integer division operations includes a magic number generator, a division iterator, and an approximation corrector, where:
[0287] The magic number generator is used to process the integer divisor based on floating-point operations to obtain an integer target magic number; the division iterator is used to generate an underestimated approximate quotient of the division operation according to the target magic number, and is also used to obtain an approximate remainder according to the underestimated approximate quotient; the approximation corrector is used to correct the underestimated approximate quotient and the approximate remainder according to the magnitude relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder.
[0288] The magic number generator includes an integer-to-floating-point unit, a numerical operation unit, and a floating-point-to-integer unit. Among them: the integer-to-floating-point unit is used to convert the integer divisor into the overestimated floating-point representation; the numerical operation unit is used to determine the underestimated floating-point representation of the reciprocal of the overestimated floating-point representation and the corresponding maximum positive calculation error; and is further used to obtain a floating-point magic number according to the underestimated floating-point representation and the maximum positive calculation error of the computer hardware reciprocal operation; the floating-point-to-integer unit is used to convert the floating-point magic number into an integer according to a first preset rounding mode to obtain a target magic number.
[0289] Among them, when the rounding mode supported by the integer-to-floating-point unit includes rounding towards positive infinity, the integer-to-floating-point unit is used to convert the integer divisor from integer type to floating-point type to obtain the overestimated floating-point representation. When the rounding mode supported by the integer-to-floating-point unit does not include rounding towards positive infinity and includes rounding to the nearest integer or rounding towards zero, the integer-to-floating-point unit is used to convert the integer divisor from integer type to floating-point type to obtain an intermediate floating-point representation; and is further used to determine the sum of the intermediate floating-point representation and the minimum precision error unit as the overestimated floating-point representation to ensure that the overestimated floating-point representation is greater than or equal to the integer divisor.
[0290] The numerical operation unit includes a reciprocal operation unit, among which:
[0291] The reciprocal operation unit is used to determine the reciprocal of the overestimated floating-point representation, and perform rounding processing on the reciprocal to obtain the corresponding underestimated floating-point representation.
[0292] The numerical operation unit further includes an integer addition operation unit, among which:
[0293] When implementing the integer addition operation unit, the difference between the bitwise operation value and the maximum positive calculation error is determined as a constant multiplication substitution value in advance; the integer addition operation unit is used to regard both the underestimated floating-point representation and the constant multiplication substitution value as corresponding integer representations, and perform bitwise integer addition on the underestimated floating-point representation and the constant multiplication substitution value to obtain a floating-point magic number.
[0294] The division iterator includes a multiplication operation unit and an integer instruction unit, among which:
[0295] The multiplication operation unit is used to determine the high-order product of the integer dividend of the division operation and the target magic number as a first iteration result;
[0296] The integer instruction unit is used to determine the product of the integer divisor and the first iteration result as the first intermediate value, and determine the difference between the integer dividend and the first intermediate value as the first remainder corresponding to the first iteration result; it is also used to determine the high-order product of the first remainder and the target magic number as the second iteration result; it is also used to determine the sum of the first iteration result and the second iteration result as the underestimated approximate quotient; it is also used to determine the product of the integer divisor and the underestimated approximate quotient as the second intermediate value, and determine the difference between the integer dividend and the second intermediate value as the approximate remainder corresponding to the underestimated approximate quotient.
[0297] The approximate corrector includes an unsigned result verification unit, where:
[0298] The unsigned result verification unit is used to, when the approximate remainder is greater than or equal to the integer divisor, add 1 to the underestimated approximate quotient to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the remainder to be checked; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the target quotient; use the approximate remainder as the target remainder; it is also used to, when the remainder to be checked is greater than or equal to the integer divisor, add 1 to the quotient to be checked to obtain the target quotient, and determine the difference between the remainder to be checked and the integer divisor as the target remainder; when the remainder to be checked is less than the integer divisor, use the quotient to be checked as the target quotient, and use the remainder to be checked as the target remainder.
[0299] The approximate corrector includes a sign check unit, a quotient check unit, and a signed result output unit, where:
[0300] The sign check unit is used to perform an exclusive OR logical operation on the integer divisor and the integer dividend to obtain an exclusive OR result;
[0301] The quotient check unit is used to, when the approximate remainder is greater than or equal to the integer divisor, add 1 to the underestimated approximate quotient to obtain the quotient to be checked, and determine the difference between the approximate remainder and the integer divisor as the target remainder; it is also used to, when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the quotient to be checked, and use the approximate remainder as the target remainder;
[0302] The signed result output unit is used to, when the exclusive OR result is less than 0, determine the negative of the quotient to be checked as the target quotient; when the exclusive OR result is greater than or equal to 0, determine the quotient to be checked as the target quotient.
[0303] Such as Figure 14As shown, it is a schematic architecture diagram of another possible implementation of the apparatus for implementing integer division operations according to an embodiment of the present invention. The apparatus for implementing integer division operations in this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 14 Take one processor 21 as an example.
[0304] The processor 21 and the memory 22 can be connected through a bus or other means, Figure 14 Take the connection through the bus as an example.
[0305] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the method for implementing integer division operations in this embodiment. The processor 21 executes the method for implementing integer division operations by running the non-volatile software programs and instructions stored in the memory 22.
[0306] The memory 22 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely provided with respect to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0307] The program instructions / modules are stored in the memory 22, and when executed by the one or more processors 21, execute the method for implementing integer division operations in the above embodiment. For example, execute the Figures 1 to 2 , Figures 5 to 6 , Figure 8 and Figures 11 to 13 each step shown.
[0308] An embodiment of the present invention also provides a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by one or more processors, for example Figure 14 one of the processors 21, it enables the above one or more processors to execute the method for implementing integer division operations in the specific embodiments of the present invention. For example, execute the Figures 1 to 2 , Figures 5 to 6 , Figure 8 and Figures 11 to 13 each step shown; it can also implement Figure 14 each module and unit described; or execute the method for implementing integer division operations in the specific embodiments of the present invention. For example, execute the Figures 1 to 2 , Figures 5 to 6 , Figure 8 andFigures 11 to 13 each of the steps shown; it can also be implemented Figure 14 each of the modules and units described above.
[0309] It should be noted that for the information interaction, execution process, etc. between the modules and units in the above-mentioned device and system, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0310] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0311] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for implementing integer division operation, characterized in that: include: Processing an integer divisor of a division operation based on a floating point operation to obtain a target magic number; wherein the target magic number is an integer; According to the target magic number, generating an underestimated approximate quotient of the division operation; and obtaining an approximate remainder according to the underestimated approximate quotient; According to the magnitude relationship between the approximate remainder and the integer divisor, the underestimated approximate quotient and the approximate remainder are corrected to obtain a target quotient and a target remainder of the division operation.
2. The method for implementing integer division operation according to claim 1, characterized in that: The method of processing the integer divisor of the division operation based on the floating point operation to obtain the target magic number includes: determining an overestimated floating point representation of the integer divisor, and determining an underestimated floating point representation of the reciprocal of the overestimated floating point representation; Obtaining a floating point magic number according to the underestimated floating point representation and the maximum forward calculation error of the computer hardware reciprocal operation; According to a first preset rounding mode, the floating point magic number is converted into an integer to obtain a target magic number.
3. The method for implementing integer division operation according to claim 2, characterized in that: Determining the overestimated floating point representation of the integer divisor and determining the underestimated floating point representation of the reciprocal of the overestimated floating point representation comprises: converting the integer divisor to the overestimated floating point representation according to a second predetermined rounding mode; determining the reciprocal of the overestimated floating point representation; The reciprocal is rounded to obtain a corresponding underestimated floating point representation.
4. The method for implementing integer division operation according to claim 3, characterized in that: The converting the integer divisor to the overestimated floating point representation according to the second preset rounding mode includes: When the second preset rounding mode is rounding towards positive infinity, converting the integer divisor from an integer type to a floating point type to obtain the overestimated floating point representation; When the second preset rounding mode is rounding to the nearest integer or rounding towards 0, the integer divisor is converted from an integer type to a floating point type to obtain an intermediate floating point representation; the sum of the intermediate floating point representation and the minimum precision error unit is determined as the overestimated floating point representation to ensure that the overestimated floating point representation is greater than or equal to the integer divisor.
5. The method for implementing integer division operation according to claim 2, characterized in that: The method of obtaining the floating point magic number according to the underestimated floating point representation and the maximum forward calculation error of the computer hardware reciprocal operation includes: Before processing the integer divisor of the division operation based on floating point operation, pre-determining the difference between the bitwise operation value and the maximum forward calculation error as a constant multiplication replacement value; The underestimated floating point representation and the constant multiplication substitute value are both regarded as corresponding integer representations, and bitwise integer addition is performed on the underestimated floating point representation and the constant multiplication substitute value to obtain a floating point magic number.
6. The method for implementing integer division operation according to any one of claims 1 to 5, characterized in that: The step of generating an underestimated approximate quotient of the division operation according to the target magic number; and obtaining an approximate remainder according to the underestimated approximate quotient comprises: Determine a high-order product of the integer dividend of the division operation and the target magic number as a first iteration result; Determine the product of the integer divisor and the first iteration result as a first intermediate value, and determine the difference between the integer dividend and the first intermediate value as a first remainder corresponding to the first iteration result; Determine the high-order product of the first remainder and the target magic number as a second iteration result; Determine the sum of the first iteration result and the second iteration result as the underestimated approximate quotient; The product of the integer divisor and the underestimated approximate quotient is determined as a second intermediate value, and the difference between the integer dividend and the second intermediate value is determined as an approximate remainder corresponding to the underestimated approximate quotient.
7. The method for implementing integer division operation according to any one of claims 1 to 5, characterized in that: When the integer dividend and the integer divisor of the division operation are unsigned integers, the step of correcting the underestimated approximate quotient and the approximate remainder according to the size relationship between the approximate remainder and the integer divisor to obtain the target quotient and target remainder of the division operation includes: When the approximate remainder is greater than or equal to the integer divisor, the underestimated approximate quotient is increased by one to obtain a quotient to be checked; the difference between the approximate remainder and the integer divisor is determined as the remainder to be checked; when the approximate remainder is less than the integer divisor, the underestimated approximate quotient is used as the target quotient; the approximate remainder is used as the target remainder; When the remainder to be checked is greater than or equal to the integer divisor, the quotient to be checked is added by one to obtain a target quotient, and the difference between the remainder to be checked and the integer divisor is determined as the target remainder; when the remainder to be checked is less than the integer divisor, the quotient to be checked is used as the target quotient, and the remainder to be checked is used as the target remainder.
8. The method for implementing integer division operation according to claim 7, characterized in that: When the integer dividend and the integer divisor are signed integers, the step of correcting the underestimated approximate quotient and the approximate remainder according to the magnitude relationship between the approximate remainder and the integer divisor to obtain the target quotient and target remainder of the division operation includes: Performing an XOR logic operation on the integer divisor and the integer dividend to obtain an XOR result; When the approximate remainder is greater than or equal to the integer divisor, the underestimated approximate quotient is increased by one to obtain a quotient to be checked; the difference between the approximate remainder and the integer divisor is determined as a target remainder; when the approximate remainder is less than the integer divisor, the underestimated approximate quotient is used as the quotient to be checked; the approximate remainder is used as the target remainder; When the XOR result is less than 0, the negative of the quotient to be checked is determined as the target quotient; when the XOR result is greater than or equal to 0, the quotient to be checked is determined as the target quotient.
9. A device for implementing integer division operation, characterized in that: The device for implementing integer division operation is used to perform a division operation of an integer dividend divided by an integer divisor to determine a target quotient and a target remainder as a result of the division operation; The device for implementing integer division operation includes a magic number generator, a division iterator and an approximation corrector, wherein: The magic number generator is used to process the integer divisor based on floating point operations to obtain an integer target magic number; The division iterator is used to generate an underestimated approximate quotient of the division operation according to the target magic number; and is also used to obtain an approximate remainder according to the underestimated approximate quotient; The approximate corrector is used to correct the underestimated approximate quotient and the approximate remainder according to the size relationship between the approximate remainder and the integer divisor to obtain the target quotient and the target remainder.
10. The device for implementing integer division operation according to claim 9, characterized in that: The magic number generator includes an integer-to-floating-point conversion unit, a numerical calculation unit, and a floating-point-to-integer conversion unit, wherein: The integer-to-floating-point conversion unit is used to convert the integer divisor into the overestimated floating-point representation; The numerical calculation unit is used to determine the underestimated floating point representation of the reciprocal of the overestimated floating point representation; and is also used to obtain a floating point magic number according to the underestimated floating point representation and the maximum forward calculation error of the reciprocal operation of the computer hardware; The floating-point to integer conversion unit is used to convert the floating-point magic number into an integer according to a first preset rounding mode to obtain a target magic number.
11. The device for implementing integer division operation according to claim 10, characterized in that: The numerical calculation unit includes a reciprocal calculation unit, wherein: The reciprocal calculation unit is used to determine the reciprocal of the overestimated floating-point representation, and round the reciprocal to obtain a corresponding underestimated floating-point representation.
12. The device for implementing integer division operation according to claim 11, characterized in that: When the rounding mode supported by the integer-to-floating-point unit includes rounding toward positive infinity, the integer-to-floating-point unit is used to convert the integer divisor from an integer to a floating point type to obtain the overestimated floating point representation; When the rounding mode supported by the integer-to-floating-point unit does not include rounding toward positive infinity and includes rounding to the nearest integer or rounding toward 0, the integer-to-floating-point unit is used to convert the integer divisor from an integer to a floating point type to obtain an intermediate floating-point representation; and determine the sum of the intermediate floating-point representation and a minimum precision error unit as the overestimated floating-point representation to ensure that the overestimated floating-point representation is greater than or equal to the integer divisor.
13. The device for implementing integer division operation according to claim 10, characterized in that: The numerical operation unit also includes an integer addition operation unit, wherein: When implementing the integer addition operation unit, the difference between the bitwise operation value and the maximum forward calculation error is determined in advance as a constant multiplication substitution value; the integer addition operation unit is used to regard the underestimated floating-point representation and the constant multiplication substitution value as corresponding integer representations, perform bitwise integer addition on the underestimated floating-point representation and the constant multiplication substitution value, and obtain a floating-point magic number.
14. The device for implementing integer division operation according to claim 9, characterized in that: The division iterator comprises a multiplication operation unit and an integer instruction unit, wherein: The multiplication operation unit is used to determine the high-order product of the integer dividend of the division operation and the target magic number as a first iteration result; The integer instruction unit is used to determine the product of the integer divisor and the first iteration result as a first intermediate value, and determine the difference between the integer dividend and the first intermediate value as a first remainder corresponding to the first iteration result; and is also used to determine the high-order product of the first remainder and the target magic number as a second iteration result; and is also used to determine the sum of the first iteration result and the second iteration result as the underestimated approximate quotient; and is also used to determine the product of the integer divisor and the underestimated approximate quotient as a second intermediate value, and determine the difference between the integer dividend and the second intermediate value as an approximate remainder corresponding to the underestimated approximate quotient.
15. The device for implementing integer division operation according to claim 9, characterized in that: The approximation corrector comprises an unsigned result verification unit, wherein: The unsigned result verification unit is used to, when the approximate remainder is greater than or equal to the integer divisor, add one to the underestimated approximate quotient to obtain the quotient to be checked; determine the difference between the approximate remainder and the integer divisor as the remainder to be checked; when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the target quotient; use the approximate remainder as the target remainder; and also to, when the remainder to be checked is greater than or equal to the integer divisor, add one to the quotient to be checked to obtain the target quotient, and determine the difference between the remainder to be checked and the integer divisor as the target remainder; when the remainder to be checked is less than the integer divisor, use the quotient to be checked as the target quotient and the remainder to be checked as the target remainder.
16. The device for implementing integer division operation according to claim 15, characterized in that: The approximation corrector comprises a sign checking unit, a quotient checking unit and a signed result output unit, wherein: The sign checking unit is used to perform an XOR logic operation on the integer divisor and the integer dividend to obtain an XOR result; The quotient checking unit is used to, when the approximate remainder is greater than or equal to the integer divisor, add one to the underestimated approximate quotient to obtain a quotient to be checked, and determine the difference between the approximate remainder and the integer divisor as a target remainder; and is also used to, when the approximate remainder is less than the integer divisor, use the underestimated approximate quotient as the quotient to be checked and the approximate remainder as the target remainder; The signed result output unit is used to determine the negative of the quotient to be checked as the target quotient when the XOR result is less than 0; and to determine the quotient to be checked as the target quotient when the XOR result is greater than or equal to 0.
17. A non-volatile computer storage medium, characterized in that: The computer storage medium stores computer executable instructions, and the computer executable instructions are executed by one or more processors to complete the method for implementing integer division operations as described in any one of claims 1 to 8.