Digit-recurrence selection constants

JP2023008865A5Active Publication Date: 2025-06-30ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022102771
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-02
Filing Date
2022-06-27
Publication Date
2025-06-30
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

Existing digit recurrence algorithms face challenges in balancing performance, circuit area, and power consumption when implementing higher radix operations for division and square root calculations, particularly in meeting the competing demands of performance and complexity in circuit design.

Method used

The approach involves subdividing higher radix operations into multiple smaller radix sub-iterations within the same processing cycle, utilizing shared circuitry for both division and square root operations, and employing speculative replication and parallel processing to reduce timing constraints and circuit area.

Benefits of technology

This method improves performance by reducing the overall time required for calculations while minimizing circuit size and power consumption, making it more efficient than traditional implementations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a data processing apparatus, method and computer-readable medium to perform a digit-recurrence operation on an input value.SOLUTION: A data processing apparatus 2 receives a remainder value of a previous iteration of the digit-recurrence operation, performs comparisons of most significant bits of the remainder value of the previous iteration of the digit-recurrence operation with each of multiple selection constants associated with available digits of a next digit of a result of the digit-recurrence operation, and outputs the next digit of the result of the digit-recurrence operation based on the comparisons. Each of the selection constants is associated with one of the available digits and an input parameter. Storage circuitry configured to store a subset of the selection constants excludes the selection constants associated with a digit excluded from the available digits.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present technology relates to the field of data processing.

[0002] Digit recursion algorithms can be used to perform processing operations such as division or square root. Digit recursion uses an iterative algorithm to perform calculations. In each iteration, the next digit of a result value is generated. Each digit is represented using several bits. In a radix r implementation of a digit recursion algorithm, each digit has log2(r) bits. For example, an implementation using a radix of 4 represents each digit with two bits, so that in each iteration, two additional bits of the result are generated, and therefore, generating a result value with a particular number of bits may require a number of iterations. Implementations using higher radixes can produce a result of a given size in fewer iterations, improving performance, but the circuitry to perform a single iteration becomes more complex. When designing a circuit to perform such a digit recursion method, meeting the competing demands of performance, circuit area, and power consumption can be challenging.

[0003] In at least some examples, a data processing apparatus is provided for performing a digit recursion operation on an input value, the data processing apparatus including: a receiving circuit configured to receive a remainder value of a previous iteration of the digit recursion operation; a comparison circuit configured to perform a comparison between each of a plurality of selection constants associated with available digits of the next digit of the result of the digit recursion operation, each selection constant being associated with one of the available digits and the input parameter, and a most significant bit of the remainder value of the previous iteration of the digit recursion operation, and outputting the next digit of the result of the digit recursion operation based on the comparison; and a storage circuit configured to store a subset of the selection constants associated with digits excluded from the available digits, excluding the excluded selection constants from the selection constants.

[0004] In at least some examples, a method of data processing is provided for performing a digit recursion operation on an input value, the method including: receiving a remainder value of a previous iteration of the digit recursion operation; performing a comparison between each of a plurality of selection constants associated with available digits of the next digit of the result of the digit recursion operation, each selection constant being associated with one of the available digits and the input parameter, and a most significant bit of the remainder value of the previous iteration of the digit recursion operation, and outputting the next digit of the result of the digit recursion operation based on the comparison; and storing a subset of the selection constants associated with digits excluded from the available digits, the subset of selection constants excluding the excluded selection constants from the selection constants.

[0005] In at least some examples, a computer-readable medium for storing computer-readable code for manufacturing a data processing apparatus for performing a digit recursion operation on an input value is provided, the computer-readable medium including: a receiving circuit configured to receive a remainder value of a previous iteration of the digit recursion operation; a comparison circuit configured to perform a comparison between each of a plurality of selection constants associated with available digits of a next digit of a result of the digit recursion operation, each selection constant being associated with one of the available digits and the input parameter, and a most significant bit of the remainder value of the previous iteration of the digit recursion operation, and outputting the next digit of the result of the digit recursion operation based on the comparison; and a storage circuit configured to store a subset of the selection constants associated with digits excluded from the available digits, excluding the excluded selection constants from the selection constants. [Brief explanation of the drawings]

[0006] Further aspects, features, and advantages of the present technology will become apparent from the following description of examples, read in conjunction with the accompanying drawings. [Figure 1]FIG. 1 illustrates a schematic diagram of an example of a data processing operation having a divide / square root processing circuit. [Figure 2] FIG. 10 illustrates schematically an example of dividing a higher base digit-recursive square root or division operation into multiple lower base sub-iterations performed in the same processing cycle. [Figure 3] FIG. 1 illustrates a circuit for performing a given base r iteration of a square root operation. [Figure 4] FIG. 10 is a diagram illustrating a remainder update circuit. [Figure 5] FIG. 1 illustrates a remainder estimation circuit. [Figure 6] FIG. 10 is a diagram illustrating a digit selection circuit. [Figure 7-1] FIG. 1 illustrates in more detail a square root processing circuit for performing a given radix-64 iteration of a square root operation by performing two radix-8 sub-iterations in the same processing cycle. [Figure 7-2] FIG. 1 illustrates in more detail a square root processing circuit for performing a given radix-64 iteration of a square root operation by performing two radix-8 sub-iterations in the same processing cycle. [Figure 8-1] A combined divide / square root processing circuit is shown that is capable of performing both division and square root operations, with the shared circuitry generating at least one output value on the same data path that is used for both the division and square root operations. [Figure 8-2] A combined divide / square root processing circuit is shown that is capable of performing both division and square root operations, with the shared circuitry generating at least one output value on the same data path that is used for both the division and square root operations. [Figure 9-1] FIG. 1 illustrates an example of a division / square root pipeline. [Figure 9-2] FIG. 1 illustrates an example of a division / square root pipeline. [Figure 10] FIG. 1 illustrates pipelining of consecutive division or square root operations, where a second operation is prohibited from starting a predetermined number of cycles after a first operation if the second operation uses a lower precision floating-point representation than the first operation. [Figure 11] FIG. 1 illustrates on-the-fly transformation. [Figure 12] FIG. 1 illustrates an example of on-the-fly transformation. [Figure 13] FIG. 10 illustrates 3X digit on-the-fly conversion. [Figure 14] FIG. 10 illustrates an example of 3X on-the-fly transformation. [Figure 15-1] FIG. 1 illustrates a circuit for performing a 3x on-the-fly conversion. [Figure 15-2] FIG. 1 illustrates a circuit for performing a 3x on-the-fly conversion. [Figure 16] FIG. 10 illustrates a selection for reconstructing partial route values. [Figure 17] FIG. 10 illustrates comparison constants for radix-8 partial iterations of a division operation. [Figure 18] FIG. 10 illustrates comparison constants for radix-8 partial iterations of square root operations. [Figure 19] FIG. 10 is a diagram illustrating an offset representing the offset of the square root comparison constant relative to the division comparison constant. [Figure 20-1] FIG. 10 illustrates a division and offset lookup table for determining comparison constants for division and square root operations. [Figure 20-2] FIG. 10 illustrates a division and offset lookup table for determining comparison constants for division and square root operations. [Figure 21-1] 1 shows a circuit for obtaining a set of comparison constants for division and square root operations. [Figure 21-2] 1 shows a circuit for obtaining a set of comparison constants for division and square root operations.

[0007] Square root processing The square root processing circuit may perform a given radix-r iteration of a radix-r square root operation by performing two or more radix-n partial iterations in the same processing cycle, where n < r. This can provide a better compromise between performance and circuit overhead compared to an implementation that does not subdivide the radix-r iteration into lower-radix partial iterations. Since the overall operation performed in one cycle is a higher-radix operation of radix r, this means that log2(r) bits of the result can be generated per processing cycle, which can provide higher performance than when a smaller radix is used. However, the radix-r iteration is divided into several radix-n partial iterations in the same processing cycle, where for each partial iteration n is smaller than r. As a result, the overall size of the circuit can be smaller than when the radix-r iteration is performed as a single operation. This is because the number of alternative options available for selection as the next digit for each partial iteration using radix n is less than the number of alternative options for the radix-r digits required when the radix-r iteration of the square root operation is performed as a single operation. However, dividing the radix-r iteration into smaller-radix partial iterations can pose timing challenges in that it may not be possible to fit those radix-n partial iterations into a single processing cycle.

[0008] For a given base-n sub-iteration, the square root processing circuit may include: a digit selection circuit that selects a next base-n result digit of the square root result based on a previous remainder estimate; a remainder update circuit that adjusts the previous remainder value based on a remainder adjustment value responsive to the next base-n result digit selected by the digit selection circuit to generate an updated remainder value; a remainder estimation circuit that generates an updated remainder estimate that indicates an estimate of a portion of the updated remainder value; and an output signal path for supplying the updated remainder value and the updated remainder estimate for use as the previous remainder value and the previous remainder estimate in a subsequent base-n sub-iteration of the given base-r iteration or in a first base-n sub-iteration of a further base-r iteration of the base-r square root operation. Because multiple sub-iterations are performed per cycle, multiple instances of the digit selection circuit, remainder update circuit, remainder estimation circuit, and output signal path may be provided for each base-n sub-iteration within the same base-r iteration of the square root operation.

[0009] In the last radix-n sub-iteration of a given radix-r iteration, the remainder estimation circuit may generate an updated remainder estimate in parallel with the remainder update circuit that generates the updated remainder value. Because the updated remainder estimate represents a portion of the updated remainder value, this is counterintuitive because one might expect the remainder value to be available first and then calculated sequentially. However, the inventors have recognized that in embodiments in which higher radix iterations are divided into smaller radix sub-iterations, it is possible to generate the updated remainder estimate for the last sub-iteration in parallel with the remainder update circuit that generates the updated remainder value for that last sub-iteration of the given radix-r iteration. This means that the delay associated with calculating the remainder estimate for the last radix-n sub-iteration can be at least partially removed from the critical timing path through the square root processing circuit, reducing the overall time it takes to perform a given radix-r iteration of the square root operation and thus improving overall performance.

[0010] The remainder update circuit may generate the updated remainder value in a redundant representation. For example, the remainder value may be represented as two terms that together represent the numerical value of the updated remainder value, but there may be two or more combinations of values ​​for the first and second terms that can represent the same numerical value. Generating the updated remainder value in a redundant representation may be useful because it may avoid computation of the updated remainder value that requires propagating a carry from one bit to another. Thus, the remainder update circuit may include a carry-save add circuit.

[0011] However, for purposes of selecting the next base-n result digit of the square root result, the digit selection circuit can perform digit selection using a non-redundant representation of the remainder, and the remainder estimation circuit can therefore generate a non-redundant representation of an updated remainder estimate that represents an estimate of at least a portion of the updated remainder value. (The non-redundant representation means that the estimate can be expressed in a single term, and for any given numeric value of the updated remainder estimate, there is a single bit pattern (and no others) in the non-redundant representation that corresponds to that numeric value.) Because the full precision of the updated remainder value may not be needed for digit selection, the updated remainder estimate can have fewer bits than the updated remainder value. (More specifically, the updated remainder estimate can have fewer bits than the number of bits of a single term of a redundantly represented remainder value that may include two redundant terms.) Limiting the number of bits in the estimate reduces delay in computing the non-redundant remainder estimate. For example, the updated remainder estimate can represent an estimate of the most significant portion of the updated remainder value, because lower bits may not significantly affect the precision of digit selection.

[0012] Therefore, computation of the remainder estimate in the non-redundant representation can use a carry propagate adder circuit, which can propagate the carry from one bit position to another and which may be slower than a carry-save adder. Thus, in typical approaches, the carry propagate adder circuit used for the remainder estimate can significantly slow down the overall processing of a particular iteration of the square root operation.

[0013] However, the inventors have recognized that in an approach in which a radix-r square root iteration is divided into multiple smaller radix-n sub-iterations that are performed within the same processing cycle, an updated remainder estimate for the last radix-n sub-iteration can be calculated in parallel with the calculation of the updated remainder value, such that the updated remainder estimate for the last radix-n sub-iteration can be calculated using information provided as input to the remainder update circuitry in the last radix-n sub-iteration and / or other information from previous sub-iterations within a given radix-r iteration, thereby avoiding the need to wait for the updated remainder value in the last radix-n sub-iteration to become available before commencing calculation of the updated remainder estimate for the last radix-n sub-iteration. This provides a relatively significant gain in performance due to the removal from the critical timing path of the relatively slow carry propagate addition for calculating the updated remainder estimate in the last radix-n sub-iteration of a given radix-r iteration.

[0014] In a remainder update, the previous remainder value is updated based on a remainder adjustment value whose value depends on the next result digit selected by the digit selection circuit. The remainder estimation circuit in the final radix-n sub-iteration can use this remainder adjustment value and the previous remainder estimate value to generate an updated remainder estimate for the final radix-n sub-iteration. Because the remainder adjustment value is used as an input to the remainder estimation circuit in the final radix-n sub-iteration, this eliminates the need to wait for the updated remainder value and allows the updated remainder estimate to be available more quickly.

[0015] The remainder estimation circuit can take advantage of the fact that the final radix-n partial iteration follows at least one previous partial iteration being executed in the same cycle, so that some information calculated in that previous partial iteration can be used by the remainder estimation circuit in the final partial iteration to calculate an updated remainder estimate sooner than if the remainder estimate were calculated consecutively after the updated remainder value was obtained.

[0016] For example, in preceding radix-n sub-iterations of a given radix-r iteration other than the final radix-n sub-iteration, the remainder estimation circuit may calculate at least one additional bit of the updated remainder estimate that is not needed to select the next radix-n result digit in the final radix-n sub-iteration of the given radix-r iteration, and in the final radix-n sub-iteration of the given radix-r iteration, the remainder estimation circuit may determine the updated remainder estimate using the at least one additional bit determined in the preceding radix-n sub-iteration. By calculating more bits than needed for the updated remainder estimate in the preceding radix-n sub-iterations, the additional bit(s) can be used to calculate the updated remainder estimate in the final radix-n sub-iteration sooner because the additional bit calculated in the preceding sub-iterations allows the updated remainder estimate in the final sub-iteration to be calculated without waiting for the updated remainder value to be available.

[0017] In the first radix-n sub-iteration of a given radix-r iteration, the remainder estimation circuit can determine an updated remainder estimate based on the updated remainder value generated by the remainder update circuit in the first radix-n sub-iteration. Thus, it is not necessary that the updated remainder estimate be calculated in parallel with the updated remainder values ​​in all sub-iterations. In the first sub-iteration of a given radix-r iteration, sufficient information may not be available to allow the remainder estimate to be calculated until the updated remainder value is available in redundant form. However, because multiple radix-n sub-iterations overlap within the same processing cycle, the circuit designer has the flexibility to change the relative timing at which portions of a subsequent sub-iteration start relative to portions of a previous sub-iteration, and can use information from the previous sub-iteration to calculate parameters in subsequent sub-iterations, making it feasible to parallelize the calculation of the updated remainder value and updated remainder estimate for at least the final sub-iteration.

[0018] In implementations in which at least three partial iterations are performed within the same cycle to perform a given base r iteration of the square root operation, the updated remainder estimate may also be calculated in parallel with the updated remainder values ​​of one or more intermediate partial iterations between the first and last partial iterations.

[0019] The square root processing circuit comprises, for a given base-n sub-iteration, one or more instances of the replica circuit, each instance of the replica circuit including two or more replica circuit units for determining, in parallel with the selection of the next base-n result digit by the digit selection circuit, two or more candidate output values ​​corresponding to different result digits that can be selected as the next base-n result digit by the digit selection circuit, and a selection circuit for selecting one of the plurality of candidate output values ​​in response to the digit selection circuit indicating which of the different result digits will be selected as the next base-n result digit, the plurality of candidate output values ​​including at least two or more candidate output values ​​generated by the two or more replica circuit units. This approach can result in faster performance because it is not necessary to wait for the next base-n result digit to be actually selected by the digit selection circuit before commencing calculations to generate the candidate output value.

[0020] Note that the number of candidate output values ​​available for selection by the selection circuit may be greater than the number of candidate output values ​​generated by two or more replicated circuit units. For example, one of the result digits available for selection may be equal to 0, and in some cases, it may not be necessary to explicitly calculate a candidate output value for the result digit 0 because the candidate output value selected if the next result digit is 0 may be identical to the input value provided to the partial iteration. Thus, the selection circuit may receive as input candidate output values ​​that are not explicitly generated by one of the replicated circuit units, as well as candidate output values ​​generated by two or more replicated circuit units.

[0021] Providing replica circuit units to speculatively calculate multiple candidate output values ​​before the time the next result digit is known may be superior in performance, but the number of replica circuit units required increases with increasing radix, which may increase circuit area cost and power consumption to support higher radix arithmetic.

[0022] One technique for limiting circuit area and power costs may be to provide at least one of the two or more replica circuit units as a shared circuit unit shared between both a positive result digit having a given magnitude and a negative result digit having the same given magnitude. The shared circuit unit is configured to output a shared candidate output value to a selection circuit on a shared signal path, and the selection circuit may select the shared candidate output value from the shared signal path when the next base-n result digit is either a positive or negative result digit having the given magnitude. This therefore eliminates the need to provide two separate replica circuit units for each of the positive and negative result digits sharing the same magnitude. This can reduce the total number of required replica circuit units, thereby saving circuit area and reducing power consumption.

[0023] For at least one instance of the replica circuit, the shared circuit unit providing the output shared between positive and negative result digits of the same magnitude can select a value to be output as a shared candidate output value on the shared signal path based on the sign of the previous remainder estimate. Thus, although a common signal path is shared between two result digit values ​​having the same magnitude but different signs, the actual numeric value output on that shared signal path can differ depending on the sign of the previous remainder estimate.

[0024] For at least one instance of the replica circuit, the shared circuit unit may include shared adding circuitry for determining shared candidate output values ​​for positive and negative result digits having a given magnitude. The technique of providing a shared circuit unit for generating shared candidate output values ​​for both positive and negative digits of the same magnitude may be particularly useful when the circuit unit includes adder circuitry, as adder circuitry is relatively costly in terms of circuit area.

[0025] For base-n partial repetition, it is usually expected that the number of candidate output values ​​available for selection by the selection circuit must be n + 1. However, by sharing a shared circuit unit between positive and negative result digits having the same magnitude, the total number of candidate output values ​​available for selection by the selection circuit can be reduced to n / 2 + 1, which means that the number of replicated circuit units can be reduced, thereby significantly reducing the circuit area.

[0026] There may be several instances of the replica circuit within the square root processing circuit, and various parts of the square root processing circuit may each use this technique, with the replica circuit unit a priori determining candidate output values ​​for multiple possible result digits, and then the correct candidate output value may be selected by the selection circuit once the next result digit is selected.

[0027] For example, the remainder update circuit may comprise one such instance of the replication circuit. If the remainder update circuit uses a speculative replication and selection technique, the candidate output value being selected by the selection circuit may be the candidate updated remainder value.

[0028] Similarly, the remainder estimation circuit may use this speculative replication and include one of the instances of the replication circuit described above. If the remainder estimation circuit includes a replication circuit, the candidate output value may be a candidate updated remainder estimate.

[0029] Another part of the digit recursion method may be performing an on-the-fly conversion. In the case of a square root operation, the adjustment of the previous remainder value to generate an updated remainder value may depend not only on the remainder adjustment value (selected based on the next result digit), but also on the partial root value, which is a numeric value corresponding to the previously selected series of result digits. Since the result digits may be selected by the digit selection circuit as signed digits, an on-the-fly conversion circuit may then be provided to convert the partial root value to a non-redundant representation to provide a partial root value in a non-redundant representation that can be used by the remainder update circuit to adjust the previous remainder value to generate an updated remainder value. As described below, it is possible to perform the on-the-fly conversion in a manner that does not require addition but can be done simply by concatenating the previous partial root value with some additional bits selected based on the latest base-n result digit.

[0030] Thus, the on-the-fly conversion circuit (for generating partial root values ​​that indicate, in a non-redundant representation, a numerical value corresponding to a previously selected series of base-n result digits) may also include an instance of the above-mentioned replica circuit, so that the replica circuit unit generates several candidate partial root values ​​and the candidate output values ​​available for selection by the selection circuit include several candidate partial root values.

[0031] Therefore, regardless of which part of the square root processing circuit performs replication, replication can help improve performance, and if implemented, sharing replication circuit units for positive and negative result digits of the same magnitude can help reduce the overall circuit size.

[0032] Some implementations may implement the replica circuit with only one or a subset of the above components of the square root processing circuit, while other components do not use the replicated approach, and performance may be maximized if the remainder update circuit, remainder estimation circuit, and on-the-fly transformation circuit each provide an instance of the replica circuit.

[0033] In general, if a given radix r iteration is divided into several back-to-back or overlapping radix n sub-iterations in the same processing cycle, the value of r may correspond to the product of the respective values ​​of n for each of the sub-iterations used in one cycle.

[0034] In the particular example described below, r=64 and n=8 for each of the partial iterations, such that there are two radix-8 partial iterations for each radix-64 iteration. This approach can provide a good balance between performance (radix-64 means that six bits can be generated per processing cycle) and circuit area and timing complexity (using radix-8 for the partial iterations means that only two partial iterations are needed, which imposes lower timing pressure compared to implementations using three or more partial iterations, but increasing the radix beyond 64 can make it infeasible to manage circuit size while meeting timing). Thus, r=64 and n=8 can be a particularly useful combination.

[0035] Nevertheless, other options are possible: for example, a radix-64 iteration of a square root operation can be performed as three sub-iterations, each of radix-4 (since 64 = 4 x 4 x 4).

[0036] Implementing each of the partial iterations with the same radix n can be useful because it can be more efficient in terms of overall circuit area and simpler in terms of design complexity to use the same radix in each partial iteration.

[0037] Nevertheless, it is also possible for different sub-iterations within the same base r iteration to use different bases. For example, a base 64 iteration of a digit-recursive square root operation can be divided into one base 4 sub-iteration, one base 8 sub-iteration, and one base 2 sub-iteration. Therefore, it is not necessary for n to be equal for each sub-iteration.

[0038] The above-described technique can be implemented in square root processing circuits of different designs. In one example, the square root processing circuit may be an iterative square root processing circuit, and the output signal path may provide an updated remainder value and an updated remainder estimate generated in the last radix-n sub-iteration from the output of the iterative square root processing circuit to the input of the same iterative square root processing circuit for use as the previous remainder value and the previous remainder estimate in the first radix-n sub-iteration of further radix-r iterations of the square root operation. Thus, to perform the square root operation as a whole, multiple passes through the iterative square root processing circuit are performed over multiple processing cycles, and the output of the iterative square root processing circuit in one cycle is fed back as an input to the same unit in a subsequent cycle.

[0039] However, as described in more detail below, the square root processing circuit may be part of a pipelined square root processing unit including several square root iteration pipeline stages, each including a respective instance of the square root processing circuit described above. In this case, the output signal path of a given pipeline stage may supply the updated remainder value and the updated remainder estimate generated in the last radix-n sub-iteration of a given radix-r iteration from the output of the square root processing circuit in one square root iteration pipeline stage to the input of the square root processing circuit (a different instance of the square root processing circuit) in a subsequent square root iteration pipeline stage for processing of the subsequent radix-r iteration in the next processing cycle. This approach allows multiple square root operations to be pipelined together, such that while a previous square root operation is being processed in a subsequent stage of the pipelined square root processing unit, a subsequent square root operation may be in a previous pipeline stage where the previous radix-r iteration is being executed, which may help improve the overall throughput of the square root operations.

[0040] Combined divide / square root circuit Commercially available processor microarchitectures typically provide separate circuit logic for division and square root operations, with these operations performed in completely separate circuit logic units and no sharing of the data path used to calculate the division result compared to the data path used to calculate the square root result. This can be simpler to implement because no extra complexity is required in the square root operation to affect the timing of the division operation. However, it may be desirable to increase the radix used for the division and square root operations to improve performance by allowing a division or square root result with a larger number of bits to be calculated per cycle. For example, a radix-64 division or square root operation, not currently available in commercially available processors, could be used to calculate a 6-bit result per cycle. However, increasing the radix means that more complex circuitry is required compared to implementations requiring lower radixes. Therefore, having separate division and square root processing circuits when operating at higher radixes can increase the circuit size and therefore the power consumption of the processor.

[0041] In the example described below, a combined divide / square root processing circuit is provided for performing a given radix-64 iteration of a radix-64 division operation in response to a division instruction and for performing a given radix-64 iteration of a radix-64 square root operation in response to a square root instruction. The combined divide / square root processing circuit has shared circuitry for generating at least one output value for the given radix-64 iteration over the same data path used for both the radix-64 division operation and the radix-64 square root operation. For example, the at least one output value may include any one or more of an updated remainder value, a selected result digit, an updated remainder estimate, and / or an on-the-fly converted partial result value. The use of shared circuitry, where the same data path is used for the output of both the division operation and the square root operation, can reduce the overall amount of circuitry compared to an implementation having a divide and square root unit. This is particularly useful for radix-64 arithmetic, given the increased circuitry required for radix-64 compared to lower radix arithmetic supported by commercially available processor microarchitectures.

[0042] The combined divide / square root processing circuitry can perform the same number of radix-64 iterations per processing cycle for both the radix-64 division operation and the radix-64 square root operation. This increases the extent to which circuitry can be shared between the square root and division operations and can help limit the overall circuit area of ​​the combined divide / square root processing circuitry.

[0043] For both radix-64 division and radix-64 square root operations, the combined divide / square root processing circuitry can perform a given radix-64 iteration by performing one or more radix-m sub-iterations in the same processing cycle, where m≦64.

[0044] In some examples, m=64, in which case the radix-64 iteration may be performed as a single, integral operation that generates the next result digit 6 bits at a time, without breaking the radix-64 iteration into separate sub-iterations. This approach may be faster, but may require additional circuit logic to accommodate the larger number of candidate result digits, as the possible result digits may extend from −32 to +32 when the radix-64 iteration is performed as a single operation.

[0045] However, in some examples, m<64, so that the combined division / square root processing circuit can perform a given radix-64 iteration by performing multiple radix-m sub-iterations in the same processing cycle. For example, in the example shown below, m equals 8, so that there are two radix-8 sub-iterations in each radix-64 iteration. Another option could be for m=4, so that there are three radix-4 sub-iterations in one radix-64 iteration per processing cycle. Although the sub-iteration radix m can take on different values ​​between different sub-iterations, as described above for the example square root processing circuit, it may be more efficient from a circuit implementation perspective if m is the same in each sub-iteration.

[0046] Thus, the term "base-m sub-iteration" is used to refer to either the entire base-64 iteration if there is no subdivision into multiple sub-iterations of smaller bases, or to an individual sub-iteration of a smaller base if such subdivision is performed.

[0047] There may be different portions of the combined divide / square root processing circuitry, which may function as the shared circuitry described above.

[0048] In one example, the shared circuitry includes a shared digit selection circuit that selects the next base-m digit of the division or square-root result based on a comparison of a previous remainder estimate with a set of comparison constants in a given base-m sub-iteration. In implementations where m=64 and thus the base-64 iteration is not divided into multiple sub-iterations, the previous remainder estimate used for digit selection may come from the previous base-64 iteration. On the other hand, if m<64 such that the base-64 iteration is divided into multiple base-64 sub-iterations, then in the first base-m sub-iteration of a given base-64 iteration, the previous remainder estimate may come from the last base-m sub-iteration of the previous base-64 iteration; and in base-m sub-iterations other than the first base-m sub-iteration of a given base-64 iteration, the shared digit selection circuitry may select the next base-m digit based on the previous remainder estimate calculated in a previous base-m sub-iteration of the given base-64 iteration.

[0049] Thus, a shared digit selection circuit can be provided to save circuit area compared to separate circuits for selecting the result digits of the division and square root operations, respectively. For example, the shared digit selection circuit can comprise the same set of comparator circuits used to perform the comparison between the previous remainder estimate and the comparison constant for both the division and square root operations.

[0050] The comparator circuitry used when performing both the division and square root operations may be the same, but the shared digit selection circuitry may use different sets of comparison constants for the base-64 division operation and the base-64 square root operation. The set of comparison constants may be selected based on the operation type.

[0051] However, one problem is that the comparison constant for a division operation may not be the same size as the comparison constant for a square root operation. Error analysis has shown that a division operation may not require as many bits in the comparison constant as the comparison constant used for a square root operation to provide sufficient precision for digit selection. Therefore, the division comparison constant may be expected to have fewer bits than the square root comparison constant. However, to facilitate circuit sharing, the comparison constant compared to the remainder estimate before a radix-64 division operation may have at least one least significant bit set to 0 to pad it to the same width as the comparison constant compared to the remainder estimate before a radix-64 square root operation. By placing at least one 0 in the least significant bit position, expanding the comparison constant for division to the same bit width as that used for the square root operation allows the same comparator in the digit selection circuit and the same data path for the remainder estimate to be used for both the square root and division operations, thereby reducing circuit area.

[0052] Another example of a shared circuit may be a shared remainder update circuit that adjusts a previous remainder value based on a remainder adjust value in a given radix-m sub-iteration to generate an updated remainder value in a redundant representation. By using a redundant representation, the remainder update may be performed using a carry-save add to avoid the increased delay of a carry-propagate add. Thus, the shared circuit may include a shared carry-save add circuit that performs a carry-save add to generate the updated remainder value. This eliminates the need for two separate carry-save adders for the division and square root operations, since the remainder value data path is shared between the division and square root operations.

[0053] However, the remainder adjust value may be different for a division operation compared to a square root operation. Thus, the shared remainder update circuit may include a selection circuit that selects as the remainder adjust value a value derived from a divisor value when performing partial iterations of a given base m as part of a base-64 division operation, and a value derived from a partial root value as a function of a series of previously selected base-m root digits when performing partial iterations of a given base m as part of a base-64 square root operation. Thus, with a small amount of additional logic in the selection circuit, a shared data path may be used for both square root and division operations when generating the remainder update.

[0054] Another example of a shared circuit may be a shared remainder estimation circuit for generating updated remainder estimates that indicate non-redundant estimates of a portion of an updated remainder value generated in a redundant representation in a given radix-m sub-iteration of a radix-64 division operation or a radix-64 square root operation. For example, the shared remainder estimation circuit may include a carry propagate adder circuit for performing carry propagate addition to generate the non-redundant estimate, thereby sharing it between the division operation and the square root operation, thereby avoiding the need for two separate carry propagate adders.

[0055] In implementations where m is less than 64, in the last radix-m sub-iteration of a given radix-64 iteration, the shared remainder estimation circuit may generate updated remainder estimates in parallel with the shared remainder update circuit that generates updated remainder values. This improves performance by reducing the latency of critical timing paths for the same reasons as described above for the square root processing circuit.

[0056] Another example of shared circuitry may be shared by on-the-fly transformation circuitry for performing on-the-fly transformations to generate partial result values ​​in a non-redundant representation in partial iterations of a given radix m. Again, on-the-fly transformation circuitry may require relatively complex hardware circuit logic, so avoiding duplicating this for division and square root operations can save a greater amount of circuit area.

[0057] However, one problem is that in typical schemes, on-the-fly conversion circuitry is implemented differently for division operations compared to square root operations. The on-the-fly conversion circuitry can insert a selected value based on the next result digit into the partial result value to generate an on-the-fly conversion value representing the partial result corresponding to the sequence of result digits selected in that cycle and any previous cycles. However, in typical schemes, the position at which the next digit is inserted into the partial result value during on-the-fly conversion differs between division and square root operations, with the division operation being implemented by inserting a value derived from the next digit into the least significant bit position with a left shift to shift up all previously inserted bits to more significant bit positions. In contrast, due to the fact that the partial result value influences the digit selection and remainder update operations in a square root operation (and therefore it is more convenient if, in each processing cycle, the most significant bit of the partial root result value remains in a consistent bit position within the stored representation of the partial result), for a square root operation, a value derived from the next result digit is inserted into a variable bit position within the partial result, with a mask used to represent the position within the partial result value where the next square root result digit is to be inserted. This mask may be adjusted between iterations or sub-iterations to gradually move the position where the next result digit is inserted towards the more significant bits of the partial result value.

[0058] Given these contrasting ways of maintaining partial result values, one might think that having shared circuit logic for on-the-fly conversion circuitry would be difficult.

[0059] However, the present inventors have recognized that it is possible to provide a shared on-the-fly conversion circuit. For a given radix-n partial iteration, the shared on-the-fly conversion circuit selects a position for inserting the next digit into the partial result value based on a mask value for both the radix-64 division operation and the radix-64 square root operation. Thus, for the division operation, instead of shifting up all digits and inserting the next digit into the least significant bit position, the shared on-the-fly conversion circuit behaves unconventionally, using a mask for the radix-64 division operation to select a position at which the next digit is inserted into the partial result value for the division operation. This allows the on-the-fly conversion for the division operation to mirror the conversion for the square root operation, using shared circuit logic and shared data paths. This helps improve overall circuit area efficiency.

[0060] Similar to the various circuit units of the square root processing circuit described above, the shared circuit in the shared division / square root circuit can include one or more instances of a replicated circuit, each instance of which includes two or more replicated circuit units for determining, in parallel with selecting the next base-m digit of the division result or square root result, two or more candidate output values ​​corresponding to different digits that can be selected as the next base-m digit, and a selection circuit for selecting one of the plurality of candidate output values ​​in response to an indication of which of the different digits has been selected as the next base-m digit, the plurality of candidate output values ​​including at least two or more candidate output values ​​generated by two or more replicated circuit units. This helps improve performance for the same reasons as described above for the square root example. Again, to reduce the total number of replicated circuit units required to process the base-m partial iterations, at least one of the replicated circuit units can be a shared circuit unit shared between positive and negative digits of equal magnitude. Various components of the combined divide / square root circuit may use any one or more of such replica circuits, for example, a remainder update circuit, a remainder estimation circuit, and an on-the-fly transformation circuit.

[0061] Similar to the square root processing circuit described above, in the case of a combined divide / square root processing circuit, this may be implemented as an iterative divide / square root processing circuit, where the output of one radix-64 iteration is input to the same iterative divide / square root processing circuit for use in further radix-64 iterations of the division or square root operation, or as a pipelined divide / square root processing unit having several pipeline stages, each having an instance of a respective one of the combined divide / square root processing circuit, with a signal path providing the output generated in one stage as input to the next stage in the pipeline.

[0062] Divide / Square Root Pipeline It is common in many programs to need to perform arithmetic operations on operands represented in floating-point format. The IEEE-754 technical standard defines various formats for floating-point representation, e.g., half-precision (HP), single-precision (SP), and double-precision (DP) (other formats are also available). The particular floating-point precision used for the operands and result of a division or square root operation can control the number of bits that need to be generated for the result, which can affect the number of iterations required for a digit-recurring division or square root operation.

[0063] Conventionally, circuit units for performing digit-recursive division or square root operations that can produce results with floating-point level precision are implemented as iterative circuit units, so that the circuit logic provided in hardware corresponds to a single iteration of the digit-recursive division or square root operation, and the output of one iteration is fed back as input to the very same circuit logic unit that performed the previous iteration, readying that same circuit unit to perform the next iteration.

[0064] In contrast, in the example described below, a division / square root pipeline is provided that includes several division / square root iterative pipeline stages, each capable of performing a respective iteration of a digit-recurring division or square root operation. A signal path is provided to supply the output generated by one pipeline stage in one iteration as an input to a subsequent pipeline stage of the division / square root pipeline to perform a subsequent iteration of the digit-recurring division or square root operation. The division / square root pipeline can perform the digit-recurring division or square root operation on floating-point operands to produce a floating-point result.

[0065] Thus, while supporting the level of precision required for floating-point formats, division or square root operations are implemented in a pipelined manner rather than as iterative units. This means that for the processing of a single division or square root operation, each iteration is performed by a different pipeline stage, with the output from one pipeline stage being input to the next pipeline stage, so that the operation can move down the pipeline and output a result until it reaches the end.

[0066] This approach can be thought of as contrary to straight pipes, and although pipelining of instructions is generally known, the much greater complexity of division / square root operations compared to other types of operations means that the overall circuit area of ​​a single circuit unit to perform a single iteration of a digit-recursive division or square root operation is relatively large, and one would therefore expect that extending the iteration unit into a pipeline containing a sufficient number of stages to produce the result precision required for floating-point processing would significantly increase the overall circuit area required for the division / square root unit by a factor corresponding to the maximum number of iterations required for the division or square root operation.

[0067] However, the present inventors have recognized that in practice, a processor microarchitecture having iterative division / square root processing circuitry may actually provide multiple parallel divide / square root units to increase the overall available bandwidth, such that, for example, there may be multiple divide functional units and / or multiple square root functional units, and more than one division or square root operation may be processed simultaneously. In a pipelined manner, the need to replicate the entire divide / square root unit is eliminated because the divide / square root pipeline may process multiple operations in a pipelined manner, with a divide / square root iterative pipeline stage after the divide / square root pipeline performing a subsequent iteration of the first digit recursive division or square root operation in parallel with a previous divide / square root iterative pipeline stage performing a first digit recursive division or square root operation and a previous iteration of the second digit recursive division / square root operation.

[0068] Thus, although pipelining appears to significantly increase circuit logic, in reality the additional circuitry may not be significant compared to commercially available processors with multiple parallel divide / square root units, and various techniques described herein for reducing circuit area can be applied, particularly using shared data paths for division and square root operations and reducing the number of replicated circuit units by sharing the same replicated circuit units for positive and negative digits of the same magnitude as described previously.

[0069] Thus, the overall pipeline may be competitive in terms of circuit area and may help improve performance, since with pipelining of operations, the pipelining scheme may avoid blocking the iterative circuit units for the total number of cycles used to perform the digit recursive division or square root operation, thereby allowing for higher throughput as successive division or square root operations may be scheduled with fewer cycles between them.

[0070] It is possible for a pipeline to perform only a division or a square root operation, such that a division / square root pipeline can perform either a division or a square root operation but not both.

[0071] However, a pipeline can be particularly useful if the combined divide / square root processing circuitry is provided with a shared data path used for both operations. Thus, each divide / square root iteration pipeline stage includes combined divide / square root processing circuitry for performing a given iteration of a digit-recurring division operation in response to a divide instruction and for performing a given iteration of a digit-recurring square root operation in response to a square root instruction. The combined divide / square root processing circuitry includes shared circuitry for generating at least one output value on the same data path used for both the given iteration of the digit-recurring division operation and the given iteration of the digit-recurring square root operation. Providing a combined divide / square root processing circuitry helps limit the overall area cost of extending a single iteration unit into a pipeline (because the area budget previously provided for the separate divide and square root units is available for the pipeline implementation) and helps the pipeline be competitive with current microarchitectures in terms of circuit area. As mentioned earlier, when a combinational divide / square root circuit is used, it may be useful for the divide / square root pipeline to perform the same number of iterations per processing cycle, in the same radix, for both the digit-recurring division operation and the digit-recurring square root operation, as this facilitates greater sharing of shared circuit units.

[0072] For a given result precision, the division / square root pipeline can process a digit-recursive division operation in the same number of processing cycles as a digit-recursive square root operation. This simplifies control of circuit timing within the pipeline and helps facilitate sharing of common circuit logic between division and square root operations.

[0073] Various floating-point formats can be supported for the operand(s) input to the division or square root operation and the floating-point result generated by the division or square root operation. For example, the operand(s) and result can be half-precision (HP), single-precision (SP), or double-precision (DP) floating-point values. The division / square root pipeline can support at least one of these formats, or it can support other types of floating-point formats. However, it is particularly useful for the division / square root pipeline to support at least one of SP and DP floating-point values. Because programs written in DP floating-point precision may be particularly common, it may be useful in some cases for the division / square root pipeline to support operations whose results are in DP floating-point representation. Pipeline stages of the division / square root pipeline can be used to process the mantissa of the floating-point operands to generate the mantissa of the floating-point result. There may be separate circuit logic for processing the exponent of the floating-point value. The exponent processing logic may be simpler than the logic for generating the mantissa and may use any known technique for generating the exponent of a division / square root result.

[0074] In some examples, the division / square root pipeline may support at least two different result precisions for digit-recurring division or square root operations. For example, the division / square root pipeline may support any two or more of HP, SP, and DP floating-point values.

[0075] For lower precision floating-point result precisions, the divide / square root pipeline can perform the division or square root operation in fewer processing cycles than when producing a higher precision result (fewer digit recursion iterations are required because fewer bits need to be generated for the result). The device can have control circuitry that controls the divide / square root pipeline to bypass at least one divide / square root iteration pipeline stage used to perform at least one iteration of the digit recursion division or square root operation when producing a higher precision result when performing the digit recursion division or square root operation to produce a lower precision result. This improves performance by making the result of the operation available sooner when fewer bits need to be calculated.

[0076] However, allowing some stages of the pipeline to be bypassed in this way can create the possibility that, when a pipelined high-precision operation is followed by a low-precision operation, both operations may conflict when they reach a post-processing stage that can perform post-processing operations on the output of the final iteration of the digit-recursive division or square root operation. For example, the post-processing stage may perform rounding of the result of the division or square root operation to provide a rounded floating-point result, and / or may perform denormal (subnormal) result processing by right-shifting to produce a result in accordance with the IEEE standard (when the result of the division or square root operation is less than the smallest number that can be represented as a regular floating-point number). To ensure that the post-processing operation receives only the output of the final iteration of a single operation per cycle, the control circuit can prevent a lower-precision digit recursive division / square root operation performed to generate a lower-precision result from starting a predetermined number of cycles after a higher-precision digit recursive division / square root operation performed to generate a higher-precision result, where the predetermined number of cycles corresponds to the difference between the number of cycles required to reach at least one post-processing stage for the higher-precision digit recursive division / square root operation and the number of cycles required to reach at least one post-processing stage for the lower-precision digit recursive division / square root operation. Thus, depending on the difference in precision between the previous high-precision operation and the subsequent low-precision operation, there may be a certain number of cycles during which the start of the low-precision operation after the high-precision operation is prohibited to avoid collisions. The predetermined number of cycles can be different for different pairs of precision formats.

[0077] Each division / square root iteration pipeline stage may include a digit selection circuit for selecting a next result digit for a partial result value of the digit recursive division or square root operation based on a comparison between the previous remainder value and a set of comparison constants, and a remainder update circuit for updating the previous remainder value based on the remainder adjustment value and the next result digit selected by the digit selection circuit. Each pipeline stage may also include other elements, such as a remainder estimation circuit for generating a non-redundant estimate of a portion of the updated remainder value generated by the remainder update circuit in a redundant representation. Each pipeline stage may also include an on-the-fly conversion circuit for maintaining on-the-fly non-redundant versions of partial result values ​​corresponding to previously selected sequences of result digits from all previous iterations of the digit recursion method.

[0078] All division / square root iteration pipeline stages of a pipeline may use the same set of comparison constants for each iteration performed within the same digit-recursive division or square root operation. While the comparison constants may be different for each operation, the same set of comparison constants may be used within each iteration of the same operation. Thus, the division / square root pipeline may perform a table lookup to obtain the set of comparison constants in a pre-processing stage of the division / square root pipeline prior to the first division / square root iteration pipeline stage of the division / square root pipeline, and the set of comparison constants may be passed from stage to stage to avoid repeating table lookups at each division / square root iteration pipeline stage within the same digit-recursive division or square root operation. This approach may reduce the timing for each individual pipeline stage because each stage does not need to perform a table lookup, reducing the overall amount of circuit logic required at each stage. There may be a set of flip-flops provided in each pipeline stage that simply capture the comparison constants received from the previous pipeline stage without the need to update their comparison constants. This greatly simplifies the pipeline and reduces the overall circuit area.

[0079] This approach may seem surprising because one might think that the comparison constants for a digit-recurring division or square root operation should not be the same for each iteration, and that a different set of comparison constants may be required compared to the constants used in later stages, particularly in the first iteration of a typical division / square root operation. However, in the example described below, the division / square root pipeline includes at least one preprocessing stage for performing operand preprocessing prior to the first division / square root iteration pipeline stage of the division / square root pipeline, where the operand preprocessing includes selecting at least one initial result digit for the result of the digit-recurring division or square root operation. By selecting at least one initial result digit for the result of the division or square root operation in the preprocessing stage so that the initial result digit is not selected within the main body of the pipeline, this means that different sets of selection criteria can be used for the result digit to avoid requiring different comparison constants in different stages of the main iteration portion of the pipeline. This means that the remaining divide / square root iterative pipeline stages can each use the same set of comparison constants within the same divide or square root operation, improving circuit timing and reducing circuit area as described above.

[0080] However, one issue in implementations where the division / square root pipeline supports both digit-recurring division and digit-recurring square root operations (where combined division / square root circuits are provided as described above) is that the number of initial digits, requiring a different set of comparison constants compared to subsequent iterations, may be different for the division and square root operations. For example, error analysis has shown that if a base of 8 is used for digit selection in a given iteration or subiteration to obtain digit selection precision sufficient for the square root operation, the selection of the first two square root digits may use different comparison constants than the selection of the remaining square root digits. If the base used is a base other than 8, the number of initial root digits selected using different comparison constants for the remaining iterations may be a number other than 2. Nevertheless, regardless of the base, in general, square root operations can use different comparison constants to select a particular number of initial root digits and use the same set of comparison constants for subsequent iterations or subiterations after those initial root digits are selected. In contrast, division operations can use the same comparison constants for selecting all result digits (regardless of the base used). However, for performance reasons, it may be desirable to select at least one of the result digits during a pre-processing stage to reduce the number of subsequent pipeline stages required for the division operation, and therefore reduce latency. For example, in the radix-8 example described below, the first division digit may be selected in a pre-processing stage.

[0081] Thus, the number of initial digits selected in the pre-processing stages may be different for the square root and division operations. For example, at least one pre-processing stage may generate a larger number of initial result digits for a digit-recursive square root operation than for a digit-recursive division operation. While this may obviously introduce some asymmetry between the two operations, in practice, this significantly aids in reducing the overall circuit area and improving pipeline performance, as it means that, for the square root operation, the comparison constants for the remaining stages can be simply latched from one stage to the next without requiring a separate table lookup at each pipeline stage.

[0082] However, because at least one pre-processing stage generates more initial result digits for the square root operation than for the division operation, this means that fewer remaining iterations are required after the pre-processing stage for the square root operation compared to the division operation, even when producing a result of the same precision, and therefore the result of the square root operation may be available to an earlier division / square root iteration pipeline stage for the square root operation compared to the division operation. To enable a shared pipeline to be used, the control circuitry can control the division / square root pipeline to cause at least one division / square root iteration pipeline stage to be used to perform at least one iteration when a digit-recurring division operation is performed, and to completely or partially skip or discard some bits of the result output when performing the digit-recurring square root operation. In some cases, an entire pipeline stage of the pipeline can be skipped for the square root operation, while in other cases, only a portion of the bits generated in a given pipeline stage may need to be discarded, depending on the floating-point precision used and the radix used for the digit-recurring operation. For example, if a given iteration of the digit recursion method is divided into multiple sub-iterations of smaller bases, as in some of the examples described above, it may be possible to skip only individual sub-iterations within a given divide / square root iteration pipeline stage for some result precisions of the square root operation, rather than skipping entire stages. Also, in some cases, if the total number of bits required for a given result precision of the square root operation is not an exact multiple of the number of bits generated per iteration or sub-iteration, truncation of the result can be achieved by fully executing a given iteration or sub-iteration, but discarding some bits of the result if other bits of the result digit generated in the most recently executed iteration or sub-iteration are still needed.

[0083] This means that, when considering the body of the pipeline, the result of the square root operation may be available earlier than the result of the division operation, but the total number of cycles taken for both the square root and division operations may be the same. For example, even if the result of the square root operation is available earlier, there may be at least one cycle as the value is passed unchanged to the next cycle, allowing the overall operation timing to reflect that of the division operation. This may, for example, make scheduling of post-processing operations easier, as the post-processing may have the same timing regardless of the operation being performed.

[0084] Another complication when using a combined divide / square root data path in a pipeline lies in maintaining partial result values ​​that provide a representation of a numerical value corresponding to a previously selected series of result digits. If a shared data path is to be used, it may be desirable to be able to insert the next result digit into the partial result value at the same bit position for both the division and square root operations when performing a given iteration of the digit recursion method in a given pipeline stage of the pipeline. However, if the pre-processing stage generates different numbers of initial result digits for the division and square root operations, this can make using shared circuit logic in the remaining pipeline stages more complicated, as the position at which the next result digit is inserted in a given iteration could differ from iteration to iteration.

[0085] Thus, when performing a digit-recurring division operation, the at least one pre-processing stage can provide a partial result value to the first division / square root iteration pipeline stage, with selected bit positions set to a dummy bit value, the selected bit positions corresponding to bit positions into which the at least one pre-processing stage inserts at least one additional result digit not generated for the digit-recurring division operation when performing the digit-recurring square root operation. This allows a given division / square root iteration pipeline stage of the division / square root pipeline to insert the next result digit into the partial result value at the same bit position for both the digit-recurring division operation and the digit-recurring square root operation. The division / square root pipeline can include a post-processing stage for removing the dummy bit value from the final result value when performing the digit-recurring division operation.

[0086] This recognizes that inserting additional dummy bit values ​​into partial results for a division operation does not affect the overall result of the division operation because the partial result values ​​are not used for remainder update or digit selection operations in the division operation. It is only for square root operations that the partial result values ​​are used to control remainder update and digit selection operations. For division operations, it does not matter if the partial result values ​​temporarily include some dummy bit values ​​that are removed in a post-processing stage, because the partial result values ​​are simply maintained "on the fly" to improve performance by not having to convert redundant representations of the result to non-redundant form at the end of the pipeline. Including dummy bit values ​​in the partial result values ​​used for the division operation allows the insertion of the next result digit to be in the same position for both operations, improving the sharing of circuit logic for both operations.

[0087] The division / square root pipeline as described above can be used for digit-recursive division or square root operations with any base.

[0088] However, since the extra number of bits of the result generated per cycle in radix-64 operations, compared to lower radices, helps to reduce the total number of pipeline stages required for the pipeline, using a division / square root pipeline can be particularly useful for radix-64 digit recurrence division or square root operations, and as a result, the pipeline can become competitive with respect to circuit area compared to iterative implementations.

[0089] In one example, each division / square root iterative pipeline stage is configured to perform each radix-r iteration of a radix-r digit recurrence division or square root operation by performing multiple radix-n partial iterations in the same processing cycle, where n < r. By splitting the higher-radix iteration into multiple lower-radix partial iterations, the amount of circuitry in each pipeline stage is reduced, and as a result, the overall circuit area of the pipeline can compete with current iterative implementations while improving performance. In a particular example, r = 64 and n = 8, but more generally, the radix-r iteration can be split into different combinations of lower-radix partial iterations as described above for the example of the square root processing circuitry.

[0090] On-the-fly conversion A data processing apparatus for converting a plurality of signed digits representing an input value into a redundant representation, comprising, in each of a plurality of iterations, a receiving circuit that receives a signed digit from the plurality of signed digits and previous intermediate data from a previous iteration, a concatenating circuit that performs a concatenation of bits corresponding to the signed digit and bits of the previous intermediate data to generate updated intermediate data, and an output circuit that provides the updated intermediate data as the previous intermediate data for a next iteration, wherein the previous intermediate data includes S3[i] in a non-redundant representation, which is at least a part of the input value multiplied by 3 in a non-redundant representation.

[0091] In these examples, each digit is signed. Thus, the input value (which can be positive or negative) is composed of individual digits, each individually signed. In this way, for example, the first digit of the input value can be positive and the second digit of the input value can be negative. This can be used to provide a representation format known as a redundant representation, in which a pair of words is used to represent the input value. This is in contrast to a non-redundant representation, in which a number is represented using a single word. Non-redundant and redundant representations are each best suited for specific types of operations, and therefore conversion between different representation formats can be useful. The conversion is performed on the fly as each digit of the input value is received, thereby avoiding the long latency that would occur if all digits were converted at once after all digits were received. The conversion process is accomplished using bit concatenation, which can be performed quickly. The concatenated bits are derived from the signed digit. A set of intermediate data is maintained between iterations and updated with each iteration. The concatenation performed depends on the newly received current digit. In particular, the intermediate data includes S3[i], which is S[i] (the partial result) multiplied by 3. The value of S3[i] is achieved without simply multiplying S[i] by 3, which would take too long to keep up with the arrival of new signed digits, not to mention be energy intensive. Note that although the term "iteration" is used here, the iteration being referenced could be a "partial iteration" as discussed above.

[0092] In some examples, the previous intermediate data includes S3[i-1]. In these examples, the value of S3 from the previous iteration, S3[i-1], is also maintained in the intermediate data. This value does not need to be calculated and can be carried over from the previous iteration. Providing such data allows adjustments to be made when carries occur during the transformation process.

[0093] In some examples, the previous intermediate data includes S3M[i], which is a non-redundant representation of at least a portion of the input value multiplied by 3 minus 1. In other words, S3M[i] = (S[i] × 3) - 1. The value of S3[i] is equal to the value of S3[i] minus 1.

[0094] In some examples, the previous intermediate data includes S3M[i-1]. In these examples, the value of S3M from the previous iteration is also maintained in the intermediate data. This value does not need to be calculated and can be carried over from the previous iteration. Providing such data allows you to adjust when carries occur during the transformation process.

[0095] In some examples, the concatenation performed by the concatenation circuitry includes concatenation for each of S3[i] and S3M[i] to generate updated intermediate data including S3[i+1] and S3M[i+1]. Thus, each of the four values ​​has concatenation performed at each iteration (or sub-iteration). The concatenation may be different for each of the four values.

[0096] In some examples, the bit corresponding to the unsigned digit is concatenated with one of S3[i] and S3M[i] to generate S3[i+1], and the other of S3[i] and S3M[i] to generate S3M[i]. One of S3[i] and S3M[i] is determined based on whether the unsigned digit is greater than or less than 0. In these examples, whether the unsigned digit is greater than, equal to, or less than 0 affects whether S3[i] or S3M[i] is used to generate S3[i+1], and the other of S3[i] and S3M[i] is used to generate S3M[i+1].

[0097] In some examples, the data processing apparatus includes an adjustment circuit configured to perform a selective adjustment on at least one of S3[i] and S3M[i] prior to concatenation based on the magnitude of the signed digit and whether the signed digit is positive or negative. The selective adjustment can be used, for example, to achieve a carry between columns of the output value.

[0098] In some instances, selective adjustment is performed when the magnitude of the signed digit multiplied by 3 exceeds the base in which the signed digit is represented. Selective adjustment can be used to handle situations where the concatenated digit multiplied by 3 is greater than the base being used for the conversion, and therefore digits need to be incremented or decremented in other positions. For example, as in base 10, if one has a partial result S[i]=512 and it is desired to add a digit to this number (thousands) 6, this can be done to achieve the number S[i+1]=6512. However, if one has S3[i]=1536 and it is desired to add a digit to this number (thousands) 6, then 3 * We need to add 6 = 18. However, because base 10 is used and 18 is greater than 10, this cannot be done by changing a single position. Instead, we add 8 to the thousands number to give us 9536, then carry "1" in as the ten thousand number to give us 19536.

[0099] In some examples, the data processing device is configured to convert multiple signed digits that represent an input value in a redundant representation without using adder circuitry. In particular, the values ​​of S3M[i] are not derived by simply taking S3[i] and subtracting 1 (e.g., using adder circuitry). Instead, by calculating these values ​​using concatenation over i iterations (concatenating a different number for each of S3[i] and SM3[i]), it is possible to determine these numbers with lower latency than would be achieved by using adder circuitry to perform the subtraction of 1.

[0100] In some examples, the data processing apparatus includes a digit recursion circuit for performing a digit recursion operation to generate a plurality of signed digits, and in each of a plurality of iterations, one of the plurality of signed numbers is provided to the receiving circuit. The digit recursion circuit can be used to provide a sequence of digits that make up the input value, with a subset of the digits being provided in an iteration (or partial iteration), e.g., each clock cycle.

[0101] In some examples, the digit recursion circuit is configured to operate in a square root operation mode in which the digit recursion operation is a square root operation. The digit recursion algorithm for calculating square roots performs a multiplication of a partial root S, where the multiplication depends on the digit being added. Because the partial root S changes with each iteration, this multiplication is performed for each iteration. Multiplying by 0 always results in 0. Multiplying by 1 is simply the identity function. On the other hand, multiplication by powers of 2 (e.g., 2 or 4) can be performed by performing bit shifts. Similarly, multiplication by −1, −2, and −4 can be performed by negating the multiplications by 1, 2, and 4, respectively. However, multiplication by 3 is significantly more complex. A multiplication circuit that performs the actual multiplication by 3 may require several processor cycles that are too slow. Even the addition of X and 2X to determine 3X requires additional circuitry, which may also take too long to execute. Therefore, by maintaining the value of S3 achieved through concatenation, square root digit recursion can be performed efficiently.

[0102] In some examples, the digit recursion circuit is configured to operate in a division operation mode in which the digit recursion operation is a division operation, the previous intermediate data includes S[i], which is at least a portion of the input value in a non-redundant representation, and SM[i], which is at least a portion of the input value in the non-redundant representation minus 1, and after multiple iterations, the output circuit is further configured to output S[i], so that the same data processing device that performs the conversion from the input value to the output value can be used for both the square root operation and the division operation. The calculation can also include generating S[i], which is at least a portion of the input value converted to a non-redundant representation, and SM[i], which is that value minus 1.

[0103] In some examples, the concatenation circuit is configured to suppress the generation of S3[i] in the division operation mode. As explained above, the value of S3 (and by extension, S3M) is particularly relevant when performing square root digit recursion. Also, when performing digit recursion division, the generation of S3 and S3M is not necessary because partial root multiplication is not required for each iteration. Therefore, suppressing the generation of S3 and S3M in the division operation mode can reduce power consumption.

[0104] In some examples, the digit recursion operation has a base of at least 8. For bases of at least 8, the available digits include at least one of, if not both, +3 and −3. Thus, during the square root digit recursion algorithm, it may be necessary to multiply the partial root by either 3 or −3, depending on the most recent digit. As previously mentioned, multiplication by 3 can be time consuming, so by maintaining S3 and S3M through the concatenation, it is possible to efficiently perform square root digit recursion for bases of 8 while meeting the timing constraints of the circuit.

[0105] In some examples, the possible values ​​of the signed digit include at least one of +3 and −3. As mentioned above, the use of signed digits can require multiplication by 3, which is more difficult to perform than multiplication involving powers of 2.

[0106] Selection Constants In some examples, a data processing apparatus for performing a digit recursion operation on an input value is provided, the data processing apparatus including: a receiving circuit configured to receive a remainder value of a previous iteration of the digit recursion operation; a comparison circuit configured to perform a comparison between each of a plurality of selection constants associated with available digits of the next digit of the result of the digit recursion operation and a most significant bit of the remainder value of the previous iteration of the digit recursion operation, and to output the next digit of the result of the digit recursion operation based on the comparison, wherein each of the selection constants is associated with one of the available digits and the input parameter; and a memory circuit configured to store a subset of the selection constants, the subset of the selection constants excluding the excluded selection constants from the selection constants associated with digits excluded from the available digits.

[0107] During the digit recursion process, a comparison is made between the most significant bit of the remainder value from the previous iteration and several selection constants to determine the next digit in the digit recursion operation, i.e., the next digit to be output. The number of selection constants corresponds to the product of the number of possible values ​​of the most significant bit of the remainder value and the number of possible values ​​the output digit can have. For example, if the six most significant bits of the remainder value are considered and there are eight possible values ​​for each output digit, the selection constant table holds 8 × 32 = 256 values. Each value may also occupy several bits. Also, typically, multiple tables are required to accommodate both square root digit recursion and division digit recursion. Therefore, the number of stored values ​​is large. In the above example, at least some of the required selection constants are not stored. That is, for the range of supported digit recursion operations (based on the considered base and number of most significant bits), at least some of the selection constants required for the digit selection process are not stored anywhere within the data processing device. This reduces the amount of storage space required, resulting in a more compact and lower-power circuit.

[0108] In some examples, the data processing device comprises a conversion circuit configured to generate selection constants that are excluded from the selection constants stored in the storage circuit, in these examples, the missing or omitted selection constants that are not stored in the data processing device are instead inferred or generated from other selection constants that are stored in the data processing device.

[0109] In some examples, the conversion circuit is configured to generate the omitted selection constants by performing a selective inversion on the sign of one of the selection constants stored in the storage circuit. In these examples, some of the omitted selection constants may be generated by taking another selection constant and inverting its sign. Inverting the sign of a number (e.g., by taking two's complement) can be performed efficiently and need not affect the time it takes to perform the selection operation.

[0110] In some examples, one of the selection constants is associated with the same input parameter and a different one of the digits available as the omitted selection constant. Thus, two columns of the selection constant table can be "merged." That is, for a given set of most significant bits of the remainder value, the selection constants of two different digits are the same (the sign changes according to the number for which the selection constant is generated). For example, the selection constant for the remainder bit 0.100010 can be "2" for the possible output digits +4 and -3. However, for the digit +4, the selection constant can be negative (-2), and for the digit -3, the selection constant can be negative (+2). Thus, these two columns can be merged into one using rules regarding whether the constants are positive or negative.

[0111] In some examples, the storage circuitry is configured to store an exception flag that indicates whether a selective inversion should be performed on the selection constant to generate an excluded selection constant. In these examples, whether to invert depends on the value of the exception flag. The inversion may also depend on other factors, such as the digit for which the selection constant is being generated. For example, considering the above example for the remainder bit 0.100010, the selection constant may be positive (+2) for one digit (+4) and negative (−2) for another digit (−3). However, the exception flag overrides this (causing both digits to have the same selection constant) or inverts it (−2 for the digit +4 and +2 for the digit +3).

[0112] In some examples, the digit recursion operation is a square root digit recursion operation, and the input parameter is a partial root.

[0113] In some examples, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor.

[0114] In some examples, in a division operation mode, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor, and in a square root operation mode, the digit recursion operation is a square root digit recursion operation and the input parameter is a partial root. Thus, in these examples, a device may be used that performs both division digit recursion and square root digit recursion depending on the operation mode.

[0115] In some examples, in a division operation mode, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor. In a square root operation mode, the digit recursion operation is a square root digit recursion operation and the input parameter is a partial root, and each selection constant is either a division digit recursion operation selection constant or a square root digit digit recursion operation selection constant. Such a data processing apparatus can perform both division and square root digit recursion, but the stored selection constants are specific to one of these two operation modes (division or square root). Storing selection constants specific to only one of the two operation modes can reduce the storage requirements of the data processing apparatus.

[0116] In some examples, each of the selection constants is a division digit recursion operation selection constant. This does not mean that all of the selection constants for division digit recursion are stored, but simply that the constants stored are division digit recursion selection constants that can be used as part of the process of generating square root digit recursion selection constants.

[0117] In some examples, the conversion circuit is configured to generate the excluded selection constant in the division operation mode by performing selective inversion of the sign of one of the division digit recursion operation selection constants, i.e., one of the division digit recursion constants is used and inverted based on some criteria (e.g., the value of the digit with which the constant is associated).

[0118] In some examples, the conversion circuit is configured to generate the selection constant to be excluded in the square root mode operation by referencing one of the division digit recursion operation selection constants.

[0119] In some examples, the memory circuitry is configured to store a plurality of mappings between the excluded selection constants in the square root operation mode and one of the division digit recursion operation selection constants, where the mappings are used to indicate which division digit recursion operation selection constant to use as a basis for creating the square root digit recursion operation selection constant and / or how to modify one of the division digit recursion operation selection constants to generate the corresponding square root digit recursion operation selection constant.

[0120] In some examples, the storage circuitry is configured to store an exception flag that indicates, for a selection constant, whether a selection inversion should be performed to generate an excluded selection constant. The exception flag may be part of a set of flags (or stored as part of a larger value) that indicate circumstances under which an inversion occurs to generate an excluded selection constant.

[0121] In some examples, the digit recursion operation is base 8. For example, the available digits may be limited to {-4, -3, -2, -1, 0, 1, 2, 3, 4}.

[0122] Examples of data processing devices FIG. 1 illustrates an example of a data processing device 2, e.g., a processor, that supports the execution of instructions defined according to a particular instruction set architecture (ISA). The device has instruction fetch circuitry 4 for fetching program instructions defined according to the architecture from an instruction cache or memory (not shown in FIG. 1). The fetched instructions are decoded by decode circuitry 6 to identify the operation to be performed. In response to a given instruction, decode circuitry 6 generates control signals that control execution unit 8 to perform the processing operation represented by the instruction. Operands for a given processing operation may be read from registers 10, and the processing results of the operation may be written back to registers 10. Execution unit 8 may include various execution units, including arithmetic units such as adder 20, multiplier 22, and divide / square root unit 24. Execution unit 8 may also include other types of functional units, such as a branch unit 26 for determining the outcome of branch instructions that may trigger a non-sequential change in program flow within the program being executed, and a load / store unit 28 for executing load instructions to load data from cache or memory into register 10 or store instructions to store data from register 10 into cache or memory.

[0123] The following example shows the circuit logic design of the division / square root execution unit 24 of the processing unit 2. When the decode stage 6 decodes a division instruction, the decode stage 6 controls the division / square root execution unit 24 to perform a digit recursion division operation. When the decode stage 6 decodes a square root instruction, the decode stage 6 controls the division / square root execution unit 24 to perform a digit recursion square root operation.

[0124] The examples that follow focus on the divide / square root execution unit 24, although it will be appreciated that the remainder of the processing unit 2 may be constructed in accordance with any known processor design techniques. It will be appreciated that Figure 1 is a simplified representation of the components of a data processor, and that in practice many other components not shown in Figure 1 may also be provided.

[0125] Theoretical foundations of digit recursive division and square root Digit recursion is the process of iterating the resulting digit p of base r. (i+1) and a class of iterative algorithms that compute the remainder rem[i]. The remainder is used to obtain the next base r digit, where base r is a power of 2, and each base r digit represents log2(r) bits of the result. Division (x / d) and Square Root

number

[0126] The partial result before iteration i is defined as:

number

number

number

number

number

[0127] For fast iteration, the remainder is held in carry-save or signed-digit redundant representation. In the implementation described below, known techniques are used to represent the remainder using a representation such as carry-save, where the remainder is represented by a positive word and a negative word (the non-redundant binary value corresponding to the remainder can be obtained by subtracting the negative word from the positive word).

[0128] On the other hand, due to the algorithm convergence condition and multiplication time r in equation (3), the remainder has several bits in the integer part. The number of integer bits depends on the base, digit set, and operation.

[0129] Then, for each iteration, the base r digits of the result are taken from the current remainder, a new remainder is calculated for the next iteration, and the partial result is updated.

[0130] The selection function for selecting the next result digit is the remainder estimate

number

number

number

number

[0131] The partial results are redundant representations of signed digits in base r, generated most significant digit first (MSDF). This is converted to a non-redundant representation at each iteration. The most efficient conversion technique is the well-known on-the-fly conversion. Essentially, on-the-fly conversion is a conversion of the digits p i+1 to the partial result P[i] (see equation (1)). However, since the digits can be negative, this addition can produce carry propagation. To prevent this slow carry propagation, another form of the result is maintained, where PM[i] has the following value: PM[i]=P[i]-r -i (6) Using this second form, the transformation algorithm for concatenation is as follows:

number

[0132] In this way, there are no arithmetic operations involved in the conversion, only the concatenation of values ​​into P[i] and PM[i], where the concatenated values ​​are the selected digit p i+1 Depends on.

[0133] The number of iterations of the digit recursion algorithm is it=[n / log2(r)] (9) n is the number of bits in the result, including the bits needed for rounding. [...] is a ceiling function, so [n / (log2(r)] is the smallest integer greater than or equal to n / (log2(r)].

[0134] The number of cycles is directly related to the number of iterations and the number of iterations performed per cycle. Now, considering m iterations per cycle, the number of cycles is: cycles=[it / m] (10)

[0135] Equations (1) through (10) can be subdivided to any base. In the next two sections, these equations are specialized for r=8, division, and square root. A higher base, r=64, is obtained by overlapping two base-8 partial iterations. The partial iteration base is 8.

[0136] Base 8 division Floating-point division of dividend x and divisor d produces the quotient q = x / d. In base 8, the partial quotient (partial result) before iteration i and the digits obtained at iteration i are Q[i] and q, respectively. i+1 and equation (1) can be rewritten as follows:

number

number

[0137] Regarding the selection function, it turns out that only the 10 most significant bits of the remainder need to be assimilated to obtain a remainder estimate accurate enough for digit selection. As mentioned before, the selection constant also depends on the divisor. The 6 most significant bits of the divisor are used to derive a set of 8 selection constants for every iteration of the current division. Different divisor values ​​can derive different sets. Note that the most significant bit of the divisor is always 1, since the operands are normalized before selecting the constants. The selection constants are stored in a look-up table (LUT).

[0138] In this implementation, it has been determined that only the 10 most significant bits (MSBs) of the remainder, 3 integer bits, and 7 fractional bits are needed to select the next quotient digit in equation (12).

[0139] base 8 square root The floating-point square root of operand x is the root

number

number

number

[0140] The initial values ​​of the remainder and partial root are rem[0]=x-1 and S[0]=1.0, respectively.

[0141] The selection function involves comparing the remainder estimate with a set of eight partial root dependent selection constants, one constant per digit value. Thus,

number

[0142] The selection constant depends on the partial root. The seven most significant bits of the partial root are used to derive a set of eight 11-bit selection constants. Different partial root values ​​can select different sets. Note that the partial roots are in the interval [0.5, 1]; the value S[i] = 1 is possible until a non-zero digit is generated. Thus, considering that the partial root has 1 integer bit (which is 0 after the first non-zero and negative digit is generated) and 6 fractional bits, and that the minimum value of the partial root is 0.5, the selection constant can be stored in a 33 × 88-bit look-up table (LUT) with 32 entries for S[i] ∈ [0.5, 1] ​​and 1 entry for S[i] = 1 (although an offset LUT can be used to reduce the storage size of the square root comparison constant, as described below in several techniques).

[0143] A simple implementation of the radix-64 square root with two radix-8 iterations Each radix-8 iteration produces 3 bits of result. Two radix-8 iterations can then be stacked to obtain 6 result bits per cycle, corresponding to the square root of the radix-64 number. A simplified implementation is shown in Figure 2. Two identical radix-8 sub-iterations are connected to obtain a radix-64 iteration. Note that only the most significant bit of the remainder is used to select the quotient digit. 11-bit remainder estimate

number

[0144] Thus, in each sub-iteration, ● The carry propagate adder 30 receives the remainder value rem[i] 31 generated in the previous partial iteration, represented in a redundant representation. The carry save adder 30 generates a non-redundant remainder estimate of a portion of the most significant bits of the remainder value 31 by performing a carry propagate addition of the most significant bits of the two words of the remainder value 31 (e.g., if the representation with positive and negative words described above is used, the negative word is subtracted from the positive word). A digit selection comparator 32 compares the remainder estimate with each of a set of comparison constants 34 to determine the next root digit 33 . ● Remainder adjust value generation circuit 36 ​​generates a remainder adjust value 39 corresponding to the "d-vector" or d[i+1] term shown in equation 17 above. Thus, for a square root operation, the remainder adjust value depends on the partial root value 37 received from the previous partial iteration and the next root digit 33 selected by digit selection comparator 32. The term "d-vector" is used simply as a label for the term d[i+1] because the number of bits in the value matches the number of bits used for vector operands in some implementations; however, it should be noted that this term does not imply that a "d-vector" is a single instruction, multiple data (SIMD) vector operand containing multiple independent data elements; a "d-vector" is a single data value, rather than a vector of multiple independent data values. ● The remainder update circuit 38 (including a 3:2 carry-save adder) updates the previous remainder 31 received from the previous sub-iteration based on the remainder adjust value 39 by adding the positive and negative words of the previous remainder 31 to the remainder adjust value 39 to generate an updated remainder 40 (still in the redundant representation) that is fed to the next sub-iteration to become the previous remainder 31 for that sub-iteration. On the path between outputting the updated remainder 40 in one sub-iteration and inputting the previous remainder 31 to the carry-save adder in the remainder update circuit 38 of the next sub-iteration, a 3-bit left shift is applied to represent the 8×rem[i] term in equation 18 above. ● On-the-fly conversion circuit 42 inserts a value determined based on selected root digit 33 into partial root value 37 to generate an updated partial root value 43 that is output to become partial root value 37 in subsequent partial iterations. The on-the-fly conversion can be performed in accordance with equations 6-8 above. Thus, although not shown in Figure 2 for simplicity, the partial root values ​​can be represented as two separate forms P and PM, as explained above, to simplify the on-the-fly conversion that can then be performed as a concatenation.

[0145] The updated remainder 40 and updated partial root value 43 from one sub-iteration become the previous remainder 31 and partial root value 37 of the next sub-iteration. Similarly, the updated remainder 40 and updated partial root value 43 from the last sub-iteration in one iteration become the previous remainder 31 and partial root value 37 of the first sub-iteration in the next iteration.

[0146] However, this simple implementation is too slow. To speed up the cycle, several techniques are used, which are described in the next section.

[0147] Base 64 Square Root Recursion 3 shows a square root processing circuit for performing digit repetition cycles corresponding to a single radix-64 square root iteration. In this example, the square root processing circuit is an iterative unit in which the output of one iteration is fed back as input to the same unit in a subsequent iteration, and flip-flop 50 latches the value passed each cycle. However, as further described below with respect to FIG. 9, the square root processing circuit can also be used in a pipelined implementation.

[0148] The square root processing circuit includes several parts: (1) remainder update circuit 34, (2) digit selection circuit (root digit calculation) 32, and (3) remainder estimation circuit 30. The connections between these components are also shown. Each of these parts will be described in detail below. The square root processing circuit also includes an on-the-fly conversion circuit 42, which will be described in more detail below. The on-the-fly partial root conversion maintains two partial root forms S[i] and SM[i], where SM[i] is the partial root S[i] minus one, and SM[i]=S[i]-8 -i (20) These two forms are used in some parts of radix-64 iteration. In addition, S3[i]=3×S[i] S3M[i]=S3[i]-8 -i They are also required for on-the-fly partial root conversion, as described in more detail below with respect to Figures 13-16. The use of S3[i] and S3M[i] simplifies the process of multiplication of ±3 root digits.

[0149] As shown in Figure 3, when a radix-64 iteration is split into two radix-8 sub-iterations, there are two instances each of remainder estimation circuit 30, digit selection circuit 32, and remainder update circuit 34, one for each radix-8 sub-iteration, although there may be some overlap between the circuits used for each sub-iteration, as explained further below. There may also be two instances of on-the-fly conversion circuit 42 for performing on-the-fly conversion using the radix-8 root digit obtained in each radix-8 sub-iteration, although for simplicity this is shown as a single block in Figure 3.

[0150] Remainder Update 4 shows in more detail the remainder update circuit 30 for performing a remainder update in a single radix-8 sub-iteration (which can be either the first or second radix-8 sub-iteration within a radix-64 iteration). The remainder update for each iteration of the cycle (see Equation 16) is done speculatively, i.e., updated remainder values ​​rem[i+1] for all possible values ​​of the root digit are calculated, and the root digit s i+1 Once the next root digit s is known, the correct remainder is selected. i+1 The system has several replica circuit units 60 that generate candidate output values ​​for each of the updated remainders corresponding to different options of s. i+1 = 0, there is no replica circuit unit 60 provided, because the above equation 18 means that the updated remainder rem[i+1] can be obtained directly from the previous remainder value rem[i] without addition. The sign of the previous remainder estimate is used to reduce the number of speculative remainders. If the remainder estimate is positive, the root digits can only be {+4, +3, +2, +1, 0}. On the other hand, if the remainder estimate is negative, the root digits can only be {-4, -3, -2, -1, 0}.

[0151] Thus, each replica circuit unit 60 has a carry-save adder 38 and a selection multiplexer 62 for selecting between alternative values ​​calculated in logic block 64 for positive and negative root digits of equal magnitude depending on the sign of the previous remainder estimate received from the previous partial iteration or iteration. This reduces the number of replica units required (four replica circuit units 60 are sufficient to accommodate digits ±1, ±2, ±3, ±4 respectively, instead of requiring eight to process each positive / negative digit separately).

[0152] The replicating circuit unit 60 constructs vectors d[i+1] (sometimes called F[i+1]) for all root digit values ​​other than 0, both positive and negative.

number

[0153] Thus, Figure 4 shows the digit bits that are concatenated in the on-the-fly calculation of each possible d[i+1] vector. The mask mask[i] signals the position at which the root digit must be concatenated (the mask is shifted by 3 bits between sub-iterations so that each successive base-8 root digit is concatenated at a position 3 bits lower than the position at which the previous base-8 root digit was inserted).

[0154] The blocks 64 labeled fda_pos and fda_neg for x=1, 2, 3, 4 are respectively |s i+1 = a| has a value corresponding to a positive or negative digit * S[i] or 2 * Perform the concatenation of SM[i] to represent the d-vector d[i+1] according to Equation 21, and also use -a×d[i+1] (the term -s in Equation 18 above) i+1 × F[i+1]) to generate d-vectors fd1, fd2, fd3, and fd4.

[0155] In the recurrence formula, d[i+1] is s i+1 To prevent 3x multiplication, i+1 The case of =±3 is handled differently, and 3×d[i+1] is constructed by blocks fd3_pos or fd3_neg using 3×S[i] directly as follows: 3×d[i+1]=2×(3×S[i])+(3×s i+1 ) x 8 -(i+1) (twenty two)

[0156] In this case, |3×s i+1Concatenate |=9, which requires 4 bits to represent. This is not a problem because a left shift of 3×S[i] by 1 bit leaves room for the additional bit. Then

number

[0157] The remainder guess code is used to select a positive or negative d[i+1] set before the 3-to-2 carry-save adder 38. This way, the result is that only 5 speculative remainders are calculated instead of 9.

[0158] The reciprocal of the remainder estimation code is placed in the least significant bit of the speculative remainder carry word, so that if the remainder estimation code is 1, the least significant bit of the speculative remainder carry word is 0, and if the remainder estimation code is 0, the least significant bit of the speculative remainder carry word is 1. This means that, as shown in equation (18), if the digit is positive (the remainder estimation code is 0), the term s i+1 ×F[i+1] must be subtracted. i+1 × F[i+1]. The two's complement is the term s i+1×F[i+1] is obtained by bit-complementing and adding 1. For example, the two's complement of 11100010 is 00011101+1=00011110. Therefore, this term is bit-complemented in the fd1_pos, fd2_pos, fd3_pos, and fd4_pos modules in Figure 4, and the "+1" is added by changing the least significant bit of the carry word, which is 0 by definition, to 1. In this way, no additional adder is required to complete the two's complement calculation. If the digit is negative (the remainder estimation code is 1), the operation in equation (18) is an addition, so there is no need to perform two's complement, and the least significant bit of the carry word is kept 0. Therefore, in summary, the reciprocal of the remainder estimation code is placed in the least significant bit of the carry word.

[0159] Among these speculative remainders provided by the replica circuit unit 60, the digit s is not included, since it does not require additional hardware. i+1 = 0 blocks fda_pos and fda_neg, and the next root digit s i+1 is determined by digit selection circuit 68, it is only an additional input in multiplexer 32 which acts as a selection circuit to select the correct candidate output value.

[0160] Each carry-save adder 38 receives two terms, the positive and negative words of the previous remainder rem[i], which are redundantly represented, and the -s in equation (18) represented by fd1-fd4. i+1 The output of each carry-save adder 38 is a candidate value for selection as the updated remainder rem[i+1], which is still in the redundant representation and therefore contains two terms, one positive and one negative. As in the case of root digit = 0, the candidate value is simply 8. * rem[i], no addition is required since there is no carry-save adder 38. The 5:1 multiplexer 68, which functions as a selection circuit, selects the root digit s selected by the root digit selection circuit 32. i+1to select among the candidate output values ​​according to , and provide the updated remainder rem[i+1].

[0161] Residual estimate 5 shows the remainder estimation circuit 30 for the first and second partial iterations. The remainder estimation is an early speculative calculation of the 11 most significant bits of the remainder for use in root digit selection. This allows for better timing because the remainder estimation is removed from the critical path through the root digit calculation.

[0162] Two different situations are shown. 1. The remainder estimate in the first sub-iteration to generate a remainder estimate to be used for digit selection in the second sub-iteration within the cycle. This is done during the first iteration based on the speculative remainder obtained by the first sub-iteration remainder update circuit 34, as shown in Figure 4. Thus, five carry propagate adders 70 calculate the remainder estimate based on the most significant bit of the sum and the speculative remainder (rem) obtained by the first sub-iteration remainder update circuit 34. d4 [i+1] to rem d1 Add the carry words of rem[i+1] and rem[i]. i+1 is known, the appropriate remainder estimate for root digit selection in the second partial iteration of the cycle is selected by multiplexer 72. Thus, this is another example of a replication circuit that includes replication circuit unit 70 and selection circuit 72. 2. A remainder estimate in the second partial iteration to generate a remainder estimate to be used for digit selection in the first partial iteration of the next cycle (the value output by remainder estimation circuit 30 in the second iteration can be flipped in flip-flop 50 ready for use in the next cycle, as shown in FIG. 3). The remainder estimate generated by remainder estimation circuit 30 in the second partial iteration is an assimilation of the most significant bit of 8×rem[i+2], which can be derived from rem[i] input as the previous remainder value in the first partial iteration as follows (based on replacing rem[i+1] with another instance of Equation 18 that relates rem[i+1] to rem[i] using the relationship from rem[i+2] to rem[i+1] using Equation 18):

number

[0163] This is calculated during the first and second iterations of the cycle as follows: msb_first=64×(8×rem[i]-s i+1 ×d[i+1]) (25) and msb_rem[i+2]=msb_first-8×s i+2 ×d[i+2] (26) where equation (25) is evaluated during the first subiteration and equation (26) is evaluated during the second subiteration. Both equations are evaluated speculatively for five possible remainders.

[0164] Note that the difference between equation (18) and equation (25) is a 64X factor, a 6-bit left shift. Then, if a 17-bit adder is used instead of two 12-bit adders, both equations can be evaluated with the same logic, with the 11 most significant bit being the remainder estimate calculated in the first partial iteration for use in digit selection in the second partial iteration of the cycle, and the 13 least significant bit being used to complete the remainder estimate calculation during the second partial iteration to obtain the remainder estimate used in digit selection in the first partial iteration of the next cycle in equation (26).

[0165] Thus, in this approach, adder 70 in the first sub-iteration calculates some additional (least significant) bits that are not actually needed in the remainder estimate used for digit selection in the second sub-iteration, but by calculating these additional bits, it allows the term msb_first shown above to be calculated in the first sub-iteration, reducing the overall circuit area compared to if a separate adder had calculated these bits in the second sub-iteration.

[0166] Adder 74 in the remainder estimation circuit for the second sub-iteration evaluates equation 26, which depends on msb_first and d-vectors 0,fd1[i+2] through fd4[i+2], which is s i+2 =0,s i+2 =±1~s i+2 = ±4, the term 8×s in the equation i+2 ×d[i+2], respectively. These vectors are generated as part of remainder update circuit 34 in the second sub-iteration of the cycle (see fd1-fd4 in FIG. 4). This approach means that there is no need to wait for carry-save adder 38 in second sub-iteration remainder update circuit 30 to perform their additions before commencing the additions by carry-propagate adder 74 in remainder estimation circuit 34 for the second sub-iteration. Instead, the calculation of the updated remainder estimate in the second sub-iteration can be performed in parallel with the remainder update in the second sub-iteration to remove latency from the critical timing path. This improves performance.

[0167] Root digit selection 6 shows the root digit calculation performed by digit selection circuit 32 (which can be either the first or second radix-8 sub-iteration within a radix-64 iteration). The calculation of the root digit was outlined above; the remainder estimate is compared to each of the eight comparison constants, and a digit is selected according to equation (19). The root digit is stored as a 1-hot 9-bit vector s[i], i=8,...,0, where s[i]=1 for digit i-4; for example, if the root digit is -1, then s[3]=1, and the 9-bit vector is s={0,0,0,0,0,1,0,0,0}.

[0168] This is shown in Figure 6. There is a set of 11-bit comparators 80 to compare the remainder estimate with each comparison constant. The carry output, ge-output, of each comparator is set to 1 if the remainder estimate is greater than the comparison constant. The ge-output and sign of the remainder estimate are then input to a set of nand and or gates to generate each bit of a 1-hot 9-bit vector.

[0169] The selection constants required for root selection are derived from values ​​stored in a look-up table (LUT). The selection constants for each radix-8 iteration depend on the partial root values ​​preceding that partial iteration, such that each partial iteration uses a different set of comparison constants. However, it has been derived that the same set of selection constants can be used for all partial iterations except the first two partial iterations. As further explained below with respect to the pipelined example of FIG. 9, the selection of the first few root digits can be done in a pre-processing stage, allowing the same selection constants to be used for each iteration, thereby avoiding a main iteration cycle that requires a separate LUT lookup.

[0170] Integrate A block diagram of the digit recursive square root processing cycle is shown in Figure 7. The different parts (remainder update circuit 34, remainder estimation circuit 30, root digit selection circuit 32, and on-the-fly root conversion 42) are identified by dotted lines, and the relationships between these parts are also shown.

[0171] As shown in more detail above, several parts of the cycle logic use speculation and duplication to meet timing constraints. Thus, duplication is used in several places to obtain speculative results for each digit value. In most cases, duplication is reduced by using the sign of the remainder to have the same logic for positive digit values ​​and their negative counterparts. In this way, the logic is duplicated five times instead of nine, resulting in a significant area reduction. Once the root digit is known, the correct value is selected from among the nine or five speculative values.

[0172] In some parts, as with the remainder update in the first and second sub-iterations and the remainder estimate in the second sub-iteration, the logic is replicated only four times, but the selection is done with a 5-to-1 mux, because one of the inputs to the mux is one of the inputs to the replicated logic (thus not requiring a replicated circuit unit to compute a new value for the speculative candidate value).

[0173] Accordingly, Figure 7 illustrates one example of a square root processing circuit that may be used in divide / square root unit 24 of Figure 1. In some examples, divide / square root unit 24 may also include a separate instance of a divide processing circuit that performs division operations in response to a divide instruction, without sharing circuitry and data paths between the square root processing circuit and the divide processing circuit.

[0174] However, as further described below with respect to FIG. 8, in some examples, the techniques described above for square root processing circuits may be used in a combined division / square root processing circuit that is also capable of performing division operations, in which case the combined division / square root processing circuit also functions as the aforementioned "square root processing circuit."

[0175] A radix-64 combined divide / square root circuit for shared divide and square root iterations. FIG. 8 shows an example of a combined divide / square root circuit for performing radix-64 division / square root iterations, which may be provided as part of the divide / square root unit 24 of FIG. 1. The combined divide / square root circuit uses shared circuitry and shared data paths to perform both division and square root operations, both with the same radix-64. The same number of radix-64 iterations are performed per cycle for both the division and square root operations (in this example, a single radix-64 iteration of digit recursion is performed per cycle for both the division and square root operations). As with the square root example above, in this example, the radix-64 iteration is divided into two overlapping radix-8 sub-iterations. The combined divide / square root circuit receives as an input a signal "div / sqrt" that indicates whether the current operation is a division or square root operation. This signal may be controlled by the instruction decoder 6 based on whether the instruction being processed is a division or square root instruction.

[0176] 3-7 for the square root example, and therefore performs the square root operation in the same manner as described above. Much of this circuitry can be reused for the division operation, so the data paths for generating the updated remainders rem[i+1], rem[i+2], remainder estimates rem_est[i+1], rem_est[i+2], and partial result values ​​S[i], SM[i] for the square root operation are also used to generate the corresponding values ​​for the division operation. (The notation Q[i], QM[i] is used for the partial result values ​​when the division operation is performed, but these are on the same data paths as the partial root values ​​S[i], SM[i] generated for the square root operation.)

[0177] Figure 8 shows the microarchitecture of a radix-64 divide / square root iteration. The two radix-8 sub-iterations that make up the radix-64 iteration are separated, the first sub-iteration at the top and the second sub-iteration at the bottom. The two iterations are very similar, but there are some differences that will be addressed later.

[0178] As stated in equations (1) and (3) above, the result after iteration i is defined by the partial result P[i] (which can be the partial quotient Q[i] or the partial root S[i]) and the remainder rem[i]. Then, each iteration includes several steps.

[0179] 1. Digit selection New result digits are generated from the remainder and divisor (in division) or partial root (in square root) using low-precision estimates instead of full-precision values ​​(see equation (2)). Thus, the combined division / square root unit 24 includes a shared digit selection circuit 32 that selects, for each radix-8 partial iteration, the next radix-8 digit of the division / square root result based on a comparison of the previous remainder estimate rem_est[i], rem_est[i+1] with a set of comparison constants. The remainder estimate word length is different for division and square root.

[0180] As already discussed above for the square root example of FIG. 6, digit selection is performed by comparing the remainder estimate to a set of eight selection constants. This set depends on the most significant bit of the divisor or partial root. The set of comparison constants is stored in a look-up table (LUT) that is addressed by the most significant bit of the divisor or partial square root (as further explained below). Error analysis of the radix-8 division and square root algorithms shows that the number of bits in the comparison constant and remainder estimate differs for the two operations: 11 bits for the square root and 10 bits for the division. However, if an 11-bit remainder estimate is used for both the division and the square root, both operations can be placed in the same logic. In this case, the comparison constant for the division is extended to 11 bits by placing a zero in the least significant bit position. In this way, the remainder estimation logic 30 and digit selection circuit 32 in the first and second partial iterations are shared between the division and the square root.

[0181] Thus, the comparisons for digit selection are performed for both the division and square root operations using the same set of comparators 80. The operation of digit selection circuit 32 is the same for both the division and square root operations (as described above with respect to FIG. 6 for the square root), except that it receives a different set of comparison constants for comparison with the 11-bit remainder estimate.

[0182] 2. Remainder Update The result digits so generated are used to update the remainder and partial results (equations (1) and (3)). Thus, a shared remainder update circuit 34 is provided in each partial iteration to adjust the previous remainder value rem[i], rem[i+1] based on the remainder adjustment value in a given radix-8 partial iteration to generate updated remainder values ​​rem[i+1], rem[i+2] in the redundant representation.

[0183] 4, a replication circuit unit is provided to generate candidate remainder values ​​for the different possible values ​​of the selected result digit (sharing circuitry between positive / negative digits of the same magnitude as described above to reduce the amount of replication required), and then a 5:1 multiplexer 68 selects one of the candidate values ​​depending on the next result digit selected by digit selection circuit 32. Carry-save adder 38 and fd calculation unit 64 are the same as in FIG.

[0184] However, as shown in equation (4), the remainder adjustment value (F[i+1] term) used in the remainder update is different for division and square root. In the case of square root, F[i+1] is the root digit s i+1 to the shifted partial root. This means that F[i+1] is calculated for each iteration by the fd calculation unit 64. However, for the division F[i+1], it is the divisor d that does not change between iterations.

[0185] Therefore, by adding the XOR gate 90, the −p in equation (3) that occurs when a division operation is performed (when F[i+1]=d as shown in equation 4) is i+1×d term. One XOR gate XORs the divisor d with the inverse of the previous remainder estimate rem_est[i], rem_est[i+1] to provide multiplication by -1. In other words, as in division, the remainder update uses multiples of +d or -d. In the case of a positive remainder, the divisor is complemented to obtain a negative multiple of the divisor. For the replicated units that calculate the candidate remainder values ​​corresponding to the ±2 and ±4 root digits, a left shift of 1 or 2 bits is applied to the paths from the XOR gate to obtain the p required in equation (3). i+1 For the square root, to avoid the need for a triple (3×d multiplications are pre-computed before the iteration to have fast iterations), a separate representation of 3 times the divisor 3×d is used, so a second XOR gate similarly XORs 3×d with the inverse of the sign of the previous remainder estimate to provide an input to the replica circuit unit which is computing candidate remainders for ±3 root digits.

[0186] 4 for the square root example is replaced with a set of 3-to-1 multiplexers 62 in FIG. 8 to select the appropriate F[i+1] value for division or square root. Each 3:1 multiplexer 62 selects the corresponding value received from XOR gate 90 based on its divisor if operation type signal div / sqrt indicates a division operation is to be performed. If operation type signal div / sqrt indicates a square root operation is to be performed, the associated one of the d-vector values ​​produced by fd1-fd4 calculation blocks 64 is selected based on the sign of the previous remainder estimate, as described above for FIG. 4. Thus, 3:1 multiplexer 62 functions as a selection circuit to select as the remainder adjust value either a value derived from the divisor value d when performing a given base-8 partial iteration as part of a base-64 division operation, or a value derived from a partial root value that depends on a previously selected series of base-8 root digits when performing a given base-8 partial iteration as part of a base-64 square root operation. The sharing of carry-save adder 38 and 5:1 multiplexer 68 between both operations provides circuit area savings.

[0187] 3. Residual estimates The remainder estimates are obtained to be used in digit calculations in the next sub-iteration. Thus, in a given radix-8 sub-iteration, there is a shared remainder estimation circuit 34 that generates updated remainder estimates rem_est[i+1], rem_est[i+2], which are non-redundant estimates of some of the updated remainder values ​​rem[i+1], rem[i+2] generated in redundant representation by remainder update circuit 30 in the given radix-8 sub-iteration. The remainder estimation circuit 30 is the same as that described above in FIG. 5 for the square root operation. Again, in the second radix-8 sub-iteration, remainder estimation circuit 30 determines the updated remainder estimate rem_est[i+2] in parallel with remainder update circuit 34, which generates the updated remainder value rem[i+2].

[0188] 4. On-the-fly conversion The partial result P[i] (quotient Q or root S) is converted from a signed digit redundant representation to a conventional binary non-redundant representation using on-the-fly conversion (equations (7) and (8)). In typical on-the-fly conversion schemes, the partial root is used in the next digit selection and remainder update for square root operations, but the fact that the partial quotient is not for division operations leads to different partial quotient update and partial root update methods. This difference is shown below (digit

number

[0189] [Table 1]

[0190] For division, each time a new digit (3 bits in base 8) is generated, the actual partial quotient is typically shifted left, with the new digits placed as the three least significant bits. In this way, the actual partial quotient is always in the left significant part. The previously inserted bits are shifted left to more significant bit positions. For square roots, on the other hand, the new root digit is concatenated to the actual partial root so that the most significant bit of the partial root is always in the most significant part of the stored data value, and masks mask[i], mask[i+1] are used to keep track of the position to which the next digit must be concatenated, as described above for square root operations.

[0191] In order to share the on-the-fly conversion logic between the division and square root, it has been decided to perform a partial quotient update as is done for a partial root update, i.e., concatenate the new quotient digits using a mask to indicate where the digits must be concatenated. This is unconventional, but means that increased sharing of datapath and circuit logic is possible.

[0192] Thus, in the first partial iteration, the shared on-the-fly transformation circuit 42 selects a position for inserting the next digit into the partial result values ​​Q[i], QM[i], S[i], SM[i] based on the mask mask[i] for both the division and square root operations. Similarly, in the second partial iteration, the shared on-the-fly transformation circuit 42 selects a position for inserting the next digit into the partial result values ​​Q[i+1], QM[i+1], S[i+1], SM[i+1] based on the mask mask[i+1] for both the division and square root operations. The mask is shifted right by three bits for each partial iteration, so that each result digit is inserted three bits to the right of the previous one.

[0193] With respect to the square root example described above with respect to FIG. 7, the combined division / square root processing circuitry can be used either in an iterative unit where the output produced in one iteration, labeled "i+2," is fed back as the input labeled "i" for the next iteration of the square root or division operation, or in a pipeline unit as described further below with respect to FIG. 9.

[0194] Divide / Square Root Pipeline The long latency of traditional division and square root implementations and the complexity of each stage, which has separate logic for division and square root, preclude the use of pipelined floating-point divide and square root units in commercial processors. Instead, commercial processors have iteration units in which parts of the logic are used across several cycles, resulting in low-bandwidth designs. In a typical scheme, the iteration logic consists of two separate parts, a division iteration and a square root iteration, with very little, if any, logic shared between both operations. To increase bandwidth, several iteration div / sqrt units are deployed operating in parallel. For example, one design has two iteration floating-point div / sqrt units performing double-, single-, and half-precision operations, and two other smaller iteration units performing single- and half-precision operations. In this way, the double-precision div / sqrt bandwidth is doubled, while the single- and half-precision division and square root bandwidth is quadrupled relative to a configuration with only div / sqrt iteration units.

[0195] In the approach shown in FIG. 9, a single pipelined div / sqrt unit 24 is instead provided. To overcome the drawbacks that prevent the use of such a unit, we developed a low-latency division and square root implementation and a common stage for division and square root, in addition to some other logic shared between both operations. Low latency is achieved by implementing a radix-64 digit-recursive division and square root algorithm with two radix-8 iterations per cycle. Such an algorithm produces a 6-bit result per cycle, as previously explained. Meanwhile, having the same algorithm for division and square root, along with a thorough stage design, allows for reduced area requirements. As a result, we have been able to design pipelined floating-point div / sqrt units for double, single, and half precision with relatively small areas. Compared to the alternative configuration described above using two double / single / half precision units and two single / half precision units, the bandwidth is significantly improved for double precision and single precision, and more modestly improved for half precision, but the circuit area of ​​the pipelined unit can be smaller than the total area of ​​the alternative configuration. Thus, the pipelined unit allows for a combination of low latency and high bandwidth to be obtained, resulting in a high performance div / sqrt unit 24.

[0196] 9, the pipeline unit 24 includes a pre-processing circuit 100, a main body of the pipeline 102 for performing digit recursion iteration, and a post-processing circuit 104. Most of the pre-processing and post-processing logic is shared between the division and square root operations, and the iterative part, the digit iteration, is spread across several pipelined radix-64 shared stages 110.

[0197] Preprocessing circuit 100 performs various preprocessing operations, including unpacking the operands, normalizing the operands (if necessary), and initializing (e.g., retrieving comparison constants and selecting one or more initial result digits).

[0198] The body of the pipeline 102 performs the digit recursion, which is the iterative portion of the digit recursion algorithm. The body of the pipeline 102 comprises several divide / square root pipeline stages 100, each of which contains an instance of the combined divide / square root processing circuitry shown in Figure 8. Thus, each pipeline stage 110 within the body 102 performs a radix-64 digit recursion floating-point division operation, q = x / d, or a radix-64 digit recursion square root operation,

number

[0199] Post-processing circuitry 104 includes rounding logic and right shifting in case of subnormal results (division only).

[0200] The pipeline unit handles three different floating-point precisions: double, single, and half (DP, SP, HP), respectively, resulting in different latencies for division or square root operations for operations of different precisions. Nevertheless, for a given precision, the latency is the same for both division and square root due to simple scheduling of the timing of the post-processing stages.

[0201] A more detailed description of the pipeline, focusing on the processing of the mantissas of the input operands x and d to produce a result, is provided below. It will be understood that the exponents of the input operands x and d are also processed. This can be done according to any known technique. For example, in the case of a division, the result exponent may correspond to the difference between the true exponents of the input operands x and d, adjusted for any right shifts in post-processing stages required for subregular expression processing. In the case of a square root operation, the result exponent may correspond to half the true exponent of the input operand x, also adjusted for any normalization applied. Here, "true exponent" refers to the significant power of two represented by the exponent of the floating-point number (removing the exponent bias applied according to the floating-point precision being used).

[0202] Preprocessing (V1, V2) Preprocessing circuit 100 performs preprocessing including unpacking floating-point operands to extract the sign, mantissa, and exponent, determining special conditions (subnormal, 0,...), normalizing the operands (e.g., handling subnormals), and look-up table (LUT) addressing to obtain selection constants needed for digit selection. When dividing by two subnormal operands, both operands are normalized in the same cycle.

[0203] Additionally, the first base-8 digit is taken. In floating-point division, the first digit can only take the values ​​{+1, +2} and is the integer digit of the quotient. In floating-point square roots, the first base-8 digit can only take the values ​​{-4, -3, -2, -1, 0} and the calculation is easily fused with the remainder and partial root initialization.

[0204] In the case of square root, the second digit is also obtained. As described above, the LUT stores the selection constants necessary for digit selection. However, in the case of square root, the selection constants for each octal iteration depend on the partial root value before that iteration such that each iteration uses a different set of comparison constants. This imposes severe timing and area constraints as the iterative logic should include the LUT and it should be read each time a new iteration starts. However, it has been derived that (by error analysis) for square root in base 8, the same set of selection constants can be used for all iterations except the first two iterations (even if the same set of selection constants is used after the first two iterations, sufficient accuracy is given to the result). Therefore, at this stage, the second root digit is obtained, then the LUT is read, and the set of selection constants thus obtained is flopped for use in digit selection in the remaining iterations.

[0205] In the case of division, several other operations are performed. To save iterations in single precision, the quotient q is forced to be such that q ∈ [1,2). Note that q < 1 only occurs when x < d. This situation is detected in the preprocessing and the dividend when q is left-shifted by 1 bit such that q = 2×x / d and q ∈ [1,2). Of course, the mantissa is the same as x / d, but the exponent needs to be decremented. Finally, 3×d = 2×d + d is calculated to be used in octal iteration, avoiding the need for 3x multiples to be calculated in each iteration and saving time.

[0206] The preprocessing stage is divided into two cycles of V1 and V2 such that unpacking, classification, and normalization of the operands, and the first root digit (square root) are done in V1. On the other hand, in V2, the following operations are performed: calculation of the second root digit (square root), calculation of the first quotient digit (division), comparison and conditional shift of x and d of the quotient (division), calculation of 3×d (division), and LUT addressing to obtain the comparison constants for the remaining iterations (division and square root).

[0207] First division digit selection and first two square root digit selections The following provides more information on how to select the first radix-8 division result digit and the first two radix-8 square root result digits in preprocessing circuit 100.

[0208] context ● Base 64 division and square root ● Each base 64 iteration consists of two base 8 iterations. ● Division: ○ The first iteration occurs before the iteration part Reason: ■ Before the iteration portion, a constant look-up table (LUT) is addressed to obtain the comparison constants needed for quotient digit selection for each radix-8 iteration. ● The LUT is addressed by the most significant bit of the divisor. ■ All iterations use the same set of comparison constants. ■ The first base-8 quotient digit can only take the values ​​+2 or +1. This means that the first iteration is much simpler than the remaining iterations. ■ In the same cycle that the LUT is addressed, there is time to perform the first divide iteration. ■ By having the first iteration of the LUT cycle, the final latency could be reduced by one cycle with some precision. ● Square root: ○ The LUT is addressed by the most significant bit of the partial route ○ The first and second iterations are performed before the iteration part Reason: ■ The radix-8 square root algorithm requires different sets of comparison constants for the first iteration, the second iteration, and the remaining iterations. ■ It has been determined that the first and second iterations are performed before the iterative portion of the square root calculation in order to have common square root iterative logic and to avoid LUT addressing in the iterative logic. ■ The first iteration occurs in the first cycle V1, along with unpacking the operands and determining special operands. ■ The second iteration occurs in the same cycle V2 as the LUT addressing to obtain the comparison constants for the remaining iterations. This cycle precedes the iterative portion of the algorithm.

[0209] Division: First base 8 digit (in V2) • The first base 8 division digit is selected using the same set of constants as the rest of the iteration, so the constants for this first digit selection and for digit selection in subsequent iterations are taken from the LUT. ● In this cycle ○ The LUT is addressed, ○ To perform the first iteration, a constant of digit = +2 is used A set of comparison constants is flopped for use in the remaining iterations. • Next, the first iteration uses the same set of constants as the remaining iterations, but due to the limited digit values, only the constant for digit = +2 is needed.

[0210] Square root: first base 8 digit (in V1) ● For base 8 iteration, the idea is the same, but the logic is not the same as for base 4. ○ Partial route is 1 (default value) The first base 8 digit can take on the values ​​-4, -3, -2, -1, or 0 o Given a partial route, the comparison constants for these five numbers are known and are wired into the first digit selection logic (only four values ​​need to be stored), so no LUT addressing is required for this. ○ These four values ​​are (comparison cte * 64, i.e., the values ​​quoted below are 64 times the actual stored constants): Digit constant = 0:-64 Digit constant = -1:-176 Digit constant = -2:-272 Digit constant = -3:-352.

[0211] Square root: second base 8 digit (in V2) ● The range of values ​​of the partial root after the first iteration is limited, only five values ​​are possible (a different partial root value for each value of the first digit): ○ First digit = 0 => Next partial root is 1.00_000 ○ First digit = -1 => Next partial root is 0.11_000 ○ First digit = -2 => Next partial root is 0.10_000 ○ First digit = -3 => Next partial root is 0.01_000 ○ First digit = -4 => Next partial root is 0.00_000 ● A small LUT is used to store these five sets of comparison constants. ● The size of this LUT is 5x88. ○ 5 rows ○ 8 bits per column to store eight 11-bit comparison constants ○ Addressing with the above partial route ○ The value stored in the LUT (again, the constant value shown is 64 times larger than the stored value) * 64): The partial root is 1.00_000=>461, 326, 191, 61, -62, -192, -317, -442.

[0212] Partial route is 0.11_000 => 406, 281, 171, 61, -62, -172, -277, -377 Partial root is 0.10_000 => 351, 241, 141, 46, -47, -142, -232, -322 Partial root is 0.01_000 => 291, 206, 121, 41, -42, -122, -192, -267 The partial root is 0.00_000 => 236, 161, 96, 31, -32, -97, -152, -212 The order of the above constants is constant: digit = +4, digit = +3, digit = +2, digit = +1, digit = 0, digit = -1, digit = -2, digit = -3.

[0213] This describes the initial digit selection of the pre-processing circuitry. Digit selection in subsequent stages is as described above in Figure 6, with reference to comparison constants shown in the LUTs described further below in Figures 17-20.

[0214] Digit repetition in pipelined divide / square root units For a general base r and a call to the number of bits in the result n, the number of iterations is

number

[0215] We detail radix-64 (r=64), two operations (division and square root), and three floating-point precisions (DS, SP, HP). The number of fractional bits per precision is 52, 23, and 10, respectively. One radix-64 iteration is performed per cycle; as mentioned above, to obtain a handy implementation, the radix-64 iteration is obtained by stacking two simpler radix-8 iterations per cycle. However, the number of iterations is still that of the radix-64 algorithm.

[0216] Floating-point division: The first digit that generates the integer bits of the final quotient is selected in preprocessing. Additionally, if the quotient is forced to [1;2), only guard bits are needed for rounding; no rounding bits are used. Then, n=53, 24, 11 for double, single, and half precision, respectively. This includes the fraction and guard bits. Then, the number of iterations for the three precisions is:

number

[0217] Floating-point square root: Since the input operands are [0:25;1), the result is [0:5;1), so the result must be left-shifted to get the final floating-point result [1;2). Like division, only one additional bit, the guard bits, needs to be rounded. Therefore, the number of bits in the root algorithm needs to be 54, 25, and 12 for DP, SP, and HP, respectively. This includes the integer bits, fractional bits, and guard bits.

[0218] Meanwhile, the first two base-8 digits are obtained in preprocessing before the iteration. The first digit selection is skipped and integrated into the initialization of the partial root of the remainder, and the second digit selection is performed in V2 so as to have a single LUT for all iterations of the remainder. These two iterations generate 6 bits of the final root, after which the number of cycles of the iteration part is

number

[0219] Therefore, several multiplexers are added to the main body of the pipeline 102, A 2:1 multiplexer 120 in stage D2 is added to select between the outputs of stages D1 and D2, allowing stage D2 to be skipped when an HP square root operation is performed. This reflects the difference between the 2 cycles required for the division and the 1 required for the square root, as shown in equations (28) and (29). ● A multiplexer (not shown in Figure 9) is added within the combined divide / square root processing circuit to allow the output of the first partial iteration of stage D4 to be selected and output as the iteration result when the SP square root operation is performed (skipping the second partial iteration of stage D4). This avoids generating the extra three bits of the second partial iteration, and the two further bits generated in the first partial iteration can also be discarded as described above. • At stage D9 a 2:1 multiplexer 122 is added to select between the outputs of stages D8 and D9, allowing stage D9 to be skipped when a DP square root operation is performed. This reflects the difference between the 9 cycles required for a division and the 8 cycles required for a square root. ● A 3:1 multiplexer 124 in stage 9 selects between the outputs received from stages D2, D4 and D9 (with or without the square root skip mentioned above), the selection by multiplexer 124 being based on a control signal indicating the floating-point precision of the current operation, which is controlled by the instruction decoder 6 depending on the type of instruction decoded to control the division / square root operation.

[0220] Thus, the instruction decoder 6 acts as a control circuit that controls the pipeline so that at least one divide / square root iteration pipeline stage used to perform at least one iteration of a digit-recursive division operation or square root operation when producing a higher precision result is bypassed when performing a digit-recursive division operation or square root operation to produce a lower precision result (controlling the multiplexer 124 to select the output of an earlier stage when the bypass is applied).

[0221] The instruction decoder 6 can also control the division / square root pipeline to cause at least one division / square root iteration pipeline stage, which is used to perform at least one iteration when a digit recursive division operation is performed, and which completely or partially skips or discards some bits of the result output when performing the digit recursive square root operation (by controlling multiplexers 120, 122 and an internal multiplexer not shown in stage D4, allowing the second partial iteration of stage D4 to be skipped and bits to be discarded).

[0222] Post-processing (W0) As mentioned above, the post-processing is a rounding and right-shift of the result in the subnormal case. Any known floating-point rounding technique can be used here. Note that the result can only be subnormal in division; there are no subnormal results in square root. The post-processing is done in one cycle for both division and square root.

[0223] Two operations and three precisions in the same pipeline - on-the-fly conversion As mentioned above, the number of digit repetition cycles in DP and HP square roots, which are one less than division, is one less than in division (see equations (28) and (29)). To maintain the same latency and collect the results in the same cycle for both operations, an empty cycle is added to the square root. That is, the inputs to D2 and D9 are passed to the output without further conversion. Furthermore, in the SP square root, the second radix-8 iteration in the D4 cycle is skipped. Also, the latency differs for each precision. The undone result for DP is available in D9, while the undone results for HP and SP are available in cycles D2 and D4, respectively. Next, the W0 cycle operation saves the signals coming out of D2, D4, or D9, depending on the precision.

[0224] To achieve an efficient digit repeat cycle, the two operations are They share most of the logic, including the on-the-fly conversion circuit 42 for updating the partial quotient or root. However, before the first digit cycle D1, preprocessing has already generated the six fractional bits for square roots or the integer digits for divisions. The shared quotient / root update logic needs to have the same new fractional digit concatenation positions for divisions and square roots.

[0225] Thus, for division, six zeros are added to the fractional part of the quotient Q[i], QM[i] in preprocessing stage V2. The new fractional bits qi generated in each subsequent iteration are then concatenated after these zeros (in the same positions where the corresponding bits are concatenated for the square root operation, as indicated by the mask). 1:000 000 q1q2q3 q4q5q6... In post-processing stage W0, these zeros are removed before rounding to have an uncrounded quotient. 1:q1q2q3 q4q5q6... The addition of these zeros does not affect the final quotient precision because partial roots are not used in the digit recursive division equation, as shown in equation (4).

[0226] Thus, for a division operation, pre-processing stage V2 provides partial result values ​​to first division / square root iteration pipeline stage D1 in which selected bit positions are set to dummy bit values ​​(0 in this example), and these selected bit positions correspond to bit positions into which at least one pre-processing stage V1, V2 inserts at least one additional result digit not generated for the digit-recursive division operation when performing the digit-recursive square root operation. In post-processing stage W0, these dummy bit values ​​are removed.

[0227] Timing Control, Latency and Throughput The microstructure of the pipeline unit is shown in Figure 9. The unit consists of 12 stages, which is the latency of the slower operation, double precision division: two pre-processing cycles (V1, V2), nine digit repetition cycles (D1-D9), and one post-processing cycle (W0). For a given floating-point precision, division and square root operations have the same latency. ● Half precision, 5 cycles: V1-V2-D1-D2-W0 ● Single precision, 7 cycles: V1-V2-D1-D2-D3-D4-W0 ● Double precision, 12 cycles: V1-V2-D1-D2-D3-D4-D5-D6-D7-D8-D9-W0 (Note that even if a cycle is skipped due to the square root at D2 or D9, the latency is still the same as the input to the 3:1 multiplexer 124 that comes after the flip-flop at the input to stage D2 or D9.) Having the same latency for both operations simplifies timing control.

[0228] Additionally, the latency is the same whether or not there are subnormal operands or results, normalization (if necessary) is performed in V1, and the subnormal quotient right shift is performed in W0 after rounding.

[0229] A timing control circuit 130 is provided to control the timing at which the division and square root operations can be initiated. Although the timing control circuit 130 is shown as a separate unit in Figure 9, in other examples the decoder 6 can function as the timing control circuit 130.

[0230] The divide / square root unit 24 is fully pipelined, meaning that a new operation can be initiated every cycle of throughput 1 if all operations are of the same precision, which is the most common case. Thus, the control circuit 130 can control the divide / square root pipeline to perform a second digit recursive division or square root operation in a divide / square root iterative pipeline stage after the divide / square root pipeline, which is performing a later iteration of the first digit recursive division or square root operation, in parallel with a previous divide / square root iterative pipeline stage performing a first digit recursive division or square root operation and a previous iteration of the second digit recursive division / square root operation.

[0231] However, when there is a mixed-precision divide or square root, a constraint appears: the two operations cannot be in the same stage at the same time. As shown in Figure 10, because the latency depends on the precision, there are some prohibited start cycles for SP and HP operations. For example, an SP div / sqrt can start 5 cycles after a DP, because in that case both operations collide at W0.

[0232] Thus, as shown in FIG. 10 , the timing control circuit 130 can control the circuitry to prevent a less precise digit recursive division / square root operation performed to produce a less precise result from starting a predetermined number of cycles after a more precise digit recursive division / square root operation performed to produce a more precise result, where the predetermined number of cycles may correspond to the difference between the number of cycles required to reach at least one post-processing stage for the more precise digit recursive division / square root operation and the number of cycles required to reach at least one post-processing stage for the less precise digit recursive division / square root operation.

[0233] The predetermined number of cycles depends on the precision used. As shown in Figure 10, the predetermined numbers are as follows: - 5 cycles when low accuracy is SP and high accuracy is DP -7 cycles when low accuracy is HP and high accuracy is DP. -2 cycles when low accuracy is HP and high accuracy is SP.

[0234] In this case, as with the case where no collision occurs in the post-processing stage W0, it is acceptable to start a low-precision operation after a high-precision operation if the number of cycles between operations is greater than a predetermined number.

[0235] This approach can provide significant bandwidth improvements by using shared pipeline divide / square root operations, and the area reduction from sharing common logic provides a better balance between performance and circuit area.

[0236] Nevertheless, pipelining can also be used for either or both of the square root and divide units in an implementation with separate square root and divide units.

[0237] Also, while FIG. 9 shows the application of the pipeline method to digit-recursive division and square root of radix 64, the pipeline method can also be used for other values ​​of radix.

[0238] Also, while FIG. 9 shows a pipeline scheme that supports all of HP, DP, and SP, other examples may support only a subset of these precisions, or may support other floating-point precisions, and therefore may use a different number of pipeline stages.

[0239] On-the-fly conversion As previously mentioned, part of the digit recursion method may involve converting from a redundant representation to a normal binary representation (non-redundant representation). Because the output digits from the digit recursion method are generated one at a time, it is useful to be able to perform the conversion one digit at a time to avoid the latency that would occur if all digits had to be converted at once. This conversion is performed using on-the-fly conversion circuitry 42.

[0240] In simple terms, the on-the-fly transformation for square roots keeps two partial root words, S[i] and SM[i] (S[0]=1.0 and SM[0]=0.0), and SM[i]=S[i]-r -i and using the update rule shown below,

number

[0241] In the formula, (X, Y) means the concatenation of X and Y, i.e., XY. In reality, SM[i] (binary number) is equivalent to S[i] (binary number) with 1 subtracted from the least significant bit position. Therefore, if S[0]=111, then SM[0]=110.

[0242] Figure 11 summarizes how S[i] and SM[i] are updated for each digit in radix-8 arithmetic. The figure {Sx[i], aaa} means the concatenation of aaa bits to the actual value of S[i] or SM[i]. Note that no arithmetic operations are performed, only concatenation.

[0243] Figure 12 shows an example of on-the-fly conversion of a root of radix 8. The sequence of digits is -1, 1, -2, -4, 2, 0, -1, where the final value of SM[i] is S[i]-1.

[0244] As shown above, in the case of square root operation, the calculation of the next remainder rem[i+1] is i+1 × S[i] multiplication (see equation (3)). In a radix-8 implementation, s i+1 ={+4,+3,+2,+1,0,-1,-2,-3,-4}, and therefore 2X, 3X, and 4X multiples of S[i] are required. While the 2X and 4X terms are easily obtained by left-shifting S[i] by 1 or 2 bits, computing 3×S[i] is much more complex, which has been a limiting factor in the practical use of radix-8 square root algorithms.

[0245] Note that in other implementations with smaller bases, the term 3X is not necessary due to the digit sets {+1, 0, -1} for base 2 and {+2, +1, 0, -1, -2} for base 4.

[0246] The present invention maintains additional partial root words representing S3[i] and S3M[i], thereby preventing the calculation from being done as 3 x S[i] by multiplying by 3, or by multiplying S by 2 and adding S. For each of S3 and S3M, the concatenation performed is as follows: 3×s i+1 ∈{+12,+9,+6,+3,0,-3,-6,-9,-12}

[0247] Figure 13 shows how the concatenation is performed. Note that 3×s i+1 To represent ={+12,+9,-9,-12}, 4 bits are needed. This means that concatenation of these digit values ​​generates a carry that is propagated to the previous digit. Therefore, 4 bits of 3×s i+1 is decomposed into 3-bit digits (3 × s[i+1]) mod 8, with values ​​of {+6,+4,+3,+1,0,-1,-3,-4,-6} and a positive or negative carry c i+1 = Take {+1,-1}.

[0248] From Figure 13, s i+1 = {+4, +3, +2, +1, 0, -1, -2, -3, -4}, then the 3-bit digits concatenated to get 3 × S[i] are (3 × s{i+1]) mod 8 = {+4, +1, +6, +3, 0, -3, -6, -1, -4}, respectively. Therefore, the concatenation process to get S3[i] and S3M[i] is as follows:

[0249] 1. |s i+1 If |={4,3}, increment / decrement the actual partial root. The actual 3X multiple S3[i] of the partial root and its decremented counterpart S3M[i] are incremented / decremented by the previous digit s depending on the carry. i s i +1 or s iIt is rebuilt by changing it to -1, S3_inc[i]=S3[i]+8 -i S3M_dec[i]=S3M[i]-8 -i Although 3 bits are used to represent each digit to be concatenated, the full range of values ​​that can be represented by these 3 bits is not used, only the maximum value of +6 is added as a digit, so the carry is added to the previous digit s i Note that the .sigma..times ...

[0250] 2. 3-bit digit concatenation. 3-bit digit concatenation is

number

[0251] FIG. 14 shows an example of an on-the-fly transformation of a 3X root multiple. The sequence of digits is -1, +1, -2, -4, +2, 0, -1. The final S3[i] result in the table is 3X times the final S[i] result in FIG. 12. In sub-iteration i=0, the initial value of S3 is 11 (3 multiplied by the initial value of S[0]=1), and the initial value of S3M is 10 (3-1=2). In sub-iteration i=1, the digit -1 is added. 3 multiplied by -1 is -3, which is equal to the concatenation of the digit -3 of S3 and the digit -2 of S3M. Referring to equations (32) and (33), it can be seen that the value of S3[i+1] is the concatenation of S3M[i] and 101 (i.e., 5), and the value of S3M[i+1] is the concatenation of S3M[i] and 100 (i.e., 4).

[0252] In the sub-iteration i=2, a digit of 1 is added. 3 multiplied by 1 is 3. Again, referring to equations (32) and (33), s i+1 We can see that S3[i+1] for i = 1 is generated by concatenating S3[i] with 011 (i.e., 3), and S3M[i+1] is generated by concatenating S3[i] with 010 (i.e., 2), resulting in S3[2] = 10.101011 and S3M[2] = 10.101010. In sub-iteration i = 3, a -2 digit is added. Multiplying 3 by -2 is -6. In the case of S3, the concatenation is performed on the previous value of S3M. Because we are operating in base 8, using S3M[i] to create S3[i+1] means that the value of S3[i+1] is 8 lower than it should be. Since we are subtracting 6, this means we must add +2 here (8 - 6 = +2). Therefore, as shown in Figure 14, the concatenation is S3M and 2 (010). Similarly, in the case of S3M, the concatenation is performed on the previous value of S3M. Therefore, as shown in Figure 14, the concatenation is S3 and 1 (001 in binary). At sub-iteration i=4, the digit to be concatenated is -4. Multiplying 3 by -4 gives -12. This is a more complicated situation because -12 cannot be represented with only 3 digits, hence the negative carry. After the negative carry, the remaining subtraction to be performed is -4 (-12 = -8 - 4). Therefore, we use the value of S3M_dec, which essentially subtracts 16 (8 is the decremented value, and 8 is derived from S3M). The resulting addition to be performed is 4 (16 - 12 = 4), so the concatenation to be performed is on the value of S3M_dec and 100 (which is 4 in binary), giving us 010 000 100. For the value of S3M, the same value is used, but the concatenation is for a value one less (i.e., 4-1=3), so the concatenation is performed between S3M_dec and 011 (which is 3 in binary). The process of digits 2, 0, and -1 used in iterations 5, 6, and 7 should be clear from the above explanation.

[0253] 15 shows an implementation of a 3X partial root multiple on-the-fly conversion that forms part of on-the-fly conversion circuit 42. Circuitry for generating the partial root values ​​S[i] and SM[i] is not shown because this can be achieved by simple adaptation (using the tables shown in the figure) of the circuitry shown in, for example, U.S. Patent Application Publication No. 2020-0293281. In each partial iteration (except the first sub-iteration), the values ​​of S3[i], S3M[i], AUX[i], and AUXM[i] from the previous partial iteration are received by receiver circuit 202. The implementation has three parts. ● Increment / decrement the actual 3X partial roots S3[i], S3M[i] using the adjustment circuit 204; ● Calculate the following 3X partial roots S3[i+1], S3M[i+1], and ● Calculation of new auxiliary 3X partial roots AUX[i+1], AUXM[i+1].

[0254] The auxiliary 3X partial route is defined as follows:

number

[0255] That is, there is carry propagation to the actual 3× partial root. According to equations (32) and (33), 3×s i+1The concatenation of produces: S3[i+1]=001 111 010 111 S3M[i+1]=001 111 010 110

[0256] Next, 3×s i+2 The concatenation of produces:

number

[0257] That is, a carry is made by the digit +3, so the set of preceding digits is incremented. However, if these digits are already saturated (in this case, the target digit of S3 is 111), then a further carry is made to the next set of bits. That is, S3[i+2] is added to the incremented S3[i+1] by (3×s i+2 ) mod 8. Note, however, that incrementing S3[i+1] not only increments the last concatenated digit value 111 → 000, but also S3M[i]_dec must increment from 001 111 010 to 001 111 011, or equivalently, S3M[i] still needs to generate S3[i+2]. Note, however, that in this example, no further carry back is required. This is because 111 is added to S[i] (digit s i+1 =-3) to get S[i+1], and the next digit s i+2 The transformation of i+2 =+4,+3). This carry propagates one digit. Theoretically, if several blocks of "111" exist in a row and a partial root must be incremented, the carry propagates more than two digits. For example, if S3[i]=0001 011 111 111 and the next digit is +3, the carry propagates three digits before. However, such a pattern cannot be generated by the concatenation process described here.

[0258] Therefore, if the carry propagated to the previous digit is +1, S3_inc[i] and S3M_inc[i] are saved for the calculation of S3[i+2] and S3M[i+2], and if carry=-1, S3_dec[i] and S3M_dec[i] are saved. This situation occurs when there is a carry of +1 or -1 in the concatenation of two consecutive root digits and for certain values ​​within the 3X partial root.

[0259] Returning to FIG. 15, the adjustment circuit 204 adjusts the output voltage from AUX[i] or AUXM[i] to S3 inc[i] , S3 dec[i]、 S3M inc[i] , and S3M dec[i] Whether AUX[i] or AUXM[i] is selected depends on the previous digit s as shown in Figure 16. i Therefore, the decoding circuit 206 depends on the previous digit s i and provides a signal to multiplexers 208a, 208b, 208c, and 208d to select between AUX[i] and AUXM[i]. Next, the previous digit s i The value of is concatenated with the output from the digit x3 circuit to provide the correction values ​​for S3_inc[i] and S3M_dec[i]. The digit x3 circuit produces four output values ​​as follows: s i If >=0: ● 3s i mod8+1 ● 3s i mod8 ● 3s i mod8-1 ● 3s i mod8-2 And, s i If < 0: ● 8-(|3s i |mod8)+1 ● 8-(|3s i |mod8) ● 8-(|3s i |mod8)-1 ● 8-(|3s i |mod8)-2

[0260] For example, s i = +1, the output is 4, 3, 2, and 1, and s i =-2, the outputs are 3, 2, 1, and 0.

[0261] Then, the new 3X partial roots S3[i+1] and S3M[i+1] are converted into the new signed digit s i The +1 to S3[i] bits are generated by concatenating the corresponding bits of S3[i], S3M[i], S3_inc[i], or S3_dec[i]. This is accomplished using concatenation circuit 210. Note that the sign of the remainder is used to reduce the number of 2:1 multiplexers whose outputs are fed to concatenation circuit 210, similar to that described with reference to FIG. 4. That is, the sign of the remainder is used to select between positive and negative digits; for example, a selection is made between digits +3 and -3 for S[i] in one multiplexer, and a selection is made between digits +3 and -3 for SM[i] in another multiplexer. A positive remainder selects a positive or 0 root digit, and a negative remainder selects a negative or 0 root digit. The digits concatenated to each digit are given by equations (32) and (33). For example, for the digit +3, 001, which is (3 × 3) modulo 8, is concatenated. On the other hand, for -1, we concatenate 111, which is 8-|3×-3|=-1 (or 111 in binary).

[0262] After performing the concatenation circuitry, output circuitry 212, in the form of a set of multiplexers, outputs the selected values ​​of S3[i+1] and S3M[i+1] along with updated AUX root values ​​AUX[i+1] and AUXM[i+1], which are generated by AUX generation circuitry 214, which decodes the most recent new digit si+1 to determine whether there is a carry or not, and then uses that information to select the appropriate value to output as AUX[i+1] and AUXM[i+1], as shown in Figure 16. Each of AUX[i+1], AUXM[i+1], S3[i+1], and S3M[i+1] is received back by receive circuitry 202 in further iterations or sub-iterations.

[0263] LUT for selection constants At each stage of the digit recursion operation, a digit selection operation SEL (see Equation 2) is performed. The digit selection function in a radix-8 division or square root digit recursion algorithm performs a comparison of the actual remainder (or a portion of it) with a set of eight selected constants or coefficients. The coefficient set is selected using the most significant portion of the divisor or partial square root. The eight coefficients in the selected set are compared with the most significant portion of the remainder, and the results of the eight comparisons are used to determine the next quotient or root digit.

[0264] These coefficient sets are stored in a look-up table (LUT) that is addressed by the most significant bit of the divisor in a division operation or the most significant part of the partial root in a square root operation. The LUT size for a radix-8 division is 32 x 72 bits, and the size for a radix-8 square root is 33 x 80 bits. In a unit that supports division and square root, two different LUTs are required, one for the division and one for the square root. Therefore, the total LUT size for such a unit is 32 x 72 + 33 x 80 = 4944 bits.

[0265] In these examples, several methods are proposed to reduce the size of the total LUT. Merging of some columns can be performed. Furthermore, the square root coefficient can be calculated by adding a small offset to the division coefficient. As a result, the square root LUT can be replaced by a smaller table and some logic. Furthermore, several optimizations are performed to further reduce the division LUT size. Therefore, the total LUT size can be reduced to 33 x 42 + 33 x 18 = 1980 bits, which represents about a 60% reduction in the required storage space.

[0266] The selection function is the remainder estimate (the most significant bit of the remainder) and the digit p i+1 This involves a comparison with a set of eight selected constants or coefficients, one constant for each possible value of .

number

[0267] In division digit recursion, the set of selected constants used to obtain the next digit depends on the divisor, while in square root it depends on the partial result. The six most significant bits of the divisor or the seven most significant bits of the partial root are used to extract a set of eight selected constants for every iteration of the current division. Different divisor or partial root values ​​extract different sets of constants.

[0268] For division, the selection constant is 10 bits wide, but the most significant bit is 0. Note that the most significant bit of the divisor is always 1, since the operands are normalized before selecting the constant. Therefore, the selection constant is stored in a 32 x 72 bit division lookup table (LUT).

[0269] For square roots, the selection constant is 11 bits wide. The partial square root is [0.5,1]. Therefore, considering that the partial root estimate has 1 integer bit and 6 fractional bits, and the minimum value of the partial root is 0.5, the selection constant is stored in a 33x80 bit square root LUT with 32 entries for R[i]∈[0.5,1) and 1 entry for R[i]=1.

[0270] Therefore, a division and square root unit (fdivsqrt unit) typically uses two LUTs: a 32x72 bit division LUT and a 33x80 bit square root LUT, for a total LUT size of 32x72 + 33x80 = 4944 bits.

[0271] This technique proposes a method to reduce the total LUT size of the fdivsqrt unit. The LUT reduction is based on the following two points:

[0272] 1. We have seen that the square root constant sqrt_ct can be derived from the division constant div_ct by adding a 4-bit offset to the base constant base_ct = [2 × div_ct / 16] × 16, where base_ct is div_ct with the four least significant bits set to 0. The 4-bit offset can be negative or positive. In this way, instead of storing the square root constant, we only need to store the offset in the offset LUT.

[0273] 2. Some symmetries in the division and offset LUTs allow further reduction in the total LUT size.

[0274] 17 and 18 show the raw division and square root LUTs. The figures show the constants set for each value of the divisor and partial root estimate. Each set is divided by the digit p i ={+4,+3,+2,+1,0,-1,-2,-3} for a total of 8 constants in the set, for division div_ct={md(4),md(3),md(2),md(1),md(0),md(-1),md(-2),md(-3)} and for square root sqrt_ct={ms(4),ms(3),ms(2),ms(1),ms(0),ms(-1),ms(-2),ms(-3)}.

[0275] The value of each comparison constant can be selected from a narrow interval. In these examples, the values ​​have been carefully chosen to make each LUT symmetric, meaning that the absolute values ​​of the constants in the columns of digits +4 and -3, +3 and -2, +2 and -1, and +1 and 0 are the same (with a few exceptions). As will be shown later, this choice helps to reduce the LUT size.

[0276] The first two divisor interval constants, md(4) and md(-3), are out of range; that is, the first two digits cannot be 4 or -3. This could be fixed by doubling the number of divisor intervals, but such an approach would be very expensive, since it would mean doubling the LUT size. Instead, the sixth fractional bit of the divisor is used to select the subinterval and adjust the two least significant bits of md(4) and md(-3).

[0277] Regarding the size of the LUT, the maximum and minimum values ​​of the division LUT are 222 and -222, respectively. Therefore, the value of the division constant is in the range [222;-222], and 9 bits are required to represent all values ​​within that range. Similarly, for the square root, the constant is [447;-446], and therefore 10 bits are required.

[0278] Offset LUT By comparing the division constants and square root comparison constants shown in FIGS. 17 and 18, the square root comparison constants can be obtained as follows:

number

[0279] That is, multiply the division constant md(k) by 2, clear the four least significant bits to 0, and add a 4-bit offset, offset(k).

number

[0280] Note that if the offset has the same sign as the base constant m_base(k), the addition involves replacing the four least significant bits of m_base(k) with the 4-bit offset. If the offset does not have the same sign as the base constant, an addition is performed.

[0281] As another example,

number

number

[0282] However, m_base(k) and offset(k) may have different signs. For example,

number

number

[0283] FIG. 19 shows the offsets for the calculation of the square root constant. It highlights the cases where the sign of the offset and the sign of the division constant are different. The square root and division comparison constants have been carefully chosen to make this table symmetrical with respect to the columns, meaning that the constants in columns +4 and -3, +3 and -2, +2 and -1, and +1 and 0 have the same absolute value (have opposite signs). There are two cases where this rule is violated: in rows 4 and 13, the offsets of digits +4 and -3 do not have the same absolute value. These cases are handled separately and can be detected, for example, via the offset correction indication circuit 252.

[0284] Symmetry Using the first division LUT: 1. The absolute value of the constant can be stored instead of the signed value, which helps reduce the LUT size. 2. Number of digits p i =+1 and p i = 0, the absolute values ​​of the constants are the same (opposite signs, specifically, the digit p i =+1 is positive and p i =0 is negative), these two columns can be replaced with just one column. 3. Number of digits p i =+2 and p i The absolute value of the constant = -1 is the same except for rows 0 and 17 (opposite sign, specifically, the digit p i =+2 is positive, and p i=-1 is negative). These two columns are stored as only one column, and the values ​​of rows 0 and 17 are corrected later, for example, in the division correction instruction circuit 250 and the division constant correction circuit 248. Note that in row 0, m(2)=50, m(-1)=-48, and in row 17, m(2)=73, m(-1)=-72. To combine these two columns, the stored values ​​are 48 in row 0 and 72 in row 17, and the final m(2) value is corrected by changing the least significant bit (row 17) or the bit to the left of the least significant bit (row 0). 4. Digit p i =+2 and p i The most significant bit of the absolute value of a constant of =-1 is 0. This bit does not need to be stored in the LUT. 5. Digit p i =+1 and p i The most significant two bits of the absolute value of a constant with .times..times.0 are 0. These bits are not stored in the LUT. 6. Digit p i =+3, p i =+2, p i =+1, p i = 0, and p i Since the constant =-1 is an even number, the least significant bit is not stored in the LUT. 7. As a result, the optimized division LUT has only 6 columns due to the column merging described above in items 2 and 3. The number of bits per column is also reduced.

[0285] The offset LUT is shown in Figure 19. This table can also be optimized. 1. Digit p i The offsets for m_base = {+2, +1, 0, -1} have the same sign as m_base, i.e., the offset is positive for digits +2 and +1, and negative for digits 0 and -1 (including 0 as negative or positive where appropriate). 2. The LUT is column-symmetric, and the offset absolute values ​​of digits +4 and -3, +3 and -2, +2 and -1, and +1 and 0 are the same except for the two cases mentioned above. As a result, only the absolute values ​​of the offsets are stored in the LUT, and when the offsets are used to obtain the square root comparison constant, their signs are set according to the digit values ​​except when the offset sign is different from the m_base sign (the value highlighted in Figure 19). 3. The signs of these exceptional values ​​are stored in a new column of the LUT. The offset LUT then has only 5 columns as a result of the column merging of entries 1 and 2, 4 columns, and an additional column for the sign.

[0286] Alternatively, it will be appreciated that a square root LUT can be provided, and the constant for the division operation is derived by looking up a value in the division LUT and performing an offset. In such situations, many of the same techniques described above can be applied to reduce the size of either the floating-point LUT or the division offset table. For example, from FIG. 18, it is clear that the magnitudes of the constants for the digits +4 and −3 are the same (the digits have opposite signs, with the +4 digit generally being positive and the −3 digit generally being negative). Similarly, the magnitudes of the constants for the digits +3 and −2 are the same (again, opposite digits, typically positive for +3 and negative for −2). Similarly, the magnitudes of the constants for the digits +2 and −1 are the same (again, opposite signs, typically positive for +2 and negative for −1).

[0287] The final division and offset table with the optimizations described in the previous section is shown in Figure 20. The table is divided into parts with the division LUT on the left and the square root offset LUT on the right. Note that the number of columns has been reduced by column merging. The resulting merging columns are labeled with the associated two-digit value. So, for example, the column labeled (+2, -1) represents the digit p in the raw table. i =+2 and pi =-1 means merging the columns corresponding to

[0288] On the other hand, note that the last row of the table in FIG. 20 is for square roots only (row 32 in FIG. 19).

[0289] The address (left-most column of the table) is accessed differently for division and square root. For division, the 6 most significant bits of the divisor form the address, but the first bit is 1. For square root, the 7 most significant bits of the partial root R[i] are used to address the table with values ​​ranging from 0.5 (0.100000 in binary) to 1.0 (1.000000 in binary). Note that the square root LUT has 33 rows, so 6 bits are used for the address.

[0290] The contents of the LUT are shown as hexadecimal values. The number of bits actually required for each column is specified in the table, and hexadecimal values ​​are given, but note that the full range of values ​​may not be possible. For example, the digit p in this division LUT i The constant value of =+3 requires only 7 bits, since it can only take on the values ​​{2,3,4}, where the most significant hexadecimal digit corresponds to the binary value {0010,0011,0100}, and therefore there is no need to store the most significant bit. Similarly for the columns (+2,-1) and (+1,0).

[0291] The offset LUT (right part) in Figure 20 stores offset absolute values ​​in columns (+4,-3), (+3,-2), (+2,-1), and (+1,0), and the 2-bit value of column sign is the offset sign of the offsets in columns (+4,-3) and (+3,-2). Note that the offsets in columns (+2,-1) and (+1,0) are positive. The sign bit 1 means that the signs of the offset and its corresponding m_base are different.

[0292] As mentioned above, the last row of the table, with address 100000, is only meaningful for square roots. Using the same base as row 011111, the comparison constant for this partial root estimate is obtained at the offset shown in the table.

[0293] Consider the following example for calculating comparison constants for division and square root: For division, the constant set is obtained from the LUT by adding leading zeros. For example, for a division operation with divisor = 1.00110x...x, the LUT address is 01_00110 and the LUT returns:

number

[0294] Note that the number of bits for each constant in the set depends on how many digits the constant is for, so taking into account the rules for reducing LUT size listed above for division, the set of comparison constants for this particular divisor value is:

number

[0295] The bits added to get the final constants are highlighted. Note that the absolute values ​​of the constants are calculated from the LUT. In a later step, the signs of m(0), m(-1), m(-2), and m(-3) are two's complemented to get the final set of constants.

[0296] Note that for the square root constants in this same row, the sign field is 01. This means that the sign of the offset for the calculation of ms(+3) and ms(-2) is different from the sign of the base constant, and therefore the calculation of these two constants requires a subtraction. From the table, LUT_offset(01_00110)={1,a,e,2,6} The offsets are as follows: Offsets with signs different from the basic constant signs are highlighted

number

[0297] The fundamental constants are

number

number

[0298] Since the positive and negative parts of the sqrt LUT are symmetric, the remaining constants are obtained by taking the two's complement of the above constants. {ms(0),ms(-1),ms(-2),ms(-3)}={-38,-114,-192,-266}

[0299] 21 shows a selection constant generator 238 used to generate the selection constant used by, for example, digit select comparator 32. The divisor and partial root bits are received by multiplexer 240. A div / sqrt select signal is provided which selects the divisor when a selection constant for division is required, and the partial root when a selection constant for square root is required. The selected bits are then used to access the associated value in storage circuit 242, which comprises a division LUT and a (square root) offset LUT.

[0300] The output from the division LUT is passed to padding circuit 246, which pads bits by adding zeros to the output constant. The padding performed is, for example, described in points 2-6 with respect to the division LUT above. The resulting constant is passed to conversion circuit 244, described below, and also to division constant correction circuit 248. Division constant correction circuit 248 receives the padded (extended) division selection constant as well as an output from division correction indication circuit 250, which indicates whether the data being obtained from the division LUT is one of the exceptional cases where the absolute values ​​of the constants are not the same (point 3 with respect to the division LUT above). That is, (i) the constants md(4) and md(-3) when the divisor estimate is 0 or 1, and (ii) the digit p when the divisor estimate is 0 or 17. i =+2 and p i = -1. These corrections require setting bits 70, 50, 1, and 0 and clearing bits 71 and 21 in the selected constant set. The corrections are performed by division constant correction circuit 248. The output from the offset LUT is passed to a conversion circuit 244 along with an output from an offset correction indication circuit 252, which indicates whether the constant being accessed is one of the exceptions where the LUT offsets do not have the same value (e.g., lines 4 and 13). If so, a correction to the correct value is made in the conversion circuit 244. The correction circuit 244 also receives a padded (extended) division constant from a padding circuit 246. A substitution circuit 254 is used to add the offset using concatenation or subtraction as described above. In particular, if the offset sign and the constant base sign are different, a subtraction is performed. The subtraction is made possible by checking the sign field in the offset LUT. Substitution of the four least significant bits of the 4-bit offset is made only if the signs are equal.

[0301] For both division constants and LUT constants, the absolute value is multiplied by digit p i A signature circuit 256 is provided to convert the signed values ​​to =0, -1, -2, -3.

[0302] Computer readable code for manufacturing The concepts described herein may be embodied in computer-readable code for the manufacture of devices embodying the described concepts. For example, the computer-readable code may be used in one or more stages of a semiconductor design and manufacturing process, including an electronic design automation (EDA) stage, to manufacture integrated circuits comprising devices embodying the concepts. Such computer-readable code may additionally or alternatively enable the definition, modeling, simulation, verification, and / or testing of devices embodying the concepts described herein.

[0303] For example, computer-readable code for producing a device embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concept. For example, the code can define a register-transfer level (RTL) abstraction of one or more logic circuits to define a device embodying the concept. The code can define an HDL representation of one or more logic circuits embodying the device in Verilog, SystemVerilog, Chisel, or VHDL (Very High Speed ​​Integrated Circuit Hardware Description Language), as well as intermediate representations such as FIRRTL. The computer-readable code can provide a definition embodying the concept using system-level modeling languages ​​such as SystemC and SystemVerilog or other behavioral representations of the concept that can be interpreted by a computer to enable simulation, functional and / or formal verification, and testing of the concept.

[0304] Additionally or alternatively, the computer-readable code may embody a computer-readable representation of one or more netlists. The one or more netlists may be generated by applying one or more logic synthesis processes to the RTL representation. Alternatively or additionally, the one or more logic synthesis processes may generate a bitstream from the computer-readable code to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for proof-of-concept and testing purposes prior to fabrication in an integrated circuit, or the FPGA may be deployed directly into a product.

[0305] The computer readable code may include a mixture of code representations for fabricating a device, including, for example, a mixture of one or more of an RTL representation, a netlist representation, or another computer readable definition used in a semiconductor design and manufacturing process to fabricate a device embodying the invention. Alternatively or additionally, a concept may be defined in a combination of a computer readable definition used in a semiconductor design and manufacturing process to fabricate a device and computer readable code that defines instructions to be executed by the defined device to be fabricated.

[0306] Such computer readable code may be disposed on any known transitory computer readable medium (such as wired or wireless transmission of code over a network) or on a non-transitory computer readable medium such as a semiconductor, magnetic disk, or optical disk. Integrated circuits manufactured using computer readable code may include components such as one or more central processing units, graphics processing units, neural processing units, digital signal processors, or other components that individually or collectively embody the concepts.

[0307] In this application, the term "configured to..." is used to mean that elements of an apparatus have a configuration that allows them to perform a defined operation. In this context, "configuration" refers to a way of arranging or interconnecting hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform the function. "Configured to" does not imply that an apparatus element must be modified in some way to provide the defined operation.

[0308] Although exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to these precise embodiments, and that various changes and modifications can be made to the embodiments by those skilled in the art without departing from the scope of the invention as defined by the appended claims.

Claims

1. A data processing apparatus for performing a digit recurrence operation on an input value, comprising: a receiving circuit configured to receive a remainder value of a previous iteration of the digit recurrence operation; each of a plurality of selection constants associated with available digits of the next digit of the result of the digit recurrence operation, each of the selection constants being associated with one of the available digits and an input parameter, and comparing each of the plurality of selection constants with the most significant bit of the remainder value of the previous iteration of the digit recurrence operation, and configured to output the next digit of the result of the digit recurrence operation based on the comparison; a storage circuit configured to store a subset of the selection constants, the subset of the selection constants excluding selection constants associated with digits excluded from the available digits; A data processing apparatus comprising the above.

2. The data processing apparatus according to claim 1, further comprising a conversion circuit configured to generate the selection constants to be excluded from the selection constants stored in the storage circuit.

3. The data processing apparatus according to claim 2, wherein the conversion circuit is configured to generate the selection constants to be excluded by performing a selective inversion on the sign of one of the selection constants stored in the storage circuit.

4. The data processing apparatus according to claim 3, wherein one of the selection constants is associated with the same input parameter and a different one of the available digits as the selection constants to be excluded.

5. The data processing apparatus according to claim 3, wherein the storage circuit is configured to store an exception flag indicating whether the selective inversion should be performed on the selection constants to generate the selection constants to be excluded.

6. The digit recurrence operation is a square root digit recurrence operation, The input parameter is a partial root, The data processing apparatus according to any one of claims 1 to 5.

7. The digit recurrence operation is a division digit recurrence operation, The input parameter is a divisor, The data processing apparatus according to any one of claims 1 to 5.

8. In the division operation mode, the digit recurrence operation is a division digit recurrence operation, the input parameter is a divisor, In the square root operation mode, the digit recurrence operation is a square root digit recurrence operation, and the input parameter is a partial root, the data processing apparatus according to any one of claims 1 to 5.

9. In the division operation mode, the digit recurrence operation is a division digit recurrence operation, and the input parameter is a divisor, In the square root operation mode, the digit recurrence operation is a square root digit recurrence operation, and the input parameter is a partial root, Each of the selection constants is a division digit recurrence operation selection constant, or each of the selection constants is a square root digit digit recurrence operation selection constant, the data processing apparatus according to claim 1 or 2.

10. Each of the selection constants is a division digit recurrence operation selection constant, the data processing apparatus according to claim 9.

11. The conversion circuit is configured to generate an exclusion selection constant in the division operation mode by performing a selective inversion of the sign of one of the division digit recurrence operation selection constants, the data processing apparatus according to claim 10 when dependent on claim 2.

12. The conversion circuit is configured to generate the excluded selection constant in the square root operation mode by referring to one of the division digit recurrence operation selection constants, the data processing apparatus according to claim 10 when dependent on claim 2.

13. The storage circuit is configured to store a plurality of mappings between the excluded selection constant in the square root operation mode and one of the division digit recurrence operation selection constants, the data processing apparatus according to claim 12.

14. The storage circuit is configured to store an exception flag indicating whether the selective inversion should be performed to generate the excluded selection constant for the selection constant, the data processing apparatus according to claim 11.

15. The digit recurrence operation is base 8, the data processing apparatus according to any one of claims 1 to 5.

16. A data processing method for performing a digit recurrence operation on an input value, Receiving a remainder value of the previous iteration of the digit recurrence operation, For each of a plurality of selection constants associated with available digits of a next digit of a result of the digit-by-digit recurrence operation, and the most significant bit of the remainder value of the previous iteration of the digit-by-digit recurrence operation, each of the selection constants is associated with one of the available digits and an input parameter, execute a comparison with each of the plurality of selection constants, and output the next digit of the result of the digit-by-digit recurrence operation based on the comparison. Store a subset of the selection constants, the subset of the selection constants excluding selection constants excluded from the selection constants, the subset of the selection constants being associated with digits excluded from the available digits. A method comprising the above. [

17. ] A computer-readable medium for storing computer-readable code for manufacturing a data processing apparatus for performing a digit-by-digit recurrence operation on an input value, A receiving circuit configured to receive a remainder value of a previous iteration of the digit-by-digit recurrence operation, For each of a plurality of selection constants associated with available digits of a next digit of a result of the digit-by-digit recurrence operation, each of the selection constants is associated with one of the available digits and an input parameter, a comparison circuit configured to execute a comparison with the most significant bit of the remainder value of the previous iteration of the digit-by-digit recurrence operation, and output the next digit of the result of the digit-by-digit recurrence operation based on the comparison, A storage circuit configured to store a subset of the selection constants, the subset of the selection constants excluding selection constants excluded from the selection constants, the subset of the selection constants being associated with digits excluded from the available digits. A computer-readable medium comprising the above.