Division / square root pipeline and method

JP2024529820A5Active Publication Date: 2025-05-27ARM LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023574634
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-07-02
Filing Date
2022-05-26
Publication Date
2025-05-27
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

Existing digit recurrence algorithms for division and square root operations face challenges in balancing performance, circuit area, and power consumption, particularly when implementing higher radixes that require complex circuitry.

Method used

The approach involves subdividing higher radix iterations into multiple sub-iterations of lower radix within the same processing cycle, utilizing shared circuitry and speculative replication to improve performance and reduce circuit area and power consumption.

Benefits of technology

This technique enhances performance by reducing critical timing paths and minimizing circuit size, making it more efficient than traditional implementations, especially for radix-64 operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

An apparatus is provided with a division / square root pipeline, the apparatus comprising: a plurality of division / square root iteration pipeline stages, each for performing a respective iteration of a digit recursive division or square root operation; and a signal path for providing an output generated by one division / square root iteration pipeline stage in one iteration as an input to a subsequent division / square root iteration pipeline stage of the division / square root pipeline for performing a subsequent iteration of the digit recursive division or square root operation, wherein the division / square root pipeline is capable of performing the digit recursive division or square root operation on floating-point operands to generate a floating-point result.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present technology relates to the data processing field.

[0002] Digit recursion algorithms can be used to perform processing operations such as division or square root. Digit recursion uses an iterative algorithm to perform the calculation. At each iteration, the next digit of the result value is generated. Each digit is represented using a number of bits. In a base r implementation of the digit recursion algorithm, each digit has log2(r) bits. For example, an implementation using a base of 4 represents each digit with 2 bits, so at each iteration, two more bits of the result are generated, so generating a result value with a particular number of bits may take a number of iterations. Implementations using higher bases may generate a result of a given size in fewer iterations, improving performance, but the circuitry to perform a single iteration is more complex. When designing circuits to perform such digit recursion methods, there may be challenges in meeting the competing demands of performance, circuit area, and power consumption.

[0003] At least some examples provide an apparatus with a division / square root pipeline comprising: a plurality of division / square root iteration pipeline stages, each for performing a respective iteration of a digit recursive division or square root operation; and a signal path for providing an output produced by one division / square root iteration pipeline stage in one iteration as an input to a subsequent division / square root iteration pipeline stage of the division / square root pipeline for performing a subsequent iteration of the digit recursive division or square root operation, wherein the division / square root pipeline can perform the digit recursive division or square root operation on floating-point operands to generate a floating-point result.

[0004] At least some examples provide a data processing method that includes performing respective iterations of a digit-recursive division or square root operation using multiple division / square root iteration pipeline stages of a division / square root pipeline, and providing an output generated by one division / square root iteration pipeline stage as an input to a subsequent division / square root iteration pipeline stage of the division / square root pipeline, wherein the division / square root pipeline can perform the digit-recursive division or square root operation on floating-point operands to generate a floating-point result.

[0005] At least some examples provide a computer-readable medium for storing computer-readable code for manufacturing an apparatus comprising a division / square root pipeline comprising: a plurality of division / square root iteration pipeline stages, each for performing a respective iteration of a digit recursive division or square root operation; and a signal path for providing an output generated by one division / square root iteration pipeline stage in one iteration as an input to a subsequent division / square root iteration pipeline stage of the division / square root pipeline for performing a subsequent iteration of the digit recursive division or square root operation, wherein the division / square root pipeline can perform the digit recursive division or square root operation on floating-point operands to generate a floating-point result. [Brief description of the drawings]

[0006] Further aspects, features, and advantages of the present technology will become apparent from the following description of examples, read in conjunction with the accompanying drawings. [Figure 1] FIG. 2 illustrates a schematic diagram of an example of a data processing operation having a division / square root processing circuit; [Diagram 2] FIG. 10 illustrates a schematic example of splitting a higher base digit-recursive square root or division operation into multiple lower base sub-iterations performed in the same processing cycle; [Diagram 3]FIG. 2 illustrates a circuit for performing a given base r iteration of a square root operation. [Figure 4] FIG. 2 illustrates a remainder update circuit. [Diagram 5] FIG. 2 illustrates a remainder estimation circuit. [Figure 6] FIG. 2 is a diagram showing a digit selection circuit. [Figure 7-1] FIG. 2 illustrates in more detail a square root processing circuit for performing a given radix-64 iteration of a square root operation by performing two radix-8 sub-iterations in the same processing cycle. [Figure 7-2] FIG. 2 illustrates in more detail a square root processing circuit for performing a given radix-64 iteration of a square root operation by performing two radix-8 sub-iterations in the same processing cycle. [Figure 8-1] A combined divide / square root processing circuit is shown that is capable of performing both division and square root operations, where the shared circuitry generates at least one output value on the same data path that is used for both the division and square root operations. [Figure 8-2] A combined divide / square root processing circuit is shown that is capable of performing both division and square root operations, where the shared circuitry generates at least one output value on the same data path that is used for both the division and square root operations. [Figure 9-1] FIG. 2 illustrates an example of a division / square root pipeline. [Figure 9-2] FIG. 2 illustrates an example of a division / square root pipeline. [Figure 10] FIG. 1 illustrates pipelining of consecutive division or square root operations where a second operation is prohibited from starting a certain number of cycles after a first operation if the second operation uses a lower precision floating-point representation than the first operation. [Figure 11] FIG. 1 illustrates on-the-fly transformation. [Figure 12] FIG. 1 illustrates an example of on-the-fly transformation. [Figure 13] FIG. 13 illustrates on-the-fly conversion of 3X digits. [Figure 14]FIG. 1 illustrates an example of 3X on-the-fly transformation. [Figure 15-1] FIG. 1 illustrates a circuit for performing a 3x on-the-fly conversion. [Figure 15-2] FIG. 1 illustrates a circuit for performing a 3x on-the-fly conversion. [Figure 16] FIG. 13 illustrates a selection for reconstructing partial route values. [Figure 17] FIG. 13 illustrates comparison constants for radix-8 sub-iterations of a division operation. [Figure 18] FIG. 13 illustrates comparison constants for radix-8 partial iterations of a square root operation. [Figure 19] FIG. 13 is a diagram illustrating an offset that represents the offset of the square root comparison constant relative to the division comparison constant. [Figure 20-1] FIG. 13 illustrates a division and offset lookup table for determining comparison constants for division and square root operations. [Figure 20-2] FIG. 13 illustrates a division and offset lookup table for determining comparison constants for division and square root operations. [Figure 21-1] 1 shows a circuit for obtaining a set of comparison constants for division and square root operations. [Figure 21-2] 1 shows a circuit for obtaining a set of comparison constants for division and square root operations.

[0007] Square root processing The square root processing circuit may perform a given radix-r iteration of the square root operation of radix r by performing two or more radix-n partial iterations in the same processing cycle, where n < r. This can provide a better compromise between performance and circuit overhead compared to an implementation that does not subdivide the radix-r iteration into lower-radix partial iterations. Since the overall operation performed in one cycle is a higher-radix operation of radix r, this means that log2(r) bits of the result can be generated per processing cycle, which can provide higher performance than when a smaller radix is used. However, the radix-r iteration is divided into several radix-n partial iterations in the same processing cycle, where for each partial iteration n is smaller than r. As a result, the overall size of the circuit can be smaller than if the radix-r iteration were performed as a single operation. This is because the number of alternative options available for selection as the next digit for each partial iteration using radix n is less than the number of alternative options for the radix-r digits required when the radix-r iteration of the square root operation is performed as a single operation. However, dividing the radix-r iteration into smaller-radix partial iterations can pose a timing challenge in that it may not be possible to fit those radix-n partial iterations into a single processing cycle.

[0008] For a given base-n partial iteration, the square root processing circuit may include a digit selection circuit that selects a next base-n result digit of the square root result based on a previous remainder estimate, a remainder update circuit that adjusts the previous remainder value based on a remainder adjustment value responsive to the next base-n result digit selected by the digit selection circuit to generate an updated remainder value, a remainder estimation circuit that generates an updated remainder estimate indicative of an estimate of a portion of the updated remainder value, and an output signal path for providing the updated remainder value and the updated remainder estimate for use as the previous remainder value and the previous remainder estimate in a subsequent base-n partial iteration of the given base-r iteration or a first base-n partial iteration of a further base-r iteration of the base-r square root operation. Since multiple partial iterations are being performed per cycle, multiple instances of the digit selection circuit, remainder update circuit, remainder estimation circuit, and output signal path may be provided for each base-n partial iteration within the same base-r iteration of the square root operation.

[0009] In the last radix-n partial iteration of a given radix-r iteration, the remainder estimation circuit may generate an updated remainder estimate in parallel with the remainder update circuit that generates the updated remainder value. This is counter-intuitive since the updated remainder estimate represents a portion of the updated remainder value, and one would expect the remainder value to be available first, and then the remainder estimate to be calculated sequentially. However, the inventors have recognized that in an embodiment in which a higher radix iteration is divided into smaller radix partial iterations, it is possible to generate an updated remainder estimate for the last partial iteration in parallel with the remainder update circuit that generates an updated remainder value for that last partial iteration of the given radix-r iteration. This means that the delay associated with the calculation of the remainder estimate for the last radix-n partial iteration can be at least partially removed from the critical timing path through the square root processing circuit, reducing the overall time it takes to perform a given radix-r iteration of the square root operation, and thus improving overall performance.

[0010] The remainder update circuit may generate the updated remainder value in a redundant representation. For example, the remainder value may be represented as two terms that together represent the numerical value of the updated remainder value, but there may be two or more combinations of values ​​of the first and second terms that can represent the same numerical value. Generating the updated remainder value in a redundant representation may be useful because it may avoid computation of the updated remainder value that requires propagating a carry from one bit to another. Thus, the remainder update circuit may comprise a carry-save add circuit.

[0011] However, for purposes of selecting the next base-n result digit of the square root result, the digit selection circuitry may perform digit selection using a representation of the remainder in a non-redundant representation, and the remainder estimation circuitry may thus generate an updated remainder estimate in a non-redundant representation that represents an estimate of at least a portion of the updated remainder value (non-redundant representation means that the estimate can be expressed in a single term, and for any given numerical value of the updated remainder estimate, there is a single bit pattern (and no other) of the non-redundant representation that corresponds to that numerical value). Because the full precision of the updated remainder value may not be required for digit selection, the updated remainder estimate may have fewer bits than the updated remainder value (more specifically, the updated remainder estimate may have fewer bits than the number of bits of a single term of a redundantly represented remainder value that may include two redundant terms), and by limiting the number of bits of the estimate, delays in computing the non-redundant remainder estimate are reduced. For example, the updated remainder estimate may represent an estimate of the most significant portion of the updated remainder value, since the lower bits may not significantly affect the precision of the digit selection.

[0012] Thus, computation of the remainder estimate in the non-redundant representation can use a carry propagate adder circuit that can propagate the carry from one bit position to another, and that may be slower than a carry-save adder. Thus, in typical implementations, the carry propagate adder circuit used for the remainder estimate can significantly slow down the overall processing of a particular iteration of the square root operation.

[0013] However, the inventors have recognized that in an approach where a radix-r square root iteration is divided into multiple smaller radix-n sub-iterations that are executed within the same processing cycle, an updated remainder estimate for the last radix-n sub-iteration may be calculated in parallel with the calculation of the updated remainder value, since the updated remainder estimate for the last radix-n sub-iteration may be calculated using information provided as input to the remainder update circuitry in the last radix-n sub-iteration and / or other information from previous sub-iterations in a given radix-r iteration, thereby avoiding the need to wait for the updated remainder value in the last radix-n sub-iteration to be available before commencing the calculation of the updated remainder estimate for the last radix-n sub-iteration. This provides a relatively significant gain in performance due to the removal from the critical timing path of the relatively slow carry propagate addition for calculating the updated remainder estimate in the last radix-n sub-iteration of a given radix-r iteration.

[0014] In the remainder update, the previous remainder value is updated based on a remainder adjustment value whose value depends on the next result digit selected by the digit selection circuit. The remainder estimation circuit in the last radix-n partial iteration can use this remainder adjustment value and the previous remainder estimate value to generate an updated remainder estimate value for the last radix-n partial iteration. Because the remainder adjustment value is used as an input to the remainder estimation circuit in the last radix-n partial iteration, this eliminates the need to wait for the updated remainder value and allows the updated remainder estimate to be available more quickly.

[0015] The remainder estimation circuit may take advantage of the fact that the final radix-n partial iteration follows at least one previous partial iteration being executed in the same cycle, so that some information calculated in that previous partial iteration may be used by the remainder estimation circuit in the final partial iteration to calculate an updated remainder estimate sooner than if the remainder estimate were calculated consecutively after the updated remainder value was obtained.

[0016] For example, in previous radix-n partial iterations of a given radix-r iteration other than the last radix-n partial iteration, the remainder estimation circuit may calculate at least one additional bit of an updated remainder estimate that is not needed to select the next radix-n result digit in the last radix-n partial iteration of the given radix-r iteration, and in the last radix-n partial iteration of a given radix-r iteration, the remainder estimation circuit may determine an updated remainder estimate using at least one additional bit determined in the previous radix-n partial iteration. By calculating more bits than are needed for the updated remainder estimate in the previous radix-n partial iteration, the additional bit(s) may be used to calculate an updated remainder estimate sooner in the last radix-n partial iteration because the additional bit calculated in the previous partial iterations allows the updated remainder estimate in the last partial iteration to be calculated without waiting for an updated remainder value to be available.

[0017] In the first radix-n sub-iteration of a given radix-r iteration, the remainder estimation circuit can determine an updated remainder estimate based on the updated remainder value generated by the remainder update circuit in the first radix-n sub-iteration. Thus, it is not necessary that the updated remainder estimate is calculated in parallel with the updated remainder value in all sub-iterations. In the first sub-iteration of a given radix-r iteration, sufficient information may not be available to be able to calculate the remainder estimate until the updated remainder value is available in redundant form. However, because multiple radix-n sub-iterations overlap within the same processing cycle, the circuit designer has the freedom to change the relative timing at which portions of a subsequent sub-iteration start with respect to portions of a previous sub-iteration, and can use information from the previous sub-iteration to calculate parameters in the subsequent sub-iteration, making it feasible to parallelize the calculation of the updated remainder value and the updated remainder estimate for at least the final sub-iteration.

[0018] In an implementation in which at least three partial iterations are performed within the same cycle to perform a given base r iteration of the square root operation, it is also possible that the updated remainder estimate value is calculated in parallel with the updated remainder values ​​of one or more intermediate partial iterations between the first and last partial iterations.

[0019] The square root processing circuit comprises, for a given base-n sub-iteration, one or more instances of the replica circuit, each instance of the replica circuit including two or more replica circuit units for determining, in parallel with the selection of the next base-n result digit by the digit selection circuit, two or more candidate output values ​​corresponding to different result digits that can be selected as the next base-n result digit by the digit selection circuit, and a selection circuit for selecting one of the plurality of candidate output values ​​in response to the digit selection circuit indicating which of the different result digits is to be selected as the next base-n result digit, the plurality of candidate output values ​​including at least the two or more candidate output values ​​generated by the two or more replica circuit units. This approach allows for faster performance since it is not necessary to wait for the next base-n result digit to actually be selected by the digit selection circuit before commencing calculations to generate the candidate output values.

[0020] It should be noted that the number of candidate output values ​​available for selection by the selection circuit may be greater than the number of candidate output values ​​generated by two or more of the replicated circuit units. For example, one of the result digits available for selection may be equal to 0, and in some cases, it may not be necessary to explicitly calculate a candidate output value for result digit 0, since the candidate output value selected when the next result digit is 0 may be identical to the input value provided to the partial iteration. Thus, the selection circuit may obtain as inputs candidate output values ​​that are not explicitly generated by one of the replicated circuit units, as well as candidate output values ​​generated by two or more replicated circuit units.

[0021] Providing replica circuit units to speculatively calculate multiple candidate output values ​​before the time the next result digit is known can be superior in performance, but the number of replica circuit units required increases with increasing radix, potentially increasing circuit area cost and power consumption to support higher radix arithmetic.

[0022] One technique for limiting circuit area and power cost may be to provide at least one of the two or more replica circuit units as a shared circuit unit shared between both a positive result digit having a given magnitude and a negative result digit having the same given magnitude. The shared circuit unit is configured to output a shared candidate output value to a selection circuit on a shared signal path, and the selection circuit may select the shared candidate output value from the shared signal path when the next base-n result digit is either a positive or negative result digit having the given magnitude. This thus eliminates the need to provide two separate replica circuit units for each of the positive and negative result digits sharing the same magnitude. This can reduce the total number of required replica circuit units, thus saving circuit area and reducing power consumption.

[0023] For at least one instance of the replica circuit, a shared circuit unit providing an output shared between positive and negative result digits of the same magnitude may select a value to be output as a shared candidate output value on a shared signal path based on the sign of a previous remainder estimate. Thus, although a common signal path is shared between two result digit values ​​having the same magnitude but different signs, the actual numerical value output on the shared signal path may differ depending on the sign of the previous remainder estimate.

[0024] For at least one instance of the replica circuit, the shared circuit unit may include a shared adding circuit for determining shared candidate output values ​​of positive and negative result digits having a given magnitude. The technique of providing a shared circuit unit for generating shared candidate output values ​​of both positive and negative digits of the same magnitude may be particularly useful when the circuit unit includes an adder circuit, because adder circuits are relatively costly in terms of circuit area.

[0025] For radix-n partial repetition, it is usually expected that the number of candidate output values ​​available for selection by the selection circuit must be n + 1. However, by sharing a shared circuit unit between positive and negative result digits having the same magnitude, the total number of candidate output values ​​available for selection by the selection circuit can be reduced to n / 2 + 1, which means that the number of replicated circuit units provided can be reduced, thereby significantly reducing the circuit area.

[0026] There may be several instances of the replicated circuit within the square root processing circuit. Different parts of the square root processing circuit may each use this technique, where the replicated circuit units a priori determine candidate output values ​​for multiple possible result digits, and then the correct candidate output value is selected by the selection circuit when the next result digit is selected.

[0027] For example, the remainder update circuit may comprise one such instance of the replication circuit. If the remainder update circuit uses a speculative replication and selection technique, the candidate output value being selected by the selection circuit may be the candidate updated remainder value.

[0028] Similarly, the remainder estimation circuit may use this speculative replication and include one of the instances of the replication circuit described above. If the remainder estimation circuit includes a replication circuit, the candidate output value may be a candidate updated remainder estimate value.

[0029] Another part of the digit recursion method may be to perform an on-the-fly conversion. In the case of a square root operation, the adjustment of the previous remainder value to generate an updated remainder value may not only depend on the remainder adjustment value (selected based on the next result digit), but also on the partial root value, which is a numerical value corresponding to the previously selected sequence of result digits. Since the result digit may be selected by the digit selection circuit as a signed digit, then an on-the-fly conversion circuit may be provided to convert the partial root value to a non-redundant representation to provide a partial root value in a non-redundant representation that may be used by the remainder update circuit to adjust the previous remainder value to generate an updated remainder value. As described below, it is possible to perform the on-the-fly conversion in a manner that does not require addition, but can be done simply by concatenating the previous partial root value with some additional bits selected based on the latest base-n result digit.

[0030] Thus, the on-the-fly conversion circuit (for generating partial root values ​​indicating, in a non-redundant representation, a numerical value corresponding to a previously selected series of base-n result digits) may also include an instance of the replica circuit described above, such that the replica circuit unit generates several candidate partial root values ​​and the candidate output values ​​available for selection by the selection circuit include several candidate values ​​for the partial root value.

[0031] Therefore, regardless of which part of the square root processing circuitry performs duplication, duplication can help improve performance, and if implemented, sharing of duplication circuit units for positive and negative result digits of the same magnitude can help reduce the overall circuit size.

[0032] Some implementations may implement the replica circuit with only one or a subset of the above components of the square root processing circuit, while other components do not use the replicated approach, and performance may be maximized when the remainder update circuit, the remainder estimation circuit, and the on-the-fly transformation circuit each provide an instance of the replica circuit.

[0033] In general, if a given radix-r iteration is divided into several back-to-back or overlapping radix-n sub-iterations in the same processing cycle, the value of r may correspond to the product of the respective values ​​of n for each of the sub-iterations used in one cycle.

[0034] In the particular example described below, r=64 and n=8 for each of the partial iterations, such that there are two radix-8 partial iterations in each radix-64 iteration. This approach can provide a good balance between performance (radix-64 means that six bits can be generated per processing cycle) and circuit area and timing complexity (using radix-8 for the partial iterations means that only two partial iterations are needed, which imposes lower timing pressures compared to implementations using three or more partial iterations, but increasing the radix beyond 64 can make it infeasible to manage circuit size while still meeting timing). Thus, r=64 and n=8 can be a particularly useful combination.

[0035] Nevertheless, other options are possible: for example, a radix-64 iteration of a square root operation can be performed as three sub-iterations, each of radix-4 (since 64=4×4×4).

[0036] Implementing each of the partial iterations in the same radix n may be useful because it may be more efficient in terms of overall circuit area and simpler in terms of design complexity to use the same radix in each partial iteration.

[0037] Nevertheless, it is also possible for different sub-iterations within the same base r iteration to use different bases: for example, a base 64 iteration of a digit recursive square root operation can be divided into one base 4 sub-iteration, one base 8 sub-iteration, and one base 2 sub-iteration. It is therefore not necessary that n be equal for each sub-iteration.

[0038] The above-mentioned technique can be implemented in square root processing circuits of different designs. In one example, the square root processing circuit may be an iterative square root processing circuit, and the output signal path may provide the updated remainder value and the updated remainder estimate value generated in the last base-n sub-iteration from the output of the iterative square root processing circuit to the input of the same iterative square root processing circuit for use as the previous remainder value and the previous remainder estimate value in the first base-n sub-iteration of the further base-r iteration of the square root operation. Thus, to perform the square root operation as a whole, multiple passes through the iterative square root processing circuit are performed over multiple processing cycles, and the output of the iterative square root processing circuit in one cycle is fed back as the input to the same unit in the subsequent cycle.

[0039] However, as described in more detail below, the square root processing circuit may be part of a pipeline square root processing unit including several square root iteration pipeline stages, each stage including a respective instance of the square root processing circuit described above. In this case, the output signal path of a given pipeline stage may provide the updated remainder value and the updated remainder estimate generated in the last radix-n sub-iteration of a given radix-r iteration from the output of the square root processing circuit in one square root iteration pipeline stage to the input of the square root processing circuit (a different instance of the square root processing circuit) in a subsequent square root iteration pipeline stage for processing of the subsequent radix-r iteration in the next processing cycle. This approach allows multiple square root operations to be pipelined with each other, so that a later square root operation may be in a previous pipeline stage where a previous radix-r iteration is being performed while a previous square root operation is being processed in a later stage of the pipeline square root processing unit, which may help improve the overall throughput of the square root operations.

[0040] Combined division / square root circuit Commercially available processor microarchitectures are typically provided with separate circuit logic for division and square root operations, respectively, so that these operations are performed in completely separate circuit logic units, and there is no sharing of the data path used to calculate the division result compared to the data path used to calculate the square root result. This may be simpler to build, since it does not require extra complexity in the square root operation to affect the timing of the division operation. However, it may be desirable to increase the radix used for the division and square root operations to improve performance by allowing a division or square root result with a greater number of bits to be calculated per cycle. For example, a 6-bit result may be calculated per cycle using a radix 64 division or square root operation that is not currently available in commercially available processors. However, the increase in radix means that more complex circuitry is required compared to implementations requiring lower radixes. Thus, having separate division and square root processing circuits when operating at higher radixes may increase the circuit size and therefore the power consumption of the processor.

[0041] In the example described below, a combined divide / square root processing circuit is provided for performing a given radix-64 iteration of a radix-64 divide operation in response to a divide instruction and for performing a given radix-64 iteration of a radix-64 square root operation in response to a square root instruction. The combined divide / square root processing circuit has shared circuitry for generating at least one output value for a given radix-64 iteration on the same data path used for both the radix-64 divide operation and the radix-64 square root operation. For example, the at least one output value may include any one or more of an updated remainder value, a selected result digit, an updated remainder estimate, and / or an on-the-fly converted partial result value. The use of shared circuitry, where the same data path is used for the output of both the divide operation and the square root operation, can reduce the overall amount of circuitry compared to an implementation having a divide and square root unit. This is particularly useful for radix-64 arithmetic, given the increased circuitry required for radix-64 compared to lower radix arithmetic supported by commercially available processor microarchitectures.

[0042] The combined division / square root processing circuitry may perform the same number of radix-64 iterations per processing cycle for both the radix-64 division operation and the radix-64 square root operation. This may increase the degree to which circuitry can be shared between the square root and division operations and may help limit the overall circuit area of ​​the combined division / square root processing circuitry.

[0043] For both radix-64 division and radix-64 square root operations, the combined division / square root processing circuitry is capable of performing a given radix-64 iteration by performing one or more radix-m sub-iterations in the same processing cycle, where m≦64.

[0044] In some examples, m=64, in which case the radix-64 iteration may be performed as a single, integral operation that generates the next result digit six bits at a time, without splitting the radix-64 iteration into separate sub-iterations. This approach may be faster, but may require additional circuit logic to accommodate a larger number of candidate result digits, since the possible result digits may be expanded from -32 to +32 when the radix-64 iteration is performed as a single operation.

[0045] However, in some examples, m<64, the combined division / square root processing circuit may perform a given radix-64 iteration by performing multiple radix-m sub-iterations in the same processing cycle. For example, m in the specific example shown below is equal to 8, so there are two radix-8 sub-iterations in each radix-64 iteration. Another option may be for m=4, so that there are three radix-4 sub-iterations in one radix-64 iteration per processing cycle. Although the sub-iteration radix m may take on different values ​​between different sub-iterations, as discussed above for the example square root processing circuit, it may be more efficient from a circuit implementation standpoint if m is the same in each sub-iteration.

[0046] Thus, the term "base-m sub-iteration" is used to refer to either the entire base-64 iteration, if there is no subdivision into multiple sub-iterations of smaller bases, or to each individual sub-iteration of the smaller base, if such subdivision is performed.

[0047] There may be different portions of the combined divide / square root processing circuitry, which may function as the shared circuitry discussed above.

[0048] In one example, the shared circuitry comprises a shared digit selection circuit that selects, in a given base-m sub-iteration, a next base-m digit of the division or square root result based on a comparison of a previous remainder estimate with a set of comparison constants. In implementations where m=64 and thus there is no division of a base-64 iteration into multiple sub-iterations, the previous remainder estimate used for digit selection may come from the previous base-64 iteration. On the other hand, if m<64 such that a base-64 iteration is divided into multiple base-64 sub-iterations, in the first base-m sub-iteration of a given base-64 iteration, the previous remainder estimate may come from the last base-m sub-iteration of the previous base-64 iteration, and in subsequent base-m sub-iterations other than the first base-m sub-iteration of a given base-64 iteration, the shared digit selection circuitry may select the next base-m digit based on the previous remainder estimate calculated in a previous base-m sub-iteration of a given base-64 iteration.

[0049] Thus, a shared digit selection circuit may be provided to save circuit area compared to separate circuits for selecting the result digits of the division and square root operations, respectively. For example, the shared digit selection circuit may comprise the same set of comparator circuits used to perform a comparison between a previous remainder estimate and a comparison constant for both the division and square root operations.

[0050] The comparator circuitry used when performing both the division and square root operations may be the same, but the shared digit selection circuitry may use different sets of comparison constants for the base 64 division operation and the base 64 square root operation. The set of comparison constants may be selected based on the operation type.

[0051] However, one problem is that the comparison constant for the division operation may not be the same size as the comparison constant for the square root operation. Error analysis has found that the division operation may not require as many bits in the comparison constant as the comparison constant used for the square root operation to provide sufficient precision of digit selection. Thus, the division comparison constant may be expected to have fewer bits than the square root comparison constant. However, to facilitate circuit sharing, the comparison constant that is compared to the remainder estimate before the radix-64 division operation may have at least one least significant bit set to 0 to pad it to the same width as the comparison constant that is compared to the remainder estimate before the radix-64 square root operation. By placing at least one 0 in the least significant bit position, expanding the comparison constant for the division to the same bit width as that used for the square root operation, this allows the same comparators in the digit selection circuitry and the same data path for the remainder estimate to be used for both the square root and division operations, reducing circuit area.

[0052] Another example of a shared circuit may be a shared remainder update circuit that adjusts a previous remainder value based on a remainder adjust value in a given radix-m sub-iteration to generate an updated remainder value in a redundant representation. By using a redundant representation, the remainder update may be performed using a carry-reserve add to avoid the increased delay of a carry-propagate add. Thus, the shared circuit may comprise a shared carry-reserve add circuit that performs a carry-reserve add to generate the updated remainder value. This eliminates the need for two separate carry-reserve adders for the division and square root operations, since the data path for the remainder value is shared between the division and square root operations.

[0053] However, the remainder adjustment value may be different for a division operation compared to a square root operation. Thus, the shared remainder update circuit may comprise a selection circuit that selects as the remainder adjustment value a value derived from a divisor value when performing a partial iteration of a given base m as part of a base-64 division operation, and a value derived from a partial root value as a function of a series of previously selected base-m root digits when performing a partial iteration of a given base m as part of a base-64 square root operation. Thus, with a small amount of additional logic in the selection circuit, a shared data path may be used for both square root and division operations when generating the remainder update.

[0054] Another example of a shared circuit may be a shared remainder estimation circuit for generating updated remainder estimates indicative of non-redundant estimates of a portion of updated remainder values ​​generated in a redundant representation in a given base-m sub-iteration of a base-64 division operation or a base-64 square root operation. For example, the shared remainder estimation circuit may comprise a carry propagate adder circuit for performing carry propagate addition to generate the non-redundant estimates, thereby sharing it between the division operation and the square root operation, thereby avoiding the need for two separate carry propagate adders.

[0055] In implementations where m is less than 64, in the last radix-m sub-iteration of a given radix-64 iteration, the shared remainder estimation circuit may generate updated remainder estimates in parallel with the shared remainder update circuit that generates the updated remainder values. This improves performance by reducing the latency of critical timing paths for the same reasons as described above for the square root processing circuit.

[0056] Another example of a shared circuit may be shared with an on-the-fly transformation circuit for performing on-the-fly transformations to generate partial result values ​​in a non-redundant representation in a given partial iteration of radix m. Again, the on-the-fly transformation circuit may require relatively complex hardware circuit logic, and thus by avoiding duplicating it for division and square root operations, a greater amount of circuit area can be saved.

[0057] However, one problem is that in typical schemes, the on-the-fly conversion circuitry is performed differently for a division operation compared to a square root operation. The on-the-fly conversion circuitry can insert a value selected based on the next result digit into the partial result value to generate an on-the-fly conversion value representing the partial result corresponding to the sequence of result digits selected in that cycle and any previous cycles. However, in typical schemes, the position at which the next digit is inserted into the partial result value during the on-the-fly conversion is different for a division operation and a square root operation, and the division operation is performed to insert a value derived from the next digit into the least significant bit position with a left shift to shift up all previously inserted bits to more significant bit positions. In contrast, due to the fact that the partial result value influences the digit selection and remainder update operations in a square root operation (and therefore it is more convenient if, in each processing cycle, the most significant bit of the partial root result value remains in a consistent bit position in the stored representation of the partial result), for a square root operation a value derived from the next result digit is inserted into a variable bit position in the partial result, with a mask used to represent the position in the partial result value where the next square root result digit is to be inserted. This mask may be adjusted between iterations or partial iterations to gradually move the position where the next result digit is inserted towards the more significant bits of the partial result value.

[0058] Given these contrasting methods of maintaining partial result values, one might think that having shared circuit logic for the on-the-fly transformation circuitry would be difficult.

[0059] However, the inventor has recognized that it is possible to provide a shared on-the-fly conversion circuit. In a given radix-n partial iteration, the shared on-the-fly conversion circuit selects a position for inserting the next digit into the partial result value based on a mask value for both the radix-64 division operation and the radix-64 square root operation. Thus, for the division operation, instead of shifting up all digits and inserting the next digit into the least significant bit position, the shared on-the-fly conversion circuit behaves unconventionally to use a mask for the radix-64 division operation to select a position where the next digit is inserted into the partial result value for the division operation. This allows the on-the-fly conversion for the division operation to mirror the conversion for the square root operation, using shared circuit logic and shared data paths. This helps improve overall circuit area efficiency.

[0060] Similar to the various circuit units of the square root processing circuit described above, the shared circuit in the shared division / square root circuit may comprise one or more instances of a replicated circuit, each instance of the replicated circuit comprising two or more replicated circuit units for determining two or more candidate output values ​​corresponding to different digits that may be selected as the next base m digit in parallel with the selection of the next base m digit of the division result or square root result, and a selection circuit for selecting one of the plurality of candidate output values ​​in response to an indication of which of the different digits has been selected as the next base m digit, the plurality of candidate output values ​​including at least two or more candidate output values ​​generated by the two or more replicated circuit units. This helps to improve performance for the same reasons as described above for the square root example. Again, to reduce the total number of replicated circuit units required to process the base m partial iterations, at least one of the replicated circuit units may be a shared circuit unit shared between positive and negative digits of equal magnitude. Various components of the combined divide / square root circuit may use any one or more of such replica circuits, for example, a remainder update circuit, a remainder estimation circuit, and an on-the-fly transformation circuit.

[0061] Similar to the square root processing circuit described above, in the case of the combined division / square root processing circuit, this may be implemented as an iterative division / square root processing circuit, where the output of one radix 64 iteration is input to the same iterative division / square root processing circuit for use in further radix 64 iterations of the division or square root operation, or as a pipelined division / square root processing unit having several pipeline stages, each having an instance of each of the combined division / square root processing circuit, and a signal path provides the output generated in one stage as input to the next stage in the pipeline.

[0062] Division / Square Root Pipeline It is common in many programs to need to perform arithmetic operations on operands represented in floating-point format. The IEEE-754 technical standard defines various formats for floating-point representation, e.g., half precision (HP), single precision (SP), and double precision (DP) (other formats are available). The particular floating-point precision used for the operands and result of a division or square root operation can control the number of bits that need to be generated for the result, which can affect the number of iterations required for a digit-recurring division or square root operation.

[0063] Conventionally, circuit units for performing digit recursive division or square root operations capable of producing results with floating-point level accuracy are implemented as iterative circuit units, such that the circuit logic provided in hardware corresponds to a single iteration of the digit recursive division or square root operation, and the output of one iteration is fed back as input to the very same circuit logic unit that performed the previous iteration, ready for that same circuit unit to perform the next iteration.

[0064] In contrast, in the example described below, a division / square root pipeline is provided that includes several division / square root iterative pipeline stages, each capable of performing a respective iteration of a digit recursive division or square root operation. A signal path is provided to supply the output generated by one pipeline stage in one iteration as an input to a subsequent pipeline stage of the division / square root pipeline to perform a subsequent iteration of the digit recursive division or square root operation. The division / square root pipeline can perform the digit recursive division or square root operation on floating-point operands to generate a floating-point result.

[0065] Thus, while supporting the level of precision required for floating-point formats, division or square root operations are implemented in a pipelined manner rather than as an iterative unit. This means that for the processing of a single division or square root operation, each iteration is performed by a different pipeline stage, and the output from one pipeline stage is input to the next pipeline stage, so that the operation can move up the pipeline and output a result until it reaches the end.

[0066] This approach can be thought of as counter to straight pipes, and although pipelining of instructions is generally known, the much greater complexity of division / square root operations compared to other forms of operations means that the overall circuit area of ​​a single circuit unit to perform a single iteration of a digit-recursive division or square root operation is relatively large, and one would therefore expect that extending the iteration unit into a pipeline containing a sufficient number of stages to produce the result precision required for floating-point processing would significantly increase the overall circuit area required for the division / square root unit by a factor corresponding to the maximum number of iterations required for the division or square root operation.

[0067] However, the inventors have recognized that in practice, a processor microarchitecture having an iterative division / square root processing circuit may actually provide multiple parallel division / square root units to increase the overall bandwidth available, such that, for example, there may be multiple division functional units and / or multiple square root functional units, and more than one division or square root operation may be processed simultaneously. In a pipelined manner, the need to replicate the entire division / square root unit is eliminated because the division / square root pipeline may process multiple operations in a pipelined manner, with a division / square root iterative pipeline stage after the division / square root pipeline performing a second digit recursive division or square root operation in parallel with a previous division / square root iterative pipeline stage performing a first digit recursive division or square root operation and a previous iteration of the second digit recursive division / square root operation.

[0068] Thus, although pipelining appears to significantly increase circuit logic, in reality the additional circuitry may not be significant compared to commercially available processors with multiple parallel divide / square root units, and various techniques described herein for reducing circuit area can be applied, particularly using shared data paths for division and square root operations and reducing the number of replicated circuit units by sharing the same replicated circuit units for positive and negative digits of the same magnitude as described previously.

[0069] Thus, the entire pipeline may be competitive in terms of circuit area and may help improve performance because with pipelining of operations, the pipelining method may avoid blocking the iterative circuit units for the total number of cycles used to perform the digit recursive division or square root operation, thereby allowing for higher throughput since successive division or square root operations may be scheduled with fewer cycles between them.

[0070] It is possible for a pipeline to perform only division or square root operations, just as a division / square root pipeline can perform either division or square root operations but not both.

[0071] However, the pipeline may be particularly useful if the combined division / square root processing circuitry is provided with a shared data path used for both operations. Thus, each division / square root iteration pipeline stage comprises a combined division / square root processing circuitry for performing a given iteration of a digit recursive division operation in response to a division instruction and for performing a given iteration of a digit recursive square root operation in response to a square root instruction. The combined division / square root processing circuitry comprises shared circuitry for generating at least one output value on the same data path used for both the given iteration of the digit recursive division operation and the given iteration of the digit recursive square root operation. Providing a combined division / square root processing circuitry helps to limit the overall area cost of extending a single iteration unit into a pipeline (since the area budget previously provided for the separate division and square root units is available for the implementation of the pipeline) and helps the pipeline to be competitive with current microarchitectures in terms of circuit area. As mentioned earlier, when a divide / square root combinational circuit is used, it may be useful for the divide / square root pipeline to perform the same number of iterations per processing cycle, in the same radix, for both the digit recursive division operation and the digit recursive square root operation, as this facilitates greater sharing of shared circuit units.

[0072] For a given result precision, the division / square root pipeline can process a digit-recursive division operation in the same number of processing cycles as a digit-recursive square root operation. This helps to simplify control of circuit timing within the pipeline and facilitates sharing of common circuit logic between division and square root operations.

[0073] A variety of floating-point formats may be supported for the operand(s) input to the division or square root operation and the floating-point result generated by the division or square root operation. For example, the operand(s) and result may be half-precision (HP), single-precision (SP), or double-precision (DP) floating-point values. The division / square root pipeline may support at least one of these formats, or may support other types of floating-point formats. However, it is particularly useful for the division / square root pipeline to support at least one of SP and DP floating-point values. Programs written in DP floating-point precision may be particularly common, so in some cases it may be useful for the division / square root pipeline to support operations whose results are in DP floating-point representation. The pipeline stages of the division / square root pipeline may be used to process the mantissa of the floating-point operands to generate the mantissa of the floating-point result. There may be separate circuit logic for processing the exponent of the floating-point value. The exponent processing logic may be simpler than the logic for generating the mantissa and may use any known technique for generating the exponent of a division / square root result.

[0074] In some examples, the division / square root pipeline may support at least two different result precisions for digit recursive division or square root operations. For example, the division / square root pipeline may support any two or more of HP, SP, and DP floating-point values.

[0075] For lower precision floating-point result precision, the division / square root pipeline can perform the division or square root operation in fewer processing cycles than when producing a higher precision result (fewer iterations of the digit recursion method are required because fewer bits need to be generated for the result). The device can have control circuitry that controls the division / square root pipeline to cause at least one division / square root iteration pipeline stage used to perform at least one iteration of the digit recursion division or square root operation when producing a higher precision result to be bypassed when performing the digit recursion division or square root operation to produce a lower precision result. This improves performance by making the result of the operation available sooner when fewer bits need to be calculated.

[0076] However, allowing some stages of the pipeline to be bypassed in this way may create the possibility that, when a pipelined high-precision operation is followed by a low-precision operation, both operations may collide when they reach a post-processing stage that may perform post-processing operations on the output of the last iteration of the digit-recursive division or square root operation. For example, the post-processing stage may perform rounding of the result of the division or square root operation to provide a rounded floating-point result, and / or may perform denormal (subnormal) result processing by right-shifting to produce a result according to the IEEE standard (when the result of the division or square root operation is less than the smallest number that can be represented as a normal floating-point number). In order to ensure that the post-processing operation receives only the output of the last iteration of a single operation per cycle, the control circuit can prevent a less precise digit recursive division / square root operation performed to generate a less precise result from starting a predetermined number of cycles after a more precise digit recursive division / square root operation performed to generate a more precise result, the predetermined number of cycles corresponding to the difference between the number of cycles required to reach at least one post-processing stage for the more precise digit recursive division / square root operation and the number of cycles required to reach at least one post-processing stage for the less precise digit recursive division / square root operation. Thus, depending on the difference in precision between the previous high-precision operation and the subsequent low-precision operation, there may be a certain number of cycles in which the start of the low-precision operation is prohibited after the high-precision operation to avoid collisions. The predetermined number of cycles may be different for different pairs of precision formats.

[0077] Each division / square root iteration pipeline stage may include a digit selection circuit for selecting a next result digit for a partial result value of the digit recursive division or square root operation based on a comparison between a previous remainder value and a set of comparison constants, and a remainder update circuit updates the previous remainder value based on the remainder adjustment value and the next result digit selected by the digit selection circuit. Each pipeline stage may also have other elements such as a remainder estimation circuit for generating a non-redundant estimate of a portion of the updated remainder value generated by the remainder update circuit in a redundant representation. Each pipeline stage may also have an on-the-fly conversion circuit for maintaining on the fly a non-redundant version of the partial result value corresponding to a previously selected sequence of result digits from all previous iterations of the digit recursion method.

[0078] All of the division / square root iteration pipeline stages of the pipeline may use the same set of comparison constants for each iteration performed within the same digit recursive division or square root operation. The comparison constants may be different for each operation, but the same set of comparison constants may be used within each iteration of the same operation. Thus, the division / square root pipeline may perform a table lookup to obtain the set of comparison constants in a pre-processing stage of the division / square root pipeline prior to the first division / square root iteration pipeline stage of the division / square root pipeline, and the set of comparison constants is passed from stage to stage to avoid repeating table lookups at each division / square root iteration pipeline stage within the same digit recursive division or square root operation. This approach may reduce timing for each individual pipeline stage since there is no need to perform table lookups at each stage, reducing the overall amount of circuit logic required at each stage. There may be a set of flip-flops provided at each pipeline stage, which simply capture the comparison constants received from the previous pipeline stage without the need to update their comparison constants. This greatly simplifies the pipeline and reduces the overall circuit area.

[0079] This approach may be surprising since one might think that the comparison constants of a digit recursive division or square root operation should not be the same for each iteration, and that a different set of comparison constants may be required compared to the constants used in later stages, especially in the first iteration of a typical division / square root operation. However, in the example described below, the division / square root pipeline comprises at least one pre-processing stage for performing operand pre-processing before the first division / square root iteration pipeline stage of the division / square root pipeline, the operand pre-processing including the selection of at least one initial result digit for the result of the digit recursive division or square root operation. By selecting at least one initial result digit for the result of the division or square root operation in the pre-processing stage such that the initial result digit is not selected within the main body of the pipeline, this means that a different set of selection criteria can be used for the result digit to avoid requiring different comparison constants in different stages of the main iteration part of the pipeline. This means that the remaining divide / square root iterative pipeline stages can each use the same set of comparison constants within the same divide or square root operation, improving circuit timing and reducing circuit area as discussed above.

[0080] However, one issue in implementations where the division / square root pipeline supports both digit-recurring division and digit-recurring square root operations (wherein a combined division / square root circuit is provided as described above) is that the number of initial digits that require a different set of comparison constants compared to subsequent iterations may be different for the division and square root operations. For example, error analysis has shown that, in order to obtain digit selections of sufficient accuracy for the square root operation, if a base of 8 is used for digit selection in a given iteration or sub-iteration, the selection of the first two square root digits may use different comparison constants than the selection of the remaining square root digits. If the base used is a base other than 8, the number of initial root digits selected using different comparison constants for the remaining iterations may be a number other than 2. Nevertheless, regardless of the base, in general, the square root operation may use different comparison constants to select a particular number of initial root digits, and use the same set of comparison constants for subsequent iterations or sub-iterations after those initial root digits have been selected. In contrast, the division operation may use the same comparison constants for the selection of all result digits (regardless of the base used). However, for performance reasons, it may be desirable to select at least one of the result digits during a pre-processing stage to reduce the number of subsequent pipeline stages required for the division operation, and therefore reduce latency. For example, in the base 8 example described below, the first division digit may be selected in a pre-processing stage.

[0081] Thus, the number of initial digits selected in the pre-processing stages may be different for the square root and division operations. For example, at least one pre-processing stage may generate a larger number of initial result digits for the digit-recursive square root operation than for the digit-recursive division operation. While this may obviously introduce some asymmetry between the two operations, in practice this goes a long way in reducing the overall circuit area and improving the performance of the pipeline, since it means that for the square root operation, the comparison constants for the remaining stages can simply be latched from one stage to the next without requiring a separate table lookup at each pipeline stage.

[0082] However, since at least one pre-processing stage generates more initial result digits for the square root operation than for the division operation, this means that fewer remaining iterations are required after the pre-processing stage for the square root operation compared to the division operation even when producing a result of the same precision, and so the result of the square root operation may be available to an earlier division / square root iteration pipeline stage for the square root operation compared to the division operation. To enable a shared pipeline to be used, the control circuitry may control the division / square root pipeline to cause at least one division / square root iteration pipeline stage, which is used to perform at least one iteration when a digit recursive division operation is performed, to completely or partially skip or discard some bits of the result output when performing the digit recursive square root operation. In some cases, an entire pipeline stage of the pipeline may be skipped for the square root operation, while in other cases, only a portion of the bits generated in a given pipeline stage may need to be discarded, depending on the floating-point precision used and the base used for the digit recursive operation. For example, if a given iteration of the digit recursion method is divided into multiple sub-iterations of smaller bases, as in some of the examples above, it may be possible to skip only individual sub-iterations within a given division / square root iteration pipeline stage for some result precisions of the square root operation, rather than skipping the entire stage. Also, in some cases, if the total number of bits required for a given result precision of the square root operation is not an exact multiple of the number of bits generated per iteration or sub-iteration, truncation of the result can be obtained by fully executing a given iteration or sub-iteration, but discarding some bits of the result if other bits of the result digits generated in the last iteration or sub-iteration executed are still required.

[0083] This means that, when considering the body of the pipeline, the result of the square root operation may be available earlier than the result of the division operation, but the total number of cycles taken for the operations may be the same for both the square root and division operations. For example, even if the result of the square root operation is available earlier, there may be at least one cycle when the value is passed unchanged to the next cycle, allowing the overall operation timing to reflect that of the division operation. This may make it easier to implement scheduling of post-processing operations, for example, since they may have the same timing regardless of the operation being performed.

[0084] Another complication when using a combined divide / square root data path in a pipeline is in maintaining the partial result values ​​that provide a representation of a numerical value corresponding to a previously selected sequence of result digits. If a shared data path is to be used, it may be desirable to be able to insert the next result digit into the partial result value at the same bit position for both the division and square root operations when performing a given iteration of the digit recursion method in a given pipeline stage of the pipeline. However, if the pre-processing stage generates a different number of initial result digits for the division and square root operations, this may make it more complicated to use shared circuit logic in the remaining pipeline stages, since the position where the next result digit is inserted in a given iteration could differ from iteration to iteration.

[0085] Thus, when performing the digit recursive division operation, the at least one pre-processing stage can provide the first division / square root iteration pipeline stage with a partial result value in which selected bit positions are set to a dummy bit value, the selected bit positions corresponding to bit positions into which the at least one pre-processing stage inserts at least one additional result digit not generated for the digit recursive division operation when performing the digit recursive square root operation. This allows a given division / square root iteration pipeline stage of the division / square root pipeline to insert a next result digit into the partial result value at the same bit position for both the digit recursive division operation and the digit recursive square root operation. The division / square root pipeline can include a post-processing stage for removing the dummy bit value from the final result value when performing the digit recursive division operation.

[0086] This recognizes that inserting additional dummy bit values ​​into the partial result for a division operation does not affect the overall result of the division operation because the partial result value is not used for remainder update or digit selection operations in the division operation. It is only for the square root operation that the partial result value is used to control the remainder update and digit selection operations. For the division operation, it is not an issue that the partial result value temporarily contains some dummy bit values ​​that are removed in a post-processing stage, since the partial result value is simply maintained "on the fly" to improve performance by not having to convert the redundant representation of the result to a non-redundant form at the end of the pipeline. Including the dummy bit values ​​in the partial result value used for the division operation allows the insertion of the next result digit to be in the same position for both operations, improving the sharing of circuit logic for both operations.

[0087] The division / square root pipeline as described above can be used for digit recursive division or square root operations with any base.

[0088] However, since the extra number of bits of the result generated per cycle in a radix 64 operation compared to a lower radix helps reduce the total number of pipeline stages required for the pipeline, using a divide / square root pipeline can be particularly useful for radix 64 digit recurrence division or square root operations, and as a result, the pipeline can become competitive with respect to circuit area compared to an iterative implementation.

[0089] In one example, each divide / square root iterative pipeline stage is configured to perform each radix r iteration of a radix r digit recurrence division or square root operation by performing multiple radix n partial iterations in the same processing cycle, where n < r. By splitting higher radix iterations into multiple lower radix partial iterations, the amount of circuitry in each pipeline stage is reduced, and as a result, the overall circuit area of the pipeline can compete with current iterative implementations while improving performance. In a specific example, r = 64 and n = 8, but more generally, a radix r iteration can be split into different combinations of lower radix partial iterations, as described above for the example of the square root processing circuit.

[0090] On-the-fly conversion A data processing apparatus for converting a plurality of signed digits representing an input value into a redundant representation, comprising, in each of a plurality of iterations, a receiving circuit that receives signed digits from the plurality of signed digits and previous intermediate data from a previous iteration, a concatenating circuit that performs a concatenation of bits corresponding to the signed digits and bits of the previous intermediate data to generate updated intermediate data, and an output circuit that provides the updated intermediate data as the previous intermediate data for a next iteration, wherein the previous intermediate data includes S3[i] in a non-redundant representation, which is the non-redundant representation of at least a part of the input value multiplied by 3.

[0091] In these examples, the individual digits are signed. Thus, the input value (which can be positive or negative) is composed of individual digits, each digit being individually signed. In this way, for example, the first digit of the input value can be positive and the second digit of the input value can be negative. This can be used to provide a form of representation known as a redundant representation, in which a pair of words are used to represent the input value. This is in contrast to a non-redundant representation, in which a number is represented using a single word. Non-redundant and redundant representations are each best suited for certain types of operations, and therefore conversion between different representation forms can be useful. The conversion is performed on the fly as each digit of the input value is received, thereby avoiding the large latency that can occur if all digits are converted at once after all digits have been received. The conversion process is achieved using a concatenation of bits, which can be performed quickly. The bits to be concatenated are derived from the signed digit. A set of intermediate data is maintained between iterations and updated with each iteration. The concatenation performed depends on the new current digit received. In particular, the intermediate data includes S3[i], which is S[i] (the partial result) multiplied by 3. The value of S3[i] is achieved without simply multiplying S[i] by 3, which would take too much time to keep up with the arrival of new signed digits, not to mention being energy intensive. Note that although the term "iteration" is used here, the iterations being referred to could be "partial iterations" as discussed above.

[0092] In some examples, the previous intermediate data includes S3[i-1]. In these examples, the value of S3 from the previous iteration, S3[i-1], is also maintained in the intermediate data. This value does not need to be calculated and can be carried over from the previous iteration. Providing such data allows adjustments to be made when carries are made during the conversion process.

[0093] In some examples, the previous intermediate data includes S3M[i], which is a non-redundant representation of at least a portion of the input value multiplied by 3 minus 1. In other words, S3M[i] = (S[i] × 3) - 1. The value of SM3[i] is equal to the value of S3[i] minus 1.

[0094] In some examples, the previous intermediate data includes S3M[i-1]. In these examples, the value of S3M from the previous iteration is also maintained in the intermediate data. This value does not need to be calculated and can be carried over from the previous iteration. Providing such data allows adjustments to be made when carries are made during the transformation process.

[0095] In some examples, the concatenation performed by the concatenation circuitry includes concatenation for each of S3[i] and S3M[i] to generate updated intermediate data including S3[i+1] and S3M[i+1]. Thus, each of the four values ​​has concatenation performed at each iteration (or sub-iteration). The concatenation may be different for each of the four values.

[0096] In some examples, the bit corresponding to the unsigned digit is concatenated to one of S3[i] and S3M[i] to generate S3[i+1], and the other of S3[i] and S3M[i] to generate S3M[i]. One of S3[i] and S3M[i] is determined based on whether the unsigned digit is greater than or less than 0. In these examples, whether the unsigned digit is greater than, equal to, or less than 0 affects whether S3[i] or S3M[i] is used to generate S3[i+1], and the other of S3[i] and S3M[i] is used to generate S3M[i+1].

[0097] In some examples, the data processing apparatus includes an adjustment circuit configured to perform selective adjustment on at least one of S3[i] and S3M[i] prior to concatenation based on the magnitude of the signed digit and whether the signed digit is positive or negative. The selective adjustment can be used, for example, to achieve a carry between columns of the output value.

[0098] In some instances, selective adjustment is performed when the magnitude of the signed digit multiplied by 3 exceeds the base in which the signed digit is represented. Selective adjustment can be used to handle situations where the concatenated digit multiplied by 3 is greater than the base being used for the conversion, and therefore it is necessary to increment or decrement a digit in another position. For example, as in base 10, if one has a partial result S[i]=512 and it is desired to add a digit to this number (thousands) 6, this can be done to achieve the number S[i+1]=6512. However, if one maintains S3[i]=1536 and it is desired to add a digit to this number (thousands) 6, then 3 * We need to add 6 = 18. However, because it is base 10 and 18 is greater than 10, this cannot be done by changing a single position. Instead, we add 8 to the thousands number to give us 9536, and then carry in a "1" as the ten thousand number to give us 19536.

[0099] In some examples, the data processing device is configured to convert multiple signed digits representing an input value in a redundant representation without using an adder circuit. In particular, the value of S3M[i] is not derived by simply taking S3[i] and subtracting 1 (e.g., using an adder circuit). Instead, by calculating these values ​​using concatenation over i iterations (concatenating a different number for each of S3[i] and SM3[i]), it is possible to determine these numbers with lower latency than would be achieved by using an adder circuit to perform the subtraction of 1.

[0100] In some examples, the data processing apparatus includes a digit recursion circuit for performing a digit recursion operation to generate a plurality of signed digits, and in each of a plurality of iterations, one of the plurality of signed numbers is provided to the receiving circuit. The digit recursion circuit can be used to provide a sequence of digits that make up the input value, with a subset of the digits being provided in each iteration (or partial iteration), e.g., each clock cycle.

[0101] In some examples, the digit recursion circuit is configured to operate in a square root operation mode where the digit recursion operation is a square root operation. The digit recursion algorithm for calculating square roots performs a multiplication of a partial root S, where the multiplication depends on the digits being added. This multiplication is performed for each iteration because the partial root S changes with each iteration. Multiplying by 0 always results in 0. Multiplying by 1 is simply the identity function. On the other hand, by performing a bit shift, it is possible to multiply by a power of 2 (e.g., 2 or 4). Similarly, multiplication by -1, -2, and -4 can be performed by negating the multiplications by 1, 2, and 4, respectively. However, multiplication by 3 is significantly more complex. A multiplication circuit that performs an actual multiplication by 3 may require several processor cycles that are too slow. Even the addition of X and 2X to determine 3X requires additional circuitry, which may also take too long to perform. Thus, by maintaining the value of S3 achieved through concatenation, an efficient square root digit recursion can be performed.

[0102] In some examples, the digit recursion circuit is configured to operate in a division operation mode in which the digit recursion operation is a division operation, the previous intermediate data includes S[i], which is at least a portion of the input value in a non-redundant representation, and SM[i], which is at least a portion of the input value in a non-redundant representation minus one, and after multiple iterations, the output circuit is further configured to output S[i], so that the same data processing device that performs the conversion from the input value to the output value can be used for both the square root operation and the division operation. The calculation can also include generating S[i], which is at least a portion of the input value converted to a non-redundant representation, and SM[i], which is its value minus one.

[0103] In some examples, the concatenation circuit is configured to suppress the generation of S3[i] in the division operation mode. As explained above, the value of S3 (and by extension, S3M) is particularly relevant when performing square root digit recursion. Also, when performing digit recursion division, the generation of S3 and S3M can be avoided because the multiplication of partial roots is not required for each iteration. Therefore, by suppressing the generation of S3 and S3M in the division operation mode, power consumption can be reduced.

[0104] In some examples, the digit recursion operation has a base of at least 8. For bases of at least 8, the available digits include at least one, if not both, of +3 and -3. Thus, during the square root digit recursion algorithm, it may be necessary to multiply the partial root by either 3 or -3 depending on the most recent digit. As previously mentioned, multiplication by 3 can be time consuming, so by maintaining S3 and S3M through the concatenation, it is possible to efficiently perform square root digit recursion for bases of 8 while still meeting the timing constraints of the circuit.

[0105] In some examples, the possible values ​​of the signed digit include at least one of +3 and -3. As mentioned above, the use of signed digits can require multiplication by 3, which is more difficult to perform than multiplication involving powers of 2.

[0106] Selection constant In some examples, a data processing apparatus for performing a digit recursion operation on an input value is provided, the data processing apparatus including: a receiving circuit configured to receive a remainder value of a previous iteration of the digit recursion operation; a comparison circuit configured to perform a comparison between each of a plurality of selection constants associated with an available digit of a next digit of a result of the digit recursion operation and a most significant bit of the remainder value of the previous iteration of the digit recursion operation, and to output the next digit of the result of the digit recursion operation based on the comparison, each of the selection constants being associated with one of the available digits and the input parameter; and a memory circuit configured to store a subset of the selection constants, the subset of the selection constants excluding the excluded selection constants from the selection constants associated with digits excluded from the available digits.

[0107] During the digit recursion process, a comparison is performed between the most significant bit of the remainder value of the previous iteration and some selection constants to determine the next digit of the digit recursion operation, i.e. the next digit to be output. The number of selection constants corresponds to the product of the number of possible values ​​of the most significant bit of the remainder value and the number of possible values ​​that the output digit can have. For example, if the 6 most significant bits of the remainder value are considered and there are 8 possible values ​​for each output digit, the selection constant table holds 8 x 32 = 256 values. Each value may also occupy several bits. Also, usually, multiple tables need to be provided to accommodate both square root digit recursion and division digit recursion. Thus, the number of values ​​to be stored is large. In the above example, at least some of the selection constants required are not stored. That is, for the range of digit recursion operations supported (based on the base considered and the number of most significant bits), at least some of the selection constants required for the digit selection process are not stored anywhere in the data processing device. This allows to reduce the amount of storage space required. This results in a more compact and lower power circuit.

[0108] In some examples, the data processing device comprises a conversion circuit configured to generate selection constants that are excluded from the selection constants stored in the storage circuit, in these examples, the missing or omitted selection constants that are not stored in the data processing device are instead inferred or generated from other selection constants that are stored in the data processing device.

[0109] In some examples, the conversion circuitry is configured to generate the omitted selection constants by performing a selective inversion on the sign of one of the selection constants stored in the storage circuitry. In these examples, some of the omitted selection constants may be generated by taking another selection constant and inverting its sign. Inverting the sign of a number (e.g., by taking two's complement) need not affect the time it takes to perform the selection operation, since this can be performed efficiently.

[0110] In some examples, one of the selection constants is associated with the same input parameter and a different one of the digits available as the excluded selection constant. Thus, two columns of the selection constant table can be "merged"; that is, for a given set of most significant bits of the remainder value, the selection constants of two different digits are the same (the sign changes according to the number for which the selection constant is generated). For example, the selection constant of the bit 0.100010 of the remainder can be "2" for possible output digits +4 and -3. However, for the digit +4, the selection constant can be negative (-2) and for the digit -3, the selection constant can be negative (+2). Thus, these two columns can be merged into one using the rules regarding whether the constants are positive or negative.

[0111] In some examples, the storage circuitry is configured to store an exception flag that indicates whether a selective inversion should be performed for the selection constant to generate the excluded selection constant. In these examples, whether to invert depends on the value of the exception flag. The inversion may also depend on other factors, for example, depending on the digit for which the selection constant is being generated. For example, considering the previous example for the remainder bit 0.100010, the selection constant may be positive (+2) for one digit (+4) and negative (-2) for another digit (-3). However, the exception flag overrides this (making both digits have the same selection constant) or inverts it (-2 for digit +4 and +2 for digit +3).

[0112] In some examples, the digit recursion operation is a square root digit recursion operation. The input parameter is a partial root.

[0113] In some examples, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor.

[0114] In some examples, in a division operation mode, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor, and in a square root operation mode, the digit recursion operation is a square root digit recursion operation and the input parameter is a partial root. Thus, in these examples, a device may be used that performs both division digit recursion and square root digit recursion depending on the operation mode.

[0115] In some examples, in a division operation mode, the digit recursion operation is a division digit recursion operation and the input parameter is a divisor. In a square root operation mode, the digit recursion operation is a square root digit recursion operation and the input parameter is a partial root, and each selection constant is a division digit recursion operation selection constant or a square root digit digit recursion operation selection constant. Such a data processing apparatus can perform both division and square root digit recursion, but the selection constants stored are specific to one of these two operation modes (division or square root). By storing selection constants specific to only one of the two operation modes, the storage requirements of the data processing apparatus can be reduced.

[0116] In some examples, each of the selection constants is a division digit recursion operation selection constant, which does not mean that all of the selection constants for division digit recursion are stored, but simply that the constants stored are division digit recursion selection constants that may be used as part of the process of generating the square root digit recursion selection constants.

[0117] In some examples, the conversion circuit is configured to generate the excluded selection constant in the division operation mode by performing a selective inversion of the sign of one of the division digit recursion operation selection constants, i.e., one of the division digit recursion constants is used and inverted based on some criteria (e.g., the value of the digit with which the constant is associated).

[0118] In some examples, the conversion circuit is configured to generate the selection constant to be excluded in the square root mode of operation by referencing one of the division digit recursion operation selection constants.

[0119] In some examples, the memory circuitry is configured to store a plurality of mappings between the excluded selection constants in the square root mode and one of the division digit recursion operation selection constants, the mappings being used to indicate which division digit recursion operation selection constant to use as a basis for creating the square root digit recursion operation selection constant and / or how to modify one of the division digit recursion operation selection constants to generate the corresponding square root digit recursion operation selection constant.

[0120] In some examples, the storage circuitry is configured to store an exception flag that indicates, for the selection constant, whether a selective inversion should be performed to generate the excluded selection constant. The exception flag may be part of a set of flags (or stored as part of a larger value) that indicate circumstances under which inversion occurs to generate the excluded selection constant.

[0121] In some examples, the digit recursion operation is base 8. For example, the available digits may be restricted to {-4, -3, -2, -1, 0, 1, 2, 3, 4}.

[0122] Examples of data processing devices FIG. 1 illustrates an example of a data processing apparatus 2, e.g., a processor, that supports the execution of instructions defined according to a particular instruction set architecture (ISA). The apparatus has instruction fetch circuitry 4 for fetching program instructions defined according to the architecture from an instruction cache or memory (not shown in FIG. 1). The fetched instructions are decoded by decode circuitry 6 to identify the operation to be performed. In response to a given instruction, decode circuitry 6 generates control signals that control execution units 8 to perform the processing operation represented by the instruction. Operands for the given processing operation may be read from registers 10, and processing results of the operation may be written back to registers 10. Execution units 8 may include various execution units, including arithmetic units such as adder 20, multiplier 22, divide / square root unit 24, etc. The execution units may also include other types of functional units, such as a branch unit 26 for determining the outcome of a branch instruction that may trigger a non-sequential change in program flow within a program being executed, and a load / store unit 28 for executing load instructions to load data from a cache or memory into register 10, or store instructions to store data from register 10 into cache or memory.

[0123] The following example shows a circuit logic design of the division / square root execution unit 24 of the processing device 2. When the decode stage 6 decodes a division instruction, the decode stage 6 controls the division / square root execution unit 24 to perform a digit recursion division operation. When the decode stage 6 decodes a square root instruction, the decode stage 6 controls the division / square root execution unit 24 to perform a digit recursion square root operation.

[0124] The examples that follow focus on the divide / square root execution unit 24, although it will be appreciated that the remainder of the processing unit 2 may be constructed in accordance with any known processor design techniques. It will be appreciated that Figure 1 is a simplified representation of the components of a data processor, and that in practice many other components not shown in Figure 1 may also be provided.

[0125] Theoretical foundations of digit recursive division and square roots Digit recursion is the process of iterating the resulting digit p of base r. (i+1) and a class of iterative algorithms that compute the remainder rem[i]. The remainder is used to obtain the next base-r digit, where base r is a power of 2, and each base-r digit represents log2(r) bits of the result. Division (x / d) and Square Root

[0126]

number

[0127] The partial result before iteration i is defined as follows:

[0128]

number

[0129]

number

[0130]

number

[0131]

number

[0132]

number

[0133] For fast iteration, the remainder is held in carry-save or signed-digit-redundant representation. In the implementation described below, known techniques are used to represent the remainder using a representation such as carry-save, where the remainder is represented by a positive word and a negative word (the non-redundant binary value corresponding to the remainder can be obtained by subtracting the negative word from the positive word).

[0134] On the other hand, due to the algorithm convergence condition in equation (3) and the multiplication time r, the remainder has some number of bits in the integer part. The number of integer bits depends on the base, digit set, and operation.

[0135] Then, for each iteration, the base r digits of the result are taken from the current remainder, a new remainder is calculated for the next iteration, and the partial result is updated.

[0136] The selection function for selecting the next result digit is the remainder estimate

[0137]

number

[0138]

number

[0139]

number

[0140]

number

[0141] The partial results are redundant representations of signed digits in base r, generated most significant digit first (MSDF). This is converted to a non-redundant representation at each iteration. The most efficient conversion technique is the well-known on-the-fly conversion. Essentially, on-the-fly conversion is a conversion of the digits p i+1 to the partial result P[i] (see equation (1)). However, because the digits can be negative, this addition can produce carry propagation. To prevent this slow carry propagation, a separate form of the result is maintained, where PM[i] has the following value: PM[i] = P[i] -r -i (6) Using this second form, the conversion algorithm for concatenation is:

[0142]

number

[0143] In this way, there is no arithmetic involved in the conversion, just a concatenation of values ​​to P[i] and PM[i], where the concatenated value is the selected digit p i+1 Depends on.

[0144] The number of iterations of the digit recursion algorithm is it=[n / log2(r)] (9) n is the number of bits in the result, including the bits needed for rounding. [...] is a ceiling function, so [n / (log2(r)] is the smallest integer greater than or equal to n / (log2(r).

[0145] The number of cycles is directly related to the number of iterations and the number of iterations performed per cycle. Now, considering m iterations per cycle, the number of cycles is: cycles=[it / m] (10)

[0146] Equations (1)-(10) can be subdivided to any base. In the next two sections, these equations are specialized for r=8, division and square root. A higher base r=64 is obtained by stacking two base-8 partial iterations. The partial iteration base is 8.

[0147] Base 8 division Floating-point division of dividend x and divisor d produces a quotient q=x / d. In base 8, the partial quotient (partial result) before iteration i and the digits obtained at iteration i are denoted Q[i] and q i+1 and equation (1) can be rewritten as

[0148]

number

[0149]

number

[0150] Regarding the selection function, it turns out that only the 10 most significant bits of the remainder need to be assimilated to obtain a remainder estimate accurate enough for digit selection. As mentioned before, the selection constant also depends on the divisor. The 6 most significant bits of the divisor are used to derive a set of 8 selection constants for every iteration of the current division. Different divisor values ​​can derive different sets. Note that the most significant bit of the divisor is always 1, since the operands are normalized before selecting the constants. The selection constants are stored in a look-up table (LUT).

[0151] In this implementation, it has been determined that only the 10 most significant bits (MSBs) of the remainder, 3 integer bits, and 7 fractional bits are needed to select the next quotient digit in equation (12).

[0152] base 8 square root The floating-point square root of the operand x is the root

[0153]

number

[0154]

number

[0155]

number

[0156] The initial values ​​of the remainder and partial root are rem[0]=x-1 and S[0]=1.0, respectively.

[0157] The selection function involves comparing the remainder estimate to a set of eight partial root dependent selection constants, one constant per digit value. Thus,

[0158]

number

[0159] The selection constants depend on the partial root. The seven most significant bits of the partial root are used to extract a set of eight 11-bit selection constants. Different partial root values ​​can select different sets. The partial roots are in the interval [0.5,1]; note that the value S[i]=1 is possible until a non-zero digit is generated. Thus, taking into account that the partial root has 1 integer bit (which is 0 after the first non-zero and negative digit is generated) and 6 fractional bits, and that the minimum value of the partial root is 0.5, the selection constants can be stored in a 33×88-bit look-up table (LUT) with 32 entries for S[i]∈[0.5,1] and 1 entry for S[i]=1 (although an offset LUT can be used to reduce the storage size of the square root comparison constants, as described below in some techniques).

[0160] A simple implementation of radix-64 square root with two radix-8 iterations Each radix-8 iteration produces 3 bits of result. Two radix-8 iterations can then be stacked to obtain 6 result bits per cycle, corresponding to the square root of radix-64. A simplified implementation is shown in Figure 2. Two identical radix-8 partial iterations are connected to obtain a radix-64 iteration. Note that only the most significant bit of the remainder is used to select the quotient digit. An 11-bit remainder estimate

[0161]

number

[0162] Thus, in each sub-iteration, ● The carry propagate adder 30 receives the remainder value rem[i] 31 generated in the previous partial iteration, represented in a redundant representation. The carry-save adder 30 generates a non-redundant remainder estimate of a portion of the most significant bits of the remainder value 31 by performing a carry-propagate addition of the most significant bits of the two words of the remainder value 31 (e.g., if a representation having positive and negative words as described above is used, the negative word is subtracted from the positive word). A digit selection comparator 32 compares the remainder estimate with each of a set of comparison constants 34 to determine the next root digit 33 . ● The remainder adjust value generation circuit 36 ​​generates a remainder adjust value 39 that corresponds to the "d-vector" or d[i+1] term shown in Equation 17 above. Thus, for a square root operation, the remainder adjust value depends on the partial root value 37 received from the previous partial iteration and the next root digit 33 selected by the digit selection comparator 32. Note that the term "d-vector" is used as a label for the term d[i+1] simply because the number of bits in the value matches the number of bits used for vector operands in some implementations, but this term does not imply that a "d-vector" is a Single Instruction Multiple Data (SIMD) vector operand containing multiple independent data elements, and that a "d-vector" is a single data value rather than a vector of multiple independent data values. ● The remainder update circuit 38 (including a 3:2 carry-save adder) updates the previous remainder 31 received from the previous sub-iteration based on the remainder adjust value 39 by adding the positive and negative words of the previous remainder 31 and the remainder adjust value 39 to generate an updated remainder 40 (still in the redundant representation) that is fed to the next sub-iteration to become the previous remainder 31 for that sub-iteration. On the path between outputting the updated remainder 40 in one sub-iteration and inputting the previous remainder 31 to the carry-save adder in the remainder update circuit 38 of the next sub-iteration, a left shift of 3 bits is applied to represent the 8×rem[i] term in equation 18 above. ● An on-the-fly conversion circuit 42 inserts a value determined based on the selected root digit 33 into the partial root value 37 to generate an updated partial root value 43 that is output to become the partial root value 37 in the subsequent partial iteration. The on-the-fly conversion may be performed in accordance with equations 6-8 above. Thus, although not shown in Figure 2 for simplicity, the partial root values ​​may be represented as two separate forms P and PM, as previously described, to simplify the on-the-fly conversion which may then be performed as a concatenation.

[0163] The updated remainder 40 and updated partial root value 43 from one sub-iteration become the previous remainder 31 and partial root value 37 of the next sub-iteration. Similarly, the updated remainder 40 and updated partial root value 43 from the last sub-iteration in one iteration become the previous remainder 31 and partial root value 37 of the first sub-iteration in the next iteration.

[0164] However, this simple implementation is too slow. To speed up the cycle, several techniques are used, which are described in the next section.

[0165] Base 64 Square Root Iteration 3 shows a square root processing circuit for performing digit iteration cycles, corresponding to a single radix-64 square root iteration. In this example, the square root processing circuit is an iterative unit where the output of one iteration is fed back as an input to the same unit in a subsequent iteration, and a flip-flop 50 latches the value passed each cycle. However, as further described below with respect to FIG. 9, the square root processing circuit may also be used in a pipelined implementation.

[0166] The square root processing circuit includes several parts: (1) remainder update circuit 34, (2) digit selection circuit (root digit calculation) 32, and (3) remainder estimation circuit 30. The connections between these components are also shown. Each of these parts will be described in detail below. The square root processing circuit also includes an on-the-fly conversion circuit 42, which will be described in more detail below. The on-the-fly partial root conversion maintains two partial root forms S[i] and SM[i], where SM[i] is the partial root S[i] minus one, SM[i] = S[i] - 8 -i (20) These two forms are used in some parts of radix-64 iteration. In addition, S3[i] = 3 × S[i] S3M[i]=S3[i]-8 -i They are also required for on-the-fly partial root conversion, as described in more detail below with respect to Figures 13-16. The use of S3[i] and S3M[i] simplifies the process of multiplication of ±3 root digits.

[0167] As shown in Figure 3, when a radix-64 iteration is split into two radix-8 sub-iterations, there are two instances each of the remainder estimation circuit 30, digit selection circuit 32, and remainder update circuit 34 corresponding to each radix-8 sub-iteration, although there may be some overlap between the circuits used for each sub-iteration, as explained further below. There may also be two instances of on-the-fly conversion circuit 42 for performing on-the-fly conversion using the radix-8 root digits obtained in each radix-8 sub-iteration, although for simplicity this is shown as a single block in Figure 3.

[0168] Remainder Update 4 shows in more detail the remainder update circuit 30 for performing a remainder update in a single radix-8 sub-iteration (which may be either the first or second radix-8 sub-iteration within a radix-64 iteration). The remainder update for each iteration of the cycle (see Equation 16) is done speculatively, i.e., updated remainder values ​​rem[i+1] for all possible values ​​of the root digit are calculated, and the remainder value rem[i+1] for the root digit s is calculated. i+1 Once the next root digit s is known, the correct remainder is selected. i+1 The system has several replica circuit units 60 each generating a candidate output value of the updated remainder corresponding to a different option of s i+1 There is no replica circuit unit 60 provided for =0, because the above equation 18 means that the updated remainder rem[i+1] can be obtained directly from the previous remainder value rem[i] without addition. The sign of the previous remainder estimate is used to reduce the number of speculative remainders. If the remainder estimate is positive, the root digits can only be {+4, +3, +2, +1, 0}. On the other hand, if the remainder estimate is negative, the root digits can only be {-4, -3, -2, -1, 0}.

[0169] Thus, each replica circuit unit 60 has a carry-save adder 38 and a selection multiplexer 62 for selecting between the alternative values ​​calculated in logic block 64 for positive and negative root digits of equal magnitude depending on the sign of the previous remainder estimate received from the previous partial iteration or iteration. This reduces the number of replicated units required (four replica circuit units 60 are sufficient to accommodate digits ±1, ±2, ±3, ±4 respectively, instead of needing eight to handle each positive / negative digit separately).

[0170] The replica circuit unit 60 constructs vectors d[i+1] (sometimes called F[i+1]) for all root digit values ​​other than 0, both positive and negative values.

[0171]

number

[0172] Thus, in Figure 4, the digit bits that are concatenated in the on-the-fly calculation of each possible d[i+1] vector are shown. The mask mask[i] signals the position to which the root digit must be concatenated (the mask is shifted by 3 bits between sub-iterations such that each successive base-8 root digit is concatenated at a position 3 bits lower than the position at which the previous base-8 root digit was inserted).

[0173] The blocks 64 labeled fda_pos and fda_neg for x=1, 2, 3, 4 are respectively i+1 = a | has a value corresponding to a positive or negative digit * S[i] or 2* Perform the concatenation of SM[i] to represent the d-vector d[i+1] according to Equation 21, and also -a×d[i+1] (the term -s in Equation 18 above) i+1 ×F[i+1]) to generate d-vectors fd1, fd2, fd3, and fd4.

[0174] In the recurrence formula, d[i+1] is s i+1 To prevent 3X multiplication, s i+1 The case = ±3 is handled differently, and 3 × d[i+1] is constructed by blocks fd3_pos or fd3_neg using 3 × S[i] directly as follows: 3×d[i+1]=2×(3×S[i])+(3×s i+1 )×8 -(i+1) (twenty two)

[0175] In this case, |3×s i+1 |=9, which requires 4 bits to represent. This is not a problem because a left shift of 3×S[i] by 1 bit leaves room for the additional bit. Then

[0176]

number

[0177] The remainder guess code is used to select the positive or negative d[i+1] set before the 3 to 2 carry-save adder 38. In this way, the result is that only 5 speculative remainders are calculated instead of 9.

[0178] The inverse of the remainder estimation code is placed in the least significant bit of the speculative remainder carry word, so that when the remainder estimation code is 1, the least significant bit of the speculative remainder carry word is 0, and when the remainder estimation code is 0, the least significant bit of the speculative remainder carry word is 1. This means that when the digit is positive (the remainder estimation code is 0), the term s i+1 This is because it is necessary to subtract s × F[i+1]. i+1 × F[i+1]. The two's complement is the term s i+1 ×F[i+1] is obtained by bit-complementing and adding 1. For example, the two's complement of 11100010 is 00011101+1=00011110. Therefore, this term is bit-complemented in the fd1_pos, fd2_pos, fd3_pos, and fd4_pos modules of FIG. 4, and the "+1" is added by changing the least significant bit of the carry word, which is 0 by definition, to 1. In this way, no additional adder is required to finish the two's complement calculation. If the digit is negative (the remainder estimation code is 1), the least significant bit of the carry word is kept 0 since there is no need to perform a two's complement since the operation of equation (18) is an addition. Therefore, in summary, the reciprocal of the remainder estimation code is placed in the least significant bit of the carry word.

[0179] Among these speculative remainders provided by the replica circuit unit 60, the digit s i+1 There is no equivalent to the blocks fda_pos and fda_neg = 0, and the next root digit s i+1 is determined by digit selection circuit 68, it is just an additional input in multiplexer 32 which acts as a selection circuit to select the correct candidate output value.

[0180] Each carry-save adder 38 receives two terms, the positive and negative words of the previous remainder rem[i], which are redundantly represented, and the -s of equation (18) represented by fd1-fd4. i+1The output of each carry-save adder 38 is a candidate value for selection as the updated remainder rem[i+1], which is still in the redundant representation and thus contains two terms, positive and negative. There are cases where the candidate value is simply 8, as in the case of root digit=0. * rem[i], no addition is required since there is no carry-save adder 38. The 5:1 multiplexer 68, which functions as a selection circuit, selects the root digit s selected by the root digit selection circuit 32. i+1 to provide an updated remainder rem[i+1].

[0181] Residual estimate 5 shows the remainder estimation circuit 30 for the first and second partial iterations. The remainder estimation is an early speculative calculation of the 11 most significant bits of the remainder for use in the root digit selection. This allows for better timing since the remainder estimation is removed from the critical path through the root digit calculation.

[0182] Two different situations are shown. 1. The remainder estimate in the first sub-iteration to generate a remainder estimate to be used for digit selection in the second sub-iteration in the cycle. This is done during the first iteration based on the speculative remainder obtained by the remainder update circuit 34 of the first sub-iteration as shown in FIG. 4. Thus, five carry propagate adders 70 calculate the most significant bit of the sum and the speculative remainder (rem) obtained by the remainder update circuit 34 of the first sub-iteration. d4 [i+1] to rem d1 Add the carry words of rem[i+1] and rem[i]. i+1 If , is known, then the appropriate remainder estimate for root digit selection in the second partial iteration of the cycle is selected by multiplexer 72. Thus, this is another example of a replica circuit including replica circuit unit 70 and selection circuit 72. 2. The remainder estimate in the second partial iteration to generate a remainder estimate to be used for digit selection in the first partial iteration of the next cycle (the value output by remainder estimation circuit 30 in the second iteration can be flopped in flip-flop 50 ready to be used in the next cycle as shown in FIG. 3). The remainder estimate generated by remainder estimation circuit 30 in the second partial iteration is an assimilation of the most significant bits of 8×rem[i+2], which can be derived from rem[i] input as the previous remainder value in the first partial iteration as follows (based on replacing rem[i+1] with another instance of Equation 18 that relates rem[i+1] to rem[i] in the relationship from rem[i+2] to rem[i+1] using Equation 18):

[0183]

number

[0184] This is calculated during the first and second iterations in the cycle as follows: msb - first = 64 × (8 × rem[i] - s i+1 ×d[i+1]) (25) and msb - rem[i+2]=msb - first-8×s i+2 ×d[i+2] (26) where equation (25) is evaluated during the first sub-iteration and equation (26) is evaluated during the second sub-iteration. Both equations are evaluated speculatively for the five remainder candidates.

[0185] Note that the difference between equation (18) and equation (25) is a 64X factor and a left shift of 6 bits. Then, if a 17-bit adder is used instead of two 12-bit adders, both equations can be evaluated with the same logic, with the 11 most significant bits being the remainder estimate calculated in the first partial iteration for use in digit selection in the second partial iteration of the cycle, and the 13 least significant bits being used to complete the remainder estimate calculation during the second partial iteration to obtain the remainder estimate value used in digit selection in the first partial iteration of the next cycle of equation (26).

[0186] Thus, in this approach, adder 70 in the first sub-iteration calculates some additional (least significant) bits that are not actually needed in the remainder estimate used for digit selection in the second sub-iteration, but by calculating these additional bits, it allows the term msb_first shown above to be calculated in the first sub-iteration, reducing the overall circuit area compared to if a separate adder had calculated these bits in the second sub-iteration.

[0187] Adder 74 in the remainder estimation circuit for the second sub-iteration evaluates Equation 26, which depends on msb_first and the d-vectors 0,fd1[i+2] through fd4[i+2], which is i+2 =0,s i+2 = ±1~s i+2 = ±4 for the term 8×s in the equation i+2 ×d[i+2], respectively. These vectors are generated as part of the remainder update circuit 34 in the second sub-iteration of the cycle (see fd1-fd4 in FIG. 4). This approach means that there is no need to wait for the carry-save adder 38 in the remainder update circuit 30 of the second sub-iteration to perform their additions before beginning the additions by the carry-propagate adder 74 in the remainder estimation circuit 34 for the second sub-iteration. Instead, the calculation of the updated remainder estimate in the second sub-iteration can be performed in parallel with the remainder update in the second sub-iteration to remove latency from the critical timing path. This improves performance.

[0188] Root Digit Selection 6 shows the root digit calculation performed by digit selection circuit 32 (which may be either the first or second base 8 sub-iteration within a base 64 iteration). The calculation of the root digit has been outlined above, the remainder estimate is compared to each of the eight comparison constants and a digit is selected according to equation (19). The root digit is stored as a 1-hot 9-bit vector s[i], i=8,...,0, where s[i]=1 for digit=i-4, e.g., if the root digit is -1 then s[3]=1 and the 9-bit vector is s={0,0,0,0,0,1,0,0,0}.

[0189] This is shown in Figure 6. There is a set of 11-bit comparators 80 to compare the remainder estimate with each comparison constant. The carry output, ge-output, of each comparator is set to 1 if the remainder estimate is greater than the comparison constant. The ge-output and the sign of the remainder estimate are then input to a set of nand and or gates to generate each bit of a 1hot 9-bit vector.

[0190] The selection constants required for root selection are derived from values ​​stored in a look-up table (LUT). The selection constants for each radix-8 iteration depend on the partial root values ​​preceding that partial iteration, such that each partial iteration uses a different set of comparison constants. However, it has been derived that the same set of selection constants can be used for all partial iterations except the first two partial iterations. As will be further explained below with respect to the pipelined example of FIG. 9, the selection of the first few root digits can be done in a pre-processing stage, which allows the same selection constants to be used for each iteration, thus avoiding a main iteration cycle that requires a separate LUT lookup.

[0191] Integrate A block diagram of the digit recursive square root processing cycle is shown in Figure 7. The different parts (Remainder Update Circuit 34, Remainder Estimation Circuit 30, Root Digit Selection Circuit 32, and On-the-Fly Root Conversion 42) are identified by dotted lines. The relationships between these parts are also shown.

[0192] As shown in more detail above, some parts of the cycle logic use speculation and duplication to meet timing constraints. Thus duplication is used in several places to obtain speculative results for each digit value. In most cases duplication is reduced by using the sign of the remainder to have the same logic for positive digit values ​​and their negative counterparts. In this way, the logic is duplicated 5 times instead of 9, resulting in significant area savings. Once the root digit is known, the correct value is selected from among the 9 or 5 speculative values.

[0193] In some parts, like the remainder updates in the first and second sub-iterations and the remainder estimate in the second sub-iteration, the logic is only replicated four times, but the selection is done with a 5-to-1 mux, because one of the inputs to the mux is one of the inputs to the replicated logic (so no replicated circuit units are needed to compute new values ​​for the speculative candidate values).

[0194] Thus, Figure 7 illustrates one example of a square root processing circuit that may be used in division / square root unit 24 of Figure 1. In some examples, division / square root unit 24 may also include a separate instance of a division processing circuit that performs division operations in response to a division instruction, without sharing circuitry and data paths between the square root processing circuitry and the division processing circuitry.

[0195] However, as described further below with respect to FIG. 8, in some examples, the techniques described above for the square root processing circuit may be used in a combined division / square root processing circuit that can also perform division operations, in which case the combined division / square root processing circuit also functions as the "square root processing circuit" discussed above.

[0196] A combined radix-64 divide / square-root circuit for shared division and square-root iterations. FIG. 8 shows an example of a combined divide / square root processing circuit for performing radix 64 division / square root iterations, which may be provided as part of the divide / square root unit 24 of FIG. 1. The combined divide / square root processing circuit performs both division and square root operations, both with the same radix 64, using shared circuitry and shared data paths. The same number of radix 64 iterations are performed per cycle for both the division and square root operations (in this example, a single radix 64 iteration of digit recursion is performed per cycle for both the division and square root operations). With respect to the square root example above, in this example, the radix 64 iteration is divided into two overlapping radix 8 sub-iterations. The combined divide / square root processing circuit receives as an input a signal "div / sqrt" indicating whether the current operation is a division or square root operation. This signal may be controlled by the instruction decoder 6 based on whether the instruction being processed is a division or square root instruction.

[0197] The combined division / square root processing circuit includes all of the components described above with respect to Figures 3-7 for the square root example, and therefore performs the square root operation in the same manner as described above. Much of this circuitry can be reused for the division operation, so that the data paths for generating the updated remainders rem[i+1], rem[i+2], remainder estimates rem_est[i+1], rem_est[i+2], and partial result values ​​S[i], SM[i] for the square root operation are also used to generate the corresponding values ​​for the division operation (the notation Q[i], QM[i] is used for the partial result values ​​when the division operation is performed, but these are on the same data path as the partial root values ​​S[i], SM[i] generated for the square root operation).

[0198] Figure 8 shows the microarchitecture of a radix-64 divide / square root iteration. The two radix-8 sub-iterations that form the radix-64 iteration are separated, the first sub-iteration at the top and the second sub-iteration at the bottom. The two iterations are very similar, but there are some differences that will be addressed later.

[0199] As stated in equations (1) and (3) above, the result after iteration i is defined by the partial result P[i] (which can be the partial quotient Q[i] or the partial root S[i]) and the remainder rem[i]. Then, each iteration includes several steps.

[0200] 1. Digit selection New result digits are generated from the remainder and divisor (in division) or partial root (in square root) using the low precision estimates instead of the full precision values ​​(see equation (2)). Thus, the combined division / square root unit 24 includes a shared digit selection circuit 32 that selects, for each radix-8 partial iteration, the next radix-8 digit of the division / square root result based on a comparison of the previous remainder estimate rem_est[i], rem_est[i+1] with a set of comparison constants. The remainder estimate word length is different for division and square root.

[0201] As already described above for the square root example of FIG. 6, digit selection is performed by comparing the remainder estimate with a set of eight selection constants. This set depends on the most significant bit of the divisor or partial root. The set of comparison constants is stored in a look-up table (LUT) that is addressed by the most significant bit of the divisor or partial square root (as further explained below). Error analysis of the radix-8 division and square root algorithms shows that the number of bits of the comparison constant and the remainder estimate differs for the two operations, 11 bits for the square root and 10 bits for the division. However, if an 11-bit remainder estimate is used for both the division and the square root, both operations can be placed in the same logic. In this case, the comparison constant for the division is extended to 11 bits by placing a zero in the least significant bit position. In this way, the remainder estimation logic 30 and the digit selection circuit 32 in the first and second partial iterations are shared between the division and the square root.

[0202] Thus, comparisons for digit selection are performed using the same set of comparators 80 for both the division and square root operations. The operation of digit selection circuit 32 is the same for both the division and square root operations (as described above with respect to FIG. 6 for the square root), except that it receives a different set of comparison constants for comparison with the 11-bit remainder estimate.

[0203] 2. Remainder Update The result digits so generated are used to update the remainder and partial result (equations (1) and (3)). Thus, a shared remainder update circuit 34 is provided in each sub-iteration to adjust the previous remainder value rem[i], rem[i+1] based on the remainder adjustment value in a given radix-8 sub-iteration to generate updated remainder values ​​rem[i+1], rem[i+2] in the redundant representation.

[0204] 4, a replica circuit unit is provided to generate candidate remainder values ​​for the different possible values ​​of the selected result digit (sharing circuitry between positive / negative digits of the same magnitude as described above to reduce the amount of replicating required), and then a 5:1 multiplexer 68 selects one of the candidate values ​​depending on the next result digit selected by the digit selection circuit 32. The carry-save adder 38 and the fd calculation unit 64 are the same as in FIG.

[0205] However, as shown in equation (4), the remainder adjustment value (F[i+1] term) used in the remainder update is different for division and square root. For square root, F[i+1] is the root digit s i+1 to the shifted partial root. This means that F[i+1] is calculated for each iteration by the fd calculation unit 64. However, for the division F[i+1], it is the divisor d that does not change between iterations.

[0206] Therefore, by adding the XOR gate 90, the −p i+1 est[i], rem_est[i+1] to produce a multiplication by -1 × d term. One XOR gate XORs the divisor d with the inverse of the sign of the previous remainder estimate rem_est[i], rem_est[i+1] to provide multiplication by -1. In other words, as in division, the remainder update uses multiples of +d or -d. In the case of a positive remainder, the divisor is complemented to obtain a negative multiple of the divisor. For the replicated units that calculate the candidate remainder values ​​corresponding to the ±2 and ±4 root digits, a left shift of 1 or 2 bits is applied to the paths from the XOR gate to obtain the p required in equation (3). i+1 For the square root, to avoid the need to do a triplet (3×d multiplications are precomputed before the iteration to have fast iterations), a separate representation of 3 times the divisor 3×d is used, so that a second XOR gate similarly XORs 3×d with the inverse of the sign of the previous remainder estimate to provide an input to the replica circuit unit which is computing candidate remainders for the ±3 root digits.

[0207] The 2-to-1 multiplexers 62 shown in Figure 4 for the square root example are replaced with a set of 3-to-1 multiplexers 62 in Figure 8 to select the appropriate F[i+1] value for division or square root. Each 3:1 multiplexer 62 selects a corresponding value received from XOR gate 90 based on its divisor if the operation type signal div / sqrt indicates that a division operation is to be performed. If the operation type signal div / sqrt indicates that a square root operation is to be performed, the associated one of the d-vector values ​​generated by fd1-fd4 calculation blocks 64 is selected based on the sign of the previous remainder estimate, as described above with respect to Figure 4. Thus, 3:1 multiplexer 62 functions as a selection circuit to select as the remainder adjust value either a value derived from the divisor value d when performing a given base 8 partial iteration as part of a base 64 division operation, or a value derived from a partial root value that depends on a previously selected series of base 8 root digits when performing a given base 8 partial iteration as part of a base 64 square root operation. The sharing of carry-save adder 38 and 5:1 multiplexer 68 between both operations provides circuit area savings.

[0208] 3. Residual estimate The remainder estimates are obtained to be used for digit calculation in the next sub-iteration. Thus, in a given radix-8 sub-iteration, there is a shared remainder estimation circuit 34 which generates updated remainder estimates rem_est[i+1], rem_est[i+2] which are non-redundant estimates of some of the updated remainder values ​​rem[i+1], rem[i+2] generated in redundant representation by the remainder update circuit 30 in the given radix-8 sub-iteration. The remainder estimation circuit 30 is the same as that described above in FIG. 5 for the square root operation. Again, in the second radix-8 sub-iteration, the remainder estimation circuit 30 determines the updated remainder estimate rem_est[i+2] in parallel with the remainder update circuit 34 which generates the updated remainder value rem[i+2].

[0209] 4. On-the-fly conversion The partial result P[i] (quotient Q or root S) is converted from a signed digit redundant representation to a conventional binary non-redundant representation using an on-the-fly conversion (equations (7) and (8)). In a typical on-the-fly conversion scheme, the partial root is used in the next digit selection and remainder update for a square root operation, but the fact that the partial quotient is not for a division operation leads to different partial quotient update and partial root update methods. This difference is shown below (digit

[0210]

number

[0211] [Table 1]

[0212] For division, each time a new digit (3 bits in base 8) is generated, in a typical scheme the actual partial quotient is left shifted and the new digit is placed as the three least significant bits. In this way the actual partial quotient is always in the left significant part. The previously inserted bit is left shifted to a more significant bit position. On the other hand, for square root, the new root digit is concatenated to the actual partial root such that the most significant bit of the partial root is always in the most significant part of the stored data value, and masks mask[i], mask[i+1] are used to keep track of the position to which the next digit must be concatenated as described above for square root operations.

[0213] In order to share the on-the-fly conversion logic between the division and square root, it has been decided to perform a partial quotient update as is done for a partial root update, i.e. concatenating the new quotient digits using a mask to indicate where the digits must be concatenated. This is unconventional but means that increased sharing of datapaths and circuit logic is possible.

[0214] Thus, in the first partial iteration, the shared on-the-fly transformation circuit 42 selects a position for inserting the next digit into the partial result values ​​Q[i], QM[i], S[i], SM[i] based on the mask mask[i] for both the division and square root operations. Similarly, in the second partial iteration, the shared on-the-fly transformation circuit 42 selects a position for inserting the next digit into the partial result values ​​Q[i+1], QM[i+1], S[i+1], SM[i+1] based on the mask mask[i+1] for both the division and square root operations. The mask is shifted right by 3 bits for each partial iteration, such that each result digit is inserted 3 bits to the right of the previous one.

[0215] With respect to the square root example described above with respect to FIG. 7, the combined division / square root processing circuitry may be used either in an iterative unit, where the output labeled "i+2" produced in one iteration is fed back as the input labeled "i" for the next iteration of the square root or division operation, or in a pipeline unit as described further below with respect to FIG. 9.

[0216] Division / Square Root Pipeline The long latency of traditional division and square root implementations and the complexity of each stage of it with separate logic for division and square root preclude the use of pipelined floating-point division and square root units in commercial processors. Instead, commercial processors have iteration units where parts of the logic are used over several cycles, resulting in low bandwidth designs. In a typical scheme, the iteration logic is composed of two separate parts, the division iteration and the square root iteration, with very little, if any, logic shared between both operations. To increase the bandwidth, several iteration div / sqrt units are arranged to operate in parallel. For example, one design has two iteration floating-point div / sqrt units performing double, single and half precision operations, and two other smaller iteration units performing single and half precision operations. In this way, the double precision div / sqrt bandwidth is doubled, while the single and half precision division and square root bandwidth is quadrupled over a configuration with only div / sqrt iteration units.

[0217] In the approach shown in FIG. 9, a single pipelined div / sqrt unit 24 is provided instead. To overcome the drawbacks that prevent the use of such units, the inventors have developed a low latency division and square root implementation as well as a common stage for division and square root, in addition to some other logic shared between both operations. Low latency is achieved by implementing a radix 64 digit recursive division and square root algorithm with two radix 8 iterations per cycle. Such an algorithm produces a 6-bit result per cycle as explained above. Meanwhile, having the same algorithm for division and square root together with a thorough stage design allows for a reduced area requirement. As a result, we have been able to design a pipelined floating-point div / sqrt unit with double, single and half precision with a relatively small area. Compared to the alternative configuration described above using two double / single / half precision units and two single / half precision units, the bandwidth is improved significantly for double precision and single precision, and more slowly for half precision, but the circuit area of ​​the pipelined unit can be smaller than the total area of ​​the alternative configuration. Thus, the pipelined unit makes it possible to combine low latency and high bandwidth to obtain a high performance div / sqrt unit 24.

[0218] 9, the pipeline unit 24 includes a pre-processing circuit 100, a main body of a pipeline 102 for performing digit recursion iteration, and a post-processing circuit 104. The pre-processing and post-processing logic is mostly shared between division and square root, and the iterative part, digit iteration, is spread out into several pipelined radix-64 shared stages 110.

[0219] Preprocessing circuit 100 performs various preprocessing operations, including unpacking the operands, normalizing the operands (if necessary), and initialization (eg, retrieving comparison constants and selecting one or more initial result digits).

[0220] The body of the pipeline 102 performs the digit recursion, which is the iterative portion of the digit recursion algorithm. The body of the pipeline 102 comprises a number of division / square root pipeline stages 100, each of which contains an instance of the combined division / square root processing circuitry shown in Figure 8. Thus, each pipeline stage 110 in the body 102 performs a radix-64 digit recursive floating point division operation, q=x / d, or a radix-64 digit recursive square root operation,

[0221]

number

[0222] Post-processing circuitry 104 includes rounding logic and right shifting in case of sub-normal results (division only).

[0223] The pipeline unit processes three different floating-point precisions: double, single, and half (DP, SP, HP), respectively, resulting in different latencies of division or square root operations for operations of different precisions. Nevertheless, for a given precision, the latency is the same for both division and square root due to simple scheduling of the timing of the post-processing stages.

[0224] A more detailed description of the pipeline is provided below, focusing on the processing of the mantissas of the input operands x, d to generate a result. It will be appreciated that the exponents of the input operands x, d are also processed. This can be done according to any known technique. For example, in the case of a division, the result exponent may correspond to the difference between the true exponents of the input operands x, d, adjusted for any right shifts in post-processing stages required for subregular expression processing. In the case of a square root operation, the result exponent may correspond to half the true exponent of the input operand x, also adjusted for any normalization applied. Here, "true exponent" refers to a significant power of two represented by the exponent of the floating-point number (removing the exponent bias applied according to the floating-point precision being used).

[0225] Pre-processing (V1, V2) Pre-processing circuit 100 performs pre-processing including unpacking floating-point operands to extract sign, mantissa and exponent, determining special conditions (sub-normal, 0,...), normalizing the operands (e.g., handling sub-normal), and look-up table (LUT) addressing to obtain selection constants needed for digit selection. When dividing by two sub-normal operands, both operands are normalized in the same cycle.

[0226] Additionally, the first base 8 digit is taken. In floating-point division, the first digit can only take values ​​{+1, +2} and is the integer digit of the quotient. In the square root of floating-point numbers, the first base 8 digit can only take values ​​{-4, -3, -2, -1, 0} and its calculation is easily blended with the remainder and partial root initialization.

[0227] In the case of square root, the second digit is also obtained. As described above, the LUT stores the selection constants necessary for digit selection. However, in the case of square root, the selection constants for each octal iteration depend on the partial root value of the previous iteration such that each iteration uses a different set of comparison constants. This imposes strict timing and area constraints since the iterative logic should include the LUT and it should be read each time a new iteration starts. However, it has been derived that (by error analysis) for all iterations of octal square root except the first two, the same set of selection constants can be used (even if the same set of selection constants is used after the first two iterations, sufficient accuracy is given to the result). Therefore, at this stage, the second root digit is obtained, then the LUT is read, and the set of selection constants thus obtained is flopped for use in digit selection in the remaining iterations.

[0228] In the case of division, several other operations are performed. To save iterations in single precision, the quotient q is forced to be q ∈ [1, 2). Note that q < 1 only when x < d. This situation is detected in the pre - processing and the dividend when q is left - shifted by 1 bit such that q = 2×x / d and q ∈ [1, 2). Of course, the mantissa is the same as x / d, but the exponent needs to be decremented. Finally, 3×d = 2×d + d is calculated to be used in octal iteration, avoiding the need for 3x multiples to be calculated in each iteration and saving time.

[0229] The pre - processing stage is divided into two cycles for V1 and V2 such that operand unpacking, classification and normalization, and the first root digit (square root) are done in V1. On the other hand, in V2, the following operations are performed: second root digit calculation (square root), first quotient digit calculation (division), comparison and conditional shift of x and d of the quotient (division), 3×d calculation (division), and LUT addressing to obtain the comparison constants for the remaining iterations (division and square root).

[0230] The first division digit selection and the first two square root digit selections The following provides more information about how to select the first radix-8 division result digit and the first two radix-8 square root result digits in the pre-processing circuit 100.

[0231] context Base 64 division and square root ● Each base 64 iteration consists of two base 8 iterations. ● Division: ○ The first iteration occurs before the iterative part Reason: ■ Before the iteration portion, a constant look-up table (LUT) is addressed to obtain the comparison constants required for quotient digit selection for each radix-8 iteration. ● The LUT is addressed by the most significant bit of the divisor. ■ All iterations use the same set of comparison constants. ■ The first base 8 quotient digit can only take on the values ​​+2 or +1. This means that the first iteration is much simpler than the remaining iterations. ■ In the same cycle that the LUT is addressed, there is time to perform the first division iteration. ■ By having the first iteration of the LUT cycle, the final latency could be reduced by one cycle with some precision. ● Square root: ○ The LUT is addressed by the most significant bit of the partial route ○ The first and second iterations are performed before the iteration Reason: ■ The radix-8 square root algorithm requires different sets of comparison constants for the first iteration, the second iteration, and the remaining iterations. ■ In order to have a common square root iteration logic in the iterative portion of the square root calculation and to avoid LUT addressing in the iterative logic, it has been determined to perform the first and second iterations before the iterative portion. ■ The first iteration is performed in the first cycle V1, along with unpacking the operands and determining whether they are special. ■ The second iteration is done in the same cycle V2 as the LUT addressing to get the comparison constants for the remaining iterations. This cycle precedes the iterative part of the algorithm.

[0232] Division: First base 8 digit (in V2) ● The first base 8 division digit is selected using the same set of constants as the rest of the iteration, so the constants for this first digit selection and for digit selection in subsequent iterations are taken from the LUT. ● In this cycle ○ The LUT is addressed, ○ A constant of digit=+2 is used to perform the first iteration o A set of comparison constants are flopped for use in the remaining iterations. • Next, the first iteration uses the same set of constants as the remaining iterations, but due to the limited digit values, only the constant for digit=+2 is needed.

[0233] Square root: first base 8 digit (in V1) ● For base 8 iterations, the idea is the same, but the logic is not the same as for base 4. ○ Partial route is 1 (default value) ○ The first base 8 digit can take on the values ​​-4, -3, -2, -1, or 0 o Given a partial route, the comparison constants for these 5 numbers are known and wired into the first digit selection logic (only 4 values ​​need to be stored), so no LUT addressing is needed for this. ○ These four values ​​are (comparison cte * 64, i.e. the values ​​quoted below are 64 times the actual stored constants): Digit constant = 0:-64 Digit constant = -1:-176 Digit constant = -2:-272 Digit constant = -3:-352.

[0234] Square root: second base 8 digit (in V2) ● The range of values ​​of the partial root after the first iteration is restricted, and only five values ​​are possible (a different partial root value for each value of the first digit): ○ First digit = 0 => next partial root is 1.00_000 ○ First digit = -1 => next partial root is 0.11_000 ○ First digit = -2 => next partial root is 0.10_000 ○ First digit = -3 => next partial root is 0.01_000 ○ First digit = -4 => next partial root is 0.00_000 ● A small LUT is used to store these five sets of comparison constants. ● The size of this LUT is 5x88. ○ 5 rows ○ 8 bits per column to store eight 11-bit comparison constants ○ Addressing with the above partial route ○ The value stored in the LUT (again, the constant value shown is a comparison cte that is 64 times larger than the stored value) * (64): The partial root is 1.00_000 => 461, 326, 191, 61, -62, -192, -317, -442.

[0235] Partial root is 0.11_000 => 406, 281, 171, 61, -62, -172, -277, -377 Partial root is 0.10_000 => 351, 241, 141, 46, -47, -142, -232, -322 Partial root is 0.01_000 => 291, 206, 121, 41, -42, -122, -192, -267 The partial root is 0.00_000 => 236, 161, 96, 31, -32, -97, -152, -212 The order of the above constants is constant: digit=+4, digit=+3, digit=+2, digit=+1, digit=0, digit=-1, digit=-2, digit=-3.

[0236] This describes the initial digit selection of the pre-processing circuitry. Digit selection in subsequent stages is as described above in Figure 6, with reference to comparison constants shown in the LUTs described further below in Figures 17-20.

[0237] Digit repetition in pipeline division / square root unit For a general base r and a call to the number of bits in the result n, the number of iterations is

[0238]

number

[0239] We detail radix 64 (r=64), two operations (division and square root), and three floating-point precisions (DS, SP, HP). The number of fractional bits per precision is 52, 23, and 10, respectively. One radix 64 iteration is performed per cycle; as mentioned before, to obtain a handy implementation, the radix 64 iteration is obtained by stacking two simpler radix 8 iterations per cycle. However, the number of iterations is still that of the radix 64 algorithm.

[0240] Floating-point division: The first digit that generates the integer bits of the final quotient is selected in preprocessing. In addition, if the quotient is forced to [1;2), only the guard bits are needed for rounding, and no round bits are used. Then, n=53, 24, 11 for double, single, and half precision, respectively. This includes the fraction and the guard bits. Then, the number of iterations for the three precisions is:

[0241]

number

[0242] Floating-point square root: Since the input operands are [0:25;1), the result is [0:5;1), so the result must be left-shifted to get the final floating-point result [1;2). Like division, only one additional bit, the guard bits, needs to be rounded. Hence, the number of bits in the root algorithm needs to be 54, 25, and 12 for DP, SP, and HP respectively. This includes the integer bits, the fractional bits, and the guard bits.

[0243] Meanwhile, the first two radix-8 digits are obtained in preprocessing before the iteration. The first digit selection is skipped and integrated into the initialization of the remainder partial root, and the second digit selection is done in V2, so as to have a single LUT for all iterations of the remainder. These two iterations generate 6 bits of the final root, and then the number of cycles of the iteration part is

[0244]

number

[0245] Therefore, several multiplexers are added to the body of the pipeline 102, A 2:1 multiplexer 120 in stage D2 is added to select between the outputs of stages D1 and D2, allowing stage D2 to be skipped when an HP square root operation is performed. This reflects the difference between the 2 cycles required for the division and the 1 required for the square root, as shown in equations (28) and (29). ● A multiplexer (not shown in FIG. 9) is added within the combined division / square root processing circuit to allow the output of the first partial iteration of stage D4 to be selected and output as the iteration result when the SP square root operation is performed (skipping the second partial iteration of stage D4). This avoids the generation of the extra three bits of the second partial iteration, and the two further bits generated in the first partial iteration can also be discarded as described above. ● At stage D9 a 2:1 multiplexer 122 is added to select between the outputs of stages D8 and D9, allowing stage D9 to be skipped when a DP square root operation is performed. This reflects the difference between the 9 cycles required for a division and the 8 cycles required for a square root. ● A 3:1 multiplexer 124 in stage 9 selects between the outputs received from stages D2, D4 and D9 (with or without the square root skip mentioned above), the selection by multiplexer 124 being based on a control signal indicating the floating-point precision of the current operation, which is controlled by the instruction decoder 6 depending on the type of instruction decoded to control the division / square root operation.

[0246] Thus, the instruction decoder 6 acts as a control circuit that controls the pipeline to bypass at least one division / square root iteration pipeline stage used to perform at least one iteration of a digit recursive division operation or a square root operation when producing a higher precision result, when performing a digit recursive division operation or a square root operation to produce a lower precision result (controlling multiplexer 124 to select the output of an earlier stage when the bypass is applied).

[0247] The instruction decoder 6 can also control the division / square root pipeline to cause at least one division / square root iterative pipeline stage, which is used to perform at least one iteration when a digit recursive division operation is performed, and to completely or partially skip or discard some bits of the result output when performing the digit recursive square root operation (by controlling multiplexers 120, 122 and an internal multiplexer not shown in stage D4, allowing the second partial iteration of stage D4 to be skipped and bits to be discarded).

[0248] Post-processing (W0) As mentioned before, the post-processing is a rounding and right-shifting of the result in the subnormal case. Any known floating-point rounding technique can be used here. Note that the result can only be subnormal in division, there are no subnormal results in square root. The post-processing is done in one cycle for both division and square root.

[0249] Two operations and three precisions in the same pipeline - on-the-fly conversion As mentioned above, the number of digit repetition cycles in DP and HP square roots is one less than in division (see equations (28) and (29)). To maintain the same latency and collect the results in the same cycle in both operations, an empty cycle is added to the square root. That is, the inputs to D2 and D9 are passed to the output without further conversion. Furthermore, in the SP square root, the second radix-8 iteration in the D4 cycle is skipped. Also, the latency differs for each precision. The undrounded result of DP is available in D9, while the undrounded results of HP and SP are available in cycles D2 and D4, respectively. Then, the W0 cycle operation saves the signals coming out of D2, D4, or D9 depending on the precision.

[0250] To achieve an efficient digit repetition cycle, the two operations are They share most of the logic, including the on-the-fly conversion circuit 42 for updating the partial quotient or root. However, before the first digit cycle D1, pre-processing has already generated the six fractional bits in the case of square root, or the integer digit in the case of division. The shared quotient / root update logic needs to have the same new fractional digit concatenation positions for division and square root.

[0251] Thus, in the case of division, six zeros are added to the fractional part of the quotient Q[i], QM[i] in pre-processing stage V2. The new fractional bits qi generated in each subsequent iteration are then concatenated after these zeros (in the same positions where the corresponding bits are concatenated for the square root operation, as indicated by the mask). 1:000 000 q1q2q3 q4q5q6... In the post-processing stage W0, these zeros are removed before rounding to have an uncrounded quotient. 1:q1q2q3 q4q5q6... The addition of these zeros does not affect the final quotient precision because partial roots are not used in the digit recursive division equation, as shown in equation (4).

[0252] Thus, for a division operation, the pre-processing stage V2 provides the first division / square root iteration pipeline stage D1 with partial result values ​​in which selected bit positions are set to dummy bit values ​​(0 in this example), which correspond to bit positions into which the at least one pre-processing stage V1, V2 inserts at least one additional result digit not generated for the digit recursive division operation when performing the digit recursive square root operation. In the post-processing stage W0, these dummy bit values ​​are removed.

[0253] Timing Control, Latency and Throughput The microstructure of the pipeline unit is shown in Figure 9. The unit consists of 12 stages, which is the latency of the slower operation, double precision division, of two pre-processing cycles (V1, V2), nine digit iteration cycles (D1-D9), and one post-processing cycle (W0). For a given floating-point precision, the division and square root operations have the same latency. ● Half precision, 5 cycles: V1-V2-D1-D2-W0 ● Single precision, 7 cycles: V1-V2-D1-D2-D3-D4-W0 ● Double precision, 12 cycles: V1-V2-D1-D2-D3-D4-D5-D6-D7-D8-D9-W0 (Note that even if a cycle is skipped due to the square root at D2 or D9, the latency is still the same as the input to 3:1 multiplexer 124 that comes after the flip-flop at the input to stage D2 or D9.) Having the same latency for both operations simplifies timing control.

[0254] In addition, the latency is the same whether or not there are subnormal operands or results; normalization (if necessary) is performed in V1, and the subnormal quotient right shift is performed in W0 after rounding.

[0255] A timing control circuit 130 is provided to control the timing at which the division and square root operations can be initiated. Although the timing control circuit 130 is shown as a separate unit in Figure 9, in other examples the decoder 6 can function as the timing control circuit 130.

[0256] The divide / square root unit 24 is fully pipelined, which means that a new operation can be started every cycle of throughput 1 if all operations are of the same precision, which is the most common case. Thus, the control circuit 130 can control the divide / square root pipeline to perform the second digit recursive division or square root operation in a later divide / square root iterative pipeline stage of the divide / square root pipeline that is performing a later iteration of the first digit recursive division or square root operation in parallel with a previous divide / square root iterative pipeline stage that is performing the first digit recursive division or square root operation and a previous iteration of the second digit recursive division / square root operation.

[0257] However, when there is a mixed-precision divide or square root, a constraint appears and the two operations cannot be in the same stage at the same time. As shown in Figure 10, there are some prohibited start cycles for SP and HP operations because the latency depends on the precision. For example, SP div / sqrt can not start 5 cycles after DP because in that case both operations collide in W0.

[0258] Thus, as shown in FIG. 10, the timing control circuit 130 can control the circuit to prevent a less precise digit recursive division / square root operation performed to generate a less precise result from starting a predetermined number of cycles after a more precise digit recursive division / square root operation performed to generate a more precise result, which may correspond to the difference between the number of cycles required to reach at least one post-processing stage for the more precise digit recursive division / square root operation and the number of cycles required to reach at least one post-processing stage for the less precise digit recursive division / square root operation.

[0259] The predetermined number of cycles depends on the precision used. As shown in FIG. - 5 cycles when low accuracy is SP and high accuracy is DP. - 7 cycles when low accuracy is HP and high accuracy is DP. -2 cycles when low accuracy is HP and high accuracy is SP.

[0260] In this case, as in the case where no collision occurs in the post-processing stage W0, it is okay to start a low-precision operation after a high-precision operation if the number of cycles between operations is greater than a predetermined number.

[0261] This approach can provide significant bandwidth improvement by using a shared pipeline divide / square root operation, and the area reduction from sharing common logic provides a better balance between performance and circuit area.

[0262] Nevertheless, pipelining can also be used for either or both of the square root and divide units in an implementation with separate square root and divide units.

[0263] Also, while FIG. 9 illustrates application of the pipeline method to radix-64 digit recursive division and square root, the pipeline method can also be used for other values ​​of radix.

[0264] Also, while FIG. 9 shows a pipeline scheme that supports all of HP, DP, and SP, other examples may support only a subset of these precisions, or may support other floating-point precisions, and therefore may use a different number of pipeline stages.

[0265] On-the-fly conversion As previously mentioned, part of the digit recursion method may involve converting from a redundant representation to a normal binary representation (non-redundant representation). Because the output digits from the digit recursion method are generated one at a time, it is useful if the conversion can be performed one digit at a time to avoid the latency that would occur if all digits had to be converted at once. This conversion is performed using on-the-fly conversion circuitry 42.

[0266] In simple terms, the on-the-fly transformation for square roots involves keeping two partial root words, S[i] and SM[i], where S[0] = 1.0 and SM[0] = 0.0, and SM[i] = S[i] - r -i and using the update rules given below,

[0267]

number

[0268] In the formula, (X, Y) means the concatenation of X and Y, i.e., XY. Note that in practice, SM[i] (binary number) is equivalent to S[i] (binary number) with 1 subtracted from the least significant bit position. Thus, if S[0]=111, then SM[0]=110.

[0269] Figure 11 summarizes how S[i] and SM[i] are updated for each digit in radix-8 arithmetic. The diagram {Sx[i], aaa} means the concatenation of aaa bits to the actual value of S[i] or SM[i]. Note that no arithmetic operations are performed, only concatenation.

[0270] Figure 12 shows an example of on-the-fly conversion of a root of radix 8. The sequence of digits is -1, 1, -2, -4, 2, 0, -1, where the final value of SM[i] is S[i]-1.

[0271] As shown above, in the case of square root calculation, the calculation of the next remainder rem[i+1] is i+1 ×S[i] multiplication (see equation (3)). In a radix-8 implementation, s i+1 ={+4,+3,+2,+1,0,-1,-2,-3,-4}, and therefore 2X, 3X and 4X multiples of S[i] are required. While the 2X and 4X terms are easily obtained by left-shifting S[i] by 1 or 2 bits, computing 3×S[i] is much more complicated, which has been a limiting factor in the practical use of radix-8 square root algorithms.

[0272] Note that in other implementations with smaller bases, the term 3X is not necessary due to the digit sets {+1, 0, -1} for base 2 and {+2, +1, 0, -1, -2} for base 4.

[0273] The present invention maintains additional partial root words representing S3[i] and S3M[i], thereby preventing the computation from being done as 3 x S[i] by multiplying by 3, or by multiplying S by 2 and adding S. For each of S3 and S3M, the concatenation performed is as follows: 3×s i+1 ∈{+12,+9,+6,+3,0,-3,-6,-9,-12}

[0274] Figure 13 shows how the concatenation is performed. Note that 3×s i+1 To represent 3×s = {+12,+9,-9,-12}, 4 bits are needed. This means that the concatenation of these digit values ​​produces a carry that is propagated to the previous digit. Thus, 3×s = {+12,+9,-9,-12} .... i+1 is decomposed into 3-bit digits (3×s[i+1]) mod 8, with values ​​{+6,+4,+3,+1,0,-1,-3,-4,-6} and a positive or negative carry c i+1 Take ={+1,-1}.

[0275] From Figure 13, s i+1 = {+4, +3, +2, +1, 0, -1, -2, -3, -4}, then the 3-bit digits concatenated to get 3 × S[i] are (3 × s{i+1]) mod 8 = {+4, +1, +6, +3, 0, -3, -6, -1, -4}, respectively. Therefore, the concatenation process to get S3[i] and S3M[i] is as follows:

[0276] 1. |s i+1 If |={4,3}, increment / decrement the actual partial root. The actual 3X multiple S3[i] of the partial root and its decremented counterpart S3M[i] are incremented / decremented by the previous digit s depending on the carry. i s i +1 or s iIt is rebuilt by changing it to -1, S3_inc[i]=S3[i]+8 -i S3M_dec[i]=S3M[i]-8 -i Although 3 bits are used to represent each digit to be concatenated, the full range of values ​​that can be represented by these 3 bits is not used, only the maximum value of +6 is added as a digit, so the carry is added to the previous digit s i Note that the metric does not need to be propagated beyond

[0277] 2. 3-bit digit concatenation. 3-bit digit concatenation is

[0278]

number

[0279] FIG. 14 shows an example of on-the-fly conversion of 3X root multiples. The sequence of digits is -1, +1, -2, -4, +2, 0, -1. The final S3[i] result in the table is 3X times the final S[i] result in FIG. 12. In sub-iteration i=0, the initial value of S3 is 11 (3 multiplied by the initial value of S[0]=1), and the initial value of S3M is 10 (3-1=2). In sub-iteration i=1, the digit -1 is added. 3 multiplied by -1 is -3, which is equal to the concatenation of the digit -3 of S3 and the digit -2 of S3M. Referring to equations (32) and (33), it can be seen that the value of S3[i+1] is the concatenation of S3M[i] and 101 (i.e., 5), and the value of S3M[i+1] is the concatenation of S3M[i] and 100 (i.e., 4).

[0280] In the sub-iteration i=2, the digit 1 is added. 3 multiplied by 1 is 3. Again, referring to equations (32) and (33), s i+1We can see that S3[i+1] for i=1 is generated by concatenating S3[i] with 011 (i.e. 3) and S3M[i+1] is generated by concatenating S3[i] with 010 (i.e. 2) which gives us S3[2]=10.101011 and S3M[2]=10.101010. In sub-iteration i=3, a -2 digit is added. Multiplying 3 by -2 is -6. In the case of S3, the concatenation is done on the previous value of S3M. Since we are operating in base 8, using S3M[i] to create S3[i+1] means that the value of S3[i+1] is 8 lower than it should be. Since we are trying to subtract 6, this means we have to add +2 here (8-6=+2). Hence, the concatenation is S3M and 2 (010) as shown in Figure 14. Similarly, in the case of S3M, the concatenation is done on the previous value of S3M. Thus, as shown in FIG. 14, the concatenation is S3 and 1 (001 in binary). At sub-iteration i=4, the digit to be concatenated is -4. 3 multiplied by -4 is -12. This is a more complicated situation since -12 cannot be represented with only 3 digits, hence the negative carry. After the negative carry, the remaining subtraction to be performed is -4 (-12=-8-4). Thus, we use the value of S3M_dec, which essentially subtracts 16 (8 is the decremented value, and 8 is derived from S3M). The resulting addition to be performed is 4 (16-12=4), so the concatenation performed is on S3M_dec and the value of 100 (which is 4 in binary), giving us 010 000 100. For the value of S3M, the same value is used, but the concatenation is to a value one less (i.e., 4-1=3), so the concatenation is performed between S3M_dec and 011 (which is 3 in binary). The process of the digits 2, 0, and -1 used in iterations 5, 6, and 7 should be clear from the above explanation.

[0281] 15 shows an implementation of a 3X partial root multiple on-the-fly conversion forming part of the on-the-fly conversion circuit 42. Circuitry for generating the partial root values ​​S[i] and SM[i] is not shown, as this can be achieved by simple adaptation (using the tables shown in the figure) of the circuitry shown in, for example, US Patent Application Publication No. 2020-0293281. In each partial iteration (except the first sub-iteration), the values ​​of S3[i], S3M[i], AUX[i], and AUXM[i] from the previous partial iteration are received by the receiver circuit 202. There are three parts to the implementation. Incrementing / decrementing the actual 3X partial roots S3[i], S3M[i] using the adjustment circuit 204; ● Calculate the next 3X partial roots S3[i+1], S3M[i+1], and ● Calculation of new auxiliary 3X partial roots AUX[i+1], AUXM[i+1].

[0282] The auxiliary 3X partial roots are defined as follows:

[0283]

number

[0284] That is, there is a carry propagation to the actual 3× partial root. According to equations (32) and (33), 3×s i+1 The concatenation of produces: S3[i+1]=001 111 010 111 S3M[i+1]=001 111 010 110

[0285] Next, 3×s i+2 The concatenation of produces:

[0286]

number

[0287] That is, the set of preceding digits is incremented because a carry is done by the digit +3. However, if these digits are already saturated (in this case, the target digit of S3 is 111), then a further carry is done to the next set of bits. That is, S3[i+2] is incremented by (3×s i+2 ) mod 8. Note, however, that incrementing S3[i+1] not only increments the final concatenated digit value 111 → 000, but also S3M[i]_dec must increment from 001 111 010 to 001 111 011, or equivalently, S3M[i] still needs to generate S3[i+2]. Note, however, that in this example, there is no need to carry back any further. This is because 111 is the digit in S[i] (the i+1 =-3) to obtain S[i+1], and the next digit s i+2 The conversion of s i+2=+4,+3). This carry propagates one digit. In theory, if there are several blocks of "111" in a row and a partial root has to be incremented, the carry propagates more than two digits. For example, if S3[i]=0001 011 111 111 and the next digit is +3. In such a case, the carry propagates to the 3 digits before. However, such a pattern cannot be generated by the concatenation process described here.

[0288] Thus, S3_inc[i] and S3M_inc[i] are saved for the calculation of S3[i+2] and S3M[i+2] if the carry propagated to the previous digit is carry=+1, and S3_dec[i] and S3M_dec[i] are saved if carry=-1. This situation occurs in the concatenation of two consecutive root digits and when there is a carry of +1 or -1 for certain values ​​in the 3X partial root.

[0289] Returning to FIG. 15, the adjustment circuit 204 receives the S3 inc[i] , S3 dec[i]、 S3M inc[i] , and S3M dec[i] Whether AUX[i] or AUXM[i] is selected depends on the previous digit s as shown in Figure 16. i Therefore, the decoding circuit 206 depends on the previous digit s i , and provides a signal to multiplexers 208a, 208b, 208c, and 208d to select between AUX[i] and AUXM[i]. Next, i The values ​​of are concatenated with the outputs from the Digit x3 circuit to provide the correction values ​​S3_inc[i] and S3M_dec[i]. The Digit x3 circuit produces four output values ​​as follows: s i If >=0: ● 3s i mod8+1 ● 3s i mod8 ● 3s i mod8-1 ● 3s i mod8-2 And s i If <0: ● 8-(|3s i |mod8)+1 ● 8-(|3s i |mod8) ● 8-(|3s i |mod8)-1 ● 8-(|3s i |mod8)-2

[0290] For example, s i = +1, the output is 4, 3, 2, and 1, and s i =-2, the outputs are 3, 2, 1, and 0.

[0291] Next, the new 3X partial roots S3[i+1] and S3M[i+1] are transformed into the new signed digit s i +1~S3[i], S3M[i] or S3_inc[i] or S3_dec[i] are generated by concatenating the corresponding bits. This is achieved using a concatenation circuit 210. Note that the sign of the remainder is used to reduce the number of 2:1 multiplexers whose outputs are fed to the concatenation circuit 210, similar to that described with reference to FIG. 4. That is, the sign of the remainder is used to select between positive and negative digits, e.g., for S[i] in one multiplexer a selection is made between digits +3 and -3, and for SM[i] in another multiplexer a selection is made between digits +3 and -3. A positive remainder selects a positive or 0 root digit, and a negative remainder selects a negative or 0 root digit. The digits to be concatenated to each digit are given by equations (32) and (33). For example, for the digit +3, concatenate 001, which is (3×3) mod 8. On the other hand, for -1, we concatenate 111, which is 8-|3×-3|=-1 (or 111 in binary).

[0292] After performing the concatenation circuitry, an output circuit 212 in the form of a set of multiplexers outputs the selected values ​​of S3[i+1] and S3M[i+1] along with updated AUX root values ​​AUX[i+1] and AUXM[i+1], which are generated by an AUX generation circuit 214, which decodes the latest new digit si+1 to determine if there is a carry and then uses that information to select the appropriate value to output as AUX[i+1] and AUXM[i+1] as shown in Figure 16. Each of AUX[i+1], AUXM[i+1], S3[i+1], and S3M[i+1] are received back by the receive circuit 202 in further iterations or sub-iterations.

[0293] LUT for selection constants At each stage of the digit recursion operation, a digit selection operation SEL (see Equation 2) is performed. The digit selection function in a radix-8 division or square root digit recursion algorithm performs a comparison of the actual remainder (or a portion of it) with a set of eight selected constants or coefficients. The coefficient set is selected using the most significant portion of the divisor or partial square root. The eight coefficients in the selected set are compared with the most significant portion of the remainder and the results of the eight comparisons are used to determine the next quotient or root digit.

[0294] These coefficient sets are stored in a look-up table (LUT) that is addressed by the most significant bit of the divisor in a division operation or the most significant part of the partial root in a square root operation. The LUT size for radix-8 division is 32×72 bits, and the size for radix-8 square root is 33×80 bits. In a unit that supports division and square root, two different LUTs are needed, one for the division and one for the square root. Thus, the total LUT size for such a unit is 32×72+33×80=4944 bits.

[0295] In these examples, several methods are proposed to reduce the size of the total LUT. Some column merging can be done. Furthermore, the square root coefficient can be calculated by adding a small offset to the division coefficient. As a result, the square root LUT can be replaced by a smaller table and some logic. Furthermore, several optimizations are performed to further reduce the division LUT size. Thus, the total LUT size can be reduced to 33×42+33×18=1980 bits, which represents about a 60% reduction in the required storage space.

[0296] The selection function is the remainder estimate (the most significant bit of the remainder) and the digit p i+1 This involves a comparison with a set of eight selection constants or coefficients, one constant for each possible value of

[0297]

number

[0298] In division digit recursion, the set of selected constants used to get the next digit depends on the divisor, whereas in square root it depends on the partial result. The six most significant bits of the divisor or the seven most significant bits of the partial root are used to extract a set of eight selected constants for every iteration of the current division. Different divisor or partial root values ​​extract different sets of constants.

[0299] For division, the selection constant is 10 bits wide, but the most significant bit is 0. Note, however, that the most significant bit of the divisor is always 1, since the operands are normalized before selecting the constant. Thus, the selection constant is stored in a 32 x 72 bit division lookup table (LUT).

[0300] For the square root, the selection constant is 11 bits wide. The partial square root is [0.5,1]. Therefore, considering that the partial root estimate has 1 integer bit and 6 fractional bits, and the minimum value of the partial root is 0.5, the selection constant is stored in a 33×80 bit square root LUT with 32 entries for R[i]∈[0.5,1) and 1 entry for R[i]=1.

[0301] Thus, a unit capable of division and square root (fdivsqrt unit) typically uses two LUTs: a 32x72 bit division LUT and a 33x80 bit square root LUT, with a total LUT size of 32x72+33x80=4944 bits.

[0302] This technique proposes a method to reduce the total LUT size of the fdivsqrt unit. The LUT reduction is based on two items:

[0303] 1. We have seen that the square root constant sqrt_ct can be derived from the division constant div_ct by adding a 4-bit offset to the base constant base_ct=[2×div_ct / 16]×16, where base_ct is div_ct with the 4 least significant bits set to 0. The 4-bit offset can be negative or positive. This way, instead of storing the square root constant, we just store the offset in the offset LUT.

[0304] 2. Some symmetries in the division and offset LUTs allow further reduction in the total LUT size.

[0305] 17 and 18 show the raw division and square root LUTs. The figures show constants set for each value of the divisor and partial root estimate. Each set contains the digit p i ={+4,+3,+2,+1,0,-1,-2,-3} for a total of 8 constants in the set, for division: div_ct={md(4),md(3),md(2),md(1),md(0),md(-1),md(-2),md(-3)}, for square root: sqrt_ct={ms(4),ms(3),ms(2),ms(1),ms(0),ms(-1),ms(-2),ms(-3)}.

[0306] The value of each comparison constant can be chosen from a narrow interval. In these examples, the values ​​have been carefully chosen to make each LUT symmetric, which means that the absolute values ​​of the constants in columns with digits +4 and -3, +3 and -2, +2 and -1, and +1 and 0 are the same (with some exceptions). As will be shown later, this choice helps to reduce the LUT size.

[0307] The first two divisor interval constants md(4) and md(-3) are out of range, i.e. the first two digits cannot be 4 or -3. This could be fixed by doubling the number of divisor intervals, but such an approach would be very expensive as it would mean doubling the LUT size. Instead, the sixth fractional bit of the divisor is used to select the sub-interval and to correct the two least significant bits of md(4) and md(-3).

[0308] Regarding the size of the LUT, the maximum and minimum values ​​of the division LUT are 222 and -222, respectively. Thus, the value of the division constant is in the range [222;-222], and 9 bits are required to represent all values ​​in such range. Similarly, for the square root, the constant is [447;-446], and therefore 10 bits are required.

[0309] Offset LUT By comparing the division constants and square root comparison constants shown in FIGS. 17 and 18, the square root comparison constants can be obtained as follows:

[0310]

number

[0311] That is, multiply the division constant md(k) by 2, clear the four least significant bits to 0, and add a 4-bit offset, offset(k).

[0312]

number

[0313] Note that if the offset has the same sign as the base constant m_base(k), then the addition involves replacing the four least significant bits of m_base(k) with the four-bit offset. If the offset does not have the same sign as the base constant, then an addition is performed.

[0314] As another example,

[0315]

number

[0316]

number

[0317] However, m_base(k) and offset(k) may have different signs. For example,

[0318]

number

[0319]

number

[0320] FIG. 19 shows the offsets for the calculation of the square root constant. We highlight the cases where the signs of the offset and the division constant are different. The square root and division comparison constants have been carefully selected to make this table symmetrical with respect to the columns, which means that the absolute values ​​of the constants in columns +4 and -3, +3 and -2, +2 and -1, and +1 and 0 are the same (have opposite signs). There are two cases where this rule is violated: in rows 4 and 13, the offsets of the digits +4 and -3 do not have the same absolute value. These cases are treated separately and can be detected, for example, via the offset correction indication circuit 252.

[0321] Symmetry Using the first division LUT: 1. The absolute value of the constant can be stored instead of the signed value, which helps reduce the LUT size. 2. Number of digits p i =+1 and p i = 0, the absolute values ​​of the constants are the same (opposite signs, specifically, the digit p i = +1 is positive, and p i =0 is negative), these two columns can be replaced with just one column. 3. Number of digits pi =+2 and p i The absolute value of the constant = -1 is the same except for rows 0 and 17 (opposite sign, specifically, the digit p i = +2 is positive, and p i =-1 is negative). These two columns are stored as only one column, and the values ​​of rows 0 and 17 are corrected later, for example, in the division correction indication circuit 250 and the division constant correction circuit 248. Note that in row 0, m(2)=50, m(-1)=-48, and in row 17, m(2)=73, m(-1)=-72. To merge these two columns, the stored values ​​are 48 in row 0 and 72 in row 17, and the final m(2) value is corrected by changing the least significant bit (row 17) or the bit to the left of the least significant bit (row 0). 4. Digit p i =+2 and p i The most significant bit of the absolute value of a constant = -1 is 0. This bit does not need to be stored in the LUT. 5. Digit p i =+1 and p i The most significant two bits of the absolute value of a constant with .DELTA.=0 are 0. These bits are not stored in the LUT. 6. Digit p i =+3, p i =+2, p i =+1, p i = 0, and p i Since the constant =-1 is an even number, the least significant bit is not stored in the LUT. 7. As a result, the optimized division LUT has only 6 columns because of the column merging shown in items 2 and 3 above. The number of bits per column is also reduced.

[0322] The offset LUT is shown in Figure 19. This table can also be optimized. 1. Digit p i The offsets for m_base = {+2, +1, 0, -1} have the same sign as m_base, i.e., the offsets are positive for digits +2 and +1, and negative for digits 0 and -1 (including 0 as negative or positive where appropriate). 2. The LUT is column symmetric, and the offset absolute values ​​of digits +4 and -3, +3 and -2, +2 and -1, and +1 and 0 are the same except for the two cases mentioned above. As a result, only the absolute values ​​of the offsets are stored in the LUT, and when the offsets are used to obtain the square root comparison constant, their sign is set according to the digit value, except when the offset sign is different from the m_base sign (the value highlighted in Figure 19). 3. The signs of these exceptional values ​​are stored in a new column in the LUT. The offset LUT then has only 5 columns as a result of the column merging of entries 1 and 2, 4 columns, and an extra column for the sign.

[0323] Alternatively, it will be appreciated that a square root LUT can be provided and the constant for the division operation is derived by looking up the value in the division LUT and performing an offset. In such circumstances, many of the same techniques as described above can be applied to reduce the size of either the floating point LUT or the division offset table. For example, from FIG. 18, it is clear that the magnitude of the constants for the digits +4 and -3 are the same (the digits have opposite signs, typically a +4 digit is positive and a -3 digit is negative). Similarly, the magnitude of the constants for the digits +3 and -2 are the same (again, opposite digits, typically positive for +3 and negative for -2). Similarly, the magnitude of the constants for the digits +2 and -1 are the same (again, opposite signs, typically positive for +2 and negative for -1).

[0324] The final division and offset table with the optimizations described in the previous section is shown in Figure 20. The table is split into parts with the division LUT on the left and the square root offset LUT on the right. Note that the number of columns has been reduced by column merging. The resulting merging columns are labeled with the associated two-digit value. So, for example, the column labeled (+2, -1) represents the digit p in the raw table. i =+2 and pi =-1 means merging the columns.

[0325] On the other hand, note that the last row of the table in FIG. 20 is for the square root only (row 32 in FIG. 19).

[0326] The address (the left-most column of the table) is accessed differently for division and square root. For division, the 6 most significant bits of the divisor form the address, but the first bit is a 1. For square root, the 7 most significant bits of the partial root R[i] are used to address a table with values ​​ranging from 0.5 (0.100000 in binary) to 1.0 (1.000000 in binary). Note that the square root LUT has 33 rows, so 6 bits are used for the address.

[0327] The contents of the LUT are shown as hexadecimal values. Note that the number of bits actually required for each column is specified in the table, and although hexadecimal values ​​are given, the full range of values ​​may not be possible. For example, the digit p in this division LUT i The constant value for =+3 only requires 7 bits since it can only take the values ​​{2,3,4} where the most significant hexadecimal digit corresponds to the binary values ​​{0010,0011,0100}, and therefore there is no need to store the most significant bit. Similarly for the columns (+2,-1) and (+1,0).

[0328] The offset LUT (right part) of Figure 20 stores offset absolute values ​​in columns (+4,-3), (+3,-2), (+2,-1), and (+1,0), and the 2-bit value of column sign is the offset sign of the offsets in columns (+4,-3) and (+3,-2). Note that the offsets in columns (+2,-1) and (+1,0) are positive. The sign bit of 1 means that the signs of the offsets and their corresponding m_base are different.

[0329] As mentioned before, the last row of the table with address 100000 is only meaningful for square roots. Using the same base as row 011111, the comparison constant for this partial root estimate is obtained at the offset shown in the table.

[0330] Consider the following example for calculating comparison constants for division and square root: For division, the constant set is obtained from the LUT by adding leading zeros. For example, for a division operation with divisor=1.00110x...x, the LUT address is 01_00110 and the LUT returns:

[0331]

number

[0332] Note that the number of bits for each constant in the set depends on how many digits the constant is for. So, taking into account the rules for reducing LUT size listed above for division, the set of comparison constants for this particular divisor value is:

[0333]

number

[0334] The bits that were added to get the final constants are highlighted. Note that the absolute values ​​of the constants are calculated from the LUT. In a later step, the signs of m(0), m(-1), m(-2), and m(-3) are two's complemented to get the final set of constants.

[0335] Note that for the square root constant in this same row, the sign field is 01. This means that the sign of the offset for the calculation of ms(+3) and ms(-2) is different from the sign of the base constant, and therefore the calculation of these two constants requires a subtraction. From the table, LUT_offset(01_00110)={1,a,e,2,6} The offsets are as follows. Offsets with signs different from the basic constant signs are highlighted.

[0336]

number

[0337] The fundamental constants are

[0338]

number

[0339]

number

[0340] Since the positive and negative parts of the sqrt LUT are symmetric, the remaining constants are obtained by taking the 2's complement of the constants above. {ms(0),ms(-1),ms(-2),ms(-3)}={-38,-114,-192,-266}

[0341] 21 shows a selection constant generator 238 which is used to generate the selection constant used by, for example, the digit selection comparator 32. The divisor and partial root bits are received by a multiplexer 240. A div / sqrt select signal is provided which selects the divisor when a selection constant for a division is required, and the partial root when a selection constant for a square root is required. The selected bits are then used to access an associated value in a storage circuit 242 which consists of a division LUT and a (square root) offset LUT.

[0342] The output from the division LUT is passed to a padding circuit 246, which pads bits by adding 0 to the constant being output. The padding performed is, for example, as described in points 2-6 with respect to the division LUT above. The resulting constant is passed to a conversion circuit 244, described below, and also to a division constant correction circuit 248. The division constant correction circuit 248 receives the padded (extended) division selection constant as well as an output from a division correction indication circuit 250, which indicates whether the data being obtained from the division LUT is one of the exceptional cases (point 3 with respect to the division LUT above) where the absolute values ​​of the constants are not the same, namely (i) the constants md(4) and md(-3) when the divisor estimate is 0 or 1, and (ii) the digit p when the divisor estimate is 0 or 17. i =+2 and p i Check the constant absolute value difference for =-1. These corrections require setting bits 70, 50, 1, and 0 in the selected constant set, and clearing bits 71 and 21. The corrections are performed by division constant correction circuit 248. The output from the offset LUT is passed to the conversion circuit 244 along with an output from an offset correction indication circuit 252, which indicates whether the constant being accessed is one of the exceptions where the LUT offsets do not have the same value (e.g., lines 4 and 13). If so, a correction to the correct value is made in the conversion circuit 244. The correction circuit 244 also receives a padded (extended) division constant from a padding circuit 246. The substitution circuit 254 is used to add the offset using concatenation or subtraction as described above. In particular, if the offset sign and the constant base sign are different, a subtraction is performed. The subtraction is made possible by checking the sign field in the offset LUT. Substitution of the four least significant bits of the 4-bit offset is only made if the signs are equal.

[0343] For both the division constant and the LUT constant, the absolute value is multiplied by digit p i A signature circuit 256 is provided to convert the input data into signed values ​​of =0, -1, -2, -3.

[0344] Computer readable code for manufacturing The concepts described herein may be embodied in computer readable code for the manufacture of devices embodying the concepts described. For example, the computer readable code may be used in one or more stages of a semiconductor design and manufacturing process, including an electronic design automation (EDA) stage, to manufacture integrated circuits comprising devices embodying the concepts. Such computer readable code may additionally or alternatively enable the definition, modeling, simulation, verification and / or testing of devices embodying the concepts described herein.

[0345] For example, computer readable code for producing a device embodying the concepts described herein can be embodied in code defining a hardware description language (HDL) representation of the concept. For example, the code can define a register transfer level (RTL) abstraction of one or more logic circuits to define a device embodying the concept. The code can define an HDL representation of one or more logic circuits embodying the device in Verilog, SystemVerilog, Chisel, or VHDL (Very High Speed ​​Integrated Circuit Hardware Description Language), as well as intermediate representations such as FIRRTL. The computer readable code can provide a definition embodying the concept using a system level modeling language such as SystemC and SystemVerilog or other behavioral representation of the concept that can be interpreted by a computer to enable simulation of the concept, functional and / or formal verification, and testing of the concept.

[0346] Additionally or alternatively, the computer readable code may embody a computer readable representation of one or more netlists. The one or more netlists may be generated by applying one or more logic synthesis processes to the RTL representation. Alternatively or additionally, the one or more logic synthesis processes may generate a bitstream from the computer readable code to be loaded into a field programmable gate array (FPGA) to configure the FPGA to embody the described concepts. The FPGA may be deployed for concept proof and testing purposes prior to manufacture in an integrated circuit, or the FPGA may be deployed directly into a product.

[0347] The computer readable code may include a mixture of code representations for fabrication of a device, including, for example, a mixture of one or more of an RTL representation, a netlist representation, or another computer readable definition used in a semiconductor design and manufacturing process to fabricate a device embodying the invention. Alternatively or additionally, a concept may be defined in a combination of a computer readable definition used in a semiconductor design and manufacturing process to fabricate a device and computer readable code that defines instructions to be executed by the defined device to be fabricated.

[0348] Such computer readable code may be disposed on any known transitory computer readable medium (such as wired or wireless transmission of code over a network) or on a non-transitory computer readable medium such as a semiconductor, magnetic disk, or optical disk. Integrated circuits manufactured using the computer readable code may include components such as one or more central processing units, graphic processing units, neural processing units, digital signal processors, or other components that individually or collectively embody the concepts.

[0349] In this application, the term "configured to..." is used to mean that elements of an apparatus have a configuration that can perform a defined operation. In this context, "configuration" refers to a manner of arrangement or interconnection of hardware or software. For example, an apparatus may have dedicated hardware that provides the defined operation, or a processor or other processing device may be programmed to perform a function. "Configured to" does not mean that the apparatus elements have to be modified in some way to provide the defined operation.

[0350] Although exemplary embodiments of the present invention are described in detail herein with reference to the accompanying drawings, it will be understood that the invention is not limited to these precise embodiments, and that various changes and modifications can be made to the embodiments by those skilled in the art without departing from the scope of the present invention as defined by the appended claims.

Claims

1. An apparatus, A division / square root pipeline, Comprising a plurality of division / square root iterative pipeline stages, each for performing respective iterations of digit-by-digit division or square root operations, and A signal path for supplying, as an input to a subsequent division / square root iterative pipeline stage of the division / square root pipeline, an output generated by one division / square root iterative pipeline stage in one iteration, for performing a subsequent iteration of the digit-by-digit division or square root operation, the division / square root pipeline comprising The apparatus, wherein the division / square root pipeline is capable of performing the digit-by-digit division or square root operation on a floating-point operand to generate a floating-point result.

2. A control circuit for controlling the division / square root pipeline to execute a subsequent iteration of the first digit-by-digit division or square root operation in a division / square root iterative pipeline stage subsequent to the division / square root pipeline, in parallel with a division / square root iterative pipeline stage prior to the division / square root pipeline for executing a previous iteration of the second digit-by-digit division / square root operation, to cause the second digit-by-digit division or square root operation to be executed, the apparatus according to claim 1, comprising the control circuit.

3. The apparatus according to claim 1 or 2, wherein each division / square root iterative pipeline stage comprises a combined division / square root processing circuit for performing a given iteration of a digit-by-digit division operation in response to a division instruction and a given iteration of a digit-by-digit square root operation in response to a square root instruction.

4. The apparatus according to claim 3, wherein the combined division / square root processing circuit comprises a shared circuit for generating at least one output value on the same data path used for both the given iteration of the digit-by-digit division operation and the given iteration of the digit-by-digit square root operation.

5. The apparatus according to claim 3, wherein the division / square root pipeline is configured to perform the same number of iterations per processing cycle using the same radix for both the digit-by-digit division operation and the digit-by-digit square root operation.

6. The apparatus according to claim 1 or 2, wherein, for a given result accuracy, the division / square root pipeline is configured to process the digit-by-digit division operation in the same number of processing cycles as the digit-by-digit square root operation.

7. The apparatus according to claim 1 or 2, wherein the division / square root pipeline is configured to support at least two different result accuracies for the digit-by-digit division or square root operation.

8. The apparatus according to claim 7, wherein the division / square root pipeline is configured to execute the digit-by-digit division or square root operation in fewer processing cycles when generating a result with lower accuracy than when generating a result with higher accuracy.

9. A control circuit that controls the division / square root pipeline to bypass at least one division / square root iteration pipeline stage used to execute at least one iteration of the digit-by-digit division operation or square root operation when generating a higher accuracy result, when executing the digit-by-digit division operation or square root operation to generate a lower accuracy result. The apparatus according to claim 7, comprising the control circuit.

10. The division / square root pipeline includes at least one post-processing stage for performing a post-processing operation on the output of the final iteration of the digit-by-digit division or square root operation. The apparatus is a control circuit that prevents a lower accuracy digit-by-digit division / square root operation, which is executed to generate a lower accuracy result, from starting for a predetermined number of cycles after a higher accuracy digit-by-digit division / square root operation, which is executed to generate a higher accuracy result. The predetermined number of cycles corresponds to the difference between the number of cycles required to reach the at least one post-processing stage for the higher accuracy digit-by-digit division / square root operation and the number of cycles required to reach the at least one post-processing stage for the lower accuracy digit-by-digit division / square root operation. The apparatus according to claim 7, comprising the control circuit.

11. Each division / square root iteration pipeline stage includes a digit selection circuit that selects the next result digit for the partial result value of the digit-by-digit division or square root operation based on a comparison between the previous remainder value and a set of comparison constants. A remainder update circuit that updates the previous remainder value based on the remainder adjustment value and the next result digit selected by the digit selection circuit, and the apparatus according to claim 1 or 2.

12. The apparatus according to claim 11, wherein the plurality of division / square root iterative pipeline stages are configured to use the same set of comparison constants for each iteration executed within the same digit-by-digit division or square root operation.

13. The division / square root pipeline is configured to perform a table lookup to obtain the set of comparison constants in a preprocessing stage of the division / square root pipeline before a first division / square root iterative pipeline stage of the division / square root pipeline, and the set of comparison constants is passed from stage to stage to avoid repeating the table lookup at each division / square root iterative pipeline stage within the same digit-by-digit division or square root operation. The apparatus according to claim 11.

14. The division / square root pipeline includes at least one preprocessing stage for performing operand preprocessing before a first division / square root iterative pipeline stage of the division / square root pipeline, and the operand preprocessing includes selecting at least one initial result digit for the result of the digit-by-digit division or square root operation. The apparatus according to claim 1 or 2.

15. The division / square root pipeline is configured to support both digit-by-digit division operations and digit-by-digit square root operations, In the operand preprocessing, the at least one preprocessing stage is configured to generate a greater number of initial result digits for the digit-by-digit square root operation than the number of initial result digits for the digit-by-digit division operation. The apparatus according to claim 14.

16. Control circuitry for controlling the division / square root pipeline to completely or partially skip or discard some bits of the result output when the digit-by-digit square root operation is performed on at least one division / square root iterative pipeline stage used to perform at least one iteration when the digit-by-digit division operation is performed. The apparatus according to claim 15.

17. When performing the digit recurrence division operation, the at least one preprocessing stage is configured to provide the first division / square root iterative pipeline stage with a partial result value in which a selected bit position is set to a dummy bit value, and the selected bit position corresponds to a bit position where at least one additional result digit that is not generated for the digit recurrence division operation is inserted when the at least one preprocessing stage performs a digit recurrence square root operation. A given division / square root iterative pipeline stage of the division / square root pipeline is configured to insert the next result digit into the partial result value at the same bit position for both the digit recurrence division operation and the digit recurrence square root operation. The apparatus according to claim 15, wherein the division / square root pipeline includes a postprocessing stage for removing the dummy bit value from the final result value when performing the digit recurrence division operation.

18. The apparatus according to claim 1 or 2, wherein the digit recurrence division or square root operation is a digit recurrence division or square root operation with a radix of 64.

19. Each division / square root iterative pipeline stage is configured to perform each radix-r iteration of a radix-r digit recurrence division or square root operation by performing a plurality of partial iterations with a radix of n in the same processing cycle, where n < r. The apparatus according to claim 1 or 2.

20. The apparatus according to claim 19, wherein r = 64 and n = 8.

21. A data processing method, comprising: using a plurality of division / square root iterative pipeline stages of a division / square root pipeline to perform each iteration of a digit recurrence division or square root operation; and feeding the output generated by one division / square root iterative pipeline stage as an input to a subsequent division / square root iterative pipeline stage of the division / square root pipeline. The data processing method, wherein the division / square root pipeline is capable of performing the digit recurrence division or square root operation on a floating-point operand to generate a floating-point result.

22. A computer-readable medium for storing computer-readable code for manufacturing an apparatus, comprising: a division / square root pipeline, ​ A plurality of division / square root iterative pipeline stages, each for performing respective iterations of digit recurrence division or square root operations, A signal path for supplying, as an input to a subsequent division / square root iterative pipeline stage of a division / square root pipeline, an output generated by one division / square root iterative pipeline stage in one iteration, for performing a subsequent iteration of the digit recurrence division or square root operation, comprising a division / square root pipeline, A computer-readable medium, wherein the division / square root pipeline is capable of performing the digit recurrence division or square root operation on a floating-point operand to generate a floating-point result.