Floating-point adder circuit with sub-normal support
By introducing a dual-path adder architecture in the DSP block, the problem that floating-point adders cannot handle sub-regular numbers is solved, effective processing of sub-regular numbers and dynamic range expansion is achieved, and calculation accuracy is improved.
Patent Information
- Application Number
- CN201810923369.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2017-09-14
- Filing Date
- 2018-08-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2038-08-14
AI Technical Summary
The existing digital signal processing block (DSP) cannot effectively process secondary regular numbers on the PLD, resulting in the floating-point adder not operating normally.
Using a dual-path adder architecture, including near and far paths, floating-point adders are dynamically configured to support the processing of regular and sub-regular numbers through circuits such as leading zero predictors, exponential value comparisons and shifters.
It realizes effective processing of secondary regular numbers, expands the dynamic range of floating-point adder and improves calculation accuracy.
Smart Images

Figure CN109508173B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to integrated circuits, and more particularly, to integrated circuits having floating point arithmetic circuits. Background Art
[0002] Programmable logic devices (PLDs) include logic circuits such as look-up tables (LUTs) and sum-of-products based logic, which can be configured to allow a user to customize the circuits according to the user's specific needs. In addition to this configurable logic, PLDs also include programmable interconnect or routing circuits for connecting the inputs and outputs of LEs and LABs. The combination of this programmable logic and routing circuits is referred to as soft logic.
[0003] In addition to soft logic, PLDs also include dedicated processing blocks that implement specific predefined functions and thus do not have to be configured by the user. Such dedicated processing blocks can include a collection of circuits on the PLD that have been partially or fully hardwired to perform one or more specific tasks (e.g., logical or mathematical operations).
[0004] A particularly useful type of dedicated processing block that has been provided on PLDs is a digital signal processing (DSP) block. Conventional DSP blocks include floating point adders that only support "normal" numbers and not "subnormal" numbers. Compared to normal numbers, subnormal numbers are numbers that are specially encoded with a predetermined minimum exponent and a mantissa component with an implied leading zero. Due to this special encoding, the floating point adder has to handle subnormal numbers differently.
[0005] The embodiments described herein arise in this context. Brief Description of the Drawings
[0006] Figure 1 is a diagram of an exemplary integrated circuit having a dedicated processing block according to an embodiment.
[0007] Figure 2 is a diagram according to an embodiment showing how a dedicated processing block can include one or more floating point adder circuits.
[0008] Figure 3 is a diagram according to an embodiment showing different floating point formats.
[0009] Figure 4 is a diagram of an exemplary dual-path floating point adder architecture according to an embodiment.
[0010] Figure 5 is a circuit diagram of an exemplary near-path arithmetic circuit according to an embodiment.
[0011] Figure 6It is a circuit diagram of an exemplary near-path normalization circuit according to an embodiment.
[0012] Figure 7 It is a circuit diagram of a far-path denormalization and arithmetic circuit according to an embodiment.
[0013] Figure 8A It is a diagram of an exemplary dedicated processing block including a floating-point adder capable of processing FP16 at its input and output according to an embodiment.
[0014] Figure 8B It is a diagram of an exemplary dedicated processing block including a floating-point adder capable of processing modified FP16 (or FP16’) at its input according to an embodiment.
[0015] Figure 8C It is a diagram of an exemplary dedicated processing block including a floating-point adder capable of processing modified FP16 at its input and output according to an embodiment.
[0016] Figure 9 It is a circuit diagram of a far-path denormalization and arithmetic circuit capable of processing modified FP16 at its input according to an embodiment. Detailed Description
[0017] The described embodiments relate to integrated circuits, and more particularly, to floating-point adders / subtractors on integrated circuits. Those skilled in the art will recognize that the described exemplary embodiments may be practiced without some or all of these specific details. In other instances, well-known operations are not described in detail so as not to unnecessarily obscure the described embodiments.
[0018] Figure 1 An exemplary embodiment of an integrated circuit is shown, such as a programmable logic device (PLD) 100 having an exemplary interconnect circuit. As Figure 1 shown, the programmable logic device (PLD) may include a two-dimensional array of functional blocks, the functional blocks including logic array blocks (LABs) 110 and other functional blocks, such as random access memory (RAM) blocks 130 and dedicated processing blocks such as dedicated processing blocks (SPBs) 120. Functional blocks such as LAB 110 may include smaller programmable regions (e.g., logic elements, configurable logic blocks, or adaptive logic modules) that receive input signals and perform custom functions on the input signals to produce output signals.
[0019] The programmable logic device 100 may include programmable memory elements. Input / output elements (IOEs) 102 may be used to load configuration data (also referred to as programming data) into the memory elements. Once loaded, the memory elements each provide corresponding static control signals that control the operation of associated functional blocks (e.g., LAB 110, SPB 120, RAM 130, or input / output elements 102).
[0020] In a typical scenario, the outputs of the loaded memory elements are applied to the gates of metal-oxide semiconductor transistors in the functional blocks to turn certain transistors on or off, and thereby configure the logic in the functional blocks including routing paths. Programmable logic circuit elements that can be controlled in this manner include portions of multiplexers (e.g., multiplexers used to form routing paths in an interconnect circuit), look-up tables, logic arrays, AND, OR, NAND, and NOR logic gates, transmission gates, etc.
[0021] The memory elements may use any suitable volatile and / or non-volatile memory structures, such as random access memory (RAM) cells, fuses, antifuses, programmable read-only memory cells, mask programming and laser programming structures, mechanical storage devices (e.g., including localized mechanical resonators), mechanically operated RAM (MORAM), combinations of these structures, etc. Since the memory elements are loaded with configuration data during programming, the memory elements are sometimes referred to as configuration memories, configuration RAMs (CRAMs), configuration memory elements, or programmable memory elements.
[0022] In addition, the programmable logic device may have input / output elements (IOEs) 102 for driving signals out of the device 100 and for receiving signals from other devices. The input / output elements 102 may include parallel input / output circuits, serial data transceiver circuits, differential receiver and transmitter circuits, or other circuits for connecting one integrated circuit to another integrated circuit. As shown, the input / output elements 102 may be positioned around the periphery of the chip.
[0023] If desired, the programmable logic device may have input / output elements 102 arranged in different ways. For example, the input / output elements 102 may form one or more columns of input / output elements that may be located anywhere on the programmable logic device (e.g., evenly distributed across the width of the PLD). If desired, the input / output elements 102 may form one or more rows of input / output elements (e.g., distributed across the height of the PLD). Alternatively, the input / output elements 102 may form islands of input / output elements that may be distributed over the surface of the PLD or may be clustered in selected areas.
[0024] The PLD may also include programmable interconnect circuitry in the form of vertical routing channels 140 (i.e., interconnects formed along the vertical axis of the PLD 100) and horizontal routing channels 150 (i.e., interconnects formed along the horizontal axis of the PLD 100), each routing channel including at least one track for routing at least one wire. If desired, the interconnect circuitry may include double data rate interconnects and / or single data rate interconnects.
[0025] If desired, the routed wires may be shorter than the entire length of the routing channel. A wire of length L may span L functional blocks. For example, a wire of length four may span four blocks. A wire of length four in a horizontal routing channel may be referred to as an "H4" wire, while a wire of length four in a vertical routing channel may be referred to as a "V4" wire.
[0026] Different PLDs may have different functional blocks connected to different numbers of routing channels. In Figure 1 A three-sided routing architecture is depicted, in which input and output connections exist on three sides of each functional block leading to the routing channels. It is also intended that other routing architectures be included within the scope of the present invention. Examples of other routing architectures include 1-sided, 1 1 / 2-sided, 2-sided, and 4-sided routing architectures.
[0027] In a direct drive routing architecture, each wire is driven by a driver at a single logical point. The driver may be associated with a multiplexer that selects the signal to be driven on that wire. In the case of a channel having a fixed number of wires along its length, the drivers may be placed at each starting point of the wires.
[0028] Note that other routing topologies in addition to Figure 1 the topology of the interconnect circuitry depicted therein are intended to be included within the scope of the present invention. For example, the routing topology may include diagonal wires, horizontal wires, and vertical wires along different portions of its extent, as well as wires perpendicular to the device plane (in the case of a three-dimensional integrated circuit), and the drivers of the wires may be located at points different from one end of the wire. The routing topology may include global wires that substantially span the entire PLD 100, fractional global wires such as wires that span a portion of the PLD 100, interleaved wires having a specific length, smaller local wires, or any other suitable arrangement of interconnect resources.
[0029] Furthermore, it should be understood that the described embodiments may be implemented in any integrated circuit. If desired, the functional blocks of such an integrated circuit may be arranged in more levels or layers, where multiple functional blocks are interconnected to form larger blocks. Other device arrangements may use functional blocks that are not arranged in rows and columns.
[0030] Figure 2is a diagram showing a dedicated processing block that can include two or more floating point (FP) addition / subtraction circuits. As Figure 2 shown, the dedicated processing block (sometimes referred to as a digital signal processing block or “DSP” block) can include one or more floating point adder / subtractor circuits 200.
[0031] It is common for floating point numbers to be used to represent real numbers in scientific notation in a computing system, and they are designed to cover a wide range of numerical values and diverse precision requirements. The IEEE 754 standard is often used for floating point numbers. A floating point number includes three different parts: (1) the sign of the floating point number; (2) the mantissa; and (3) the exponent. Each of these parts can be represented by a binary number, and in the IEEE 754 format, they can have different bit sizes, depending on the precision (e.g., see Figure 3 ).
[0032] As Figure 3 shown, the single-precision floating point format (sometimes simply referred to as “FP32”) requires 32 bits, which are distributed as follows: one sign bit (bit 32), eight exponent bits (bits [31:24]), and 23 mantissa bits (bits [23:1]). The double-precision floating point format (sometimes simply referred to as “FP64”) requires 64 bits, including: one sign bit (bit 64), 11 exponent bits (bits [63:53]), and 52 mantissa bits (bits [52:1]). The half-precision floating point format (sometimes simply referred to as “FP16”) requires 16 bits, which are distributed as follows: one sign bit (bit 16), five exponent bits (bits [15:11]), and ten mantissa bits (bits [10:1]). The modified or alternative half-precision floating point format (sometimes simply referred to as modified FP16, FP16’, or FP16+++) requires 19 bits, which are distributed as follows: one sign bit (bit 19), eight exponent bits (bits [18:11]), and ten mantissa bits (bits [10:1]). The three additional exponent bits in modified FP16’ can help the floating point adder better handle subnormal numbers with an extended dynamic range and improved accuracy.
[0033] The sign of a floating point number according to the IEEE 754 standard is represented using a single bit, where “0” represents a positive number and “1” represents a negative number. Since IEEE 754 floating point numbers are defined in this sign-magnitude format, addition and subtraction operations are essentially the same.
[0034] The exponent of a floating point number is preferably an unsigned binary number and, for the FP32 single-precision format, it is in the range of 0 to 255. To represent very small numbers, negative exponents must be used. Thus, the exponent preferably has a bias that should be subtracted from the unsigned exponent value. For single-precision floating point numbers, the bias is preferably 127. For example, the exponent value 140 actually represents (140 - 127) = 13, and the value 101 represents (101 - 127) = -26. For the FP16 half-precision number format, the exponent bias is 15. For the FP16' modified half-precision format, the exponent bias is preferably 127 (i.e., the same as FP32 because the number of exponent bits for both is the same). For simplicity, the floating point adder architecture described in the text uses the FP16 format, which is merely illustrative. In general, circuit techniques and improvements can be applied to any suitable IEEE 754 floating point format.
[0035] As discussed above, according to IEEE 754, the mantissa is typically a normalized number (i.e., it does not have leading zeros and represents the precision or fractional component of the floating point number). Since the mantissa is stored in binary format, the leading bit can be either 0 or 1. For a normalized number, the leading bit will always be 1. Thus, a "normal" floating point value is encoded as follows:
[0036] x = (-1) s 2 e-偏置量 1.F if e ≥ 1 (1)
[0037] where s represents the sign bit, e represents the unsigned exponent value, and F represents the mantissa bits. Thus, the smallest normal number or "non-denormal" number in FP16 is equal to:
[0038] x reg_min = (-1) s 2 1-偏置量 1.0 = (-1) s 2 -14 *1.0 (2)
[0039] Compared to normal floating point numbers, there is another type of floating point number that is smaller than normal floating point numbers and is sometimes referred to as a "denormal" or "un-normalized" number. A denormal number is a floating point value that is encoded as follows:
[0040] x = (-1) s 2 e_min 0.F if e = 0 (3)
[0041] Among them, assuming a binary FP16 format, e_min is set to be equal to (1 - bias) = -14. Although the biased exponents stored in the exponent field are all zero (e.g., 00000), their actual value is 1. Therefore, a denormal input is considered to have an exponent of 1, and its leading bit before the decimal point is 0. In other words, when considering the bias, an exponent encoded as zero actually represents -14. The fractional part F is not all zero; otherwise, the value would represent the encoding of zero. Therefore, the largest subnormal number in FP16 is equal to:
[0042] x sub_max = (-1) s 2 e_min 0.1111111111 = (-1) s 2 -14 * (1 – 2 -10 ) (4)
[0043] And the smallest subnormal number in FP16 is equal to:
[0044] x sub_min = (-1) s 2 e_min 0.0000000001 = (-1) s 2 -24 (5)
[0045] The floating-point adder circuit 200 can be configured to be capable of processing not only normal (non-denormal) numbers but also denormal numbers. Figure 4 is a diagram showing a suitable implementation of the adder circuit 200 that can support denormal numbers. As Figure 4 shown, the adder 200 may include a setting circuit, such as an exponent / fraction comparison and near / far path routing circuit 40 that receives inputs A and B (i.e., the exponents and mantissas of numbers A and B). The circuit 40 compares the exponents and mantissas of inputs A and B and splits the numbers into a "near" path and a "far" path.
[0046] If the difference between the exponents is equal to zero or one (i.e., if the magnitude of the difference between the exponent of A and the exponent of B is less than or equal to one), then the near path can be used (where the value 1 is set as a predetermined threshold). In the near path, generally only subtraction occurs. The larger number, which can be either A or B, will be routed to the "near larger" path 42, while the smaller number will be routed to the "near smaller" path 44.
[0047] On the other hand, if the magnitude of the difference between the exponents is greater than one or for a true addition operation, the far path can be taken (e.g., the far path can also handle addition for near values). The larger number, which can be either A or B, will be routed to the "far larger" path 46, while the smaller number will be routed to the "far smaller" path 48.Figure 4 An exemplary architecture in which the median value is split into a near path and a far path can sometimes be referred to as a "dual-path" adder architecture.
[0048] To illustrate the operation of the near path, consider an example where A equals +(2 5 )*(1.xxxxxxxxxx) and B equals -(2 4 )*(1.yyyyyyyyyy). To match the exponents, B can be shifted one position to the right (i.e., B is divided by the factor 2 since the exponent is base 2). Since B is negative, the magnitude of B can be subtracted from A, resulting in 0.00001zzzzz, which has a corresponding exponent of 5. To normalize 0.00001zzzzz, the number of leading zeros can be determined, and the result can be shifted left by an appropriate number of bits. As shown in the example above, the near path can involve a subtraction operation (which can be performed using the near path arithmetic logic unit 50) and a normalization operation (which can be performed using the normalization circuit 52).
[0049] To illustrate the operation of the far path, consider an example where A equals +(2 5 )*(1.xxxxxxxxxx) and B equals (-1) s (2 2 )*(1.yyyyyyyyyy). To match the exponents, B can be shifted three positions to the right (i.e., B is divided by 8 since the exponent is base 8). Depending on whether B is positive or negative, the resulting valid value can be: 10.zzzzzzzzzz, 1.zzzzzzzzzz, or 0.lzzzzzzzzz, which has a corresponding exponent of 5. In the first case, a one-bit right shift operation is required. In the second case, no shift is needed. In the third case, a one-bit left shift operation is required. As shown in the example above, the far path can involve shifting the far smaller number to the right (which can be performed using the denormalization / align circuit 54) and addition / subtraction operations (which can be performed using the circuit 56).
[0050] Still referring to Figure 4 , the floating-point adder circuit 200 can also include a multiplexing circuit, such as a multiplexer 58, which has a first data input that receives the output signal from the normalization circuit 52 in the near path, a second data input that receives the output signal from the + / - circuit 56 in the far path, and an output that can route either the near path result or the far path result. Subsequently, the output of the multiplexer 58 can be fed through the exception / error handling circuit 60 to generate the final output signal OUT.
[0051] Figure 5is a circuit diagram of an exemplary near-path arithmetic logic unit (ALU) with subnormal support, such as arithmetic circuit 50. As Figure 5 shown, circuit 50 includes right-shift circuits 210 and 213, multiplexer circuits 211, 214, and 216, subtraction circuits 212 and 215, and a leading zero anticipator (LZA) circuit 217. The mantissa X of the near larger value is received at input 201, while the mantissa Y of the near smaller value is received at input 202.
[0052] Multiplexer 211 has a first input that receives the mantissa (Y) directly from input 202, a second input that receives a right-shifted version of the mantissa (Y) from input 202 via right-shifter 210, and an output. Subtraction circuit 212 has a first input that receives the mantissa (X) directly from input 201, a second input that receives a signal from the output of multiplexer 211, and an output. Right-shifter 210 performs a 1-bit right alignment for the case where the difference between the exponents is equal to 1.
[0053] Multiplexer 214 has a first input that receives the mantissa (X) directly from input 201, a second input that receives a right-shifted version of the mantissa (X) from input 201 via right-shifter 213, and an output. Subtraction circuit 215 has a first input that receives the mantissa (Y) directly from input 202, a second input that receives a signal from the output of multiplexer 214, and an output. Right-shifter 213 performs a 1-bit right alignment for the case where the difference between the exponents is equal to 1.
[0054] Multiplexer 216 has a first (0) input that receives a signal from the output of subtraction circuit 212, a second (1) input that receives a signal from the output of subtraction circuit 215, a control input that receives the most significant bit (MSB) of the output of subtraction circuit 212, and an output 203. By being configured in this way, the left subtraction circuit 212 calculates a first difference equal to (mantissa(X) - mantissa(Y)), while the right subtraction circuit 215 calculates a second difference equal to (mantissa(Y) - mantissa(X)). Thus, two subtraction operations are calculated in parallel. The MSB of the output of subtraction circuit 212 will indicate whether (mantissa(X) - mantissa(Y)) is positive or negative. If the MSB is 0, then multiplexer 216 routes the result from subtractor 212 to its output 203. If the MSB is 1, then multiplexer 216 routes the result from subtractor 215 to its output 203.
[0055] The leading zero anticipator (LZA) 217 can receive the mantissa (X) from input 201 and the mantissa (Y) from input 202. Based on the received mantissas, the LZA 217 can be configured to predict or estimate the number of leading zeros that may occur at output 203. The LZA 217 can produce an estimated count with a deviation of at most one unit.
[0056] Figure 6 is a circuit diagram of an exemplary near-path normalization circuit 52. As Figure 6 shown, the near-path normalization circuit 52 may include a left-shift circuit 325, a one-bit left-shift circuit 307, multiplexers 308 and 324, subtraction circuits 320, 309, and 310, a comparison circuit 321, logic AND gates 326 and 327, and a logic OR gate 328. The left shifter 325 may receive the positive mantissa difference from output 203 (i.e., Figure 5 the output of multiplexer 216 in ), may receive the count from multiplexer 324 that determines how many left shifts will be performed, and may generate a corresponding shifted value at its output.
[0057] The circuit 52 may also receive an estimated count from the LZA via path 204 and receive a base exponent e_base. The base exponent represents the larger exponent of A and B, which may be output by the setting circuit 40 ( Figure 4 ). The subtraction circuit 320 may be configured to compute an output equal to e_base minus a predetermined value 1. The comparison circuit 321 may be configured to compare e_base with the LZA count value. If e_base is greater than the LZA count, then the comparison circuit 321 may generate an output m equal to 1, or if e_base is less than or equal to the LZA count, then the comparison circuit 321 may generate an output m equal to 0. The multiplexer 324 may have a first (0) input receiving (e_base - 1) from the subtractor 320, a second (1) input directly receiving the LZA count value, a control input receiving the control signal m from the comparison circuit 321, and an output that controls how many shifts the shifter 325 performs.
[0058] The exponent e_base represents the maximum amount of left shift allowed for a normalized number. As long as e_base is greater than the LZA count (i.e., whenever m is asserted), the left-shift circuit 325 will perform the shift by the LZA count. However, when e_base is less than or equal to the LZA count (i.e., whenever m is de-asserted, which represents the subnormal case), the left-shift circuit 325 should only be allowed to perform (e_base - 1) shifts to ensure that the exponent is not less than 1. In Figure 6 the arrangement configured, the comparison circuit 321 and the multiplexer 324 ensure that these criteria are met.
[0059] The subtraction circuit 309 can be configured to compute the difference between e_base and the LZA count and output a corresponding intermediate exponent e_inter. The logic AND gate 326 can receive e_inter and can use the control signal m to enable the logic AND gate 326 to act as a gating circuit that selectively allows e_inter to pass through and reach its output. In response to receiving m = 1, e_inter can pass through unchanged and reach the output of gate 326. In response to receiving m = 0, gate 326 can force e_inter = 1.
[0060] For the near path, the shifted value at the output of circuit 325 can be 1.z or 0.1z. Thus, sometimes it is necessary to use the shifter 307 to shift the output of circuit 325 one additional bit to the left. The multiplexer 308 can have a first (0) input that receives a signal directly from shifter 325, a second (1) input that receives a signal from shifter 307, a control input, and an output 304 that provides a final tail value. The control of multiplexer 308 depends at least in part on the MSB of the output of shifter 325. If the MBS is 1, then the first (0) input of multiplexer 308 should be selected. If the MBS is 0, then the second (1) input of multiplexer 308 should be selected.
[0061] An additional consideration for selecting input (0) or input (1) of multiplexer 308 depends on whether e_inter is greater than one. The logic OR gate 328 monitors the high bits of e_inter (e.g., gate 328 looks at e_inter[5:2], ignoring the least significant bit [1]) and outputs a high value only when e_inter is greater than a predetermined value of one (i.e., whenever at least one of the high bits is non-zero). This example of using a logic OR gate to detect the greater-than-one condition is illustrative only; other logic implementations can be used if desired.
[0062] As described above, the intermediate exponent e_inter is set to 1 to handle subnormal cases. Thus, in subnormal situations, gate 328 will output 0. The logic AND gate 327 has an inverted input that receives the MSB of the output of shifter 325 and a non-inverted input that receives the output of gate 328. Configured in this way, AND gate 327 will output 0 when e_inter is forced to 1 and will output 1 only when e_inter is greater than 1 and the MSB received at its inverted input is low. The output of AND gate 327 is coupled to the control input of multiplexer 308.
[0063] The subtraction circuit 310 can be configured to compute the difference between e_inter provided at the output of gate 326 and the bit provided at the output of gate 327, and generate a corresponding output 305. By being configured in this way, whenever a normal number is being processed and the MSB at the output of 325 is low, the final exponent at output 305 will be equal to (e_inter - 1), or whenever a subnormal number is being processed, the final exponent at output 305 will be equal to 1. Selecting the (1) input of mux 308 to perform an additional one-bit left shift is sometimes referred to as “fine” normalization. Whenever fine normalization is performed, the subtraction circuit 310 will be used to decrement the near-path exponent at output 305 by one.
[0064] Figure 5 circuit 217 and Figure 6 circuits 320, 321, 324, 326, 327, and 328 contribute to subnormal normalization and are thus sometimes collectively referred to as the near-path subnormal control circuit. Figure 5 and Figure 6 The exemplary arrangement of and is merely exemplary. If desired, other suitable ways of implementing similar concepts can be used to support subnormal number processing in the near path.
[0065] Figure 7 is a circuit diagram of a far-path denormalization and arithmetic circuit 700 including a denormalization / align circuit 54 and a far-path arithmetic logic unit 56. A far larger value mantissa (i.e., mantissa(X)) is received at input 401, while a far smaller value mantissa (i.e., mantissa(Y)) is received at input 402. Circuit 700 can also receive a base exponent e_base from a setup circuit 40 and an operation type “op” indicating whether the current operation is addition or subtraction. Circuit 54 can include a right shift circuit 407 for shifting mantissa(Y), where the amount of shift is the difference in exponent values.
[0066] As Figure 7 shown, the far-path ALU 56 can include a first adder / subtractor 408, a second adder / subtractor 409, a third adder / subtractor 415, multiplexing circuits 410 and 414, a rounding condition selection circuit 411, a one-bit right shift circuit 412, a one-bit left shift circuit 413, a logic OR gate 430, and a control circuit 422. Circuit 408 can be configured to compute the addition or subtraction of mantissa(X) and a left-shifted version of mantissa(Y). At the same time, circuit 409 can be configured to compute a rounded version of the addition or subtraction of mantissa(X) and a left-shifted version of mantissa(Y).
[0067] The rounding condition selector 411 can receive the results from circuits 408 and 409, the output from circuit 407, and the signal op for determining whether to use the result from circuit 408 or the result from circuit 409. If no rounding is required, then selector 411 can output the signal r = 0, and if rounding is required, then the output signal r = 1. The multiplexer 410 can have a first (0) input receiving the result from circuit 408, a second (1) input receiving the result from circuit 409, a control input receiving the signal r from selector 411, and an output providing the intermediate mantissa.
[0068] In the far path, in the subnormal case, the intermediate mantissa generated at the output of multiplexer 410 can be 01.z, 10.z, or 00.1z. In the first case, the intermediate mantissa is already normalized and thus no shift is required. In the second case, the intermediate mantissa must be shifted right by one bit, which can be performed using the shifter 412. In the third case, the intermediate mantissa must be shifted left by one bit, which can be performed using the shifter 413. The multiplexer 414 has a first (0) input receiving the right-shifted version from circuit 412, a second (1) input receiving the unaltered intermediate mantissa directly from the output of multiplexer 410, a third (2) input receiving the left-shifted version from circuit 413, a control input, and an output 405 providing the far path mantissa.
[0069] The multiplexer 414 can be configured using the control circuit 422 (sometimes referred to as the multiplexer control circuit). The control circuit 422 can receive the two MSBs of the intermediate mantissa and also the signal e* generated by the logic OR gate 430. The gate 430 can receive the high-order bit of e_base and can output the signal e* indicating whether e_base is greater than a predetermined value of 1. This example of using a logic OR gate to detect the greater-than-one condition is merely illustrative; other logic implementations can be used if desired. If e_base is greater than 1, then at least one of the high-order bits of e_base will be non-zero, thereby forcing e* to be high. In the subnormal case, the signal e* will be 0 only when e_base is equal to 1.
[0070] First, the control circuit 422 can configure the multiplexer 414 to select its (0) input when the two MSBs = 10, and can configure the circuit 415 to increment e_base by 1 to generate the output 406. Second, the control circuit 422 can configure the multiplexer 414 to select its (1) input when the two MSBs = 01 or when the two MSBs = 00 and e* = 0 (i.e., in the subnormal case where no further left shift of the intermediate mantissa can be done), and can also configure the circuit 415 to keep e_base unchanged at the output 406. Third, the control circuit 422 can configure the multiplexer 414 to select its (2) input when the two MSBs = 00 and e* = 1, and can configure the circuit 415 to decrement e_base by 1 to generate the output 406.
[0071] Figure 7 The circuits 422, 415, and 430 contribute to subnormal number calculations and are sometimes collectively referred to as the far-path subnormal calculation circuits. Figure 7 The exemplary arrangements in are merely exemplary. If desired, other suitable ways of implementing similar concepts can be used to support subnormal number handling in the far path.
[0072] Using Figures 4 - 7 The floating-point adder circuit 200 implemented using the circuits shown in is often included in a dedicated processing block that implements a dot product. This type of dedicated processing block can include multiplication circuits 810 and 812 that respectively calculate a first product and a second product, and a floating-point adder 200 for combining the first product and the second product (e.g., see Figure 8A ). As Figure 8A shown, the adder 200 using the near-path and far-path implementations shown in Figures 5 - 7 supports subnormal processing for inputs and outputs having the FP16 format. Figure 8A
[0073] In another suitable embodiment, the adder 200' with the configuration 802 shown in Figure 8B processes inputs having the FP16' format while outputting values having the FP16 format. In yet another suitable embodiment, the adder 200'' with the configuration 803 shown in Figure 8C processes both inputs and outputs having the FP16' format. The configuration 803 can be used where downstream components perform operations using the floating-point single-precision format FP32, since the number of exponents is the same between FP16' and FP32. In this case, the conversion from FP16' to FP32 involves mantissa expansion, which is achieved by padding zeros to the right of the FP16' mantissa.
[0074] The dynamic range of Fp16 is smaller than that of FP16'. Thus, the normal range result of the floating-point number at the input of adder 200' can be converted to the subnormal range at the output of adder 200'. In addition, the normal range result of the floating-point number at the input of adder 200' can be converted to the exceptional condition range at the output of adder 200'. If desired, the exceptional result of the floating-point number at the input of adder 200' can be converted to the normal range at the output of adder 200'. Optionally, the overflow result of the floating-point number at the input of adder 200' can be converted to the normal range at the output of adder 200' without loss of the information contained in the input floating-point number. In configuration 802, the input of multiplier 810 can also be FP16. Thus, the subnormal range result of the floating-point number at the input of multiplier 810 can be converted to the normal range at the output of multiplier 810.
[0075] In configuration 802, since the output format is FP16, the minimum allowable exponent value is -14. An exponent value lower than -14 (e.g., if in the range of -15 to -25) will require subnormal processing. If the intermediate exponent value after mantissa rounding is less than -25 (i.e., when the intermediate result is less than the smallest subnormal number), then an underflow condition is detected and the result is flushed to zero. Typically, the near path processes cases where the operation is subtraction and the relative shift between the two operands is at most 1. For configuration 802, the near path can be triggered only when the smaller exponent is greater than -14. Thus, for configuration 802, the near path circuit does not have to be changed.
[0076] If the exponent value is less than -14, then the far path will be used. In one case, this can be implemented by making the input in the near path zero while the far path defaults to processing the input value. In another case, both the near path and the far path can receive the input value, but Figure 4 the multiplexer 58 in Figure 4 circuit 40 will select the far path output. The control circuit of multiplexer 58 can be in
[0077] Figure 9 is the circuit diagram of the far path denormalization and arithmetic circuit 900, which is capable of processing FP16' at its input (as Figure 8Bas shown in configuration 802). Circuit 900 includes a denormalization / alignment circuit 54 and an arithmetic circuit 56. Circuit 900 may also include a logical OR gate 930 for generating a signal e*. The gate 930 may receive the high bits of e_base (e.g., e_base[8:2] since FP16’ has 8 exponent bits), and may output an output signal e* indicating whether e_base is greater than 1. This example of using a logical OR gate to detect the greater-than-one condition is merely illustrative; other logical implementations may be used if desired. If e_base is greater than 1, then at least one of the high bits of e_base will be non-zero, thereby forcing e* to be high. In the subnormal case, the signal e* will only be 0 when e_base is equal to 1.
[0078] Circuit 900 must be modified to process FP16’ at the input of circuit 900. As Figure 9 shown, a much larger value of the mantissa (i.e., mantissa(X)) is received at input 901, while a much smaller value of the mantissa (i.e., mantissa(Y)) is received at input 902. Circuit 900 may receive the base exponent e_base at inputs 904 and 905, and may receive the difference of the exponents (i.e., exponent(X) - exponent(Y)) at input 903.
[0079] Circuit 52 may further include a subtraction circuit 924, an addition circuit 925, a first right shift circuit 908, a second right shift circuit 909, a first comparison circuit 910, a second comparison circuit 911, and gating circuits 923, 912, and 913 (e.g., logical AND gates). The subtraction circuit 924 may be configured to compute the difference between e_min (which is equal to -14 for half precision) and e_base at its output (e.g., circuit 924 computes -14 minus e_base).
[0080] The gating circuit 923 (e.g., a logical AND gate) receives the computed difference and selectively passes the difference depending on whether the computed difference is positive or negative (e.g., by monitoring the MSB of the computed difference). An MSB of 0 indicates a positive difference, meaning that e_base is less than -14 (i.e., a subnormal number is detected), and the difference is passed through as an output signal s. An MSB of 1 indicates a negative difference, meaning that e_base is equal to or greater than -14, and the signal s will be set to 0. Circuit 908 may be configured to denormalize the mantissa(X) received at input 901 by shifting the mantissa(X) s bits to the right. Thus, no shift will occur in the normal case when e_base is equal to or greater than -14, while a shift of a certain amount (-14 minus e_base) will occur in the subnormal case when e_base is less than -14.
[0081] In addition, if the shift amount exceeds 13 (which can be verified using the comparison circuit 910), then the final shifted value will be forced to zero using the gating circuit 912 (e.g., a logic AND gate). The threshold of 13 is chosen because, for FP16, only 10 mantissa bits are actually used along with 3 additional bits, such as guard bits, rounding bits, and sticky bits. Thus, any shift beyond 13 will only result in a zero output. Circuits 910 and 912 are optional and may not be used because any shift beyond 13 bits will naturally output zero. The final output of the much smaller path in circuit 54 is represented as output a, which can be 15 bits: 1 bit used as an overflow protection bit, 1 sign bit, and 10 mantissa bits followed by 1 guard bit, 1 rounding bit, and 1 sticky bit.
[0082] In the subnormal case, the much smaller path may also be affected by the signal s. The adder circuit 925 can be configured to compute the sum of the difference between the exponents (X) and (Y) received at input 903 and the signal s. The circuit 909 can be configured to denormalize the mantissa (Y) by shifting it to the right. In the normal case, when s = 0, the circuit 909 will shift the mantissa (Y) to the right by the difference in exponents. In the subnormal case, when s is non-zero, the circuit 909 will shift the mantissa (Y) not only by the difference in exponents but also by the position of an additional s bits. In addition, if the value at the output of the circuit 925 is greater than 13 (which can be verified using the comparison circuit 911), then the final shifted value will be forced to zero using the gating circuit 913 (e.g., a logic AND gate). Circuits 911 and 913 are optional and may not be used because any shift beyond 13 bits will naturally output zero. The final output of the much smaller path in circuit 54 is represented as output b, which can have the same number of bits as output a.
[0083] The circuit 56 of the circuit 900 can include logic XOR gates 972 and 974, adder circuits 914, 915, 971, and 973, selector circuit 917, multiplexers 916 and 922, shift circuits 918 and 919, control circuit 921, and adder / subtractor 920. The logic XOR gate 972 can have a first input receiving the signal op, a second input receiving the 3 LSBs of output b (i.e., b[3:1]), and an output. The logic XOR gate 974 can have a first input receiving the signal op, a second input receiving the remaining MSBs of output b (i.e., b[15:4]), and an output.
[0084] The adder 971 may have a first input that receives the 3 LSBs of output a (i.e., a[3:1]), a second input that receives the signal from the output of gate 972, a third carry-in input that receives the signal op, and an output that provides a 4-bit signal (e.g., carry output bit C, guard bit G, round bit R, and sticky bit S). By being configured in this way, the guard, round, and sticky (GRS) bits of two operands are added using the smaller adder 971. When the operation is subtraction, the binary complement of the GRS bits of output b is performed by logic gate 972. The carry of adder 971 is also the signal op that completes the two's complement of the GRS bits of output b.
[0085] The adder 914 has a first input that receives the 12 MSBs of output a (i.e., a[15:4]), a second input that receives the signal from the output of gate 974, and an output that generates the signal p[12:1]. The adder 915 has a first input that receives a[15:4], a second input that receives the signal from the output of gate 974, a third carry input that receives one, and an output that generates the signal q[12:1]. The adder 973 has a first input that receives a[15:4], a second input that receives the signal from the output of gate 974, a third carry input that receives two, and an output that generates the signal t[12:1]. By being configured in this way, adders 914, 915, and 973 add portions a[15:4] and b[15:4] with corresponding 0, 1, and 2 carry values (optionally via binary complement).
[0086] The selector circuit 917 may receive and analyze the 4-bit output from adder 917, the two most significant bits of p, q, and t, op, and e_base, so as to distinguish between different alignment cases. The logic of circuit 921 is related to the logic of circuit 917. The multiplexer 916 may output a corresponding control signal r, which may determine which one of the three input signals (i.e., p, q, or t) is routed to the output of multiplexer 916.
[0087] The remainder of the arithmetic circuit 56 of circuit 900 is substantially equivalent to that of circuit 700. Specifically, Figure 9 circuits 918, 919, 922, and 920 respectively correspond to Figure 7 circuits 412, 413, 414, and 415, and exhibit the same structure and functional capabilities.
[0088] When the operands are both aligned to the e = -14 exponent, the logic behind circuits 917 and 921 can be explained by first looking at addition, then subtraction, and then addition / subtraction.
[0089] For the normal case, when the larger input exponent e_base is greater than or equal to -14 and when the operation (op) is addition, there are three cases to consider. In the first case where p[12:11] = 01 and q[12:11] = 01 (i.e., when both p and q are directly normalized), the relevant bits to be checked are p[1] and the GRS bits. Bit p[1] can be the bit immediately to the left of the guard bit G. If GRS = 100 and p[1] = 1, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q. Otherwise, p will be selected.
[0090] In the second case where p[12:11] = 01 and q[12:11] = 10 (i.e., when p = 01.1111111111), the relevant bits to be checked are still p[1] and the GRS bits. If GRS ≥ 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output from shifter 918. Otherwise, circuit 917 configures multiplexer 916 to select p, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916.
[0091] In the third case where p[12:11] = 10 and q[12:11] = 10, the relevant bits to consider are p[1] (sometimes denoted as bit "L" in the text) and p[2] (sometimes denoted as "L1" in the text). If L1 = 1 and LGRS = 1000, or if LGRS > 1000, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output from shifter 918.
[0092] When the larger input exponent e_base is less than -14, this means that a[3:1] ≠ 000. When the operation (op) is addition, there are also three cases to consider. In the first case where p[12:11] = 00 and q[12:11] = 00, a subnormal may be generated in FP16. If C = 1 (i.e., if the carry output bit of adder 971 is high), then q should be selected before rounding. Applying rounding to q will involve checking the GRS bits and q[1]. If q[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select t, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. If C = 0, then we check the GRS bits and p[1]. If p[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select p, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916.
[0093] In the second case where p[12:11] = 00 and q[12:11] = 01, a subnormal may also be generated. If C = 1, then we check the GRS bits and q[1]. If q[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select t, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. If C = 0, then we check the GRS bits and p[1]. If p[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select p, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916.
[0094] In the third case where p[12:11]=01 and q[12:11]=01, subnormals will not be generated. If C = 1, then we select q before rounding and check the GRS bit and q[1]. If q[1]=1 and GRS = 100, or if GRS>100, then circuit 917 configures multiplexer 916 to select t. Additionally, if t[12:11]=01, then circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. However, if t[12:11]=10, then circuit 921 configures multiplexer 922 to select the output from shifter 918, which should not occur in this case. Otherwise, circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. If C = 0, then we select p before rounding and check the GRS bit and q[1]. If p[1]=1 and GRS = 100, or if GRS>100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select p, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916.
[0095] For the normal case, when the larger input exponent e_base is greater than or equal to -14 and the operation (op) is subtraction, there are also three cases to consider. The first case is when p[12:11]=00 and q[12:11]=00. If e_base equals -14, then subnormals will be generated. If C = 0, then let L = p[1]. Otherwise, let L = q[1]. If GRS = 100 and L = 1, or if GRS>100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select p. If e_base is greater than -14, then normal numbers will be generated. If GRS = 110, then circuit 917 configures multiplexer 916 to select q, and G will be flipped to 0. If RS>10 and if G = 1, then circuit 917 configures multiplexer 916 to select q and flips G to 0; if G = 0, then circuit 917 configures multiplexer 916 to select p and flips G to 1. Otherwise circuit 917 configures multiplexer 916 to select p.
[0096] The second case is when p[12:11] = 00 and q[12:11] = 01, which means p = 0.1111111111. If C = 0, then before rounding, we should return p. Specifically, if e_base > -14, then rounding will occur because p will first be normalized and then rounded. Thus, if G = 1 and RS ≥ 10, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. If G = 0 and RS > 10, then circuit 917 configures multiplexer 916 to select p and flips G, and circuit 921 configures multiplexer 922 to select the output from shifter 919. However, if e_base equals -14, then a subnormal output is assumed, and GRS and p[1] are checked. If p[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916 to establish a subnormal output. Otherwise, circuit 917 configures multiplexer 916 to select p. If C = 1, then before rounding, we will return q, and the result will be normalized. If GRS > 100 and q[1] = 0, then circuit 917 configures multiplexer 916 to select t, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise circuit 917 configures multiplexer 916 to select q.
[0097] The third case is when p[12:11] = 01 and q[12:11] = 01. In this case, the result is normalized before rounding, so there is no distinct case for e_base equal to -14. If C = 1, then before rounding, q should be selected. Rounding involves checking q[1] and GRS. If q[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select t, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise circuit 917 configures multiplexer 916 to select q. If C = 0, then before rounding, p should be selected. Rounding involves checking p[1] and GRS. If p[1] = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise circuit 917 configures multiplexer 916 to select p.
[0098] Exceptional cases where e_base is less than -14 and the operation (op) is subtraction will also be handled on the far path. If p[12:11] = 00 and q[12:11] = 00, then the GRS bit will be checked. If C = 0, then let L = p[1]. If L = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select q, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select p. If C = 1, then let L = q[1]. If L = 1 and GRS = 100, or if GRS > 100, then circuit 917 configures multiplexer 916 to select t, and circuit 921 configures multiplexer 922 to select the output directly from multiplexer 916. Otherwise, circuit 917 configures multiplexer 916 to select q.
[0099] Compared with configuration 802, configuration 803 processes both the input and output using FP16'. The limited range of the data at the outputs of multipliers 810 and 812 (see Figure 8C ) ensures that subnormal numbers cannot be established at the output of adder 200". For example, the smallest non-zero input to the multiplier is 2 -14 *2 -10 = 2 -24 . The product of two such numbers has the value 2 -48 . If these values have different signs and are processed on the near path, then the minimum value after possible cancellation will be 2 -58 . The value of 2 -58 is definitely within the normal range of 2 -126 in FP16' format, so subnormals will not be established. Thus, the adder architecture is similar to an adder architecture where subnormal numbers are flushed to zero at both the input and output.
[0100] If desired, programmable integrated circuit 10 can be dynamically configured to support three different configuration modes 801, 802, or 803. Specifically, for configuration 801, the exponent processing circuit should be configured to process a 5-bit exponent field, and for configurations 802 and 803, the exponent processing circuit should be configured to process an 8-bit exponent field. For the far path, the infrastructure of Figure 9 can be used. To support mode 801, the shifter 908 can be deactivated by setting the shift value to zero (e.g., by forcing the strobe logic 923 to output zero). In other words, the signal s can be forced to 0. These changes can also apply to mode 803.
[0101] For the near path, the infrastructure of Figure 6The infrastructure, which can support both modes 801 and 802. For mode 803, the multiplexer 324 can be forced to always shift the LZA count value received at input 204. Since the exponent cannot become negative in mode 803, the output m of the comparison circuit 321 should be forced to 1, which will also allow the gating circuit 326 to transmit (e_base - LZA count). Since the leading zero predictor is still used, the downstream logic is similar.
[0102] So far, embodiments have been described with respect to integrated circuits. The methods and devices described herein can be incorporated into any suitable circuit. For example, they can be incorporated into many types of devices such as programmable logic devices, application specific standard products (ASSPs), application specific integrated circuits (ASICs), microcontrollers, microprocessors, central processing units (CPUs), graphics processing units (GPUs), etc. Examples of programmable logic devices include programmable array logic (PAL), programmable logic array (PLA), field programmable logic array (FPLA), electrically programmable logic device (EPLD), electrically erasable programmable logic device (EEPLD), logic cell array (LCA), complex programmable logic device (CPLD), and field programmable gate array (FPGA), which are just a few of the examples listed.
[0103] The programmable logic devices described in one or more embodiments herein can be part of a data processing system that includes one or more of the following components: a processor; a memory; IO circuitry; and peripherals. Data processing can be used in a wide range of applications, such as computer networking, data networking, instrumentation, video processing, digital signal processing, or any other suitable application where the advantages of programmable or reprogrammable logic are desired. Various different logic functions can be performed using programmable logic devices. For example, a programmable logic device can be configured as a processor or controller that works in cooperation with a system processor. A programmable logic device can also be used as an arbiter to arbitrate access to shared resources in a data processing system. In yet another example, a programmable logic device can be configured as an interface between a processor and one of the other components in the system.
[0104] Example:
[0105] The following examples relate to other embodiments.
[0106] Example 1 is an integrated circuit that includes: a floating-point adder circuit that receives a first floating-point number and a second floating-point number and outputs a corresponding third floating-point number, where: the first floating-point number and the second floating-point number have a first format, and the third floating-point number has a second format different from the first format; the first floating-point number and the second floating-point number have a first dynamic range, and the third floating-point number has a second dynamic range smaller than the first dynamic range; and the floating-point adder circuit includes a first shifter for denormalizing the first floating-point number and a second shifter for denormalizing the second floating-point number, so that the normalized range results of the first floating-point number and the second floating-point number are converted to the subnormal range of the third floating-point number.
[0107] Example 2 is the integrated circuit of Example 1, which further includes: a first floating-point multiplier that receives a first set of floating-point numbers and outputs the first floating-point number; and a second floating-point multiplier that receives a second set of floating-point numbers and outputs the second floating-point number.
[0108] Example 3 is the integrated circuit of Example 2, where the first and second sets of floating-point numbers have a third format different from the first format.
[0109] Example 4 is the integrated circuit of Example 3, where the third format is the same as the second format.
[0110] Example 5 is the integrated circuit of Example 3, where the subnormal range result of the first set of floating-point numbers is converted to the normalized range of the first floating-point number, and where the subnormal range result of the second set of floating-point numbers is converted to the normalized range of the second floating-point number.
[0111] Example 6 is the integrated circuit of Example 5, where the normalized range results of the first and second floating-point numbers are converted to the exception condition range of the third floating-point number.
[0112] Example 7 is the integrated circuit of Example 6, where the exception results of the first floating-point number and the second floating-point number are converted to the normalized range of the third floating-point number.
[0113] Example 8 is the integrated circuit of Example 7, where the overflow results of the first floating-point number and the second floating-point number are converted to the normalized range of the third floating-point number without losing the information contained in the first and second floating-point numbers.
[0114] Example 9 is an integrated circuit of any one of Examples 1-8, wherein the floating-point adder circuit includes: a near-path circuit that operates on first and second floating-point numbers having an exponent difference equal to zero or one; and a far-path circuit that operates on first and second floating-point numbers having an exponent difference greater than one, wherein the far-path circuit is further configured to operate on first and second floating-point numbers having an exponent difference equal to zero or one while performing an addition operation, and wherein the first and second shifters are part of the far-path circuit.
[0115] Example 10 is an integrated circuit of Example 9, wherein the far-path circuit includes: a rounding circuit; and a selection circuit, wherein the far-path circuit is further configured to operate on first and second floating-point numbers having an exponent difference equal to zero or one while performing a subtraction operation and when the selection circuit determines that left-shift normalization is not required and a rounding operation is possible.
[0116] Example 11 is an integrated circuit of Example 10, wherein paths in the floating-point adder circuit that are not used to process a set of numbers are flushed to zero.
[0117] Example 12 is a method of operating an integrated circuit, the method including: receiving a first floating-point number and a second floating-point number by means of a floating-point adder on the integrated circuit; outputting a corresponding third floating-point number by means of the floating-point adder, wherein the first and second floating-point numbers have a first format that has a first dynamic range, and wherein the third floating-point number has a second format that has a second dynamic range different from the first dynamic range.
[0118] Example 13 is the method of Example 12, further including: converting a normal range result of the first and second floating-point numbers to a subnormal range of the third floating-point number.
[0119] Example 14 is the method of Example 12, further including: receiving a first set of floating-point numbers and outputting the first floating-point number by means of a first multiplier on the integrated circuit; receiving a second set of floating-point numbers and outputting the second floating-point number by means of a second multiplier on the integrated circuit, wherein the first and second sets of floating-point numbers have the first format; converting a subnormal range result of the first set of floating-point numbers to a normal range of the first floating-point number; and converting a subnormal range result of the second set of floating-point numbers to a normal range of the second floating-point number.
[0120] Example 15 is a method of any one of Examples 12-14, further comprising: converting the normal range results of the first and second floating-point numbers to the subnormal range of the third floating-point number; converting the subnormal results of the first and second floating-point numbers to the normal range of the third floating-point number; and converting the overflow results of the first and second floating-point numbers to the normal range of the third floating-point number without loss of information contained in the first and second floating-point numbers.
[0121] Example 16 is an integrated circuit comprising: a dedicated processing block including: a first multiplier configured to generate a first product; a second multiplier configured to generate a second product; and a floating-point adder circuit that receives the first and second products and generates a corresponding output, wherein the floating-point adder circuit is capable of operating in multiple modes to support subnormal numbers.
[0122] Example 17 is the integrated circuit of Example 16, wherein when the floating-point adder is placed in the first of the multiple modes, the floating-point adder is configured to receive the first and second products in a half-precision format using a 5-bit exponent field.
[0123] Example 18 is the integrated circuit of Example 17, wherein when the floating-point adder is placed in the second of the multiple modes, the floating-point adder is configured to receive the first and second products in a modified half-precision format using an 8-bit exponent field.
[0124] Example 19 is the integrated circuit of Example 18, wherein when the floating-point adder is placed in the second of the multiple modes, the floating-point adder is configured to generate an output in the half-precision format.
[0125] Example 20 is the integrated circuit of any one of Examples 18-19, wherein when the floating-point adder is placed in the third of the multiple modes, the floating-point adder is configured to receive the first and second products in the modified half-precision format and also generate an output in the modified half-precision format.
[0126] Example 21 is an integrated circuit comprising: a module for receiving a first floating-point number and a second floating-point number and outputting a corresponding third floating-point number, wherein the first and second floating-point numbers have a first format having a first dynamic range, and wherein the third floating-point number has a second format having a second dynamic range less than the first dynamic range.
[0127] Example 22 is the integrated circuit of Example 21, further comprising: a module for converting the normal range results of the first and second floating-point numbers to the subnormal range of the third floating-point number.
[0128] Example 23 is the integrated circuit of Example 21, further comprising: a module for receiving a first set of floating-point numbers and outputting the first floating-point number; a module for receiving a second set of floating-point numbers and outputting the second floating-point number, wherein the first and second sets of floating-point numbers have a first format; a module for converting a subnormal range result of the first set of floating-point numbers to a normal range of the first floating-point number; and a module for converting a subnormal range result of the second set of floating-point numbers to a normal range of the second floating-point number.
[0129] Example 24 is the integrated circuit of any one of Examples 21-23, further comprising: a module for converting a normal range result of the first and second floating-point numbers to an exception condition range of the third floating-point number; a module for converting an exception result of the first and second floating-point numbers to a normal range of the third floating-point number.
[0130] Example 25 is the integrated circuit of any one of Examples 21-23, further comprising: a module for converting an overflow result of the first and second floating-point numbers to a normal range of the third floating-point number without losing information contained in the first and second floating-point numbers.
[0131] For example, all optional features of the devices described above can also be implemented with respect to the methods or processes described herein. The foregoing is merely illustrative of the principles of the present disclosure, and those skilled in the art can make various modifications. The above embodiments can be implemented individually or in any combination.
Claims
1. An integrated circuit, comprising: A floating - point adder circuit that receives a first floating - point number and a second floating - point number and outputs a corresponding third floating - point number, wherein: The first floating - point number and the second floating - point number have a first format, and the third floating - point number has a second format different from the first format; The first floating - point number and the second floating - point number have a first dynamic range, and the third floating - point number has a second dynamic range smaller than the first dynamic range; and The floating - point adder circuit includes a first shifter for denormalizing the first floating - point number and a second shifter for denormalizing the second floating - point number, so that the normalized range results of the first floating - point number and the second floating - point number are converted to the sub - normal range of the third floating - point number.
2. The integrated circuit according to claim 1, further comprising: A first floating - point multiplier that receives a first set of floating - point numbers and outputs the first floating - point number; And A second floating - point multiplier that receives a second set of floating - point numbers and outputs the second floating - point number.
3. The integrated circuit according to claim 2, wherein, The first set of floating - point numbers and the second set of floating - point numbers have a third format different from the first format.
4. The integrated circuit according to claim 3, wherein The third format is the same as the second format.
5. The integrated circuit according to claim 3, wherein, The sub - normal range results of the first set of floating - point numbers are converted to the normal range of the first floating - point number, and wherein, the sub - normal range results of the second set of floating - point numbers are converted to the normal range of the second floating - point number.
6. The integrated circuit according to claim 5, wherein, The normal range results of the first floating - point number and the second floating - point number are converted to the exceptional condition range of the third floating - point number.
7. The integrated circuit according to claim 6, wherein, The exceptional results of the first floating - point number and the second floating - point number are converted to the normal range of the third floating - point number.
8. The integrated circuit according to claim 7, wherein, The overflow results of the first floating - point number and the second floating - point number are converted to the normal range of the third floating - point number without loss of the information contained in the first floating - point number and the second floating - point number.
9. The integrated circuit according to any one of claims 1-8, wherein, The floating - point adder circuit includes: A near - path circuit that operates on the first floating - point number and the second floating - point number having an exponent difference equal to zero or one; and A far - path circuit that operates on the first floating - point number and the second floating - point number having an exponent difference greater than one, wherein the far - path circuit is further configured to operate on the first floating - point number and the second floating - point number having an exponent difference equal to zero or one while performing an addition operation.
10. The integrated circuit according to claim 9, wherein, The far - path circuit includes: A rounding circuit; and A selection circuit, wherein the far - path circuit is further configured to operate on the first floating - point number and the second floating - point number having an exponent difference equal to zero or one while performing a subtraction operation and when the selection circuit determines that left - shift normalization is not required and a rounding operation is possible.
11. The integrated circuit according to claim 10, wherein, The paths in the floating - point adder circuit that are not used to process a set of numbers are flushed to zero.
12. A method of operating an integrated circuit, the method comprising: Receiving a first floating - point number and a second floating - point number by means of the floating - point adder circuit on the integrated circuit; With the aid of the floating-point adder circuit, a corresponding third floating-point number is output, including: denormalizing the first floating-point number by means of a first shifter in the floating-point adder circuit and denormalizing the second floating-point number by means of a second shifter in the floating-point adder circuit, so as to convert the normalized range results of the first floating-point number and the second floating-point number to the subnormal range of the third floating-point number, wherein the first floating-point number and the second floating-point number have a first format, the first format has a first dynamic range, and wherein the third floating-point number has a second format, the second format has a second dynamic range different from the first dynamic range.
13. The method according to claim 12, further comprising: Receiving a first set of floating-point numbers and outputting the first floating-point number by means of a first multiplier on the integrated circuit; Receiving a second set of floating-point numbers and outputting the second floating-point number by means of a second multiplier on the integrated circuit, wherein the first set of floating-point numbers and the second set of floating-point numbers have the first format; Converting the subnormal range result of the first set of floating-point numbers to the normalized range of the first floating-point number; and Converting the subnormal range result of the second set of floating-point numbers to the normalized range of the second floating-point number.
14. The method according to any one of claims 12-13, further comprising: Converting the normalized range results of the first floating-point number and the second floating-point number to the exceptional condition range of the third floating-point number; Converting the exceptional results of the first floating-point number and the second floating-point number to the normalized range of the third floating-point number; And Converting the overflow results of the first floating-point number and the second floating-point number to the normalized range of the third floating-point number without losing the information contained in the first floating-point number and the second floating-point number.
15. An integrated circuit, comprising: A dedicated processing block, which includes: A first multiplier configured to generate a first product; A second multiplier configured to generate a second product; and A floating-point adder circuit that receives the first product and the second product and generates a corresponding output, wherein the floating-point adder circuit can operate in multiple modes to support subnormal numbers, and wherein the floating-point adder circuit receives a first floating-point number and a second floating-point number and outputs a corresponding third floating-point number, wherein: The first floating-point number and the second floating-point number have a first format, and the third floating-point number has a second format different from the first format; The first floating-point number and the second floating-point number have a first dynamic range, and the third floating-point number has a second dynamic range smaller than the first dynamic range; and The floating-point adder circuit includes a first shifter for denormalizing the first floating-point number and a second shifter for denormalizing the second floating-point number, so that the normalized range results of the first floating-point number and the second floating-point number are converted to the subnormal range of the third floating-point number.
16. The integrated circuit according to claim 15, wherein, When the floating-point adder circuit is placed in the first of the multiple modes, the floating-point adder circuit is configured to receive the first product and the second product in a half-precision format using a 5-bit exponent field.
17. The integrated circuit according to claim 16, wherein, When the floating-point adder circuit is placed in the second of the multiple modes, the floating-point adder circuit is configured to receive the first product and the second product in a modified half-precision format using an 8-bit exponent field.
18. The integrated circuit according to claim 17, wherein, When the floating-point adder circuit is placed in the second of the multiple modes, the floating-point adder circuit is configured to generate the output in the half-precision format.
19. The integrated circuit according to any one of claims 17 - 18, wherein, When the floating-point adder circuit is placed in the third of the multiple modes, the floating-point adder circuit is configured to receive the first product and the second product in the modified half-precision format and also generate the output in the modified half-precision format.