Arithmetic circuit and related method
Patent Information
- Application Number
- CN202211118581.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2042-09-13
AI Technical Summary
由于不同的舍入模式有不同的舍入条件,因此在先前技术中,浮点数的舍入运算电路会对浮点数进行一系列的移位运算以及复杂的条件判断,且在部分模式中还需要进行多位的加法,导致运算效率不佳,且运算电路也需要较大的面积来实作
[0005]由于本申请的运算电路及运算方法可以依据浮点数的阶码产生掩码,并可利用掩码产生浮点数的保护位、最低有意义位及黏滞位以进行舍入位的判断,而无须对尾数进行多次的移位,因此可以简化舍入运算的过程。此外,由于本申请的运算电路及相关方法还可将尾数的小数部分截断而改以填入舍入位,而仅需要将待舍入尾数与一位舍入位相加就能得到浮点整数,无须完整的多位加法器,因此还能够减少硬件所需的电路面积。
Smart Images

Figure CN117742655B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an arithmetic unit, and more specifically to an arithmetic unit for rounding floating-point numbers to floating-point integers. Background Technology
[0002] In some applications, the processor must convert floating-point numbers to floating-point integers during computation. This means the processor must discard the fractional part of the floating-point number and determine whether to carry over the integer part based on the rounding mode. Common rounding modes include round to nearest even (RNE), round to positive infinity (RTP), round to negative infinity (RTN), and round to zero (RTZ). Because different rounding modes have different rounding conditions, in previous technologies, floating-point rounding circuits performed a series of shift operations and complex condition checks on the floating-point number. In some modes, multi-bit addition was also required, resulting in poor computational efficiency and requiring a large area for implementation. Therefore, how to more efficiently convert floating-point numbers to floating-point integers has become a problem to be solved. Summary of the Invention
[0003] One embodiment of this application provides an arithmetic circuit for rounding a floating-point number to a floating-point integer. The floating-point number includes an N-bit mantissa and a P-bit exponent, where N and P are positive integers. The arithmetic circuit includes a mask generation unit, a rounding bit generation unit, a mantissa truncation unit, and a rounding calculation unit. The mask generation unit generates a first mask and generates a second and a third mask based on the first mask. The first mask distinguishes the integer and fractional parts of the mantissa, the second mask marks the first digit after the decimal point, and the third mask marks all digits below the second decimal point. The rounding bit generation unit obtains the guard bit and least significant bit of the mantissa based on the second mask, obtains the sticky bit of the mantissa based on the third mask, and determines the rounding bit based on the rounding mode of the arithmetic circuit, the guard bit, the least significant bit, and the sticky bit. The mantissa truncation unit fills the fractional part of the mantissa with the rounding bits to generate the mantissa to be rounded. The rounding calculation unit is used to generate the floating-point integer based on the mantissa to be rounded and the rounding digit.
[0004] Another embodiment of this application provides an arithmetic method for rounding a floating-point number to a floating-point integer. The floating-point number includes an N-bit mantissa and a P-bit exponent, where N and P are positive integers. The method includes: generating a first mask, wherein the first mask is used to distinguish the integer part and the fractional part of the mantissa; generating a second mask based on the first mask, wherein the second mask is used to mark the first digit after the decimal point of the mantissa; generating a third mask based on the first mask, wherein the third mask is used to mark all bits below the second digit after the decimal point of the mantissa; obtaining the guard bit of the mantissa based on the second mask; obtaining the least significant bit of the mantissa based on the second mask; obtaining the sticky bit of the mantissa based on the third mask; determining the rounding bit based on the rounding mode, the guard bit, the least significant bit, and the sticky bit; filling the fractional part of the mantissa with the rounding bit to generate a mantissa to be rounded; and generating the floating-point integer based on the mantissa to be rounded and the rounding bit.
[0005] Since the arithmetic circuit and method of this application can generate a mask based on the exponent of the floating-point number, and can use the mask to generate the protection bit, least significant bit, and sticky bit of the floating-point number for rounding judgment, without needing to shift the mantissa multiple times, the rounding operation process can be simplified. In addition, since the arithmetic circuit and related method of this application can also truncate the decimal part of the mantissa and fill it into the rounding bit, and only the mantissa to be rounded needs to be added to one rounding bit to obtain the floating-point integer, without the need for a complete multi-bit adder, the required circuit area of the hardware can also be reduced. Attached Figure Description
[0006] The aspects of this disclosure will be better understood from the following embodiments when read in conjunction with the accompanying drawings. It should be noted that, in accordance with standard industry practice, the various structures are not drawn to scale. In fact, the dimensions of the various structures may be arbitrarily increased or decreased for clarity of explanation.
[0007] Figure 1 This is a schematic diagram of an embodiment of the operational circuit of this application.
[0008] Figure 2 This is a diagram illustrating floating-point numbers.
[0009] Figure 3 This is a flowchart of the method for rounding floating-point numbers to floating-point integers in this application.
[0010] Figure 4 This is a diagram showing the comparison between the first mask and the mantissa.
[0011] Figure 5 This is a diagram showing the comparison between the second mask and the mantissa.
[0012] Figure 6This is a diagram showing the comparison between the third mask and the mantissa.
[0013] Figure 7 This is a schematic diagram illustrating the operation of generating protection bits based on the second mask.
[0014] Figure 8 This is a schematic diagram illustrating the operation of generating the least meaningful bit based on the second mask.
[0015] Figure 9 This is a schematic diagram of the operation of generating sticky bits based on the third mask.
[0016] Figure 10 This is a schematic diagram illustrating the operation of generating the mantissa to be rounded.
[0017] Figure 11 This is a schematic diagram illustrating the operation of generating floating-point integers. Detailed Implementation
[0018] The following disclosure provides numerous different embodiments or examples of various components for implementing the provided subject matter. Specific examples of components and arrangements are described below to simplify this disclosure. Of course, such examples are merely illustrative and are not intended to be limiting. For example, in the following description, a first component formed above or on a second component may include embodiments in which the first and second components are in direct contact, and may also include embodiments in which additional components may be formed between the first and second components such that the first and second components are not in direct contact. Additionally, reference numerals and / or letters may be repeated in various instances of this disclosure. This repetition is for simplicity and clarity and does not in itself indicate a relationship between the various embodiments and / or configurations discussed.
[0019] Furthermore, for ease of description, spatial relative terms such as “below,” “under,” “below,” “above,” “above,” and similar terms may be used herein to describe the relationship of one component or member to another component or member illustrated in the figures. In addition to the orientations depicted in the figures, spatial relative terms are intended to cover different orientations of the device during use or operation. The device may be oriented in other ways (rotated 90 degrees or otherwise) and therefore the spatial relative descriptors used herein may be interpreted in the same way.
[0020] As used herein, terms such as “first,” “second,” and “third” describe various components, parts, areas, layers, and / or sections, but such components, parts, areas, layers, and / or sections should not be limited by such terms. These terms are used only to distinguish one component, part, area, layer, or section from another. For example, the use of terms such as “first,” “second,” and “third” herein does not imply sequence or order unless explicitly indicated by the context.
[0021] The singular forms "a," "an," and "the" may also include plural forms unless the context explicitly indicates otherwise. The term "connection," along with its derivatives, may be used herein to describe structural relationships between components. "Connection" may be used to describe two or more elements in direct physical or electrical contact with each other. "Connection" may also be used to indicate two or more elements in direct or indirect physical or electrical contact with each other (with intervening elements between them), and / or two or more elements cooperating or interacting with each other.
[0022] Figure 1 This is a schematic diagram of an embodiment of the operational circuit of this application. For example... Figure 1 As shown, the arithmetic circuit 100 may include a mask generation unit 110, a rounding bit generation unit 120, a mantissa truncation unit 130, and a rounding calculation unit 140. In this embodiment, the arithmetic circuit 100 can be used to round the floating-point number FP1 into a floating-point format integer INTF1.
[0023] Figure 2 This is a schematic diagram of the floating-point number FP1. In this embodiment, the floating-point number FP1 may include a sign bit SGN, an N-bit mantissa MTS1, and a P-bit exponent EXP, where N and P are positive integers. The sign bit SGN indicates whether the floating-point number FP1 is positive or negative, the exponent EXP represents the exponent of the floating-point number FP1, and the mantissa MTS1 is the number of significant bits in the floating-point number FP1. For example, the floating-point number FP1 may be a 32-bit floating-point number as defined by IEEE-754. Figure 2 As shown, a 32-bit floating-point number FP1 can include a one-bit sign bit SGN, an 8-bit exponent EXP, and a 23-bit mantissa MTS1, that is, N is 23 and P is 8. Figure 2 For example, the value of the floating-point number FP1 can be represented as shown in equation (1).
[0024] FP1 = (-1) SGN ×1.MTS1×2 EXP-B =1.01001010100001100100001×2 10 Equation (1)
[0025] In equation (1), B is the preset bias, and the exponent EXP minus the preset bias B is the exponent of the floating-point number FP1. In this embodiment, the exponent EXP is 137, and the preset bias B can be, for example, about half the maximum value of the 8-bit exponent EXP, which is 127. In this case, the decimal point of the floating-point number FP1 will fall at the default position PP and shifted 10 bits lower. Therefore, the integer part of the floating-point number FP1 may include the hidden bit 1 of the highest bit and the highest 10 bits of the mantissa MTS1, that is, 10100101010, while the fractional part is the lowest 13 bits of the mantissa MTS1, that is, 0001100100001.
[0026] In the arithmetic circuit 100, the mask generation unit 110 can generate a first mask MSK1 based on the exponent EXP of the floating-point number FP1 to distinguish the integer part and the fractional part of the mantissa MTS1, and can generate a second mask MSK2 and a third mask MSK3 based on the first mask MSK1. The second mask MSK2 can be used to mark the first digit after the decimal point in the mantissa MTS1, and the third mask MSK3 can be used to mark all digits after the second decimal point in the mantissa MTS1. Next, the rounding bit generation unit 120 can generate a guard bit G, a least significant bit L, and a sticky bit S based on the second mask MSK2 and the third mask MSK3, and can generate a rounding bit R based on the guard bit G, the least significant bit L, and the sticky bit S to determine whether the fractional part should be rounded up. After determining the rounding bit R, the mantissa truncation unit 130 can use the rounding bit R to fill the fractional part of the mantissa MTS1 to generate the rounded mantissa MTS2, and the rounding calculation unit 140 can generate the floating-point integer INTF1 based on the mantissa MTS2 to be rounded and the rounding bit R.
[0027] Since the arithmetic circuit 100 can generate masks MSK1, MSK2, and MSK3 based on the exponent EXP, and can use masks MSK2 and MSK3 to generate the protection bit G, the least significant bit L, and the sticky bit S for determining the rounding bit R, without needing to perform multiple shifts on the mantissa MTS1, the rounding operation process can be simplified compared to the prior art. In addition, since the mantissa truncation unit 130 truncates the fractional part of the mantissa MTS1 and fills it into the rounding bit R, the rounding calculation unit 140 only needs to add the mantissa MTS2 to be rounded to one rounding bit R to obtain the floating-point integer INTF1, without needing a complete N-bit adder, thereby reducing the circuit area required by the arithmetic circuit 100.
[0028] Figure 3 This is a flowchart illustrating the method for rounding floating-point numbers to floating-point integers as described in this application. Figure 3 As shown, method M1 may include steps S110 to S190. In this embodiment, method M1 may be executed by the arithmetic circuit 100.
[0029] In step S110, the mask generation unit 110 can generate a first mask MSK1 based on the exponent EXP of the floating-point number FP1. In this embodiment, the first mask MSK1 can be used to separate the integer part and the fractional part of the mantissa MTS1. For example, the first mask MSK1 can have the same N bits as the mantissa MTS1, and in the first mask MSK1, each bit corresponding to the integer part of the mantissa MTS1 is 1, and each bit corresponding to the fractional part of the mantissa MTS1 is 0. Figure 4 This is a schematic diagram showing the comparison between the first mask MSK1 and the mantissa MTS1.
[0030] In this embodiment, the mask generation unit 110 can perform a shift operation on an N-bit 1 value (N'b1) to generate a first mask MSK1. The mask generation unit 110 can determine the number of bits in the integer and fractional parts of the mantissa MTS1 based on the exponent EXP minus a preset bias B, thereby determining the shift number. For example, when the exponent is greater than or equal to 0 and less than N, it indicates that the mantissa MTS1 should have a fractional part; in this case, the mask generation unit 110 can subtract the exponent from N to obtain the shift number. Figure 2 Taking the floating-point number FP1 as an example, the exponent EXP minus the preset bias B is 10, indicating that the decimal point will be shifted 10 bits to the right (lower bit). In this case, since 10 is greater than 0 and less than N (N=2^3), the mask generation unit 110 will obtain a shift value of 13, and shift the value of N bits of 1 to the higher bit by 13 bits to generate the first mask MSK1 (the lower bit part is padded with zeros), indicating that the lowest 13 bits of the mantissa MTS1 belong to the fractional part, while the remaining higher 10 bits are the integer part.
[0031] However, when the exponent is greater than or equal to N, it means that the decimal point will be shifted to the right (lower digit) by at least N positions. Therefore, all bits of the mantissa MTS1 are integers. In this case, the mask generation unit 110 can set the shift number to 0 and shift the N-bit 1 value to the left by 0 positions. That is, the mask generation unit 110 does not need to perform a shift operation at this time, so the first mask MSK1 is the value of N-bit 1s.
[0032] Conversely, if the exponent is less than 0, it means that the decimal point will be shifted to the left (higher bit) by at least 1 bit from its original position. Therefore, all bits of the mantissa MTS1 belong to the decimal. At this time, the mask generation unit 110 can set the shift number to N, thereby shifting the value of N bits of 1 to the higher bit by N bits (filled with zeros) to obtain the value of the first mask MSK1 as N bits of 0 (N'b0).
[0033] After obtaining the first mask MSK1, the mask generation unit 110 can generate a second mask MSK2 in step S120 based on the first mask MSK1. In this embodiment, the second mask MSK2 can be used to mark the first decimal place of the mantissa MTS1. For example, in the second mask MSK2, the bit corresponding to the first decimal place in the mantissa MTS1 is 1, and the other bits are 0. Figure 5 This is a schematic diagram showing the comparison between the second mask MSK2 and the mantissa MTS1.
[0034] In this embodiment, when the exponent is not equal to 0, the mask generation unit 110 can shift the first mask MSK1 down by 1 bit and pad it with 1 in the high bits to generate the intermediate mask MSKA, such as... Figure 5As shown, performing a bitwise XOR operation between the intermediate mask MSKA and the first mask MSK1 generates the second mask MSK2. Furthermore, if the exponent is 0, it indicates that the most significant bit of the mantissa MTS1 is the first decimal place. In this case, the mask generation unit 110 can directly shift the N bits of 1 to the higher position by (N-1) bits to generate the second mask MSK2. In this scenario, the second mask MSK2 has only its most significant bit set to 1, while all other bits are 0.
[0035] In this embodiment, the mask generation unit 110 can also generate a third mask MSK3 based on the first mask MSK1 in step S130. In this embodiment, the third mask MSK3 can be used to mark all bits after the second decimal place of the mantissa MTS1. For example, in the third mask MSK3, all bits corresponding to the second decimal place of the mantissa MTS1 are 1, and the other bits are 0. Figure 6 This is a schematic diagram showing the comparison between the third mask MSK3 and the mantissa MTS1.
[0036] Since the integer and fractional parts of the mantissa MTS1 can be determined based on the first mask MSK1, in step S130, the mask generation unit 110 can shift the complement of the first mask MSK1 to obtain the third mask MSK3. However, when the exponent is less than 0, the fractional part of the floating-point number FP1 will also include the previously preset hidden bit 1, so the most significant bit of the mantissa MTS1 should be after the second decimal place. In this case, to avoid misjudging the most significant bit of the mantissa MTS1, during the shift, it is also necessary to determine whether the most significant value of the third mask MSK3 should be 1 or 0 based on whether the exponent is less than 0.
[0037] In this embodiment, when the exponent is less than 0, the mask generation unit 110 can generate a supplementary bit AB with a value of 1; otherwise, the mask generation unit 110 can generate a supplementary bit AB with a value of 0. Then, the mask generation unit 110 can set the supplementary bit AB in the high-order bits and combine it with the highest (N-1) bits MSK1'[N-1:1] of the complement of the first mask MSK1 to form the third mask MSK3. Figure 6 Since the exponent is greater than 0, the supplementary bit AB is 0; in this case, the highest 11 bits of the third mask MSK3 are 0, while the lowest 12 bits of the third mask MSK3 are 1.
[0038] After obtaining the masks MSK2 and MSK3, the rounding bit generation unit 120 can generate a protection bit G in step S140 based on the second mask MSK2. In this embodiment, the protection bit G is the first decimal place of the floating-point number FP1. Figure 7This is a schematic diagram illustrating the operation of generating the protection bit G based on the second mask MSK2. Since the second mask MSK2 has marked the position of the first digit after the decimal point in the mantissa MTS1, the rounding bit generation unit 120 can perform a bitwise AND operation between the mantissa MTS1 and the second mask MSK2 to generate the result RT1, as shown below. Figure 7 As shown, a bitwise OR operation is then performed on each bit of the result RT1 to generate the protection bit G. It should be noted that if the exponent is -1, it means that the highest bit of the mantissa MTS1 corresponds to the second decimal place of the floating-point number FP1, while the first decimal place of the floating-point number FP1 is the preset hidden bit 1 of the floating-point number FP1; in this case, the rounding bit generation unit 120 can set the value of the protection bit G to 1.
[0039] In step S150, the rounding bit generation unit 120 can also generate the least significant bit L according to the second mask MSK2. In this embodiment, the least significant bit L is the least significant bit of the integer part of the floating-point number FP1. Figure 8 This is a schematic diagram illustrating the operation of generating the least significant bit L based on the second mask MSK2. For example... Figure 8 As shown, the rounding bit generation unit 120 can set the value of 1 bit 1 in the high bit and combine it with the highest (N-1) bit MTS1[N-1:1] of the mantissa MTS1, and then perform a bitwise AND operation with the second mask MSK2 to generate the operation result RT2. Then, it performs an OR operation on each bit of the operation result RT2 to generate the least meaningful bit L.
[0040] In step S160, the rounding bit generation unit 120 can generate a sticky bit S based on the third mask MSK3. In this embodiment, the sticky bit S is the result of a logical OR operation on all bits below the second decimal place of the floating-point number FP1. Figure 9 This is a schematic diagram illustrating the operation of generating the sticky bit S based on the third mask MSK3. Since the third mask MSK3 has marked the positions of all bits below the second decimal point of the mantissa MTS1, under normal circumstances, the rounding bit generation unit 120 can perform a bitwise AND operation on the mantissa MTS1 and the third mask MSK3 to produce the result RT3, and then perform an OR operation on each bit of the result RT3 to generate the sticky bit S, such as... Figure 9 As shown. It should be noted that when the exponent is less than -1, the default hidden bit 1 of the floating-point number FP1 will be after the second decimal place of FP1, therefore the sticky bit S should be 1. However, according to IEEE specifications, when every bit of the floating-point number FP1 is 0, it can be represented as 0. That is, when the exponent EXP is 0, the hidden bit of the floating-point number FP1 can be 0 instead of 1. In this case, the sticky bit S cannot be directly set to 1, but it can still be determined whether the sticky bit S is 0 by performing a bitwise AND operation on the mantissa MTS1 and the third mask MSK3.
[0041] After obtaining the protection bit G, least significant bit L, and sticky bit S of the floating-point number FP1, the rounding bit generation unit 120 can further determine the rounding bit R in step S170 based on the rounding mode, protection bit G, least significant bit L, and sticky bit S. When the rounding mode is RNE, if the protection bit G is 0, the fractional part can be directly discarded without carrying over the integer part. Conversely, if the protection bit G is 1 and the sticky bit S is also 1, it means that the floating-point number FP1 should be greater than the midpoint of its two closest integers, and a carry is required. In addition, if the protection bit G is 1 and the sticky bit S is 0, according to the RNE rule, the least significant bit L must be considered. If the least significant bit is 1, a carry is required; otherwise, no carry is required. In this case, the rounding bit generation unit 120 can perform an OR operation on the least significant bit L and the sticky bit S, and then perform an AND operation on the protection bit G to generate the rounding bit R, as shown in equation (2).
[0042] R = G && (L||S) Equation (2)
[0043] Furthermore, in the case of RTP rounding mode, if the sign bit is 0, it indicates that the floating-point number FP1 is positive. In this case, a carry is required as long as the protection bit G or the sticky bit S is 1. Conversely, if the sign bit is 1, it indicates that the floating-point number FP1 is negative. In this case, no carry is required regardless of whether the protection bit G and the sticky bit S are 1. In this case, the rounding bit generation unit 120 can perform an OR operation on the protection bit G and the sticky bit S, and then perform an AND operation on the complement of the sign bit SGN to generate the rounding bit R, as shown in equation (3).
[0044] R = ! SGN && (G||S) Equation (3)
[0045] Similarly, in the rounding mode RTN, if the sign bit is 1, it indicates that the floating-point number FP1 is negative. In this case, a carry is required as long as the protection bit G or the sticky bit S is 1. Conversely, if the sign bit is 0, it indicates that the floating-point number FP1 is positive. In this case, no carry is required regardless of whether the protection bit G and the sticky bit S are 1. In this case, the rounding bit generation unit 120 can perform an OR operation on the protection bit G and the sticky bit S, and then perform an AND operation on the sign bit SGN to generate the rounding bit R, as shown in equation (4).
[0046] R = SGN && (G||S) Equation (4)
[0047] Furthermore, in the case of RTZ rounding mode, since there is no need to carry, the rounding bit generation circuit 120 can directly set the rounding bit R to 0.
[0048] After obtaining the rounding bit R, the mantissa truncation unit 130 can use the rounding bit R to fill the decimal part of the mantissa MTS1 in step S180 to generate the mantissa MTS2 to be rounded, and the rounding calculation unit 140 can then generate the floating-point integer INTF1 based on the mantissa MTS2 to be rounded and the rounding bit R in step S190. Figure 10 This is a schematic diagram illustrating the operation of generating the mantissa MTS2 to be rounded; Figure 11 This is a schematic diagram of the operation that generates the floating-point integer INTF1.
[0049] like Figure 10 As shown, the mantissa truncation unit 130 can first perform a bitwise AND operation on the mantissa MTS1 and the first mask MSK1 to generate the first number A1, and then perform a bitwise AND operation on the complement MSK1' of the first mask MSK1 and the N rounding bits N'bR to generate the second number A2. Then, it can perform a bitwise OR operation on the first number A1 and the second number A2 to generate the mantissa MTS2 to be rounded.
[0050] After generating the mantissa MTS2 to be rounded, as follows Figure 11 As shown, in step S190, the rounding calculation unit 140 can set the exponent EXP in the high-order bits and combine it with the mantissa MTS2 to be rounded, then add it to the rounding bit R to obtain the exponent and mantissa of the floating-point integer INTF1. In this way, when the rounding bit R is 1, the rounding calculation unit 140 only needs to add 1 to the mantissa MTS2 to obtain the carry-over result, thus simplifying the hardware requirements. Furthermore, since the exponent EXP of the floating-point number FP1 can be set in the high-order bits and combined with the mantissa MTS2 to be rounded, if the carry-over results in an increase in the number of bits in the integer part, the exponent of the floating-point integer INTF1 will also be directly carried over, thus obtaining the correct exponent.
[0051] In summary, the arithmetic circuit and method of this application can round floating-point numbers into floating-point integers according to different rounding modes. Furthermore, since the arithmetic circuit and method of this application can generate a mask based on the exponent, and can use the mask to generate the guard bit, least significant bit, and sticky bit of the floating-point number for rounding determination, it eliminates the need for multiple shifts of the mantissa, thus simplifying the rounding process. In addition, the arithmetic circuit and method of this application can also truncate the decimal part of the mantissa and fill it into the rounding bit. Therefore, only the mantissa to be rounded needs to be added to a single rounding bit to obtain a floating-point integer, eliminating the need for a complete multi-bit adder, thereby reducing the required circuit area.
[0052] The foregoing outlines the structures of several embodiments to enable those skilled in the art to better understand aspects of this disclosure. Those skilled in the art will understand that this disclosure can be readily used as a basis for designing or modifying other manufacturing processes and structures for carrying out the same purposes and / or achieving the same advantages of the embodiments described herein. Those skilled in the art should also recognize that such equivalent constructions do not depart from the spirit and scope of this disclosure, and that various changes, substitutions, and modifications can be made herein without departing from the spirit and scope of this disclosure.
Claims
1. An arithmetic circuit for rounding floating-point numbers to floating-point integers, characterized in that, The floating-point number includes an N-bit mantissa and a P-bit exponent, where N and P are positive integers. The arithmetic circuit includes: A mask generation unit is configured to generate a first mask based on the integer part and the fractional part of the mantissa, and to generate a second mask and a third mask based on the first mask, wherein the first mask is used to distinguish the integer part and the fractional part of the mantissa, the second mask is used to mark the first digit after the decimal point of the mantissa, and the third mask is used to mark all digits below the second digit after the decimal point of the mantissa. The rounding bit generation unit is used to obtain the protection bit and the least significant bit of the mantissa according to the second mask, obtain the sticky bit of the mantissa according to the third mask, and determine the rounding bit according to the rounding mode of the arithmetic circuit, the protection bit, the least significant bit and the sticky bit. A mantissa truncation unit is used to fill the fractional part of the mantissa with the rounding bits to generate a mantissa to be rounded; and The rounding calculation unit is used to set the exponent in the high bit, combine it with the mantissa to be rounded, and then add it to the rounding bit to generate the floating-point integer.
2. The operational circuit according to claim 1, characterized in that, The mask generation unit is used to: When the exponent obtained by subtracting the preset bias from the exponent is greater than or equal to 0 and less than N, the difference is subtracted from N to obtain the shift number. When the exponent is greater than or equal to N, the shift number is set to 0; When the exponent is less than 0, the shift number is N; and The N-bit 1 value is shifted up by the shift number by one bit to generate the first mask.
3. The arithmetic circuit according to claim 2, wherein The mask generation unit is used to: When the exponent is equal to 0, the N-bit 1 value is shifted up by (N-1) bits to generate the second mask; and When the exponent is not equal to 0, the first mask is shifted 1 bit to the lower position and 1 is added to the higher position to generate an intermediate mask. Then, a bitwise XOR operation is performed between the intermediate mask and the first mask to generate the second mask.
4. The arithmetic circuit according to claim 2, wherein The mask generation unit is used to: When the exponent is less than 0, a supplementary bit with a value of 1 is generated; otherwise, a supplementary bit with a value of 0 is generated. The supplementary bit is set in the high bit and combined with the complement of the highest (N-1) bits of the first mask to form the third mask.
5. The operational circuit according to claim 1, characterized in that: In the first mask, each bit corresponding to the integer part of the mantissa is 1, and each bit corresponding to the fractional part of the mantissa is 0; In the second mask, the first bit after the decimal point of the mantissa is 1, and the other bits are 0; and In the third mask, all bits below the second decimal place corresponding to the mantissa are 1, and the other bits are 0.
6. The operational circuit according to claim 5, characterized in that, The rounding bit generation unit sets the protection bit to 1 when the exponent obtained by subtracting the preset bias from the exponent is -1; otherwise, it performs a bitwise AND operation on the mantissa and the second mask to generate a first operation result, and then performs an OR operation on each bit of the first operation result to generate the protection bit.
7. The arithmetic circuit according to claim 5, wherein The rounding bit generation unit sets the 1-bit value in the high bit and combines it with the highest (N-1) bits of the mantissa, then performs a bitwise AND operation with the second mask to generate the second operation result, and performs an OR operation on each bit of the second operation result to generate the lowest meaningful bit.
8. The arithmetic circuit according to claim 5, wherein The rounding bit generation unit sets the sticky bit to 0 when the exponent obtained by subtracting a preset bias from the exponent is less than -1 and the exponent is not 0; otherwise, it performs a bitwise AND operation on the mantissa and the third mask to generate a third operation result, and then performs an OR operation on each bit of the third operation result to generate the sticky bit.
9. The operational circuit according to claim 1, characterized in that, The floating-point number also includes a sign bit, and the rounding bit generation unit is further used for: When the rounding mode is RNE, the least significant bit and the sticky bit are ORed together, and then the bit is ANDed with the guard bit to generate the rounded bit. When the rounding mode is RTP, the guard bit and the sticky bit are ORed together, and then the complement of the sign bit is ANDed together to generate the rounding bit. When the rounding mode is RTN, the guard bit and the sticky bit are ORed together, and then ANDed with the sign bit to generate the rounding bit; and When the rounding mode is RTZ, the rounding bit is set to 0.
10. The arithmetic circuit according to claim 1, wherein The mantissa truncation unit performs a bitwise AND operation on the mantissa and the first mask to generate a first number, performs a bitwise AND operation on the complement of the first mask and N rounding bits to generate a second number, and performs a bitwise OR operation on the first number and the second number to generate the mantissa to be rounded.
11. An arithmetic method for rounding a floating-point number to a floating-point integer, characterized by, The floating-point number includes an N-bit mantissa and a P-bit exponent, where N and P are positive integers. The method includes: A first mask is generated based on the integer part and the fractional part of the mantissa, wherein the first mask is used to distinguish the integer part and the fractional part of the mantissa; A second mask is generated based on the first mask, wherein the second mask is used to identify the first digit after the decimal point of the mantissa; A third mask is generated based on the first mask, wherein the third mask is used to identify all bits below the second decimal place of the mantissa; The protection bit of the mantissa is obtained based on the second mask; The least significant bit of the mantissa is obtained based on the second mask; The sticky bit of the mantissa is obtained based on the third mask; The rounding bit is determined based on the rounding mode, the guard bit, the least significant bit, and the sticky bit. The decimal part of the mantissa is filled with the rounding bits to generate the mantissa to be rounded; and The exponent is set in the high bit and combined with the mantissa to be rounded, and then added to the rounding bit to generate the floating-point integer.
12. The method of claim 11, wherein, The steps for generating the first mask include: When the exponent obtained by subtracting the preset bias from the exponent is greater than or equal to 0 and less than N, the exponent is subtracted from N to obtain the shift number. When the exponent is greater than or equal to N, the shift number is set to 0; When the exponent is less than 0, the shift number is N; and The N-bit 1 value is shifted up by the shift number by one bit to generate the first mask.
13. The method of claim 12, wherein, The steps for generating the second mask based on the first mask include: When the exponent is equal to 0, the N bits of 1 are shifted up by (N-1) bits to generate the second mask; and When the exponent is not equal to 0, the first mask is shifted 1 bit to the lower position and 1 is added to the higher position to generate an intermediate mask. Then, a bitwise XOR operation is performed between the intermediate mask and the first mask to generate the second mask.
14. The method of claim 12, wherein, The steps for generating the third mask based on the first mask include: When the exponent is less than 0, a supplementary bit with a value of 1 is generated; otherwise, a supplementary bit with a value of 0 is generated. The supplementary bit is set in the high bit and combined with the complement of the highest (N-1) bits of the first mask to form the third mask.
15. The method according to claim 11, characterized in that: In the first mask, each bit corresponding to the integer part of the mantissa is 1, and each bit corresponding to the fractional part of the mantissa is 0; In the second mask, the first bit after the decimal point of the mantissa is 1, and the other bits are 0; and In the third mask, all bits below the second decimal place corresponding to the mantissa are 1, and the other bits are 0.
16. The method according to claim 15, characterized in that, The step of obtaining the protection bit of the mantissa based on the second mask includes: When the exponent obtained by subtracting the preset bias from the exponent is -1, the protection bit is set to 1; otherwise, a bitwise AND operation is performed on the mantissa and the second mask to produce a first operation result, and then an OR operation is performed on each bit of the first operation result to produce the protection bit.
17. The method according to claim 15, characterized in that, The step of obtaining the least significant bit of the mantissa based on the second mask includes: The value of 1 bit is set in the high bit and combined with the highest (N-1) bits of the mantissa. Then, a bitwise AND operation is performed with the second mask to produce the second operation result. And a OR operation is performed on each bit of the second operation result to produce the lowest meaningful bit.
18. The method of claim 15, wherein, The step of obtaining the sticky bit of the mantissa based on the third mask includes: When the exponent obtained by subtracting the preset bias from the exponent is less than -1 and the exponent is not 0, the sticky bit is set to 0; otherwise, a bitwise AND operation is performed on the mantissa and the third mask to produce a third operation result, and then an OR operation is performed on each bit of the third operation result to produce the sticky bit.
19. The method according to claim 11, characterized in that, The floating-point number also includes a sign bit. The step of determining the rounding bit based on the rounding mode, the protection bit, the least significant bit, and the sticky bit of the arithmetic circuit includes: When the rounding mode is RNE, the least significant bit and the sticky bit are ORed together, and then the bit is ANDed with the guard bit to generate the rounded bit. When the rounding mode is RTP, the guard bit and the sticky bit are ORed together, and then the complement of the sign bit is ANDed together to generate the rounding bit. When the rounding mode is RTN, the guard bit and the sticky bit are ORed together, and then ANDed with the sign bit to generate the rounding bit; and When the rounding mode is RTZ, the rounding bit is set to 0.
20. The method according to claim 11, characterized in that, The step of filling the decimal part of the mantissa with the rounding bits to generate the mantissa to be rounded includes: Perform a bitwise AND operation between the mantissa and the first mask to generate a first number; Perform a bitwise AND operation on the complement of the first mask and N rounding bits to produce a second number; and Perform a bitwise OR operation on the first number and the second number to produce the mantissa to be rounded.
Citation Information
Patent Citations
Apparatus and method for floating-point multiplication
CN106970776A