Order-matching shifting method and device and floating-point number operation method

By using the circuit structure based on the transmission gate array and the control signal of the transmission gate gate gate shift array in the floating-point number addition and subtraction circuit for the order shift operation, the problem of low computing performance caused by the complexity of the order shift operation in the prior art is solved, and the effect of improving the calculation frequency, energy efficiency and computing power density is achieved.

CN120215877APending Publication Date: 2025-06-27SHANGHAI JIANQI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375447.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The implementation method of order shift operations in the prior art is complex, resulting in low calculation performance of floating-point numbers, which has become a key bottleneck affecting the computing performance of the entire machine.

Method used

The circuit structure based on the transmission gate array is adopted, and the control signal of the transmission gate gate shift array is used, which is absolutely worth the difference between the order bits of the two floating point numbers, and the data logic level and path are reduced by compensating the order shift remainder.

Benefits of technology

The working frequency of floating-point number addition and subtraction circuits is significantly improved, the calculation energy efficiency is improved, the computing power density is improved, and the calculation accuracy is guaranteed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215877A_ABST
    Figure CN120215877A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of data processing, and discloses an order shifting method and device and a floating-point number operation method. In the invention, a circuit structure based on a transmission gate array is adopted in order shift operation, and the transmission gate array is controlled by combining a control signal of a transmission gate strobe shift array obtained by utilizing an absolute value of a difference between order digits of two floating-point numbers, so that data logic levels and paths are remarkably reduced; by performing precision compensation on the order shift remainder, the calculation precision can be ensured under the condition that the bit width of the adder is not increased. The effects of reducing the level of computational logic, further improving the working frequency of the floating-point number addition and subtraction circuit, improving the computational energy efficiency and improving the computational power density are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of data processing, and in particular, to a method and device for exponent alignment and shifting, and a floating-point number operation method. Background Art

[0002] A floating-point number consists of a sign bit, an exponent, and a mantissa. It provides a much larger value range and higher calculation accuracy than a fixed-point number, and is widely used in the fields of processors supporting floating-point calculations and computing power chips for artificial intelligence. The addition and subtraction calculation circuits of floating-point numbers often become important unit circuits that seriously affect key indicators such as calculation performance, calculation energy efficiency, and computing power density in processor chips and large computing power chips.

[0003] According to the definition of the IEEE-754 standard, the traditional floating-point number format is defined as Figure 1 shown. Whether it is double-precision (abbreviated as FP64), single-precision (abbreviated as FP32), or half-precision (abbreviated as FP16), it consists of a sign bit, an exponent segment, and a mantissa segment. Among them, S is the sign bit, E is the exponent segment, and M is the mantissa segment. The encodings of all 0s and all 1s in the exponent segment are not used as normal exponents, but are special identifiers reserved for operands to label absolute 0, subnormal numbers (the integer part of the significand is 0), infinity, and non-numeric data (abbreviated as NaN). The true value formula of this floating-point number is: N = (-1) s ×2 E-k ×(I + M), where: k = 2 w-1 -1 is the exponent offset value, which is related to the bit width W of the exponent code E, and is used to represent the actual positive and negative exponents with a k-bit unsigned number. I is the integer part of the significand, which takes 1 when the floating-point number is a normal number (the exponent bit is neither all 1s nor all 0s), and takes 0 when it is a subnormal number (the exponent bit is all 0s and the mantissa is not 0); M is the fractional part of the mantissa (significand), and the bit [x] from left to right represents 1 / 2 x , and the detailed definition of its standard can be found in the IEEE 754 technical standard specification.

[0004] The mantissas of traditional floating-point numbers are transmitted and stored in the form of a sign bit + original code. Therefore, in the traditional floating-point addition and subtraction calculation architecture, the mantissa is first complemented, and after the binary addition calculation is completed, the original code of the calculation result is obtained again. The conversion between these original codes and complement codes requires an operation of inverting each bit and then adding 1.

[0005] Compared with the addition and subtraction operations of fixed-point numbers, the mantissa segment of floating-point numbers needs to be exponent-aligned before addition and subtraction operations can be performed. This exponent alignment and shifting circuit is very complex and may greatly reduce the performance of floating-point addition and subtraction calculations.

[0006] The inventors found that there are at least the following problems in the related art: The implementation methods of alignment shift in the related art all adopt a multiplexer for one - of - many selection, or a cascaded manner of multi - stage selectors, or a multi - stage multi - input selector between the two. Since the range of shift bits involved in the alignment shift operation is relatively large, and the parallel operation of multi - bit mantissas is performed simultaneously, the data logic level of this circuit is very high, and the circuit logic is also very large, which often becomes the key bottleneck affecting the overall machine computing performance. Summary of the Invention

[0007] The purpose of the embodiments of the present invention is to provide an alignment shift method, device and floating - point operation method, which significantly reduce the data logic level and path. By using the alignment shift remainder for precision compensation, the calculation precision can be guaranteed without increasing the bit width. The effect of reducing the level of calculation logic is achieved, thereby increasing the working frequency of the floating - point addition and subtraction circuit, improving the calculation energy efficiency, and increasing the computing power density.

[0008] To solve the above - mentioned technical problems, an embodiment of the present invention provides an alignment shift method, including: obtaining a control signal for gating a shift array by using the absolute value of the difference between the exponent bits of two floating - point numbers; wherein, the control signal includes: a shift gating signal and a zero - filling signal for indicating a zero - filling operation; controlling the shift array gated by the transmission gate to perform exponent alignment on the mantissa bits of the floating - point number according to the control signal; wherein, the shift array gated by the transmission gate is composed of a plurality of the transmission gates; the shift gating signal is used to indicate that the transmission gate enters a conducting state or a closing state to control the shift operation of the mantissa bits of the floating - point number by the shift array gated by the transmission gate, and the zero - filling signal is used to perform a zero - filling operation on the empty bit positions formed after the shift operation.

[0009] An embodiment of the present invention further provides a floating-point operation method, including: aligning the exponent of the smaller mantissa for two floating-point numbers to be operated by the above-mentioned exponent alignment and shifting method to obtain a shifted remainder and a mantissa with aligned exponents; wherein, the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; obtaining a mantissa operation type control signal and a final calculation result sign bit through logical operations according to the sign bits of the two floating-point numbers, the floating-point operation operator type, and the absolute value size relationship between the two floating-point numbers; the mantissa operation type control signal is used to indicate whether to perform an inversion operation on the mantissa with aligned exponents; if the mantissa operation type control signal indicates to perform an inversion operation, then perform an unsigned fixed-point full addition operation using the larger mantissa and the mantissa after the inversion operation, otherwise directly perform an unsigned fixed-point full addition operation on the mantissa with aligned exponents and the larger mantissa; wherein, the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value size relationship between the two floating-point numbers; performing a normalization shift operation on the result of the unsigned fixed-point full addition operation and the shifted remainder according to the number of leading zeros in the result of the unsigned fixed-point full addition operation to obtain a normalized result; performing an overflow judgment and result normalization process on the normalized result by using the exponential adjustment result and the final calculation result sign bit based on the number of leading zeros for the larger exponent bit among the two floating-point numbers.

[0010] An embodiment of the present invention further provides a device for exponent alignment and shifting, including: a transmission gate control signal decoding circuit, configured to obtain a control signal for selecting and enabling a transmission gate array for shifting by using the absolute value of the difference between the exponent bits of two floating-point numbers; wherein, the control signal includes: a shift selection signal for indicating a shift operation and a zero-padding signal for indicating a zero-padding operation; a transmission gate selection and enabling shift array, configured to control the transmission gate selection and enabling shift array to perform exponent alignment on the mantissa bits of the floating-point number according to the control signal; wherein, the transmission gate selection and enabling shift array is composed of a plurality of the transmission gates; the shift selection signal is used to indicate that the transmission gate enters a conducting state or a non-conducting state to control the transmission gate selection and enabling shift array to perform a shift operation on the mantissa bits of the floating-point number, and the zero-padding signal is used to perform a zero-padding operation on the empty bit positions formed after the shift operation.

[0011] In the embodiment of the present invention, by adopting a circuit structure based on a transmission gate array in the exponent alignment and shifting operation, and combining the control of the transmission gate array by using the control signal of the transmission gate selection and enabling shift array obtained from the absolute value of the difference between the exponent bits of two floating-point numbers, the data logic levels and paths are significantly reduced. By using the exponent alignment and shifting remainder for precision compensation, the calculation precision can be guaranteed without increasing the bit width of the adder. The effect of reducing the level of calculation logic is achieved, thereby improving the working frequency of the floating-point addition and subtraction circuit, enhancing the calculation energy efficiency, and increasing the computing power density. Brief Description of the Drawings

[0012] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements. Unless otherwise stated, the figures in the drawings do not constitute a scale limitation.

[0013] Figure 1 is a schematic diagram of the definition of a traditional floating - point format provided according to the related art;

[0014] Figure 2 is a schematic diagram of the overall structure of a floating - point adder - subtractor provided according to the related art;

[0015] Figure 3 is a schematic diagram of the structure for implementing exponent alignment and shifting using a multiplexer provided according to the related art;

[0016] Figure 4 is a schematic diagram of the structure for implementing exponent alignment and shifting using a multi - stage two - way selector provided according to the related art;

[0017] Figure 5 is a flowchart of the exponent alignment and shifting method provided according to an embodiment of the present invention;

[0018] Figure 6 is a schematic diagram of the hardware structure of the exponent alignment and shifting and negation module provided according to another embodiment of the present invention;

[0019] Figure 7 is a schematic diagram of the hardware structure of the transmission - gate gated shift array provided according to an embodiment of the present invention;

[0020] Figure 8 is a flowchart of the floating - point operation method provided according to another embodiment of the present invention;

[0021] Figure 9 is an architecture diagram of a floating - point adder - subtractor using sign - bit processing provided according to another embodiment of the present invention;

[0022] Figure 10 is a flowchart of the hardware operation data provided according to another embodiment of the present invention;

[0023] Figure 11 is a block diagram of the absolute - value comparison and decision module provided according to another embodiment of the present invention;

[0024] Figure 12 is a schematic diagram of restoring the integer part of the mantissa provided according to another embodiment of the present invention;

[0025] Figure 13aIt is a schematic diagram of the alignment shift and negation module shift operation according to another embodiment of the present invention;

[0026] Figure 13b It is a schematic diagram of the alignment shift and negation module shift operation according to another embodiment of the present invention;

[0027] Figure 14 It is a circuit structure diagram of an unsigned fixed-point full adder according to another embodiment of the present invention;

[0028] Figure 15a It is a schematic diagram of the result mantissa normalization shift operation according to another embodiment of the present invention;

[0029] Figure 15b It is a schematic diagram of the result mantissa normalization shift operation according to another embodiment of the present invention;

[0030] Figure 15c It is a schematic diagram of the result mantissa normalization shift operation according to another embodiment of the present invention;

[0031] Figure 16 It is a structure diagram of the alignment shift device according to another embodiment of the present invention. Specific embodiments

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will elaborate on each embodiment of the present invention with reference to the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present invention, many technical details are provided to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented. The division of the following embodiments is for convenience of description and should not constitute any limitation to the specific implementation of the present invention. The various embodiments can be combined and cross-referenced with each other on the premise of no contradiction.

[0033] An embodiment of the present invention relates to a method for alignment shifting, which can be implemented in a shifting alignment module of a floating-point adder / subtractor using sign bit processing (an example of the architecture of a floating-point adder / subtractor using sign bit processing will be described in detail later in this article). In this embodiment, by adopting a circuit structure based on a transmission gate array in the alignment shifting operation, and combining the control of the transmission gate array with the control signal of the transmission gate selection shifting array obtained from the absolute value of the difference between the exponent bits of two floating-point numbers, the data logic levels and paths are significantly reduced. By using the alignment shifting remainder for precision compensation, the calculation precision can be guaranteed without increasing the bit width of the adder. The effect of reducing the level of calculation logic is achieved, thereby increasing the operating frequency of the floating-point adder / subtractor circuit, improving the calculation energy efficiency, and increasing the computing power density.

[0034] To facilitate the understanding of the application scenario of this application, an example of the overall structure of a conventional floating-point adder / subtractor is given as Figure 2 As shown, assuming that the exponent segments of two floating-point operands A and B are the A exponent bit and the B exponent bit respectively, it is necessary to compare the values of the A exponent bit and the B exponent bit through a shifting decision maker to determine the magnitude relationship of their exponents, and calculate the absolute difference between the two exponent values as the shifting bit. The shifting bit generates a shifting bit control signal through the alignment shifter; the first MUX selects the mantissa of the operand with the smaller exponent according to the comparison result "eA < eB" of the shifting decision maker and sends it to the alignment shifter for alignment shifting; the mantissa after alignment shifting is sent to the first two's complement obtaining module for transformation from the original code to the two's complement; the second MUX sends the mantissa of the operand with the larger exponent to the second two's complement obtaining module to complete the transformation from the original code to the two's complement; the operator control signal add / sub is sent to the two two's complement obtaining modules for the calculation of two's complement obtaining in combination with the sign bit; the mantissas of the operands in the two's complement format output from the two two's complement obtaining modules are sent to the fixed-point adder for signed binary fixed-point addition operation; the leading 0 / 1 counter counts the leading 0 (if the sign bit of the adder calculation result is 0) or leading 1 (if the sign bit of the adder calculation result is 1) of the output of the fixed-point adder, and sends the counting result to the mantissa left shift operation module for normalization shifting operation. After the result is converted to the original code format through the mantissa original code recovery module, it is sent to the result normalization module; the exponent adjustment module uses the result of the leading 0 / 1 counter to perform adjustment calculation on the larger exponent of the input operand, and the exponent adjustment result is sent to the result normalization module; the result normalization module performs overflow judgment processing on the adjusted result exponent and the result mantissa in the original code form according to the IEEE-754 specification requirements, and finally outputs the calculation result in the normalized floating-point format. Figure 2The fixed-point adder in the example shown uses a fixed-point adder for signed numbers to support binary addition or subtraction operations on mantissas where the sign bit may be negative. Since the numerical format of the mantissa segment in the IEEE-754 format is defined as the original code rather than the complement code, the mantissa segments in the input and output floating-point numbers of the floating-point adder / subtractor in the example are unsigned original codes. It is necessary to complete the conversion from the original code to the complement code before performing fixed-point addition and subtraction processing, and to complete the restoration calculation from the complement code to the original code before sending it for overflow judgment and output normalization processing. The series of operations of converting from the original code to the complement code before performing fixed-point addition and subtraction processing and restoring from the complement code to the original code before sending it for overflow judgment and output normalization processing require two levels of a total of three original code-complement code conversion (or vice versa) circuits.

[0035] At the same time, the implementation methods of the alignment shift module in the existing solutions all use a multiplexer for selection (attached Figure 3 - The existing structure that uses a multiplexer for selection to implement alignment shift), or the method of cascading multiple-level selectors (attached Figure 4 -- The existing structure that uses multiple-level two-way selectors to implement alignment shift), or the implementation using a multi-input selector between the two. Since the range of the shift bits involved in this module is relatively large and parallel operations are performed on multi-bit mantissas, the data logic level of this circuit is very high and the circuit logic is also very large, which often becomes the key bottleneck affecting the overall computing performance of the machine. For ease of understanding, this is further explained below.

[0036] Attached Figure 3It is the circuit structure of the alignment shifter in the floating-point adder / subtractor in related calculations. This circuit uses a multiplexer to implement the right shift operation. Each bit of the output mantissa Y[i:0] uses an independent multiplexer to select one bit from the possible bits of the input mantissa for output. The selection control signal is the input shift control signal SFT (in this figure, the maximum right shift range is 0 to 15 as an example, and the SFT bit width is 4 bits, that is, represented by SFT[3:0]); each multiplexer is a 16-to-1 selector; the input of the selector is 16-bit data, coming from 16 higher-order left-side data (including this position) starting from this position in the input mantissa. For example: the output mantissa Y[0] uses X[15:0] as the input of its selector, the output mantissa Y[1] uses X[16:1] as the input of its selector, and so on. The output mantissa Y[i] uses X[i + 15:i] as the input of its selector; when the bit selected by the selector input exceeds the actual data bit width boundary of the input mantissa X (including the high-order limit on the left and the low-order limit on the right), 0 is directly connected to the input end of this selector. For example, if the actual input mantissa X has a bit width of 24 bits, that is, X[23:0], all bits other than X[23:0] are directly connected to the selector with 0, that is, when i < 0 or i > 23, X[i] is directly replaced with 0.

[0037] Figure 4 In order to solve the fan-in and fan-out problems of the multi-input MUX in related calculations, a multi-stage 2-to-1 selector is used to implement the right shift operation. Figure 4 Each small circle in it represents a 2-input selector. The dotted line represents each bit of the shift control signal SFT, and this bit controls whether the output of each 2-to-1 selector connected to it selects the input from the left or the right; and so on. Each bit of the output data Y has passed through the path of 4 stages of 2-to-1 selectors. Different SFT values determine the data of each output bit to pass through the path of the multi-stage selector from a specific bit of the original input mantissa. As Figure 4 As shown by the arrow connection line in it, under the control of SFT[3:0] = 0101 (the control bit being 1 represents selecting the input on the left, and being 0 represents selecting the input on the right), Xi + 7 is right-shifted by 5 bits and output from Yi + 2. When the bit selected by the selector input exceeds the actual data bit width boundary of the input mantissa X (including the high-order limit on the left and the low-order limit on the right), 0 is directly connected to the input end of this selector. For example, if the actual input mantissa X has a bit width of 24 bits, that is, X[23:0], all bits other than X[23:0] are directly connected to the selector with 0, that is, when i < 0 or i > 23, X[i] is directly replaced with 0.

[0038] The implementation details of the alignment shift method in this embodiment will be specifically described below. The following content is only the implementation details provided for convenience of understanding and is not necessary for implementing this solution.

[0039] In step 501, the control signal for gating the shift array is obtained by using the absolute value of the difference between the exponent bits of two floating-point numbers. The control signal includes: a shift gating signal and a zero-padding signal for indicating a zero-padding operation.

[0040] In step 502, according to the control signal, the transmission gate gating shift array is controlled to perform exponent alignment on the mantissa bits of the floating-point number. The transmission gate gating shift array is composed of multiple transmission gates. The shift gating signal is used to indicate whether the transmission gate enters the conducting state or the off state, so as to control the transmission gate gating shift array to perform a shift operation on the mantissa bits of the floating-point number. The zero-padding signal is used to perform a zero-padding operation on the empty bit positions formed after the shift operation.

[0041] In an example, the transmission gates in the transmission gate gating shift array are divided into multiple groups. According to the control signal, the transmission gate gating shift array is controlled to perform exponent alignment on the mantissa bits of the floating-point number. For example, it can be: by sequentially indicating each group of transmission gates to enter the conducting state or the off state through the shift gating signal to indicate the mantissa bits to shift, obtaining the mantissa to be zero-padded and the remainder to be zero-padded. The transmission gates corresponding to the empty bit positions in the remainder to be zero-padded and the transmission gates corresponding to the empty bit positions in the mantissa to be zero-padded are all in the off state. The zero-padding signal for the mantissa in the zero-padding signal indicates that the zero-padding transmission gate queue in the transmission gate gating shift array performs a zero-padding operation on the empty bit positions in the mantissa to be zero-padded, obtaining the exponent-aligned mantissa bits. The zero-padding signal for the remainder in the zero-padding signal indicates that the zero-padding transmission gate queue in the transmission gate gating shift array performs a zero-padding operation on the empty bit positions in the remainder to be zero-padded, obtaining the shifted remainder.

[0042] Since there are various ways of dividing groups and the corresponding shift operations also vary, in an example, a more energy-saving shift method is provided. That is, the above method of sequentially indicating each group of transmission gates to enter the conducting state or the off state through the shift gating signal to indicate the mantissa bits to shift can be: by indicating a group of transmission gates corresponding to the shift number of bits to enter the conducting state through the shift gating signal and indicating all the remaining groups of transfer gates to enter the off state to indicate the mantissa bits to shift. The shift number of bits is obtained by the absolute value of the difference between the exponent bits of two floating-point numbers.

[0043] The number of transmission gates in the transmission gate gating shift array can be determined by itself. To achieve better results, in an example, the number of transmission gates in the transmission gate gating shift array, the bit width of the shifted remainder, and the bit width of the exponent-aligned mantissa bits are all related to the bit width of the absolute value of the difference between the exponent bits of two floating-point numbers and the bit width of the mantissa bits.

[0044] To better understand this embodiment, attached Figure 6 discloses a hardware structure of a mantissa shifting and negating module. The mantissa shifting and negating module is composed of sub-module 201 - transmission gate control signal decoding circuit module, sub-module 202 - transmission gate gated shifting array circuit, and a post-processing circuit for negating the output data and outputting the complement correction flag of CIN. The content of the hardware structure is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.

[0045] Figure 6 In the sub-module 201 - transmission gate control signal decoding circuit module, the shift number control signal eSHIFT is decoded to obtain the control signals of the transmission gate array: sel, yclr, and tclr. Their bit-width parameters are related to the shift range defined by eSHIFT. The logical relationship between these groups of control signals and the input eSHIFT is shown in Table 1 below:

[0046]

[0047] As an example, the shift control signal eSHIFT supported in Table 1 has a 4-bit width, and the supported right shift range is 0 to 15; the corresponding sel and yclr have a 16-bit width, while tclr has a 15-bit width.

[0048] Figure 6 In the sub-module 202 - transmission gate gated shifting array module, it is used to complete the specific shifting operation. This module is mainly implemented by forming an array of transmission gate units. The input operand with a smaller absolute value, the complete mantissa LIT, is sent into the transmission gate array as an operand. After passing through the path selected by the transmission gate control signal, it is output from the mantissa output ports Y and T. Among them, Y is the data output after the shifting operation for mantissa addition and subtraction operations, and T is used as the shift remainder to participate in the result mantissa normalization shifting operation for compensating the calculation accuracy. To simplify the operation of obtaining the complement code of the output data LITTLE before fixed-point addition, this module provides an optional data reverse output and the carry-in flag CIN required in the complement code calculation. The input signal sub comes from the sign bit operator processing module and is used to control whether to perform the complement code operation on the output data, as attached Figure 6As shown, the bitwise inversion operation is implemented using an exclusive-OR gate. When sub = 0, it means adding the original number, and the exclusive-OR gate directly outputs the output of the selector. When sub = 1, it means subtracting the original number, and the exclusive-OR gate outputs the inverted result of the output of the selector. To simplify the calculation logic, the increment operation for two's complement processing in this module only takes effect when eSHIFT = 0, that is, when the exponents of the two input operands are equal and no alignment shift of LIT is required. That is: when eSHIFT = 0 and sub = 1, CIN outputs 1, otherwise it outputs 0. The deviation in two's complement calculation in other cases is directly absorbed into the error of the remainder compensation. In an example, as a preferred design for improving calculation accuracy, the alignment shift and inversion module can output the remainder TAIL generated by the right shift operation for accuracy compensation during the mantissa normalization shift operation of the addition result. The data width of the TAIL output is determined by the bit width of eSHIFT. In an example, for the alignment shift and inversion module, when the LIT bit selected by the selector input exceeds the actual data bit width boundary of the input mantissa LIT (including the high-order limit on the left and the low-order limit on the right), 0 is directly used to access the input end of the selector. For example, if the bit width of the smaller mantissa LIT actually input to the module is 24 bits, that is, LIT[23:0], all bits other than LIT[23:0] are directly used 0 to access the selector, that is, when x < 0 or x > 23, LIT[x] is directly replaced by 0. In an example, for the alignment shift and inversion module, to keep the output remainder TAIL and the aligned mantissa LITTLE having the same two's complement attribute, the output of the remainder TAIL adopts the same inversion output control logic as LITTLE. In an example, for the alignment shift and inversion module, as one of the optimization options, when the overall floating-point calculation accuracy requirement does not require the remainder for accuracy compensation, the remainder TAIL is not output, and its corresponding generation circuit can be optimized to save circuit resources and area power consumption.

[0049] Furthermore, Figure 7 A circuit example of a transmission gate gated shift array is disclosed. The bit width of the input mantissa LIT supported by this example is 24 bits, and the supported alignment shift range is 0 to 15; the following description is only for the implementation details provided for easy understanding and is not necessary for implementing this solution. This circuit mainly includes two parts of circuit functions: a shift transmission gate array (with Figure 7 area A marked by the virtual box in the attachment) and an output zero-padding transmission gate queue (with Figure 7The B1, B2, and B3 regions in it). The data lines in the horizontal direction of the shift transmission gate array are the bits of the mantissa LIT of the input operand. Each row inputs 1 bit, and the bits are arranged from top to bottom from high to low. Its highest bit is the most significant bit LIT

[23] of LIT, and the lowest bit is LIT[0]; in the vertical direction of the shift transmission gate array is the output data after the shift operation. Each column outputs 1 bit. From left to right are the mantissa output Y and the remainder output T. Among them, Y is the higher significant bit and T is the lower significant bit. The leftmost highest bit is the most significant bit

[23] of Y, and the rightmost lowest bit is T[0]; the diagonal line from the upper left to the lower right in the shift transmission gate array is the control signal line sel of the transmission gate that controls the shift operation. Each diagonal line is a control bit, and the bits are arranged from left to right from high to low; its highest bit is the most significant bit

[15] of sel, corresponding to the selection of the transmission gate without shifting, and its lowest bit is the most significant bit [0] of sel, corresponding to the selection of the transmission gate shifted 15 bits to the right; the shift transmission gate array is used to complete the selection and connection of the input data line in the horizontal direction to the output data line in the vertical direction to form an output. This array includes a total of 16×24 = 384 transmission gates (taking the shift in the range of 0 to 15 for 24-bit data as an example), divided into 16 groups, and each group of 24 transmission gates is uniformly controlled by the same control signal sel[i] (the diagonal line in area A); when this control signal is 1, the 24 transmission gates it controls are simultaneously turned on, and all the other transmission gates in area A are turned off;

[0050] In one example, the output zero-padding transmission gate queue is used to make the control zero-padding transmission gate enter the conducting state through the corresponding zero-padding control signal when all the transmission gates in area A connected to the output data in each column are in the off state, and transmit the low level (zero state) to the output data line for output. According to the different output data types in the B1, B2, and B3 regions, the corresponding zero-padding transmission gate control signals are also different. The data output in the B1 region is the high significant bit segment Y[23:8] of the mantissa output, and its zero-padding transmission gate control signal is yclr[15:0]. The data output in the B3 region is the remainder output bit segment T[14:0], and its zero-padding transmission gate control signal is tclr[14:0]. The data output in the B2 region is the low significant bit segment Y[7:0] of the mantissa output, and its zero-padding transmission gate control signal is fixed at 0 to ensure that this zero-padding transmission gate always remains in the off state.

[0051] In one example, the bit widths of the high and low segments of the mantissa Y output in the B1 and B2 regions, the bit width of the remainder T output in the B3 region, and the size of the transmission gate array are all related to the bit width [n - 1:0] of the input mantissa LIT and the bit width [k:0] of the shift control signal eSHIFT. n and k can be other values. Attached Figure 7 Only take n = 24 and k = 3 as an example for description.

[0052] As one of the optimization options, when the overall floating-point calculation accuracy does not require precision compensation using remainders, the transmission gate circuit for generating remainders can be optimized to save circuit resources, area, and power consumption.

[0053] Appendix Figure 7 In the disclosed structural diagram, the high and low bit arrangements of the data lines in the horizontal and vertical directions and the control line data in the upper left to lower right direction are just one arrangement implementation of this circuit structure, and other arrangement orders can be adopted on the basis of maintaining the corresponding logical relationships.

[0054] In this embodiment, a circuit structure based on a transmission gate array is adopted in the alignment shift operation, and the control signal of the transmission gate selection shift array obtained by using the absolute value of the difference between the exponent bits of two floating-point numbers is combined to control the transmission gate array, thus significantly reducing the data logic levels and paths. By using the alignment shift remainder for precision compensation, the calculation accuracy can be guaranteed without increasing the bit width. It achieves the effect of reducing the level of calculation logic, thereby increasing the operating frequency of the floating-point addition and subtraction circuit, improving the calculation energy efficiency, and increasing the computing power density.

[0055] Another embodiment of the present invention relates to a floating-point operation method, which can be applied to electronic devices capable of performing floating-point operations such as mobile phones, computers, chips, etc., or can also be an aggregate composed of multiple electronic devices, and can also be carried on an architecture of a floating-point adder and subtractor using sign bit processing. In this embodiment, the sign bit processing is optimized by using the magnitude relationship information, sign bit information, and floating-point type of the floating-point number, simplifying the processing logic of the mantissa part operation in the floating-point addition and subtraction operation. A circuit structure based on a transmission gate array is adopted in the alignment shift operation, significantly reducing the data logic levels and paths. By using the alignment shift remainder to directly participate in the result mantissa normalization shift, the calculation accuracy is guaranteed without increasing the bit width of the adder. These improvements and optimizations significantly reduce the level of the overall logic, achieving the effects of increasing the operating frequency of the floating-point addition and subtraction circuit, improving the calculation energy efficiency, reducing the scale of the logic circuit, and increasing the computing power density. The implementation details of the floating-point operation method in this embodiment are specifically described below. The following content is only the implementation details provided for convenient understanding and is not necessary for implementing this solution.

[0056] The implementation details of the floating-point operation method in this embodiment are specifically described below. The flow of the floating-point operation method in this embodiment is as Figure 8As shown, in step 801, the two floating-point numbers to be operated are subjected to exponent alignment of the smaller mantissa through the above-mentioned exponent alignment and shifting method to obtain the shifted remainder and the mantissa with exponent alignment; where the smaller mantissa is the complete mantissa of the smaller floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; in step 802, according to the sign bits of the two floating-point numbers, the type of floating-point operation operator, and the magnitude relationship of the absolute values of the two floating-point numbers, a mantissa operation type control signal and the sign bit of the final calculation result are obtained through logical operations; the mantissa operation type control signal is used to indicate whether to perform a negation operation on the mantissa with exponent alignment; in step 803, if the mantissa operation type control signal indicates to perform a negation operation, an unsigned fixed-point full addition operation is performed using the larger mantissa and the negated smaller mantissa, otherwise the mantissa with exponent alignment and the larger mantissa are directly subjected to an unsigned fixed-point full addition operation; where the larger mantissa is the complete mantissa of the larger floating-point number obtained according to the magnitude relationship of the absolute values of the two floating-point numbers; in step 804, a normalization shift operation is performed on the result of the unsigned fixed-point full addition operation based on the number of leading zeros and the shifted remainder of the result of the unsigned fixed-point full addition operation to obtain a normalized result; in step 805, the exponent of the larger exponent bit in the two floating-point numbers is adjusted based on the number of leading zeros, and an overflow judgment and result normalization process are performed on the normalized result using the exponent adjustment result and the sign bit of the final calculation result.

[0057] To better understand the execution principle of the floating-point number operation method of the present invention, this embodiment provides an optimized architecture of a floating-point adder / subtractor as a carrier of the floating-point number operation method. The optimized architecture of the floating-point adder / subtractor is as Figure 9 shown. The adder / subtractor in this example can better complete the addition or subtraction operation of two floating-point numbers A and B compared to Figure 2 the example shown. It should be noted that Figure 9 this is an example of a relatively complete and detailed optimized architecture provided for easy understanding. The floating-point number operation method involved in the present invention can act on this structure, but it does not mean that the floating-point number operation method can only act on this structure. Figure 9The structure in [the above] consists of an absolute value comparison and decision module (labeled as Module 1), a mantissa alignment shift and negation module (labeled as Module 2, where the mantissa alignment shift method in the above embodiment operates, and step 801 in this embodiment can be correspondingly executed by this module), a sign bit processing module (labeled as Module 3, and step 802 in this embodiment can be correspondingly executed by this module), a unsigned fixed-point full adder (labeled as Module 4, and the fixed-point full addition operation in step 803 of this embodiment can be correspondingly executed by this module), a leading zero counter (labeled as Module 5), a result mantissa normalization shift module (labeled as Module 6, and step 804 in this embodiment can be correspondingly executed by this module), an exponent adjustment module (labeled as Module 7), and a result normalization processing module (labeled as Module 8, and step 805 in this embodiment can be correspondingly executed by Module 7 and Module 8). Each of these modules has corresponding functions and corresponding expansion methods, which will be discussed in detail separately in the following text.

[0058] As Figure 9 shown, the adder / subtractor can perform the addition or subtraction operation of two floating-point numbers A and B: Y = A + B or Y = A - B, where the input operands A, B, and the output result Y are all floating-point numbers that meet the IEEE-754 specification. SignA and SignB are the sign bits of A and B respectively, eA and eB are the exponent segments of A and B respectively, and MA and MB are the mantissa segments of A and B respectively. The data operation process corresponding to this architecture is as Figure 10As shown, for the input operands A and B, the operands A and B will be split and then sent to the absolute value comparison and decision module. The absolute value comparison and decision module receives the exponent bits eA / eB and mantissa bits MA / MB data of the two input operands, and after internal processing, outputs the complete mantissas MantA and MantB of the two operands with the hidden integer bits restored, as well as the comparison result ALTB representing the absolute value size relationship between the two operands A and B. In addition to directly comparing the absolute value sizes of two floating-point numbers, in one example, before step 801, the exponent bits of the two floating-point numbers can be compared first. When the exponent bits of the two floating-point numbers are equal, the absolute value size relationship between the two floating-point numbers is obtained by comparing the absolute values of the mantissa bits of the two floating-point numbers. This is to minimize the consumption of comparison operations as much as possible. The internal processing of the absolute value comparison and decision module includes, for example: comparing the sizes of the exponent bits, comparing the absolute values of the mantissas when eA = eB, restoring the complete mantissas MantA / MantB, calculating the larger exponent value, and calculating the absolute value of the exponent difference. The comparison result of the size relationship is used together with the operator for sign bit processing to output the final result sign bit and the flag indicating that the smaller number needs to be inverted. The smaller mantissa output by the absolute value comparison and decision module is used together with the final result sign bit output by the sign bit processing and the flag indicating that the smaller number needs to be inverted for the alignment shifter operation. Thus, the mantissa of the smaller operand after exponent alignment and the shift remainder are output through the alignment shifter operation, and it is determined whether to invert according to the sign bit processing result. The alignment shifter operation will obtain the shift remainder and the mantissa after alignment. The mantissa after alignment is used for unsigned addition of the mantissas with the larger mantissa output by the absolute value comparison and decision module, and then the leading 0 count of the result is performed, and the result normalization shift is performed with the shift remainder. The leading 0 count of the result can also be used for result exponent adjustment with the larger exponent obtained by the absolute value comparison and decision module. Finally, combining the results of sign bit processing, result normalization shift, and result exponent adjustment, the overflow judgment and result normalization processing are jointly completed.

[0059] The absolute value comparison and decision module (such as Figure 9 the module 1 shown), is mainly used to receive the exponent bits eA / eB and mantissa bits MA / MB data of the two input operands, and after internal processing, outputs the complete mantissas MantA and MantB of the two operands with the hidden integer bits restored, outputs the comparison result ALTB representing the absolute value size relationship between the two operands A and B, outputs the larger value eMAX = Max(eA - k, eB - k) of the exponents after subtracting the offset value k, and at the same time, this module also outputs an unsigned number eSHIFT representing the absolute value of the difference between the two exponent bits |eA - eB|. For example, the internal structure of the absolute value comparison and decision module is as Figure 11As shown, the absolute value comparison discriminator consists of sub-module 101 - shift discriminator, sub-module 102 - mantissa restorer, sub-module 103 - mantissa comparator, and a multiplexer MUX, and is used to complete the calculation and processing of the exponent bits of two input operands, the restoration and comparison of the complete mantissas of two input operands, output the larger exponent value eMAX after removing the exponent offset, the control signal eSHIFT for the exponent shift register, the gating control signal ALTB for the mantissa data path, and the complete mantissas MantA and MantB of the two operands for addition calculation.

[0060] Figure 11 The sub-module 101 - shift discriminator in it is used to complete the comparison and calculation operations of the exponent bits eA and eB of two operands, including: comparing whether eA is equal to eB, if equal, output the signal eq = 1, otherwise output eq = 0; comparing the magnitude relationship between eA and eB, if eA < eB, output the signal eaLTeb = 1, otherwise output eaLTeb = 0; calculating and outputting the absolute value of the difference between eA and eB, eSHIFT = |eA - eB|; outputting the larger value eMAX = MAX(eA - k, eB - k) after removing the exponent offset k from eA and eB, etc.

[0061] Figure 11 The sub-module 102 - mantissa restoration module in it is used to complement the hidden mantissa integer bit information according to the eA and eB values of the exponent segment. When the exponent bits are all 0, the integer bit of the corresponding restored complete mantissa is 0, otherwise the integer bit is 1. Taking operand A as an example, the operation schematic of this sub-module is as shown in the appendix Figure 12 shown;

[0062] Figure 11 The sub-module 103 - mantissa comparator in it is used to directly compare the magnitudes of the mantissa segment values MA and MB. If MA is less than MB, output 1, otherwise output 0. This comparison result will only be gated and output from port ALTB when the control signal eq = 1. When eq = 0, ALTB directly gates and outputs the exponent comparison result eaLTeb of sub-module 101, and the comparison result of sub-module 103 will be ignored. This path selection function is implemented by a multiplexer MUX.

[0063] The absolute value comparison and discrimination module (such as Figure 9The output signal ALTB of the module 1) shown is used to control the first MUX to send the complete mantissa LIT = Min(MantA, MantB) of the smaller operand to module 2 for alignment shift operation; and control the second MUX to send the complete mantissa BIG = Max(MantA, MantB) of the larger number to module 4 for addition operation. For example, when ALTB = 1, it means the absolute value of operand A (including the exponent bit and mantissa segment data) is less than the absolute value of operand B. This signal controls the first MUX to send the complete mantissa MantA of the smaller absolute value number A to the alignment shifter, and controls the second MUX to send the complete mantissa MantB of the larger absolute value number B to the fixed-point negation adder; when ALTB = 0, the data selection path is exactly the opposite; the ALTB signal is also directly used by module 3 to participate in the sign bit calculation to simplify its internal calculation logic; the output signal eSHIFT of module 1 is sent to module 2 to control the operation of the alignment shifter, and its value represents the number of shift bits for the right shift operation of the alignment shifter; the eMAX output by module 1 is sent to module 7 for finally calculating the exponent code of the output result.

[0064] The alignment shift and negation module (such as Figure 9 module 2) in it. This module completes the right shift operation of the input mantissa LIT according to the value of eSHIFT, and then performs the negation preprocessing of taking the two's complement on the output result according to the sub signal status output by module 3. The alignment shift operation is performed based on the transmission gate array structure, which can be: through the decoding circuit in the transmission gate array structure, decoding the shift bit control signal generated based on the difference between the large exponent and the small exponent to obtain the control signal of the transmission gate array; according to the path selected by the control signal of the transmission gate array, controlling the small mantissa to pass through the transmission gate selection shift array in the transmission gate array structure to obtain the initial shift remainder and the initially aligned small mantissa; using the processing result of the sign bit processing to perform the negation preprocessing of taking the two's complement on the initial shift remainder and the initially aligned small mantissa to obtain the shift remainder and the aligned small mantissa.

[0065] It is not difficult to find that the alignment shift and negation module in this embodiment is the corresponding module in the above alignment shift method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0066] In one example, as a preferred design for improving calculation accuracy, the exponent alignment shift and inversion module can output the remainder TAIL generated by the right shift operation, which is used for precision compensation during the mantissa normalization shift operation on the addition result. The data width of the TAIL output is determined by the bit width of eSHIFT. In one example, for the exponent alignment shift and inversion module, when the selected LIT bit of the selector input exceeds the actual data bit width boundary of the input mantissa LIT (including the high-order limit on the left and the low-order limit on the right), 0 is directly connected to the input end of the selector. For example, if the bit width of the smaller mantissa LIT of the actual input module is 24 bits, i.e., LIT[23:0], all bits other than LIT[23:0] are directly connected to the selector with 0. That is, when x < 0 or x > 23, LIT[x] is directly replaced with 0. In one example, for the exponent alignment shift and inversion module, to keep the output remainder TAIL and the mantissa LITTLE after exponent alignment have the same two's complement property, the output of the remainder TAIL adopts the same inversion output control logic as LITTLE. In one example, for the exponent alignment shift and inversion module, as one of the optimization options, when the overall floating-point calculation accuracy requirement does not need the remainder for precision compensation, the remainder TAIL is not output, and its corresponding generation circuit can be optimized to save circuit resources and area power consumption.

[0067] In one example, the data output by the exponent alignment shift and inversion module 2 is divided into two segments: high and low. Figures 13a to 13b Disclosed is the shift data correspondence relationship between the input mantissa LIT, the output high significant bit data segment LITTLE, and the low significant bit data segment TAIL; the mantissa LITTLE[n - 1:0] output after alignment shift maintains the same format as the input mantissa LIT[n - 1:0], and its data bit width is the same as that of the input mantissa; while the bit width of the output remainder TAIL[m - 1:0] is determined by the bit width k of eSHIFT[k:0], that is, m = 2^k - 1, and the bits exceeding the input bit width are filled with 0; for example: when the value of eSHIFT is 0, as Figure 13a shown, there is no shift and it is determined whether to invert bit by bit according to the state of the sub signal. When the bit width of eSHIFT is 4 bits, as Figure 13b shown, the corresponding maximum right shift amplitude is from 0 to 15, the data bit width of TAIL is 15, that is, m = 15, and it is determined whether to invert bit by bit according to the state of the sub signal.

[0068] In one example, the exponent alignment shift and inversion module (such as Figure 9The high significant bit segment LITTLE output by module 2) therein is sent to module 4 for unsigned fixed-point addition calculation; as an optional item to improve calculation accuracy, the low significant bit data segment TAIL output by module 2 can skip the adder and be directly sent to module 5 to participate in the leading zero count operation and be sent to module 6 to participate in the left shift operation of the mantissa of the calculation result.

[0069] In one example, the single-bit output signal CIN output by the exponent shifting and negation module (such as Figure 9 module 2 therein) is connected to module 4 as the two's complement plus 1 flag bit, and is used to correct the two's complement deviation of the input LITTLE participating in the fixed-point addition calculation when the exponent bits of the two operands are exactly equal (eA = eB) and sub = 1 (subtraction operation).

[0070] The sign bit processing module (such as Figure 9 module 3 therein) outputs the sign bit sign of the final calculation result and the mantissa operation type control signal sub through logical operations according to the sign bits SA and SB of the input operands A and B, the operation operator (add_sub), and the ALTB signal output by module 1. sub = 1 represents subtracting the mantissa of the smaller operand from the mantissa of the larger operand; sub = 0 represents adding the mantissa of the smaller operand to the mantissa of the larger operand. The input-output logical operation truth table of module 3 is as follows

[0071] shown in Table 2.

[0072]

[0073]

[0074] In one example, the output signal sub of module 3 is sent to module 2 to control the negation preprocessing of the two's complement conversion before the data after the exponent right shift operation is output; another output signal sign is the sign bit of the entire addition and subtraction operation result. sign = 0 represents that the final result is positive, and sign = 1 represents that the final output result is negative. This signal is sent to module 8 to be combined with the exponent bit and mantissa bit of the final calculation result to output in floating-point format.

[0075] In one example, the unsigned fixed-point full addition operation using the large mantissa and the negated small mantissa in step 804 can be: when the exponent bits of the two floating-point numbers are equal, perform an increment operation on the negated small mantissa, and perform an unsigned fixed-point full addition operation on the large mantissa and the incremented small mantissa; when the exponent bits of the two floating-point numbers are not equal, perform an unsigned fixed-point full addition operation on the large mantissa and the negated small mantissa, specifically as follows:

[0076] The unsigned fixed-point full adder (such as Figure 9Module 4) in it realizes the function of adding two complete mantissas BIG and LITTLE after order alignment, and the output result SUM[n:0] = BIG[n-1:0] + LITTLE[n-1:0], where the highest bit SUM[n] is the carry output bit; in one example, the adder is implemented by cascaded fixed-point full adders, and its internal logic architecture is as shown in the appendix Figure 14 As shown, the CIN signal output from Module 2 is connected to the lowest carry input of the full adder, supporting the addition operation of adding 1 after bitwise inversion for obtaining the two's complement of the smaller operand mantissa after alignment shift when the two exponents are equal; when the exponents of the two operands are not equal, the addition operation during the two's complement obtaining operation is absorbed as an error into the remainder TAIL output from the alignment shift.

[0077] In one example, obtaining the floating-point operation result according to the result of the mantissa operation, the larger exponent bit of the two floating-point numbers, and the sign bit processing result in step 805 can be: performing a normalization shift operation on the result of the unsigned fixed-point full addition operation based on the number of leading zeros in the result of the unsigned fixed-point full addition operation to obtain a normalization result; performing an exponent adjustment on the larger exponent bit of the two floating-point numbers based on the number of leading zeros, and performing an overflow judgment and result normalization processing on the normalization result using the final calculated result sign bit in the exponent adjustment result and the sign bit processing result of the two floating-point numbers. Specifically, it is manifested as:

[0078] Leading 0 counter (such as Figure 9 Module 5) in it is used to count the number of leading zeros in the output result SUM of Module 4, that is, to count how many "0"s there are before the first "1" appears in SUM[n:0] from the high bit to the low bit. For example, taking the 24-bit wide fixed-point full adder used in the architecture as an example, the output result SUM[24:0] of Module 4 is 0_00011100_11000010_10001101, and its highest bit SUM

[24] is the carry output bit, then the leading 0 count result LZCNT = 4, indicating that there are 4 zeros in front of the first 1 in SUM[24:0].

[0079] In one example, performing a normalization shift operation on the result of the unsigned fixed-point full addition operation based on the number of leading zeros in the result of the unsigned fixed-point full addition operation to obtain a normalization result can be: performing a left shift operation on the result of the unsigned fixed-point full addition operation using the number of leading zeros; performing a precision compensation on the result of the unsigned fixed-point full addition operation after the left shift operation using the shift remainder generated during the exponent alignment process to obtain a normalization result.

[0080] In one example, while the leading zero count result LZCNT is used by module 6 to control the left shift operation of the result mantissa, it is also sent to module 7 for correcting the exponent (order) of the final output result. That is, before performing the overflow judgment and result normalization processing, the result exponent of the large order is adjusted through the result of the leading zero count.

[0081] The result mantissa normalization shift module (such as Figure 9 module 6 in

[0082] is used to perform a normalization shift operation on the calculation result of module 4 to ensure that the most significant bit in the shifted output mantissa result MantiY is fixed at 1 and is the only integer bit of the mantissa, achieving the purpose of normalization.

[0083] In one example, Figure 8 step 304 in

[0084] can be: According to the large order and the standard order offset, convert the mantissa in the result normalization shift result into the form of a subnormal number, calculate the order segment represented by the floating-point number specification, hide the integer bit in the mantissa of the subnormal number form to form the mantissa segment represented by the floating-point number specification, and output the calculation result floating-point number of the overall floating-point adder / subtractor to implement the overflow judgment and result normalization processing. LZCNT

[0085] In one example, the shift bit number of the normalization shift operation is determined by the output signal LZCNT of module 5, that is, MantiY = SUM * 2 LZCNT Figure 15a Figure 15b Figure 15c Figure 15a Figure 15b Figure 15a Figure 15a shows the case where SUM[n] = 1, LZCNT = 0, and no shift is performed. Figure 15bIndicates the case where SUM[n, n-1]=01, LZCNT=1, and it is shifted left by 1 bit. Figure 15c Indicates the case where SUM[n:n-3]=0001, LZCNT=3, and it is shifted left by 3 bits.

[0086] The exponent adjustment module (such as Figure 9 Module 7 in) is used to adjust and correct the eMAX output from Module 1 to form an exponent value that matches the result mantissa MantY output after the result mantissa normalization shift operation. The output result eOUT = eMAX - LZCNT + 1.

[0087] The result normalization processing module (such as Figure 9 Module 8 in), this module is used to complete the overflow judgment and normalization processing of the calculation result, including judging that when eOUT <= -k (k is the exponent offset defined by the IEEE-754 standard), converting the mantissa into the form of a subnormal number, calculating the exponent segment eY = eOUT + k represented by the floating-point number specification, hiding the integer bits in the mantissa to form the mantissa segment represented by the floating-point number specification, and finally outputting the calculation result floating-point number Y of the overall floating-point adder / subtractor from this module to ensure that the processed result meets the requirements of the IEEE-754 specification.

[0088] In this embodiment, by using the size relationship information of floating-point numbers, the floating-point sign bit information, and the floating-point type, an optimized sign bit processing is performed, which simplifies the processing logic of the mantissa part operation in the floating-point addition and subtraction operation. In the alignment shift operation, a circuit structure based on a transmission gate array is adopted, which significantly reduces the data logic level and path. By using the alignment shift remainder to directly participate in the result mantissa normalization shift, the calculation accuracy is guaranteed without increasing the adder bit width. These improvements and optimizations significantly reduce the overall logic level, achieving the effects of increasing the working frequency of the floating-point addition and subtraction circuit, improving the calculation energy efficiency, reducing the scale of the logic circuit, and increasing the computing power density.

[0089] The step division of the above method is only for clear description. When implemented, it can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of this application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0090] Another embodiment of the present invention relates to an alignment shift device, such as Figure 16As shown in the figure, it includes: a transmission gate control signal decoding circuit 1601, which is used to obtain the control signal for gating the shift array by using the absolute value of the difference between the exponent bits of two floating-point numbers; wherein, the control signal includes: a shift gating signal for indicating a shift operation and a zero-padding signal for indicating a zero-padding operation; a transmission gate gating shift array 1602, which is used to control the transmission gate gating shift array to perform exponent alignment on the mantissa bits of the floating-point number according to the control signal; wherein, the transmission gate gating shift array is composed of multiple transmission gates; the shift gating signal is used to indicate that the transmission gate enters a conducting state or a closed state to control the transmission gate gating shift array to perform a shift operation on the mantissa bits of the floating-point number, and the zero-padding signal is used to perform a zero-padding operation on the empty bit positions formed after the shift operation.

[0091] In one example, the transmission gates in the transmission gate gating shift array are divided into multiple groups, and the above-mentioned controlling the transmission gate gating shift array to perform exponent alignment on the mantissa bits of the floating-point number according to the control signal can be: successively indicating each group of transmission gates to enter a conducting state or a closed state through the shift gating signal to indicate the mantissa bits to be shifted, so as to obtain the mantissa to be zero-padded and the remainder to be zero-padded; indicating the zero-padding transmission gate queue in the transmission gate gating shift array to perform zero-padding on the empty bit positions in the mantissa to be zero-padded through the mantissa zero-padding signal in the zero-padding signal to obtain the mantissa bits with exponent alignment; indicating the zero-padding transmission gate queue in the transmission gate gating shift array to perform zero-padding on the empty bit positions in the remainder to be zero-padded through the remainder zero-padding signal in the zero-padding signal to obtain the shifted remainder; wherein, the transmission gates corresponding to the columns of the empty bit positions in the remainder to be zero-padded and the transmission gates corresponding to the columns of the empty bit positions in the mantissa to be zero-padded are all in a closed state.

[0092] In one example, successively indicating each group of transmission gates to enter a conducting state or a closed state through the shift gating signal to indicate the mantissa bits to be shifted can be: indicating a group of transmission gates corresponding to the shift number of bits to enter a conducting state through the shift gating signal, and indicating all the remaining groups of transfer gates to enter a closed state to indicate the mantissa bits to be shifted; wherein, the shift number of bits is obtained by using the absolute value of the difference between the exponent bits of two floating-point numbers.

[0093] In one example, the number of transmission gates in the transmission gate gating shift array, the bit width of the shifted remainder, and the bit width of the mantissa bits with exponent alignment are all related to the bit width of the absolute value of the difference between the exponent bits of two floating-point numbers and the bit width of the mantissa bits.

[0094] In this embodiment, a circuit structure based on a transmission gate array is adopted in the alignment shift operation, and the control of the transmission gate array is combined with the control signal of the transmission gate selection and shift array obtained by using the absolute value of the difference between the exponent bits of two floating-point numbers. As a result, the data logic levels and paths are significantly reduced. By using the alignment shift remainder for precision compensation, the calculation precision can be guaranteed without increasing the bit width. The effect of reducing the level of the calculation logic is achieved, thereby increasing the operating frequency of the floating-point addition and subtraction circuit, improving the calculation energy efficiency, and increasing the computing power density.

[0095] It is not difficult to find that this embodiment is a device embodiment corresponding to the above alignment shift method embodiment, and this embodiment can be implemented in cooperation with the above method embodiment. The relevant technical details mentioned in the above method embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above method embodiment.

[0096] It is worth mentioning that each module involved in this embodiment is a logic module. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovative part of the present invention, units that are not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other units in this embodiment.

[0097] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made in form and details without departing from the spirit and scope of the present invention.

Claims

1. A method for shifting the order, characterized in that: include: The absolute value of the phase difference between the order bits of two floating point numbers is used to obtain a control signal for a transmission gate to select a shift array; wherein the control signal includes: a shift selection signal and a zero-filling signal for indicating a zero-filling operation; Control the transmission gate to select the shift array to perform exponent alignment on the mantissa bits of the floating point number according to the control signal; Among them, the transmission gate selection shift array is composed of a plurality of the transmission gates; the shift selection signal is used to instruct the transmission gate to enter an on state or an off state, so as to control the transmission gate selection shift array to perform a shift operation on the mantissa bits of the floating point number, and the zero-filling signal is used to perform a zero-filling operation on the empty bits formed after the shift operation.

2. The rank shifting method according to claim 1, characterized in that: The transmission gates in the transmission gate gating shift array are divided into a plurality of groups, and the transmission gate gating shift array is controlled according to the control signal to perform exponent alignment on the mantissa bits of the floating point number, including: Instructing each group of the transmission gates to enter a conducting state or a closing state in sequence through the shift selection signal to instruct the mantissa bits to shift, thereby obtaining a mantissa to be padded with zeros and a remainder to be padded with zeros; The mantissa zero-filling signal in the zero-filling signal instructs the transmission gate to select the zero-filling transmission gate queue in the shift array to fill the empty bits in the mantissa to be filled with zero, so as to obtain the mantissa bits with the order code aligned; The remainder zero-filling signal in the zero-filling signal instructs the transmission gate to select the zero-filling transmission gate queue in the shift array to fill the empty bits in the remainder to be filled with zeros with zeros, so as to obtain a shift remainder; Among them, each column of transmission gates corresponding to the empty bit positions in the to-be-filled zero remainder and each column of transmission gates corresponding to the empty bit positions in the to-be-filled zero mantissa are both in a closed state.

3. The rank shifting method according to claim 2, characterized in that: Instructing each group of transmission gates to enter a conducting state or a closing state in sequence through the shift selection signal to instruct the mantissa bits to shift, including: Instructing a group of transmission gates corresponding to the number of shifted bits to enter a conducting state through the shift selection signal, and instructing all the remaining groups of transmission gates to enter a closed state, so as to instruct the mantissa bits to shift; The shift bit number is obtained by the absolute value of the difference between the order bits of the two floating point numbers.

4. The rank shifting method according to claim 1, characterized in that: The number of transmission gates in the transmission gate selection shift array, the bit width of the shift remainder and the bit width of the mantissa bits aligned with the exponent are all related to the bit width of the absolute value of the difference between the exponent bits of the two floating point numbers and the bit width of the mantissa bits.

5. A floating point operation method, characterized in that: include: The two floating-point numbers to be calculated are subjected to exponent alignment of the little mantissas by the exponent shifting method described in any one of claims 1 to 4 to obtain a shift remainder and a little mantissa with the exponent aligned; wherein the little mantissa is the complete mantissa of the smaller floating-point number obtained according to the absolute value relationship of the two floating-point numbers; According to the sign bits of the two floating-point numbers, the floating-point operation operator type and the absolute value relationship of the two floating-point numbers, a mantissa operation type control signal and a final calculation result sign bit are obtained through logical operation; the mantissa operation type control signal is used to indicate whether to perform a negation operation on the little mantissa aligned with the exponent; If the mantissa operation type control signal indicates to perform a negation operation, an unsigned fixed-point full addition operation is performed using the large mantissa and the small mantissa after the negation operation, otherwise an unsigned fixed-point full addition operation is performed directly on the small mantissa and the large mantissa whose exponents are aligned; wherein the large mantissa is the complete mantissa of the larger floating-point number obtained according to the absolute value relationship of the two floating-point numbers; Performing a normalization shift operation on the result of the unsigned fixed-point full addition operation and the shift remainder according to the number of leading zeros of the result of the unsigned fixed-point full addition operation to obtain a normalized result; Based on the number of leading zeros, the larger exponent bit of the two floating point numbers is exponentially adjusted, and the normalized result is overflow-checked and normalized using the exponential adjustment result and the sign bit of the final calculation result.

6. The floating point operation method according to claim 5, wherein: The method of performing an unsigned fixed-point full addition operation using the big endian and the little endian after the inversion operation comprises: When the exponents of the two floating-point numbers are equal, performing an addition operation on the inverted little endian, and performing an unsigned fixed-point full addition operation on the big endian and the little endian that has been added; When the exponents of the two floating-point numbers are not equal, an unsigned fixed-point full addition operation is performed on the large mantissa and the small mantissa that has been inverted.

7. The floating point operation method according to claim 5, wherein: According to the result of the mantissa operation, the larger exponent bit of the two floating point numbers and the sign bit processing result of the two floating point numbers, a floating point operation result is obtained, including: Based on the number of leading zeros of the result of the unsigned fixed-point full addition operation, performing a normalization shift operation on the result of the unsigned fixed-point full addition operation and the shift remainder to obtain a normalized result; The larger exponent bit of the two floating point numbers is exponentially adjusted based on the number of leading zeros, and the normalized result is overflow-checked and normalized using the exponential adjustment result and the final calculation result sign bit in the sign bit processing results of the two floating point numbers.

8. The floating point number operation method according to claim 7, wherein: The step of performing a normalization shift operation on the result of the unsigned fixed-point full addition operation based on the number of leading zeros of the result of the unsigned fixed-point full addition operation to obtain a normalized result includes: Using the leading zero number, a left shift operation is performed on the result of the unsigned fixed-point full addition operation; The shift remainder generated in the process of exponent alignment is used to perform precision compensation on the result of the unsigned fixed-point full addition operation after the left shift operation to obtain a normalized result.

9. The floating point operation method according to any one of claims 5 to 8, characterized in that: The method further comprises: Before aligning the exponents of the little endian numbers using the absolute value of the difference in the exponents of the two floating point numbers, the exponents of the two floating point numbers are compared; When the exponents of two floating-point numbers are equal, the absolute value relationship of the two floating-point numbers can be obtained by comparing the absolute values ​​of the mantissas of the two floating-point numbers.

10. A step shifting device, characterized in that: include: A transmission gate control signal decoding circuit is used to obtain a control signal of a transmission gate selection shift array by using the absolute value of the phase difference between the order bits of two floating point numbers; wherein the control signal includes: a shift selection signal for indicating a shift operation and a zero padding signal for indicating a zero padding operation; A transmission gate enable shift array is used to control the transmission gate enable shift array to perform exponent alignment on the mantissa bits of a floating point number according to the control signal; wherein the transmission gate enable shift array is composed of a plurality of the transmission gates; the shift enable signal is used to instruct the transmission gate to enter an on state or an off state to control the transmission gate enable shift array to perform a shift operation on the mantissa bits of the floating point number, and the zero padding signal is used to perform a zero padding operation on the empty bits formed after the shift operation.