Floating-point arithmetic unit, floating-point processing device, method, chip and electronic device
By combining the logic processing module and the delay control module, the delay of the floating-point arithmetic unit can be controlled and the hardware resources optimized. This solves the flexibility problem of the floating-point arithmetic unit in different application scenarios in the prior art, and improves the computing efficiency and hardware processing capabilities.
Patent Information
- Application Number
- CN202510719274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing floating-point arithmetic units are difficult to adjust flexibly in different application scenarios, resulting in problems such as precision loss and high hardware complexity, which limit the improvement of the overall performance of floating-point arithmetic.
By introducing a logic processing module and a delay control module, the delay of the floating-point arithmetic unit can be made controllable. Combined with the calling of multiple floating-point arithmetic units and register optimization, a controllable pipeline structure and custom register design are adopted to optimize the utilization of hardware resources.
It effectively balances floating-point operation latency, improves operational efficiency, reduces power consumption, saves circuit area, enables mixed-precision design, and significantly enhances hardware processing capabilities.
Smart Images

Figure CN120233979B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of floating-point arithmetic technology, and in particular to a floating-point arithmetic unit, a floating-point processing device, a method, a chip, and an electronic device. Background Technology
[0002] With the rapid development of scientific computing and machine learning training, multiply-accumulate units capable of supporting floating-point data processing have emerged, such as floating-point multipliers, floating-point adders, and floating-point multiply-accumulators. They are widely used in scientific computing, digital signal processing, image processing, machine learning, and other fields.
[0003] In related technologies, floating-point arithmetic units typically employ a fixed latency structure, making it difficult to flexibly adjust for different application scenarios. Although some latency-controllable design schemes have been proposed, these schemes still have some drawbacks, such as precision loss and high hardware complexity, thus limiting the overall performance improvement of floating-point arithmetic. Summary of the Invention
[0004] This application proposes a floating-point arithmetic unit, a floating-point processing device, a method, a chip, and an electronic device, which can effectively balance the latency of floating-point operations, improve computational efficiency, and thus enhance the overall performance of floating-point operations.
[0005] To achieve the above objectives, the technical solution of this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a floating-point arithmetic unit, which includes:
[0007] The logic processing module is used to perform logical operations on the operands to be processed in the input floating-point arithmetic unit to obtain the operation results;
[0008] The delay control module is used to control the delay between the input logic processing module and the input floating-point arithmetic unit; and / or, to control the delay between the output calculation result of the logic processing module and the output calculation result of the floating-point arithmetic unit.
[0009] Secondly, embodiments of this application provide a floating-point processing apparatus, which includes a plurality of floating-point arithmetic units as described in the first aspect, wherein:
[0010] A floating-point processing unit is used to call a target number of floating-point arithmetic units among multiple floating-point arithmetic units to perform floating-point operations on the operands to be processed, based on the precision of the operands to be processed, and to obtain the calculation result; wherein, the target number is related to the precision of the operands to be processed.
[0011] Thirdly, embodiments of this application provide a floating-point processing apparatus, which includes registers and a plurality of floating-point arithmetic units as described in the first aspect, wherein:
[0012] A register is used to store at least two floating-point numbers in the same clock cycle;
[0013] The floating-point arithmetic unit is used to retrieve operands to be processed from registers.
[0014] Fourthly, embodiments of this application provide a floating-point processing method, the method comprising:
[0015] The logic processing module performs logical operations on the input operands to obtain the results.
[0016] The delay control module controls the delay between the input logic processing module and the input floating-point arithmetic unit for the operands to be processed; and / or controls the delay between the output calculation result of the logic processing module and the output calculation result of the floating-point arithmetic unit.
[0017] Fifthly, embodiments of this application provide a floating-point processing method, the method comprising:
[0018] Depending on the precision of the operand to be processed, a target number of floating-point arithmetic units are called from multiple floating-point arithmetic units, such as those in the first aspect, to perform floating-point operations on the operand to be processed and obtain the operation result; wherein, the target number is related to the precision of the operand to be processed.
[0019] Sixthly, embodiments of this application provide a floating-point processing method, the method comprising:
[0020] Registers store at least two floating-point numbers in the same clock cycle;
[0021] The floating-point arithmetic unit obtains the operand to be processed from the register so that the floating-point arithmetic unit can perform the floating-point processing method as described in the fourth aspect.
[0022] In a seventh aspect, embodiments of this application provide a chip, which includes: a floating-point arithmetic unit as described in the first aspect, or a floating-point processing device as described in the second aspect, or a floating-point processing device as described in the third aspect.
[0023] Eighthly, embodiments of this application provide an electronic device including a processor, wherein the processor includes: a floating-point arithmetic unit as described in the first aspect, or a floating-point processing device as described in the second aspect, or a floating-point processing device as described in the third aspect.
[0024] This application provides a floating-point arithmetic unit, floating-point processing device, method, chip, and electronic device. In the floating-point arithmetic unit, a logic processing module performs logical operations on the operands input to the floating-point arithmetic unit to obtain the operation result. A delay control module controls the delay between the input logic processing module and the input floating-point arithmetic unit for the operands to be processed; and / or controls the delay between the output operation result of the logic processing module and the output operation result of the floating-point arithmetic unit. In this way, for the floating-point arithmetic unit (e.g., a multiplication unit or an addition unit), the delay of floating-point operations can be effectively balanced, and the input and / or output delays throughout the processing can be controlled. Based on this delay-controllable strategy, hardware resource waste can be avoided, computational efficiency can be improved, and power consumption can be reduced. Furthermore, depending on the precision of the operands to be processed, a target number of floating-point arithmetic units can be called from multiple floating-point arithmetic units to perform floating-point operations on the operands to be processed to obtain the operation result. The target number is related to the precision of the operands to be processed. This approach also enables mixed-precision design, significantly improving hardware processing power while maintaining accuracy requirements. Furthermore, by reusing basic modules such as multiplication and addition units, power consumption and resource consumption are further reduced, saving circuit area. In addition, register design can be optimized, replacing standard registers in related technologies with custom devices. This not only enables more efficient data processing but also solves the power waste caused by excessive and unnecessary timing operations in related technologies, thereby improving the overall performance of floating-point operations. Attached Figure Description
[0025] Figure 1 A schematic diagram of the composition structure of a floating-point arithmetic unit provided in this application embodiment. Figure 1 ;
[0026] Figure 2 A schematic diagram of the composition structure of a floating-point arithmetic unit provided in this application embodiment. Figure 2 ;
[0027] Figure 3 A schematic diagram of the composition structure of a multiplication unit provided in an embodiment of this application;
[0028] Figure 4 A schematic diagram of the composition structure of an addition unit provided in an embodiment of this application;
[0029] Figure 5 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 1 ;
[0030] Figure 6A A schematic diagram of the multi-precision parameter configuration interface provided in the embodiments of this application. Figure 1 ;
[0031] Figure 6B A schematic diagram of the multi-precision parameter configuration interface provided in the embodiments of this application. Figure 2 ;
[0032] Figure 6C A schematic diagram of the multi-precision parameter configuration interface provided in the embodiments of this application. Figure 3 ;
[0033] Figure 6D A schematic diagram of the multi-precision parameter configuration interface provided in the embodiments of this application. Figure 4 ;
[0034] Figure 7 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 2 ;
[0035] Figure 8 A schematic diagram of the application framework for a floating-point processing device provided in the embodiments of this application. Figure 1 ;
[0036] Figure 9 A schematic diagram of the application framework for a floating-point processing device provided in the embodiments of this application. Figure 2 ;
[0037] Figure 10 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 3 ;
[0038] Figure 11 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 4 ;
[0039] Figure 12 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 5 ;
[0040] Figure 13 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment is shown in Figure 6.
[0041] Figure 14 A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 1 ;
[0042] Figure 15 A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 2 ;
[0043] Figure 16 A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 3 . Detailed Implementation
[0044] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.
[0045] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0046] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0047] It should also be noted that the terms "first, second, and third" used in the embodiments of this application are only used to distinguish similar objects and do not represent a specific order of objects. It is understood that "first, second, and third" can be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0048] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant terms and technologies of the embodiments of this application are described below. The following related technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and they all fall within the protection scope of the embodiments of this application.
[0049] The Floating Point Unit (FPU) is the structure that performs floating-point operations. The FPU is primarily responsible for performing mathematical operations involving floating-point numbers, such as basic arithmetic operations like addition, subtraction, multiplication, and division. Compared to the integer arithmetic unit, the FPU directly supports numerical calculations including decimal points through hardware circuitry, significantly improving the accuracy and speed in scenarios such as scientific computing and graphics rendering. Furthermore, the FPU can also handle transcendental function operations (such as trigonometric functions and logarithmic functions), further expanding its application scope.
[0050] Multiply Accumulate (MAC) is a special operation in digital signal processors or some microprocessors. The hardware circuit unit that implements this operation is called a "multiplier accumulator". This operation adds the product of the multiplication to the value of the accumulator A and then stores the result in the accumulator.
[0051] Not a Number (NAN) is a type of numeric data type in computer science that represents undefined or unrepresentable values. It is commonly used in floating-point arithmetic.
[0052] A leading zero counter (LZC) is a counter used to count the number of leading zeros in a binary number. In a leading zero counter, the counter increments by 1 when the most significant bit of the input binary number is 0, and stops counting when the most significant bit is 1.
[0053] It should be understood that floating-point numbers are mainly composed of three parts: the sign bit, the exponent (exp), and the mantissa, and their encoding format is shown in Table 1.
[0054] Table 1
[0055]
[0056] Taking single-precision floating-point numbers as an example, a single-precision floating-point number occupies 32 bits and is divided into the following three parts:
[0057] The sign bit (1 bit) is used to indicate whether the floating-point number is positive or negative; 0 indicates a positive number and 1 indicates a negative number.
[0058] The exponent bit, also known as the "exponent code" (8 bits), is used to represent the exponent part of a floating-point number; it uses offset representation, where the actual exponent value is the stored value minus 127 (the offset).
[0059] The mantissa (23 bits) represents the fractional part of the floating-point number.
[0060] In this application embodiment, floating-point numbers include multiple formats, namely: normal numbers, subnormal numbers, and special numbers. Special numbers can include positive and negative zero, positive and negative infinity, and NOT numbers. Specifically, the exponent and mantissa bits of positive and negative zero are all 0; the exponent bit of positive and negative infinity is all 1 and the mantissa bit is all 0; the exponent bit of NOT numbers is all 1 and the mantissa bit is not 0; the exponent bit of subnormal numbers is all 0 and the mantissa bit is not 0; and the rest are represented as normal numbers.
[0061] It should also be understood that in floating-point arithmetic, units that simultaneously perform multiplication and addition operations, such as multiply-accumulate units (or "floating-point multiply-accumulate units"), are core components for processing real number operations in modern computer systems. They are widely used in scientific computing, digital signal processing, and image processing. The design and implementation of these units must comply with the IEEE 754 standard, which defines the representation and operation rules for floating-point numbers. With the development of scientific computing and machine learning, the demand for multi-precision floating-point operations is increasing. Traditional fixed-point multipliers have a fixed number of input bits, making it difficult to meet the requirements of multi-precision computation. Therefore, methods supporting multi-precision floating-point multiplication have emerged, such as reconfigurable floating-point multiply-accumulate units. These units can dynamically adjust precision as needed, improving hardware utilization and reducing bit redundancy. In terms of hardware implementation, the design of floating-point arithmetic units needs to strike a balance between speed, area, and power consumption. In other words, the following main challenges need to be addressed in implementing high-speed, high-precision floating-point multipliers:
[0062] 1. Timing Optimization and Pipeline Design: The greater the pipeline depth, the less computation is required at each stage. However, pipeline buffers and data dependencies can also lead to performance bottlenecks. Balancing pipeline depth and throughput is a key design issue.
[0063] 2. Power Consumption and Area Control: High-performance multipliers often employ complex parallel adder structures, such as Wallace trees, but this results in a larger silicon area and higher power consumption. Low-power design is particularly important in mobile terminals and embedded systems, thus requiring optimized logic circuits and clock gating techniques.
[0064] 3. Error Control and Rounding Strategy: Floating-point operations require results to meet the precision requirements of the IEEE 754 standard. Rounding errors and inappropriate rounding methods may occur during mantissa multiplication and normalization. To ensure numerical accuracy, a multi-level rounding verification mechanism is typically employed, with sophisticated control circuitry implemented in hardware.
[0065] It should also be understood that in the field of modern high-performance computing, the design of floating-point multipliers and adders faces stringent requirements regarding latency, power consumption, and area. Traditional floating-point arithmetic units typically employ fixed-latency structures, making it difficult to flexibly adjust them for different application scenarios. To address this challenge, researchers have proposed various latency-controllable floating-point multiplier and adder design schemes in recent years.
[0066] For example, a multi-channel floating-point multiply-accumulator architecture divides the processing into four data paths by analyzing the relationships between operands and the type of operations. This adapts to different situations, avoids unnecessary processing steps, and thus improves computational speed and reduces power consumption. However, this design may increase hardware complexity, thereby affecting area and power consumption. As another example, researchers have proposed a high-precision, low-power approximate floating-point multiplier design based on partial product probability analysis to optimize the latency of floating-point multipliers. This design analyzes the probability of a partial product being 1 and proposes an approximate 4-2 compressor and low-bit OR gate compression method, effectively reducing hardware resource consumption, power consumption, and latency. However, this design may suffer from error accumulation in some high-precision applications.
[0067] In other words, although some floating-point multiply-accumulate unit designs with controllable latency already exist, these designs still have some drawbacks, such as precision loss and high hardware complexity, which limit the overall performance improvement of floating-point operations.
[0068] Based on this, embodiments of this application provide a floating-point arithmetic unit, a floating-point processing device, a method, a chip, and an electronic device. For floating-point arithmetic units (e.g., multiplication units or addition units), the latency of floating-point operations can be effectively balanced, and the latency of input and / or output during the entire processing can be controlled. This latency-controllable strategy can also avoid wasting hardware resources, improve computational efficiency, and reduce power consumption. Furthermore, depending on the precision of the operands to be processed, a target number of floating-point arithmetic units can be called from multiple floating-point arithmetic units to perform floating-point operations on the operands to be processed. This also enables mixed-precision design, thereby significantly improving the hardware's processing power while ensuring precision requirements. Moreover, based on the reuse of basic modules such as multiplication and addition units, power consumption and resource consumption are further reduced, saving circuit area. In addition, the design of registers can be optimized, replacing standard registers in related technologies with custom devices. This not only enables more efficient data processing but also solves the power waste caused by excessive unnecessary timing operations in related technologies, thereby improving the overall performance of floating-point operations.
[0069] The embodiments of this application will now be described in detail with reference to the accompanying drawings.
[0070] In one embodiment of this application, Figure 1 A schematic diagram of the composition structure of a floating-point arithmetic unit provided in this application embodiment. Figure 1 .like Figure 1 As shown, the floating-point arithmetic unit 10 may include a logic processing module 101 and a delay control module 102, wherein:
[0071] The logic processing module 101 is used to perform logical operations on the operands to be processed in the input floating-point arithmetic unit 10 to obtain the operation results;
[0072] The delay control module 102 is used to control the delay between the input logic processing module 101 and the input floating-point arithmetic unit 10; and / or, to control the delay between the output calculation result of the logic processing module 101 and the output calculation result of the floating-point arithmetic unit 10.
[0073] In this embodiment, the logic processing module 101 performs logical operations, such as multiplication or addition, on the operands input to the floating-point arithmetic unit 10 to obtain the corresponding results. The delay control module 102 can control the delay before the operands are input to the logic processing module 101, and / or control the delay before the operands are output to the floating-point arithmetic unit 10; thereby effectively balancing the delay of the logical operation, improving the operational efficiency of the floating-point arithmetic unit 10, and ensuring accurate result output.
[0074] In this embodiment, the floating-point unit 10 can adopt a pipelined structure, dividing the entire floating-point processing into several pipeline registers, which can effectively avoid performance bottlenecks caused by excessive pipeline delays. For example, the entire floating-point processing can be divided into: an input pipeline register, at least one intermediate pipeline register, and an output pipeline register, thereby improving pipeline efficiency and saving time.
[0075] In some embodiments, see Figure 2 The delay control module 102 may include an input pipeline register 1021 and an output pipeline register 1022, wherein:
[0076] Input pipeline register 1021 is used to control the computation path delay of at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed; and / or,
[0077] Output pipeline register 1022 is used to control the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0078] It should be noted that in the pipelined structure, the input pipeline register 1021 and the output pipeline register 1022 are controllable. That is, whether the input pipeline register 1021 and the output pipeline register 1022 are enabled can be controlled according to the corresponding configuration parameters in order to achieve delay balance.
[0079] In one possible implementation, the input pipeline register 1021 is used to control whether at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed is timed according to the first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, initial exponent bit, and initial mantissa bit meets the first delay requirement.
[0080] In another possible implementation, for the output pipeline register 1022, the output pipeline register 1022 is used to control whether at least one of the target sign bit, target exponent bit and target mantissa bit corresponding to the operation result is timed according to the second configuration parameter, so that the output path delay of at least one of the target sign bit, target exponent bit and target mantissa bit meets the second delay requirement.
[0081] In this embodiment, the first delay requirement and the second delay requirement are preset and used to measure whether the path delay of each segment meets the requirements. Additionally, the first configuration parameter and the second configuration parameter can be designed according to the overall requirements of the floating-point unit 10 to control whether the corresponding pipeline register is enabled.
[0082] For example, the input pipeline register 1021 can be controlled to be enabled or disabled according to the first configuration parameter, that is, to control whether at least one of the initial sign bit, initial exponent bit and initial mantissa bit of the input is timed according to the first configuration parameter; the output pipeline register 1022 can be controlled to be enabled or disabled according to the second configuration parameter, that is, to control whether at least one of the target sign bit, target exponent bit and target mantissa bit of the output is timed according to the second configuration parameter.
[0083] It should also be noted that, in the embodiments of this application, for the controllable input pipeline register 1021 and output pipeline register 1022, both pipeline registers can be turned on, or both can be turned off, or only one pipeline register can be turned on (for example, turn on input pipeline register 1021 and turn off output pipeline register 1022; or turn on output pipeline register 1022 and turn off input pipeline register 1021). No limitation is made here.
[0084] In other words, in this embodiment, the input pipeline register and output pipeline register of the floating-point arithmetic unit 10 can be controllably designed according to corresponding configuration parameters. Here, the input pipeline register can control whether to press the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, and the output pipeline register can control whether to press the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result according to the second configuration parameter. In this way, when the delay of a certain path is large, the delay balance is achieved by pressing the other paths.
[0085] In some embodiments, see continue to see Figure 2 At least one intermediate pipeline register 1011 is provided between the input pipeline register 1021 and the output pipeline register 1022, wherein:
[0086] At least one intermediate pipeline register 1011 is used to perform logical operations on the initial sign bit, initial exponent bit and initial mantissa bit corresponding to the operand to be processed, and to determine the target sign bit, target exponent bit and target mantissa bit corresponding to the operation result.
[0087] In this embodiment, the operands to be processed include a first operand and a second operand. Here, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit; the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit. The target sign bit can be obtained by performing logical operations on the first and second initial sign bits, the target exponent bit can be obtained by performing logical operations on the first and second initial exponent bits, and the target mantissa bit can be obtained by performing logical operations on the first and second initial mantissa bits.
[0088] Thus, in this embodiment of the application, the logical processing can be divided into at least one intermediate pipeline register, and computational tasks can be allocated between each intermediate pipeline register, thereby optimizing the execution time of floating-point processing, avoiding performance bottlenecks caused by excessive pipeline delays, and improving computational efficiency.
[0089] Understandably, in the embodiments of this application, the floating-point arithmetic unit 10 can be a multiplication unit (or "floating-point multiplication unit", "multiplier", etc.), and / or, the floating-point arithmetic unit 10 can also be an addition unit (or "floating-point addition unit", "adder", etc.). The floating-point arithmetic unit 10 will be described in detail below with the multiplication unit and the addition unit as examples.
[0090] In one possible implementation, the floating-point unit 10 can be a multiplication unit. In this case, at least one intermediate pipeline register can include a pipeline register (i.e., a first intermediate pipeline register). That is, the multiplication unit adopts a three-stage pipelined design, such as an input pipeline register (first-stage pipeline), a first intermediate pipeline register (second-stage pipeline), and an output pipeline register (third-stage pipeline). Moreover, the input pipeline register and the output pipeline register can be controlled to perform timing according to corresponding configuration parameters. Therefore, the multiplication unit can be called a "controllable three-stage pipelined multiplier".
[0091] In other words, in the multiplication unit, an intermediate pipeline register is set between the input pipeline register and the output pipeline register, and the input pipeline register and the output pipeline register are controllable. That is, the input pipeline register can control whether to press the clock on at least one of the initial sign bit, initial exponent bit and initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, and the output pipeline register can control whether to press the clock on at least one of the target sign bit, target exponent bit and target mantissa bit corresponding to the operation result according to the second configuration parameter.
[0092] Furthermore, in this embodiment of the application, the operands to be processed include a first operand and a second operand, wherein the first operand can be a multiplicand and the second operand can be a multiplier; or, the first operand can be a multiplier and the second operand can be a multiplicand; no limitation is made here. For example, assuming the first operand is the multiplicand x and the second operand is the multiplier y, then the result of the operation C = x × y.
[0093] For example, Figure 3 This is a schematic diagram illustrating the structural composition of a multiplication unit provided in an embodiment of this application. For example... Figure 3 As shown, from a pipeline perspective, this can include: an input pipeline register (first-stage pipeline) A1, a first intermediate pipeline register (second-stage pipeline) A2, and an output pipeline register (third-stage pipeline) A3.
[0094] In some embodiments, see continue to see Figure 3 The logic processing module 101 may include a first logic module (such as...) Figure 3 301-1 or 301-2 in the second logic module 302, wherein:
[0095] The first logic module is used to perform logical operations on the first initial sign bit and the second initial sign bit, as well as the first initial exponent bit and the second initial exponent bit, respectively, to determine the target sign bit and the target exponent bit.
[0096] The second logic module is used to perform a multiplication operation on the first initial mantissa and the second initial mantissa to determine the target mantissa.
[0097] In this embodiment, the second logic module 302 is located in the first intermediate pipeline register (i.e., the second-level pipeline). For the first logic module (such as... Figure 3 For example, in the case of 301-1 or 301-2, when the input pipeline register (i.e., the first-level pipeline) is enabled, the first logic module 301-1 is located in the input pipeline register (i.e., the first-level pipeline); or when the input pipeline register (i.e., the first-level pipeline) is disabled, the first logic module 301-2 is located in the first intermediate pipeline register (i.e., the second-level pipeline). In other words, the position of the first logic module can be related to whether the input pipeline register (i.e., the first-level pipeline) is enabled.
[0098] In some embodiments, taking the first logic module 301-1 as an example, see further. Figure 3 The first logic module 301-1 may include a sign bit processing unit 3011 and an exponent bit processing unit 3012. The sign bit processing unit 3011 is used to perform a sign XOR operation on the first initial sign bit and the second initial sign bit to obtain the target sign bit; the exponent bit processing unit 3012 is used to perform an exponent addition operation on the first initial exponent bit and the second initial exponent bit to obtain the target exponent bit.
[0099] In this embodiment, the target sign bit can be obtained by performing a sign XOR operation on the initial sign bits of the two operands (e.g., the first operand and the second operand). If the two sign bits are the same, the sign XOR result is false (0), indicating that the target sign bit is positive; otherwise, if the two sign bits are different, the sign XOR result is true (1), indicating that the target sign bit is negative.
[0100] In this embodiment, the target exponent can be obtained by performing exponential addition on the initial exponents of two operands (e.g., the first operand and the second operand). Here, the exponent is obtained by shifting based on an offset (bias). For example, bias = 2. (k-1) -1, where k is the number of bits corresponding to the exponent. Thus, if the operand is a single-precision floating-point number, k=8, and the bias is 127; if the operand is a half-precision floating-point number, k=5, and the bias is 15; if the operand is a double-precision floating-point number, k=11, and the bias is 1023.
[0101] In other words, in this embodiment of the application, for the exponent bit of each operand, if the operand is a single-precision floating-point number, then the exponent bit is used for 8-bit storage, and its storage format is the sum of the value and 127; if the operand is a half-precision floating-point number, then the exponent bit is used for 5-bit storage, and its storage format is the sum of the value and 15; otherwise, if the operand is a double-precision floating-point number, then the exponent bit is used for 11-bit storage, and its storage format is the sum of the value and 1023.
[0102] For example, assuming the initial exponent bits corresponding to the two operand cells are E1 and E2 (i.e., the actual values of the binary exponent fields), the exponent addition operation during multiplication is: Actual exponent = (E1 - bias) + (E2 - bias) = E1 + E2 - 2 × bias. Finally, this calculation result needs to be added back with the single offset to obtain the target stored exponent value (i.e., the target exponent bit), which can be E1 + E2 - bias. It is important to note that in the multiplication operation, the two actual exponents are actually added together, not the stored exponents.
[0103] It should also be noted that, for each operand, in addition to the initial sign bit and initial exponent bit, a corresponding flag bit can be introduced to indicate whether the operand is a special number, such as zero, NAN, positive or negative infinity, or overflow. For example, the first operand may also include a first initial flag bit to indicate whether the first operand is a special number; the second operand may also include a second initial flag bit to indicate whether the second operand is a special number.
[0104] In some embodiments, taking the first logic module 301-1 as an example, see below. Figure 3 The first logic module 301-1 may further include a flag processing unit 3013. The flag processing unit 3013 is used to perform logical operations based on the first initial flag and the second initial flag to determine the target flag corresponding to the operation result. The target flag is used to indicate whether the operation result is a special number.
[0105] In other words, in this embodiment of the application, for each operand, in addition to the initial sign bit, initial exponent bit and initial mantissa bit, an initial flag bit (flag(zero, nan, overflow)) is also included to determine whether the corresponding operand is a special number.
[0106] For example, when at least one of the first and second operands is a special number, the following explanations illustrate several cases:
[0107] If the first operand is NAN and the second operand is any value, then the result C is NAN, meaning that any value and NAN will result in NAN.
[0108] If the first operand is zero and the second operand is a non-zero finite number, then the result C is zero, that is, the result of multiplying zero by a finite number is still 0. The target sign bit is determined by XORing the sign bits of the two operands.
[0109] If the first operand is zero and the second operand is positive or negative infinity, the result C is NAN, meaning that multiplying zero by infinity is an undefined operation and returns NAN.
[0110] If the first operand is positive or negative infinity and the second operand is a non-zero finite number, then the result C is positive or negative infinity, that is, the product of infinity and a finite non-zero number is infinity. The target sign bit is determined by XORing the sign bits of the two operands.
[0111] If the first operand is positive or negative infinity, and the second operand is positive or negative infinity, then the result C is positive or negative infinity. That is, multiplying infinity with the same sign results in positive infinity, and multiplying infinity with opposite signs results in negative infinity.
[0112] It can also be understood that in the second logic module 302, the target mantissa can be obtained by performing mantissa operations on the corresponding mantissas of the multiplicand and multiplier. In some embodiments, the second logic module 302 may include a mantissa product unit 3021 and a rounding operation unit 3022. The mantissa product unit 3021 is used to perform mantissa multiplication on the first initial mantissa and the second initial mantissa to obtain a first mantissa product; the rounding operation unit 3022 is used to perform a rounding operation on the first mantissa product to determine the target mantissa.
[0113] In this embodiment, for mantissa multiplication, the implicit bits of the mantissa bits corresponding to each operand need to be restored first. For example, if the operand is a normalized number, an implicit leading 1 is added to the corresponding mantissa bits to form the actual mantissa 1.M (e.g., if the stored mantissa M=101, the actual mantissa value is 1.101); if the operand is a non-normalized number, an implicit leading 0 is added to the corresponding mantissa bits to form the actual mantissa 0.M. Then, the actual mantissas of the two operands (e.g., 1.M1 and 1.M2) are multiplied using an unsigned binary multiplication operation, resulting in a result twice the length of the mantissa bits; for example, if it is single precision (23-bit mantissa), the first mantissa product after multiplication is 46 bits; if it is double precision (52-bit mantissa), the first mantissa product after multiplication is 104 bits.
[0114] For example, suppose the mantissa of the single-precision multiplicand is 1.0100... and the mantissa of the single-precision multiplier is 1.1000..., then the first mantissa product is 1.0100... × 1.1000... = 1.1110...; at this point, further rounding operation is needed on the first mantissa product.
[0115] In the embodiments of this application, rounding operations may include truncation rounding, rounding up, rounding down, or rounding to the nearest number. Truncation rounding, also known as "rounding towards 0," directly truncates extra digits regardless of their value and without changing the sign of the number. Rounding up, also known as "rounding towards positive infinity," rounds up for positive numbers if the extra digits are non-zero, and truncates negative numbers. Rounding down, also known as "rounding towards negative infinity," adjusts negative numbers to a smaller value if the extra digits are non-zero, and truncates positive numbers. Rounding to the nearest number, also known as "rounding towards even numbers," prioritizes the closest rounding value; if the value is in the middle (e.g., extra digits are 1000...), it rounds towards the nearest even number. For example, for the floating-point number 2.5 (10.1 in binary), the value of the rounding operation is 2; for the floating-point number 1.375 (1.011 in binary, with 2 mantissas retained), the value of the rounding operation is 1.4.
[0116] In some embodiments, the rounding operation unit 3022 is further configured to perform a rounding operation on the first mantissa product to obtain a second mantissa product when the operation result is an irregular number; and to perform normalization processing on the second mantissa product to obtain a target mantissa digit when the second mantissa product exceeds a preset range, and to perform a carry operation on the target exponent digit so that the multiplication result is transformed from an irregular number to a normal number.
[0117] It should be noted that, in this embodiment, whether the calculation result is a non-standard number can be determined based on whether the target exponent is all zeros, and whether the second mantissa product exceeds a preset range can refer to whether the second mantissa product overflows. Specifically, if the second mantissa product does not exceed the preset range, then no exponent adjustment is needed, i.e., no carry operation is needed to the target exponent; if the second mantissa product exceeds the preset range, then the exponent needs adjustment, i.e., a carry operation is needed to the target exponent. In other words, after rounding, the result of multiplying the two initial mantissas may have the following two outcomes:
[0118] Case 1: The product of the second mantissa after rounding does not overflow (e.g., 1.111...111 → 1.000...000), in which case there is no need to adjust the exponent.
[0119] Case 2: The product of the second mantissa after rounding overflows (e.g., 1.111...111 + 1 → 10.000...000). In this case, it is necessary to shift right by 1 bit, that is, add 1 to the exponent (carry).
[0120] It should also be noted that, in this embodiment, the rounding operation can be performed according to the IEEE-754 rule. For the first mantissa product, the processing of normalized and non-normalized numbers are two different branches and cannot be combined. For example, in one possible implementation, if the result is a normalized number, carry-over (specifically, carry-over to the target exponent, i.e., incrementing the target exponent by 1) is performed when the guard bit (G), round bit (R), and sticky bit (S) satisfy preset combination conditions. In another possible implementation, if the result is a non-normalized number, after rounding the first mantissa product, if carry-over to the target exponent (i.e., incrementing the target exponent by 1) is performed, the target exponent will not be all zeros, thus transforming the result from a non-normalized number to a normalized number. Therefore, after performing mantissa multiplication on the initial mantissa bits of the two operands, when dealing with underflow numbers, the above situation needs to be taken into account, and the case of non-standard carry needs to be handled separately when assigning the final target exponent bit.
[0121] It should also be noted that, in this embodiment, the multiplication unit can adopt a controllable three-stage pipelined structure. Specifically, the input pipeline register (i.e., the first-stage pipeline) can be enabled or disabled based on a first configuration parameter, i.e., the input timing is controlled based on the first configuration parameter; the output pipeline register (i.e., the third-stage pipeline) can be enabled or disabled based on a second configuration parameter, i.e., the output timing is controlled based on the second configuration parameter. In one possible implementation, if both the input and output pipeline registers are disabled, then the multiplication unit can be considered a single-stage pipelined structure.
[0122] It should also be noted that in this embodiment, the multiplication unit adopts a controllable three-stage pipelined structure. This is because the result of mantissa multiplication requires one cycle to obtain. Therefore, a first intermediate pipelined register is inserted between the input pipelined register and the output pipelined register, allowing the two initial mantissa bits to be multiplied directly. At this time, the result of mantissa multiplication arrives simultaneously with the target flag bit, target sign bit, and target exponent bit. This is because multiplication involves complex computational logic, requiring multiple levels of logic gates and a longer computational path compared to addition and XOR gates, resulting in greater latency. Here, timing can be applied to other paths to achieve latency balance.
[0123] In another possible implementation, the floating-point unit 10 can be an adder. In this case, at least one intermediate pipeline register can include three pipeline registers (i.e., a first intermediate pipeline register, a second intermediate pipeline register, and a third intermediate pipeline register). That is, the adder unit adopts a five-stage pipelined design, such as an input pipeline register (first-stage pipeline), a first intermediate pipeline register (second-stage pipeline), a second intermediate pipeline register (third-stage pipeline), a third intermediate pipeline register (fourth-stage pipeline), and an output pipeline register (fifth-stage pipeline). Moreover, the input and output pipeline registers can be controlled by whether to use timing according to corresponding configuration parameters. Therefore, the adder unit can be called a "controllable five-stage pipelined adder".
[0124] In other words, in the addition unit, three intermediate pipeline registers are set between the input pipeline register and the output pipeline register, and the input pipeline register and the output pipeline register are also controllable. In this embodiment, the input pipeline register and the output pipeline register are the same as those in the multiplication unit, and whether they are enabled can be determined by the corresponding configuration parameters. For example, the input pipeline register can control whether at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed is clocked according to the first configuration parameter, and the output pipeline register can control whether at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result is clocked according to the second configuration parameter.
[0125] Furthermore, in the embodiments of this application, the operands to be processed include a first operand and a second operand, wherein the first operand can be the augend and the second operand can be the addend; or, the first operand can be the addend and the second operand can be the augend; however, which operand is the augend can be determined by the size of the exponent (i.e., the "initial exponent bit") of the operand. For example, assuming the first operand is the augend x and the second operand is the addend y, then the operation result C = x + y.
[0126] For example, Figure 4 This is a schematic diagram illustrating the structural composition of an addition unit provided in an embodiment of this application. For example... Figure 4 As shown, from a pipeline perspective, this can include: input pipeline register (first stage pipeline) B1, first intermediate pipeline register (second stage pipeline) B2, second intermediate pipeline register (third stage pipeline) B3, third intermediate pipeline register (fourth stage pipeline) B4 and output pipeline register (fifth stage pipeline) B5.
[0127] In some embodiments, see continue to see Figure 4The logic processing module 101 may include a first logic module 401, a second logic module 402, a third logic module 403, and a fourth logic module 404, wherein:
[0128] The first logic module 401 is used to determine the exponent difference between the first initial exponent bit and the second initial exponent bit and the smaller exponent of the two, and to perform a shift operation on the mantissa bit corresponding to the smaller exponent bit according to the exponent difference so that the first initial mantissa bit and the second initial mantissa bit are aligned.
[0129] The second logic module 402 is used to determine the sign XOR result of the first initial sign bit and the second initial sign bit, and to perform an addition operation on the aligned first initial mantissa bit and the second initial mantissa bit according to the sign XOR result to obtain the first mantissa sum and the target sign bit.
[0130] The third logic module 403 is used to perform leading zero processing on the first mantissa sum to determine the leading zero statistical result and the shifted second mantissa sum;
[0131] The fourth logic module 404 is used to perform an exponent shift operation based on the statistical results of leading zeros to obtain the target exponent; and to perform a rounding operation on the second mantissa to obtain the target mantissa.
[0132] It should be noted that, for the first and second operands, in addition operations, the augend and addend must first be determined based on the exponents of the operands. For example, if the first initial exponent is greater than the second initial exponent, i.e., the exponent of the first operand is larger, then the first operand can be determined as the augend, and the second operand as the addend; otherwise, if the second initial exponent is greater than the first initial exponent, i.e., the exponent of the second operand is larger, then the second operand can be determined as the augend, and the first operand as the addend. Additionally, embodiments of this application may also set register flags to indicate whether the augend and addend need to be interchanged.
[0133] It should also be noted that the slight difference in parsing two operands between the addition and multiplication units is that operand parsing can only proceed after the operand positions have been determined in the previous step, but no operations are performed on the individual parts of the two operands at this stage. In other words, the addition unit needs to determine the operand positions by comparing the exponents before parsing, while the multiplication unit does not require comparing the exponents before parsing.
[0134] It should also be noted that the first logic module 401 is located in the input pipeline register (i.e., the first-level pipeline), the second logic module 402 is located in the first intermediate pipeline register (i.e., the second-level pipeline), the third logic module 403 is located in the second intermediate pipeline register (i.e., the third-level pipeline), and the fourth logic module 404 is located in the third intermediate pipeline register (i.e., the fourth-level pipeline). If the input pipeline register (i.e., the first-level pipeline) is closed, then the first logic module 401 is located in the first intermediate pipeline register (i.e., the second-level pipeline). In other words, the position of the first logic module 401 is related to whether the input pipeline register (i.e., the first-level pipeline) is enabled.
[0135] In some embodiments, see continue to see Figure 4 The first logic module 401 may include an exponent subtraction unit a1, a first control unit a2, a first selection unit (MUX1) a3, a second selection unit (MUX2) a4, a third selection unit (MUX3) a5, and a mantissa shifting unit a6, wherein:
[0136] The exponent subtraction unit a1 is used to perform a subtraction operation on the first initial exponent bit and the second initial exponent bit to determine the exponent difference and transmit the exponent difference to the first control unit a2.
[0137] The first control unit a2 is used to receive the exponential difference and generate a first control signal to be transmitted to the first selection unit a3, a second control signal to be transmitted to the second selection unit a4, and a third control signal to be transmitted to the third selection unit a5 based on the exponential difference.
[0138] The first selection unit a3 is used to receive a first control signal and select a larger exponent from the first initial exponent bit and the second initial exponent bit according to the first control signal.
[0139] The second selection unit a4 is used to receive the second control signal and select the mantissa corresponding to the smaller exponent according to the second control signal in the first initial mantissa and the second initial mantissa.
[0140] The third selection unit a5 is used to receive the third control signal and select the mantissa corresponding to the larger exponent according to the third control signal in the first initial mantissa and the second initial mantissa.
[0141] The mantissa shifting unit a6 is used to shift the mantissa bit corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa bit and the second initial mantissa bit.
[0142] In this embodiment, the first logic module 401 is used to align the initial mantissa bits of two operands based on the exponent difference between them. Specifically, the exponent subtraction unit a1 determines the exponent difference (i.e., "exponent difference") between the first and second initial exponent bits. Then, the first control unit a2 generates a first control signal, a second control signal, and a third control signal based on the exponent difference. The first control signal selects the larger exponent from the first and second initial exponent bits; the second control signal selects the mantissa bit corresponding to the smaller exponent from the first and second initial mantissa bits; and the third control signal selects the mantissa bit corresponding to the larger exponent from the first and second initial mantissa bits. Thus, based on the exponent difference, the mantissa bit corresponding to the smaller exponent is right-shifted by the mantissa shift unit a6, thereby aligning the first and second initial mantissa bits.
[0143] In other words, in the addition unit, the smaller exponent must first be determined, and the mantissa corresponding to the smaller exponent must be right-shifted to align the two initial mantissas. During this process, the smaller exponent in the two operands must also be adjusted to the larger one, meaning the two exponents are equal in size.
[0144] For example, suppose the exponent of the first operand is 3 and the mantissa is 1.101 (the actual value is 1.101 × 2). 3 (Converted to decimal, this is 13); the exponent of the second operand is 1, and the mantissa is 1.110 (the actual value is 1.110 × 2). 1 (Converted to decimal, this is 3.5). Since the exponent of the second operand is smaller, and the difference between the two exponents is 3-1=2, it is necessary not only to add 2 to the exponent of the second operand to align it with the larger exponent 3, but also to shift the mantissa of the second operand to the right by 2 bits, that is, to make the mantissa of the second operand 0.0111, so as to align the decimal points of the two mantissas.
[0145] In some embodiments, the second logic module 402 may include an XOR unit b1, a second control unit b2, and a mantissa addition unit b3, wherein:
[0146] The XOR unit b1 is used to perform a sign XOR operation on the first initial sign bit and the second initial sign bit to obtain the sign XOR result, and then transmit the sign XOR result to the second control unit b2.
[0147] The second control unit b2 is used to receive the sign XOR result and generate a fourth control signal to be transmitted to the mantissa addition unit b3 based on the sign XOR result.
[0148] The mantissa addition unit b3 is used to perform addition operations on the aligned first initial mantissa bits and second initial mantissa bits according to the fourth control signal to obtain the first mantissa sum and the target sign bit.
[0149] In this embodiment, the second logic module 402 mainly calculates the first mantissa sum between the aligned first initial mantissa bits and the second initial mantissa bits. Specifically, the XOR unit b1 performs a signed XOR operation based on the initial sign bits of the two operands to determine the signed XOR result. This result is then transmitted to the second control unit b2, which in turn transmits a fourth control signal to the mantissa addition unit b3. The fourth control signal determines whether the aligned first initial mantissa bits and the second initial mantissa bits are to be added or subtracted to obtain the first mantissa sum. Based on this first mantissa sum, the target sign bit corresponding to the operation result is determined.
[0150] It should also be noted that for two operands (e.g., the first operand and the second operand), if the first initial sign bit and the second initial sign bit are the same, then the sign XOR result can be determined to be false (0); otherwise, if the two sign bits are different, then the sign XOR result can be determined to be true (1).
[0151] In this embodiment, if the XOR result is false (0), the fourth control signal is used to instruct the aligned first initial mantissa and second initial mantissa to perform an addition operation, and the target sign bit remains unchanged, that is, the target sign bit is the same as the sign bit of the first operand or the second operand. If the XOR result is true (1), the fourth control signal is used to instruct the aligned first initial mantissa and second initial mantissa to perform a subtraction operation, and the target sign bit is determined by the sign bit with the larger absolute value among the aligned first initial mantissa and second initial mantissa. For example, if the first initial sign bit is positive and the second initial sign bit is negative, and the absolute value of the first initial mantissa is larger among the aligned first initial mantissa and second initial mantissa, then the target sign bit can be determined to be positive; otherwise, if the absolute value of the second initial mantissa is larger among the aligned first initial mantissa and second initial mantissa, then the target sign bit can be determined to be negative.
[0152] In other words, in this embodiment of the application, if the sign XOR result is false (0), then the first initial mantissa and the second initial mantissa after alignment can be added to obtain the first mantissa sum, and the target sign bit remains unchanged; if the sign XOR result is true (1), then the first initial mantissa and the second initial mantissa after alignment can be subtracted, specifically by subtracting the initial mantissa with the smaller absolute value from the initial mantissa with the larger absolute value, and the target sign bit is determined by the sign bit with the larger absolute value among the first initial mantissa and the second initial mantissa after alignment.
[0153] In some embodiments, the third logic module 403 may include a leading zero processing unit c1. The leading zero processing unit c1 is used to predict leading zeros in the first mantissa sum to determine the leading zero statistical result; and to shift the first mantissa sum according to the leading zero statistical result to obtain a second mantissa sum.
[0154] In this embodiment of the application, for the leading zero processing unit c1, a strategy of step-by-step shifting and judging zero is adopted to determine the case where the high bits of the corresponding stage are all zero, so as to obtain the leading zero statistical result; then, the first mantissa sum is shifted according to the leading zero statistical result to obtain the second mantissa sum, and the highest bit of the mantissa of the second mantissa sum is 1.
[0155] In other words, if the sum of the first mantissas after addition has many zeros at the beginning, then LZC (Least Normalized Conversion) needs to be introduced, and the first mantissa sum needs to be left-shifted. For example, this applies to the addition of denormalized numbers, or when the sign bits of the two operands are different, resulting in a very small first mantissa sum. For instance, the first operand has an exponent of 2 and a mantissa of 1.1; the second operand has an exponent of 1 and a mantissa of 1.1. After alignment, the exponent of the second operand is adjusted to 2, and the mantissa is shifted one bit to the right, meaning the mantissa of the second operand becomes 0.11. When adding the two mantissas, in binary, 1.1 plus -0.11 equals 0.11, which is 0.75 in decimal. Since the first mantissa sum is 0.11, normalization is needed, i.e., shifting it left by one bit to become 1.1, and simultaneously decreasing the exponent by 1, i.e., changing it from 2 to 1.
[0156] In this process, the number of leading zeros (i.e., the leading zero count) determines the number of bits to shift left. For example, if the first mantissa sum is 0.00101, the leading zero count is 3. In this case, the first mantissa sum needs to be shifted left by three bits to get 1.01 (the second mantissa sum), while the exponent is reduced by 3.
[0157] In some embodiments, the fourth logic module 404 may include an exponent shift unit d1 and a rounding operation unit d2. The exponent shift unit d1 is used to receive the larger exponent sent by the first selection unit and the leading zero statistics sent by the leading zero processing unit c1, and to perform an exponent shift operation on the larger exponent based on the leading zero statistics to obtain the target exponent digit; the rounding operation unit d2 is used to perform a rounding operation on the second mantissa to obtain the target mantissa digit.
[0158] In this embodiment, the exponent shifting unit d1 mainly performs exponent shifting based on the leading zero statistical result sent by the leading zero processing unit c1. Here, "exponent shifting" refers to subtracting the exponent size. For example, if the first mantissa sum is 0.00101, the leading zero statistical result is 3. In this case, not only does the first mantissa sum need to be shifted left by three bits to obtain 1.01 (the second mantissa sum), but the larger exponent also needs to be subtracted, that is, the larger exponent is reduced by 3 to obtain the target exponent.
[0159] In the embodiments of this application, the rounding operation may include truncation rounding, rounding up, rounding down, or rounding to the nearest integer. The rounding operation is similar to the rounding operation in the multiplication unit, and will not be described in detail here.
[0160] In one possible implementation, the rounding operation unit d2 is also used to round the second mantissa sum to obtain the third mantissa sum when the result of the operation is a non-standard number; when the third mantissa sum exceeds a preset range, the third mantissa sum is normalized to obtain the target mantissa digit, and the target exponent digit is adjusted by adding 1 so that the result of the operation is transformed from a non-standard number to a standard number.
[0161] It should be noted that, in this embodiment, whether the calculation result is a non-standard number can be determined based on whether the target exponent is all zeros, and whether the third mantissa sum exceeds a preset range can refer to whether the third mantissa sum overflows. Specifically, if the third mantissa sum does not exceed the preset range, then no exponent adjustment is needed, i.e., no carry-over operation to the target exponent is required; if the third mantissa sum exceeds the preset range, then the exponent needs to be adjusted, i.e., a carry-over operation to the target exponent is required. In other words, after rounding the second mantissa sum, the following two situations may exist:
[0162] Case 1: The third digit after rounding does not overflow (e.g., 1.111...111 → 1.000...000), in which case no adjustment of the exponent is needed.
[0163] Case 2: The third mantissa after rounding overflows (e.g., 1.111...111 + 1 → 10.000...000). In this case, it is necessary to shift right by 1 bit, that is, add 1 to the exponent (carry).
[0164] It should also be noted that in the embodiments of this application, the processing of the result obtained from the addition operation and the processing of the normalized number and the non-normalized number are also two different branches and cannot be combined. For example, in one possible implementation, if the result is a normalized number, when the guard bit (G), round bit (R), and sticky bit (S) respectively meet the preset combination conditions, carry-over (specifically, carry-over to the target exponent bit, i.e., increment the target exponent bit by 1). In another possible implementation, if the result is a non-normalized number, after rounding the second mantissa, if carry-over to the target exponent bit is performed, i.e., incrementing the target exponent bit by 1, then the target exponent bit will not be all zeros, thereby transforming the result from a non-normalized number to a normalized number.
[0165] It should also be noted that, for Figure 4 The addition unit shown has the following connections: the output of the exponent subtraction unit a1 is connected to the first control unit a2; the first output of the first control unit a2 is connected to the control of the first selection unit a3; the second output of the first control unit a2 is connected to the control of the second selection unit a4; the third output of the first control unit a2 is connected to the control of the third selection unit a5; the fourth output of the first control unit a2 is connected to the control of the mantissa shift unit a6; the output of the second selection unit a4 is connected to the input of the mantissa shift unit a6; the outputs of the mantissa shift unit a6 and the third selection unit a5 are respectively connected to the two inputs of the mantissa addition unit b3; the output of the XOR unit b1 is connected to the input of the second control unit b2; the output of the second control unit b2 is connected to the control of the mantissa addition unit b3; the output of the mantissa addition unit b3 is connected to the input of the leading zero processing unit c1; and the output of the leading zero processing unit c1 is connected to the exponent shift unit d1 and the rounding operation unit d2, respectively, to obtain the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0166] Understandably, in the embodiments of this application, using custom devices to improve performance is a very effective strategy, especially with significant advantages in power consumption optimization. Specifically, by employing custom registers, the design can be precisely adjusted according to actual needs, thereby reducing system power consumption.
[0167] In some embodiments, the input pipeline register is used to merge and store the data to be pressed in the initial sign bit, initial exponent bit and initial mantissa bit corresponding to the operand to be processed within the same clock cycle; the output pipeline register is used to merge and store the data to be pressed in the target sign bit, target exponent bit and target mantissa bit corresponding to the operation result within the same clock cycle; and the intermediate pipeline register is used to merge and store the intermediate results of logical operations on the initial sign bit, initial exponent bit and initial mantissa bit corresponding to the operand to be processed within the same clock cycle.
[0168] In other words, the registers (such as input pipeline registers, intermediate pipeline registers, and output pipeline registers) in both multiplication and addition units can be replaced with custom-designed devices. This represents a precise optimization to address timing bottlenecks. Custom-designed registers not only process data more efficiently but also combine multiple data points requiring timing within the same clock cycle, making register inputs more flexible. During register output, precise reallocation is performed according to the needs of different modules, avoiding power waste caused by excessive and unnecessary timing operations in related technologies.
[0169] The floating-point arithmetic unit provided in this application embodiment can effectively balance the latency of floating-point operations and control the latency of input and / or output during the entire processing. Based on this latency controllable strategy, it can also avoid the waste of hardware resources, improve the operation efficiency, and reduce power consumption. Moreover, it can optimize the design of registers, replacing the standard registers in related technologies with custom devices. This not only enables more efficient data processing but also solves the power consumption waste caused by too many unnecessary timing operations in related technologies, thereby improving the overall performance of floating-point operations.
[0170] In another embodiment of this application, Figure 5 A schematic diagram of the composition structure of a floating-point processing device provided in this application embodiment. Figure 1 .like Figure 5 As shown, the floating-point processing device 50 may include a configuration module 501 and a plurality of floating-point arithmetic units 10 as described in any of the foregoing embodiments, wherein:
[0171] The configuration module 501 is used to call a target number of floating-point arithmetic units 10 from multiple floating-point arithmetic units 10 according to the precision of the operand to be processed.
[0172] The floating-point arithmetic unit 10 is used to perform floating-point operations on the operands to be processed and obtain the operation results; wherein, the target number is related to the precision of the operands to be processed.
[0173] In this embodiment, the precision of the operand to be processed can include half-precision, single-precision, and double-precision. For example, double-precision can be represented using FP64, which includes 1 sign bit, 11 exponent bits, and 52 mantissa bits; single-precision can be represented using FP32, which includes 1 sign bit, 8 exponent bits, and 23 mantissa bits; half-precision can be represented using BF16, which includes 1 sign bit, 8 exponent bits, and 7 mantissa bits. BF16 is similar to FP32, but the mantissa bits are shortened to 7 bits.
[0174] Here, for floating-point multiply-accumulate units, MAC (Mixed-Precision) design is a common optimization strategy aimed at balancing the performance, power consumption, and accuracy requirements of hardware design. In high-performance computing, especially in the design of deep learning accelerators, how to improve computational efficiency without sacrificing accuracy is a critical problem that urgently needs to be solved. By adopting mixed-precision arithmetic, the processing power of the hardware can be significantly improved while ensuring accuracy requirements. Specifically, through parametric design and module reuse, multiply-accumulate units with various precisions and ratios can be constructed to meet the needs of different application scenarios.
[0175] In some embodiments, the configuration module 501 is further configured to truncate the operand to be processed according to the precision of the operand to be processed, obtain a target number of data segments, and input the target number of data segments into a target number of floating-point arithmetic units.
[0176] In this embodiment of the application, the configuration module 501 can be used according to... Figures 6A-6D The configuration interface allows for parameter configuration, enabling the truncation of operands based on their precision, thus facilitating precision switching. For example... Figures 6A-6D As shown, this illustration provides four parameter positions for truncating operands, for example... Figure 6A This corresponds to the configuration interface where parameter position = 0. Figure 6B This corresponds to the configuration interface when parameter position = 1. Figure 6C This corresponds to the configuration interface where parameter position = 2. Figure 6D This corresponds to the configuration interface where parameter position = 3.
[0177] In one possible implementation, if the operand is single-precision, then two data segments can be extracted, such as data segments b[31:0] and b[63:32]. For example, for parameter positions = 0, 1, the extracted data segment is b[31:0]; for parameter positions = 2, 3, the extracted data segment is b[63:32].
[0178] In another possible implementation, if the operand is half-precision, then four data segments can be extracted, such as data segments b[15:0], b[31:16], b[47:32], and b[63:48]. For example, for parameter position = 0, the extracted data segment is b[15:0]; for parameter position = 1, the extracted data segment is b[31:16]; for parameter position = 2, the extracted data segment is b[47:32]; and for parameter position = 3, the extracted data segment is b[63:48].
[0179] In this way, after inputting the target number of data segments into the target number of floating-point arithmetic units (e.g., multiplication units), the outputs of each multiplication unit can be placed at the top level, merged, and then input into the addition unit in the MAC. This not only achieves signal integration and unified management—by aggregating the outputs of each multiplier at the top level, signal flow can be uniformly managed and controlled, simplifying design complexity, reducing signal transmission paths, lowering latency, and improving overall system performance—but also optimizes timing and latency balance. In multi-channel designs, the calculations of each multiplication unit may have different latency; by merging signals at the top level, the latency of each channel can be balanced, ensuring that the input signals received by the addition unit are consistent in timing, avoiding timing errors caused by latency mismatch. Furthermore, it enhances modularity and maintainability, facilitating modular design and making each function more independent. This improves system maintainability and scalability, facilitating subsequent debugging and optimization; simultaneously, it reduces power consumption and resource consumption, such as reducing unnecessary intermediate storage and transmission, thus lowering power consumption. This helps achieve efficient computation in resource-constrained environments.
[0180] In some embodiments, the target number is set to m when the precision of the operand to be processed is double precision; the target number is set to n when the precision of the operand to be processed is single precision; and the target number is set to k when the precision of the operand to be processed is half precision.
[0181] In this embodiment, m, n, and k are all positive integers, and the ratio m:n:k satisfies a preset ratio requirement. For example, the preset ratio requirement can be 1:2:4, but it is not specifically limited.
[0182] In other words, mixed-precision MAC units can be designed to support computations with different precision ratios, such as a typical FP64:FP32:BF16 design of 1:2:4. In this case, the precision can be flexibly adjusted according to different computational needs to adapt to a wider range of computational tasks. FP64 precision computation is mainly used for applications requiring extremely high precision, while FP32 and BF16 are used for applications with relatively lower precision requirements but extremely high computational demands. Especially in machine learning and deep learning tasks, the data formats of FP32 and BF16 can significantly improve computational efficiency while ensuring the effectiveness of model training and inference.
[0183] In some embodiments, the floating-point processing device 50 may further include a mode selection unit, wherein multiple input terminals of the mode selection unit are respectively connected to the output terminals of floating-point arithmetic units called at different precisions, and the control terminal of the mode selection unit is used to receive an operation mode signal, wherein:
[0184] The mode selection unit is used to output the operation results corresponding to the target number of floating-point arithmetic units called based on the precision of the operand to be processed, according to the operation mode signal.
[0185] In this embodiment, the operation mode signal can be represented by op_mode. Based on this operation mode signal, the calculation result corresponding to the appropriate precision can be switched according to the requirements, thereby achieving more efficient calculation.
[0186] In some embodiments, the multiple floating-point arithmetic units may include multiple multiplication units and multiple addition units. The configuration module 501 is configured to call a target number of multiplication units and a target number of addition units among the multiple multiplication units, based on the precision of the operands to be processed. The multiplication units are configured to perform multiplication operations on the first and second operands among the operands to be processed to obtain a multiplication result. The addition units are configured to perform addition operations on the multiplication result and the third operand among the operands to be processed to obtain an operation result.
[0187] In this embodiment, if both multiplication and addition units exist, then there are two mode selection units, such as a first mode selection unit and a second mode selection unit. Thus, the first mode selection unit is used to select and output the multiplication result corresponding to the precision of the operands to be processed in the multiplication operation based on the operation mode signal; the second mode selection unit is used to select and output the addition result corresponding to the precision of the operands to be processed in the addition operation based on the operation mode signal.
[0188] It should also be noted that the number of multiplication units and addition units invoked are related to the precision of the operands to be processed. For example, if the operands are half-precision, then 4 multiplication units and 4 addition units can be invoked; if the operands are single-precision, then 2 multiplication units and 2 addition units can be invoked, but there is no specific limitation here.
[0189] In one possible implementation, taking multiple floating-point devices including multiple multiplication units and multiple addition units as an example, such as... Figure 7 As shown, the floating-point processing device 50 may include multiple multiplication units 701, multiple addition units 702, a first mode selection unit (MUX_md_1) 703, a first trigger unit 704, a second trigger unit 705, and a second mode selection unit (MUX_md_2) 706.
[0190] It should be noted that the first trigger unit 704 is used to input the multiplication result to the target number of addition units in a preset clock cycle; the second trigger unit 705 is used to input the flag information of the multiplication result to the target number of addition units in a preset clock cycle.
[0191] It should also be noted that the operands to be processed here include the first operand d_i_x, the second operand d_i_y, and the third operand d_i_z. Each operand also includes flag information. First, the precision of the operands to be processed is determined, and according to the precision of the operands to be processed, a target number of multiplication units 701 are called. The multiplication results corresponding to the target number of multiplication units are selected and output to the input of the first trigger unit 704 through the first mode selection unit 703. The flag information of the multiplication results is input to the input of the second trigger unit 705. The clock signal clk controls the first trigger unit 704 to output the obtained multiplication results, and the multiplication results and the third operand d_i_z are input to the target number of addition units. Finally, the addition results (i.e., the operation result d_o) corresponding to the target number of addition units are selected and output through the second mode selection unit 706. During this process, the clock signal clk can also control the second trigger unit 705 to output the flag information corresponding to the multiplication result, so that in the addition operation, the flag information can be used to directly determine whether the multiplication result is a special number.
[0192] In other words, in this embodiment, multiple multiplication units and multiple addition units are multiplexed for different precisions, thereby saving circuit area. Furthermore, the first trigger unit 704 is a data trigger, used to input the obtained multiplication result to the target number of addition units corresponding to a preset clock cycle; the second trigger unit 705 is a flag trigger, used to input the flag information of the obtained multiplication result to the target number of addition units corresponding to a preset clock cycle, so that the addition units no longer need to determine whether the received multiplication result is a special number, further saving circuit area.
[0193] For example, Figure 8 A schematic diagram of the application framework for a floating-point processing device provided in the embodiments of this application. Figure 1 , Figure 9 A schematic diagram of the application framework for a floating-point processing device provided in the embodiments of this application. Figure 2 .like Figure 8 and Figure 9 As shown, the floating-point processing device 50 can be compatible with both FP32 and BF16 precision. Figure 7 Based on this, it may also include a third trigger unit 801, a fourth trigger unit 802, a fifth trigger unit 803, a sixth trigger unit 804, a fourth selection unit 805, a fifth selection unit 806, and a seventh trigger unit 807. The third trigger unit 801, the fourth trigger unit 802, and the fifth trigger unit 803 are used to sample and output the first operand d_i_x, the second operand d_i_y, and the third operand d_i_z in a preset clock cycle. Then, the fourth selection unit 805 selects data from the first operand d_i_x and 0 to input into the corresponding multiplication unit, and the fifth selection unit 806 selects data from the second operand d_i_y and 0 to input into the corresponding multiplication unit. The sixth trigger unit 804 is used to output the operation mode signal op_mode to the control terminal of the first mode selection unit 703 or the control terminal of the second mode selection unit 706 in a preset clock cycle, so as to select whether the calculation result is FP32 precision or BF16 precision. Finally, the seventh trigger unit 807 samples and outputs the calculation result d_o in a preset clock cycle.
[0194] It is important to note that, in Figure 8 or Figure 9 In this configuration, multiple multiplication units 701 and multiple addition units 702 can be multiplexed at different precisions. For example, if the ratio of FP32 to BF16 is 2:4, then in... Figure 9In the diagram, multiple multiplication units 701 include six multiplication units, and multiple addition units 702 include six addition units. At FP32 precision, only two multiplication units and two addition units are needed; at BF16 precision, only four multiplication units and four addition units are needed. To further save circuit area, such as... Figure 8 As shown, the multiple multiplication units 701 include four multiplication units, and the multiple addition units 702 include four addition units. At FP32 precision, only two multiplication units and two addition units are needed; at BF16 precision, only four multiplication units and four addition units are needed. Furthermore, if the ratio of FP32 to BF16 is adjusted to 2:5, then in... Figure 9 In this case, only two multiplication units and two addition units are needed under FP32 precision, and only five multiplication units and five addition units are needed under BF16 precision. No restrictions are imposed here.
[0195] In other words, the choice of data format is particularly important in the design of mixed-precision MAC units. Suppose the data to be processed has a lower numerical range but higher precision requirements, such as TF32 (Tensor Float 32) or BF16 (BFloat 16). For these data formats, embodiments of this application can also use zero-padding to pass the input data into the module. Specifically, when the input data has a low bit width (e.g., 16 bits) but its numerical range requirement is not high, zero-padding can be used to convert the data to 32 bits for operation. This not only effectively improves computational precision but also increases computational speed and power efficiency by reducing hardware resource usage.
[0196] For addition operations, especially addition operations at BF16 precision, similar optimization strategies also apply. Before performing the addition operation, the BF16 data can be extended to 32 bits by padding with zeros at the lower bits (e.g., ...). Figure 9 The BF16_to_FP32 module is used to perform the addition calculation. This not only improves the accuracy of the addition but also ensures that the numerical precision is not lost due to insufficient bit width during the entire operation. As a result, the final output result is more accurate, meeting the application scenarios with higher precision requirements, while maintaining low computing cost and power consumption.
[0197] In short, designing a mixed-precision MAC unit presents certain challenges at the hardware implementation level. A balance needs to be found between hardware resources, timing, power consumption, and computational precision. Therefore, the efficiency of multipliers and adders must be fully considered in the aforementioned processes, especially their performance when processing data of different precisions. The introduction of mixed precision requires ensuring efficient data transmission and processing while avoiding unnecessary computational redundancy. In the multiply-accumulate unit, multiple data points of different precisions need to undergo precise scheduling and optimization to ensure efficient data flow and overall system stability.
[0198] In another embodiment of this application, a floating-point processing apparatus is also provided, see [link to previous application]. Figure 10 The floating-point processing device 50 may include a register 1001 and a plurality of floating-point arithmetic units 10 as described in any of the foregoing embodiments, wherein:
[0199] Register 1001 is used to store at least two floating-point numbers in the same clock cycle;
[0200] The floating-point arithmetic unit 10 is used to obtain operands to be processed from registers.
[0201] It should be noted that the standard registers in related technologies cannot merge multiple data sets that need to be timed within the same clock cycle, resulting in unnecessary power consumption waste. However, in this embodiment, the register 1001 can be a custom device, precisely optimized for sequential logic; it can not only process data more efficiently, but also merge multiple data sets that need to be timed within the same clock cycle, making the register input more flexible.
[0202] In some embodiments, see Figure 11 The floating-point arithmetic unit 10 can be a multiplication unit, and the register 1001 can include the first register r1, wherein:
[0203] The first register r1 is used to store the first operand and the second operand in the same clock cycle, and to input the operand output by the first register r1 into at least one multiplication unit;
[0204] At least one multiplication unit 1002 is used to perform multiplication operations on the operands output by the first register r1 to determine at least one multiplication result.
[0205] In this embodiment of the application, if the floating-point processing device 50 is only used to implement multiplication operations, then it may include at least one multiplication unit 1002. After merging and storing the first operand x and the second operand y input to the first register r1, the operand output from the first register r1 can be correspondingly input to at least one multiplication unit 1002, thereby obtaining the multiplication result.
[0206] In some embodiments, see Figure 12 The floating-point arithmetic unit 10 can be an addition unit, and register 1001 can include a second register r2, wherein:
[0207] The second register r2 is used to store the third operand and the fourth operand in the same clock cycle, and to input the operand output by the second register r2 into at least one addition unit;
[0208] At least one addition unit 1003 is used to perform addition operations on the operands output by the second register r2 to determine at least one addition result.
[0209] In this embodiment, if the floating-point processing device 50 is only used to perform addition operations, it may include at least one addition unit 1003. After merging and storing the third operand z and the fourth operand u to be added in the second register r2, the operands output from the second register r2 can be correspondingly input into at least one addition unit 1003 to obtain the addition result.
[0210] In some embodiments, see Figure 13 The floating-point arithmetic unit 10 may include a multiplication unit and an addition unit, and the register 1001 may include a first register r1 and a second register r2, wherein:
[0211] The first register r1 is used to store the first operand and the second operand in the same clock cycle, and to input the operand output by the first register r1 into at least one multiplication unit;
[0212] At least one multiplication unit 1002 is used to perform multiplication operations on the operands output by the first register r1 to determine at least one multiplication result;
[0213] The second register r2 is used to store the third operand and at least one multiplication result in the same clock cycle, and to input the operand output by the second register r2 into at least one addition unit;
[0214] At least one addition unit 1003 is used to perform addition operations on the operands output by the second register r2 to determine at least one addition result.
[0215] In some embodiments, see continue to see Figure 13 The floating-point processing device 50 may further include a redistribution unit 1004. The redistribution unit 1004 is used to redistribute inputs to at least one multiplication unit when the first register r1 is output, and to input the redistributed operands to at least one multiplication unit 1002.
[0216] It should be noted that in this embodiment, the first register r1 and the second register r2 can also be custom-made devices. This not only enables more efficient data processing but also allows for the merging of multiple data sets requiring timing within the same clock cycle, making register inputs more flexible. During register output, precise reallocation is performed according to the needs of different modules, avoiding power waste caused by excessive and unnecessary timing operations in related technologies.
[0217] It should also be noted that, in the embodiments of this application, the multiplication unit and the addition unit can be located in the MAC unit. For example... Figure 13 As shown, the three input operands also include flags (zero, nan, overflow) to indicate whether the corresponding input operand is a special number.
[0218] It should also be noted that, with Figure 13 For example, the third operand z, due to design requirements, will arrive one step later than the first operand x and the second operand y. The input can be processed according to different design requirements. The redistribution unit allocates input for each MAC (multiplication). The second register r2 simultaneously adds at least one multiplication result to the third operand z, and then inputs it to each MAC for addition to obtain the final result = x × y + z.
[0219] This application provides a floating-point processing device comprising multiple floating-point arithmetic units. Depending on the precision of the operands to be processed, a target number of floating-point arithmetic units can be invoked from among the multiple units to perform floating-point operations on the operands and obtain the result. This also enables mixed-precision design, significantly improving hardware processing power while ensuring accuracy requirements. Furthermore, by reusing basic modules such as multiplication and addition units, power consumption and resource consumption are further reduced, saving circuit area. In addition, register design can be optimized, replacing standard registers in related technologies with custom devices. This not only enables more efficient data processing but also solves the power waste caused by excessive unnecessary timing operations in related technologies, thereby improving the overall performance of floating-point operations.
[0220] In yet another embodiment of this application, Figure 14 A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 1 .like Figure 14 As shown, the method may include:
[0221] S1401, the logic processing module performs logical operations on the input operands to obtain the operation results.
[0222] S1402, the delay control module controls the delay between the input logic processing module and the input floating-point arithmetic unit for the operand to be processed; and / or, controls the delay between the output calculation result of the logic processing module and the output calculation result of the floating-point arithmetic unit.
[0223] In this embodiment, the floating-point processing method is applied to the floating-point arithmetic unit described in the foregoing embodiments. This effectively balances the latency of the logical operation, improves the operational efficiency of the floating-point arithmetic unit 10, and ensures accurate result output.
[0224] In this embodiment, the floating-point unit can adopt a pipelined structure, dividing the entire floating-point processing into several pipelined registers, which can effectively avoid performance bottlenecks caused by excessive pipeline delays. For example, the entire floating-point processing can be divided into: an input pipelined register, at least one intermediate pipelined register, and an output pipelined register, thereby improving pipeline efficiency and saving time.
[0225] In some embodiments, the delay control module includes an input pipeline register and an output pipeline register, and the method may further include:
[0226] The input pipeline register controls the computation path delay of at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed; and / or,
[0227] The output pipeline register controls the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0228] In this embodiment, the input pipeline register 1021 and the output pipeline register 1022 are controllable in the pipeline structure. That is, whether the input pipeline register 1021 and the output pipeline register 1022 are enabled can be controlled according to the corresponding configuration parameters in order to achieve delay balance.
[0229] In some embodiments, the method may further include: an input pipeline register controlling whether to clock at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed according to a first configuration parameter, so that the computation path delay of at least one of the initial sign bit, initial exponent bit, and initial mantissa bit meets a first delay requirement; and / or, an output pipeline register controlling whether to clock at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result according to a second configuration parameter, so that the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit meets a second delay requirement.
[0230] In this embodiment, the first delay requirement and the second delay requirement are preset and used to measure whether the path delay of each segment meets the requirements. Furthermore, the first configuration parameter and the second configuration parameter can be designed according to the overall requirements of the floating-point arithmetic unit to control whether the corresponding pipeline register is enabled.
[0231] In some embodiments, at least one intermediate pipeline register is provided between the input pipeline register and the output pipeline register, and the method may further include:
[0232] At least one intermediate pipeline register performs logical operations on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, and determines the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0233] In this embodiment, the operands to be processed include a first operand and a second operand. Here, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit; the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit. The target sign bit can be obtained by performing logical operations on the first and second initial sign bits, the target exponent bit can be obtained by performing logical operations on the first and second initial exponent bits, and the target mantissa bit can be obtained by performing logical operations on the first and second initial mantissa bits.
[0234] In this way, the logical processing can be divided into at least one intermediate pipeline register, and the computation task can be allocated between each intermediate pipeline register. This optimizes the execution time of floating-point processing, avoids performance bottlenecks caused by excessive pipeline delays, and improves computational efficiency.
[0235] Those skilled in the art should understand that the description of the floating-point processing method in the embodiments of this application can be understood with reference to the relevant description of the floating-point arithmetic unit in the foregoing embodiments, and will not be described in detail here.
[0236] In yet another embodiment of this application, Figure 15A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 2 .like Figure 15 As shown, the method may include:
[0237] S1501, the configuration module calls the target number of floating-point arithmetic units from multiple floating-point arithmetic units according to the precision of the operand to be processed.
[0238] S1502, the floating-point arithmetic unit performs floating-point operations on the operands to be processed and obtains the calculation result; wherein, the target number is related to the precision of the operands to be processed.
[0239] In this embodiment, the floating-point processing method is applied to the floating-point processing device described in the foregoing embodiments. The precision of the operands to be processed can include half-precision, single-precision, and double-precision. Here, by employing mixed-precision arithmetic, the processing power of the hardware can be significantly improved while ensuring the required precision. Specifically, through parametric design and module reuse, multiply-accumulate units with various precisions and ratios can be constructed to meet the needs of different application scenarios.
[0240] In some embodiments, the method may further include: a configuration module truncating the operands to be processed according to their precision to obtain a target number of data segments, and inputting the target number of data segments into a target number of floating-point arithmetic units.
[0241] In this embodiment of the application, the configuration module can be used according to... Figures 6A-6D The configuration interface allows you to configure parameters to truncate operands based on their precision, thus enabling you to switch precision.
[0242] In some embodiments, the method may further include: a configuration module calling a target number of multiplication units and a target number of addition units in a plurality of multiplication units according to the precision of the operands to be processed; the multiplication units performing multiplication operations on the first operand and the second operand in the operands to be processed to obtain a multiplication result; and the addition units performing addition operations on the multiplication result and the third operand in the operands to be processed to obtain an operation result.
[0243] In the embodiments of this application, the number of multiplication units and addition units invoked are related to the precision of the operands to be processed. For example, if it is a half-precision operand, then 4 multiplication units and 4 addition units can be invoked; if it is a single-precision operand, then 2 multiplication units and 2 addition units can be invoked, but no specific limitation is made here.
[0244] In other words, in this embodiment, after inputting the target number of data segments into the target number of floating-point arithmetic units (e.g., multiplication units), the outputs of each multiplication unit can be merged and input into the addition unit in the MAC after being placed at the top level. This not only achieves signal integration and unified management—by aggregating the outputs of each multiplier at the top level, signal flow can be uniformly managed and controlled, simplifying design complexity, reducing signal transmission paths, lowering latency, and improving overall system performance—but also optimizes timing and latency balance. In multi-channel designs, the calculations of each multiplication unit may have different latency; by merging signals at the top level, the latency of each channel can be balanced, ensuring that the input signals received by the addition unit are consistent in timing, avoiding timing errors caused by latency mismatch. Furthermore, it enhances modularity and maintainability, facilitating modular design and making each function more independent. This improves system maintainability and scalability, facilitating subsequent debugging and optimization; simultaneously, it reduces power consumption and resource consumption, such as reducing unnecessary intermediate storage and transmission, thus lowering power consumption. This helps achieve efficient computation in resource-constrained environments.
[0245] Furthermore, those skilled in the art should understand that the description of the floating-point processing method in the embodiments of this application can be understood with reference to the relevant description of the floating-point processing device described in the foregoing embodiments, and will not be elaborated here.
[0246] In yet another embodiment of this application, Figure 16 A flowchart illustrating a floating-point processing method provided in this application embodiment. Figure 3 .like Figure 16 As shown, the method may include:
[0247] The S1601 register stores at least two floating-point numbers in the same clock cycle.
[0248] S1602, the floating-point arithmetic unit obtains the operand to be processed from the register so that the floating-point arithmetic unit can execute: the logic processing module performs logical operations on the operand to be processed to obtain the operation result; the delay control module controls the delay between the input logic processing module and the input floating-point arithmetic unit of the operand to be processed; and / or controls the delay between the output operation result of the logic processing module and the output operation result of the floating-point arithmetic unit.
[0249] It should be noted that, for step S1602, after the floating-point unit obtains the operand to be processed from the register, the floating-point unit executes... Figure 14 The floating-point processing method shown.
[0250] It should also be noted that the standard registers in related technologies cannot merge multiple data sets that need to be timed within the same clock cycle, resulting in unnecessary power consumption waste. However, in the embodiments of this application, the registers here can be custom devices, precisely optimized for timing logic; they can not only process data more efficiently, but also merge multiple data sets that need to be timed within the same clock cycle, making the register input more flexible.
[0251] In some embodiments, the floating-point arithmetic unit includes a multiplication unit and an addition unit, and the register includes a first register and a second register. Accordingly, the method further includes: the first register storing a first operand and a second operand in the same clock cycle, and inputting the operand output by the first register into at least one multiplication unit; the at least one multiplication unit performing a multiplication operation on the operand output by the first register to determine at least one multiplication result; the second register storing a third operand and at least one multiplication result in the same clock cycle, and inputting the operand output by the second register into at least one addition unit; the at least one addition unit performing an addition operation on the operand output by the second register to determine at least one addition result.
[0252] In some embodiments, the method further includes: the reallocation unit reallocates input to at least one multiplication unit when the first register outputs, and inputs the reallocated operands to at least one multiplication unit.
[0253] It should be noted that in the embodiments of this application, the first register and the second register can be custom-made devices. This not only enables more efficient data processing but also allows for the merging of multiple data sets requiring timing within the same clock cycle, making register inputs more flexible. During register output, precise reallocation is performed according to the needs of different modules, avoiding power waste caused by excessive and unnecessary timing operations in related technologies.
[0254] For example, with Figure 13 For example, the third operand z, due to design requirements, will arrive one step later than the first operand x and the second operand y. The input can be processed according to different design requirements. The redistribution unit allocates input for each MAC (multiplication). The second register r2 simultaneously adds at least one multiplication result to the third operand z, and then inputs it to each MAC for addition to obtain the final result = x × y + z.
[0255] Furthermore, those skilled in the art should understand that the description of the floating-point processing method in the embodiments of this application can also be understood with reference to the relevant description of the floating-point processing device described in the foregoing embodiments, and will not be detailed here.
[0256] In another embodiment of this application, a latency-balanced, high-performance floating-point arithmetic unit is proposed, which can be applied to the field of high-performance computing, such as large model training, artificial intelligence (AI), and intelligent computing. The floating-point arithmetic unit may include a multiplication unit and / or an addition unit, which will be described in detail below.
[0257] In one possible implementation, a delay-balanced, controllable three-stage pipelined multiplication unit design is provided. The floating-point multiplication unit employs a delay-balanced, pipelined design. Inputs are equipped with flags to indicate whether the operand is a special number, such as zero, NaN, or overflow. Figure 3 As shown, the first-stage pipeline can select whether to pass the input, timed register data down the pipeline based on configured parameters. After inserting the second-stage pipeline, the mantissa bits are multiplied directly.
[0258] In this implementation, the multiplication result requires one cycle to obtain. At this time, the multiplication result arrives simultaneously with the flag bit, sign bit, and exponent bit. Considering that multiplication involves complex computational logic and requires multiple levels of logic gates and a longer computational path compared to addition and XOR gates, resulting in greater latency, this embodiment of the application can synchronize other paths to achieve latency balance.
[0259] The third-level flow also controls whether the output is timed by configuring parameters.
[0260] It should be noted that rounding operations can be performed according to IEEE-754 rules. Feedback from testing revealed that the processing of denominated and normalized numbers are two separate branches and cannot be combined. Specifically, for normalized numbers, the rounding rules for floating-point numbers can be summarized as follows: when the reserved bits (G), approximation bits (R), and sticky bits (S) meet specific combination conditions, a carry-over occurs (normal carry). For denominated numbers, there is a situation where, after rounding, the mantissa carries over to the exponent, causing the exponent to not be all zeros, meaning the denominated number becomes a normalized number after the carry-over. Therefore, when handling underflow numbers, this situation needs to be considered, and the carry-over case for denominated numbers needs to be handled separately when assigning the final exponent value.
[0261] In another possible implementation, a delay-balanced, controllable five-stage pipelined adder unit design is presented. In high-performance computing, the adder unit is one of the fundamental arithmetic units, and its design performance directly affects the overall system efficiency. To optimize the performance of the adder unit, especially when dealing with floating-point calculations, designing a delay-balanced, controllable pipelined adder unit has become a crucial technical solution. This effectively balances the latency of the adder operation, improves the computational efficiency of the adder unit, and ensures accurate result output.
[0262] In this implementation, the two-stage pipeline of the addition unit's input and output is the same as that of the multiplication unit, both of which can be determined by configuration parameters. The addition operation first selects the larger operand based on its exponent and uses this operand as the addend. Here, a register flag can be set to indicate whether the addend and addend need to be swapped. Slightly different from the multiplication unit's operand parsing, operand parsing can only proceed after the operand positions have been determined in the previous step, but no operations are performed on the individual parts of the two operands. Specifically, the exponent alignment operation aligns the decimal point of the mantissa based on the difference in the exponents of the two operands. Here, a shift alignment operation is performed based on the exponent difference, and the result is inserted into the pipeline register after shifting. After obtaining the mantissa sum, the sign bit of the operand determines whether the result needs to be inverted. Next, a shift operation is performed to adjust the exponent size appropriately, normalizing the result. After normalization, the result is inserted into the pipeline register. This process also incorporates LZC, which employs a step-by-step shifting and zero-checking strategy. The register records the state where the most significant bits are all zeros at each stage. The final register value determines the number of leading zeros in the mantissa and performs the shift accordingly. This also ensures that the most significant bit of the final mantissa is 1. See details... Figure 4 Furthermore, the rounding rules are similar to those for the multiplication unit mentioned above, and will not be detailed here.
[0263] Finally, the core advantage of the entire five-stage pipelined addition unit design lies in its controllable latency balancing. Through the five-stage pipeline structure, the addition unit can distribute computational tasks across each stage, thereby optimizing the execution time of the addition operation. The latency of each stage is carefully designed to ensure that the final addition result does not contain violations and to avoid performance bottlenecks caused by excessive pipeline latency. Furthermore, the latency balancing strategy employed here effectively avoids wasting hardware resources, improves computational efficiency, and reduces power consumption.
[0264] In another possible implementation, a mixed-precision multiply-accumulate unit design is presented. MAC unit design is a common optimization strategy aimed at balancing the performance, power consumption, and accuracy requirements of hardware design. In high-performance computing, especially in the design of deep learning accelerators, improving computational efficiency without sacrificing accuracy is a critical problem that urgently needs to be solved. By employing mixed-precision operations, the processing power of the hardware can be significantly improved while maintaining accuracy requirements. Based on fundamental modules such as multiplication and addition units, through parametric design and module reuse, we can construct multiply-accumulate units with various precisions and ratios to meet the needs of different application scenarios.
[0265] For example, such as Figures 6A-6D As shown, at the top level of MAC, operands can be truncated according to the configured parameters to achieve the purpose of switching precision.
[0266] In addition, placing the outputs of each multiplication unit at the top level and merging them before inputting them back to the addition unit of the MAC unit has the following important implications: (1) Signal integration and unified management. By aggregating the outputs of each multiplier at the top level, the signal flow can be managed and controlled in a unified manner, simplifying the design complexity. It reduces signal transmission paths, lowers latency, and improves the overall performance of the system. (2) Optimize timing and delay balance. In multi-channel designs, the calculations of each multiplier may have different delays. By merging the signals at the top level, the delays of each channel can be balanced, ensuring that the input signals received by the adder are consistent in timing, and avoiding timing errors caused by delay mismatch. (3) Enhance modularity and maintainability. It helps with modular design, making each part more independent. It improves the maintainability and scalability of the system, facilitating subsequent debugging and optimization. (4) Reduce power consumption and resource consumption. It reduces unnecessary intermediate storage and transmission, lowering power consumption. It helps achieve efficient computation in resource-constrained environments.
[0267] Specifically, mixed-precision MAC units can be designed to support computations with different precision ratios, such as a typical FP64:FP32:BF16 design of 1:2:4. This design allows for flexible adjustment of precision to suit a wider range of computational tasks. In this design, FP64 precision is primarily used for applications requiring extremely high precision, while FP32 and BF16 are used for applications with relatively lower precision requirements but extremely high computational demands. Particularly in machine learning and deep learning tasks, the FP32 and BF16 data formats can significantly improve computational efficiency while ensuring the effectiveness of model training and inference.
[0268] Furthermore, embodiments of this application can also select an appropriate precision ratio according to specific needs. For example, in some applications, the ratio of FP32 to BF16 can be adjusted to 2:4 to optimize resource utilization and further improve computing speed. This flexible precision selection enables the mixed-precision MAC unit to perform well in different tasks, meeting precision requirements while improving computing efficiency and reducing hardware power consumption and resource consumption. Figures 7-9 It provides an illustrative example of mixed precision applications.
[0269] In mixed-precision MAC unit design, the choice of data format is particularly important. This involves assuming the data needs to be processed in a low-precision range but with high accuracy requirements, such as TF32 or BF16. For these data formats, this embodiment can use zero-padding to pass the input data into the module. Specifically, when the input data has a low bit width (e.g., 16 bits) but its numerical range requirement is not high, zero-padding can be used to convert the data to 32 bits for computation. This method not only effectively improves computational accuracy but also increases computational speed and power efficiency by reducing hardware resource usage.
[0270] For addition operations, especially addition operations at BF16 precision, similar optimization strategies also apply. Before performing the addition operation, the BF16 data can be extended to 32 bits by padding with zeros at the lower bits (e.g., ...). Figure 9 The BF16_to_FP32 module in the algorithm is used to perform the addition calculation. This method not only improves the accuracy of the addition but also ensures that the numerical precision is not lost due to insufficient bit width during the entire operation. Through this technique, the final output result will be more accurate, meeting the application scenarios with higher precision requirements, while maintaining low computational cost and power consumption.
[0271] Understandably, designing a mixed-precision MAC unit presents certain challenges at the hardware implementation level. A balance needs to be found between hardware resources, timing, power consumption, and computational precision. Therefore, the efficiency of the multiplication and addition units must be fully considered during the design process, especially their performance when processing data of different precisions. The introduction of mixed precision requires ensuring efficient data transmission and processing while avoiding unnecessary computational redundancy. In the multiply-accumulate unit, multiple data points of different precisions need to undergo precise scheduling and optimization to ensure efficient data flow and overall system stability.
[0272] Furthermore, the design of hardware accelerators needs to fully consider flexibility. The design of mixed-precision MAC units is not merely an optimization for a specific task, but rather requires dynamic adjustment of precision and resource allocation based on the needs of different tasks. For example, in deep learning training, the proportion of precision used may be adjusted according to different levels of the model and the computational load to maximize hardware performance. Designers also need to provide sufficient control mechanisms, such as an operation mode signal (op_mode), to switch precision and operation modes at runtime according to task requirements, thereby achieving more efficient computation.
[0273] In another possible implementation, performance enhancement is achieved using custom components, which may require negotiation with the backend. Using custom components is a highly effective strategy, particularly in terms of power consumption optimization. Specifically, by employing custom registers, precise adjustments can be made according to actual needs, thereby reducing system power consumption. The motivation for implementing this technical solution stems from the high proportion of sequential logic identified in the comprehensive test report. As the complexity of sequential logic increases, power consumption and performance bottlenecks become increasingly apparent, especially in register data transfer and storage. Therefore, optimizing register design can significantly reduce unnecessary power consumption, resulting in a substantial improvement in power efficiency.
[0274] In this embodiment, standard registers in the module can be replaced with custom devices, representing a precise optimization to address timing bottlenecks. Custom registers not only process data more efficiently but also merge multiple data sets requiring timing within the same clock cycle, making register inputs more flexible. During register output, precise reallocation is performed according to the needs of different modules, avoiding power waste caused by excessive unnecessary timing operations in related technologies. This approach not only improves performance but also significantly optimizes overall power consumption, effectively enhancing the overall system performance, especially in complex systems.
[0275] For example, taking the top level as an example, in Figure 13 In this process, the input operands also include a flag (zero, nan, overflow) to indicate whether the operand is a special number. The third operand z is input because, due to design requirements, it arrives one clock cycle later than the second operand x and the third operand y. In this case, the input can be processed according to different design requirements. The reallocation unit allocates input to each MAC. The second register r2 simultaneously multiplies the results of the multiplications of each MAC with the third operand z, and then inputs them to each MAC for addition to obtain the final result.
[0276] It should also be noted that the delay-balanced multiply-add unit (MAC) provided in this application plays a key role in digital signal processing. Its efficient multiplication and addition operations have led to its widespread application in many fields, such as intelligent computing and high-performance computing.
[0277] For example, in the field of Digital Signal Processing (DSP), MAC units are used to perform filtering, convolution, and other signal processing tasks. Their efficient multiply-accumulate operations make them widely used in audio, video, and communication systems. In the field of High Performance Computing (HPC), MAC units are used for matrix multiplication and linear algebra operations. Their efficient computing power makes them widely used in scientific computing, engineering simulation, and data analysis. In the field of AI and Machine Learning (ML), MAC units are used for forward and backward propagation of neural networks. Their efficient computing power makes them widely used in the training and inference of deep learning models. In the field of Graphics Processing Units (GPUs), MAC units are used for graphics rendering and image processing. Their efficient computing power makes them widely used in games, virtual reality, and computer vision. In the field of communication systems, MAC units are used for modulation / demodulation, encoding / decoding, and signal processing. Their efficient computing power makes them widely used in wireless communication, satellite communication, and fiber optic communication. In the field of the Internet of Things (IoT), MAC units are used for sensor data processing and edge computing. Its high computing power has led to its widespread application in smart homes, smart cities, and industrial automation. In the field of medical devices, the MAC unit is used for image processing and signal analysis. Its high computing power has led to its widespread application in medical imaging, diagnosis, and monitoring.
[0278] In another embodiment of this application, an embodiment of this application provides a chip that may include: a floating-point arithmetic unit 10 as described in any of the foregoing embodiments, or a floating-point processing device 50 as described in any of the foregoing embodiments.
[0279] In another embodiment of this application, an electronic device is provided, which includes a processor, wherein the processor includes: a floating-point arithmetic unit 10 as described in any of the foregoing embodiments, or a floating-point processing device 50 as described in any of the foregoing embodiments.
[0280] In summary, the above embodiments have provided a detailed explanation of the specific implementation of the aforementioned embodiments. It can be seen that the technical solutions of these embodiments can effectively balance the latency of floating-point operations for floating-point arithmetic units (such as multiplication or addition units), and control the input and / or output latency throughout the processing. This controllable latency strategy can also avoid wasting hardware resources, improve computational efficiency, and reduce power consumption. Furthermore, through mixed-precision design, the processing power of the hardware can be significantly improved while ensuring accuracy requirements. Moreover, the reuse of basic modules such as multiplication and addition units further reduces power consumption and resource consumption, saving circuit area. In addition, the design of registers can be optimized, replacing standard registers in related technologies with custom devices. This not only enables more efficient data processing but also solves the power waste caused by excessive unnecessary timing operations in related technologies, thereby improving the overall performance of floating-point operations.
[0281] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0282] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0283] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0284] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0285] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium.
[0286] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0287] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0288] The methods disclosed in the several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0289] The features disclosed in the several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0290] The features disclosed in the several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0291] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A floating-point arithmetic unit, characterized in that, The floating-point arithmetic unit includes: The logic processing module is used to perform logical operations on the operands to be processed input to the floating-point arithmetic unit to obtain the operation results; The delay control module is used to control the delay between the input of the operand to be processed to the logic processing module and the input to the floating-point arithmetic unit; and / or, to control the delay between the output of the calculation result by the logic processing module and the output of the calculation result by the floating-point arithmetic unit; The delay control module includes an input pipeline register and an output pipeline register, wherein: The input pipeline register is used to control whether to clock at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, according to a first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit meets the first delay requirement; and, The output pipeline register is used to control whether to clock at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the calculation result according to the second configuration parameter, so that the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit meets the second delay requirement; The first configuration parameter and the second configuration parameter are designed according to the overall requirements of the floating-point arithmetic unit so that the floating-point arithmetic unit can satisfy the delay balance when outputting the arithmetic result.
2. The floating-point arithmetic unit according to claim 1, characterized in that, At least one intermediate pipeline register is provided between the input pipeline register and the output pipeline register, wherein: The at least one intermediate pipeline register is used to perform logical operations on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, and to determine the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result; wherein, the operand to be processed includes a first operand and a second operand, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit, and the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit.
3. The floating-point arithmetic unit according to claim 2, characterized in that, When the floating-point arithmetic unit is a multiplication unit, the at least one intermediate pipeline register includes a pipeline register; The logic processing module includes a first logic module and a second logic module, wherein: The first logic module is configured to perform logical operations on the first initial sign bit and the second initial sign bit, as well as the first initial exponent bit and the second initial exponent bit, respectively, to determine the target sign bit and the target exponent bit; The second logic module is used to perform a multiplication operation on the first initial mantissa and the second initial mantissa to determine the target mantissa.
4. The floating-point arithmetic unit according to claim 3, characterized in that, The second logic module includes a mantissa multiplication unit and a rounding operation unit, wherein: The mantissa multiplication unit is used to perform mantissa multiplication on the first initial mantissa and the second initial mantissa to obtain the first mantissa product; The rounding operation unit is used to perform a rounding operation on the first mantissa product to determine the target mantissa digits.
5. The floating-point arithmetic unit according to claim 4, characterized in that, The rounding operation unit is further configured to, when the calculation result is a non-standard number, perform a rounding operation on the first mantissa product to obtain a second mantissa product; when the second mantissa product exceeds a preset range, perform normalization processing on the second mantissa product to obtain the target mantissa digit, and perform a carry operation on the target exponent digit, so that the calculation result is transformed from a non-standard number to a standard number.
6. The floating-point arithmetic unit according to claim 3, characterized in that, The first logic module includes a sign bit processing unit and an exponent bit processing unit, wherein: The sign bit processing unit is used to perform a sign XOR operation on the first initial sign bit and the second initial sign bit to obtain the target sign bit. The exponent processing unit is used to perform exponential addition on the first initial exponent and the second initial exponent to obtain the target exponent.
7. The floating-point arithmetic unit according to claim 6, characterized in that, The first operand includes a first initial flag bit, and the second operand includes a second initial flag bit; the first logic module further includes a flag bit processing unit, wherein: The flag processing unit is used to perform logical operations based on the first initial flag and the second initial flag to determine the target flag corresponding to the operation result. The target flag is used to indicate whether the operation result is a special number.
8. The floating-point arithmetic unit according to claim 2, characterized in that, When the floating-point arithmetic unit is an addition unit, the at least one intermediate pipeline register includes three pipeline registers; The logic processing module includes a first logic module, a second logic module, a third logic module, and a fourth logic module, wherein: The first logic module is used to determine the exponent difference between the first initial exponent and the second initial exponent and the smaller exponent of the two, and to perform a shift operation on the mantissa corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa and the second initial mantissa. The second logic module is used to determine the sign XOR result of the first initial sign bit and the second initial sign bit, and to perform an addition operation on the aligned first initial mantissa bit and the second initial mantissa bit according to the sign XOR result to obtain the first mantissa sum and the target sign bit; The third logic module is used to perform leading zero processing on the first mantissa sum to determine the leading zero statistical result and the shifted second mantissa sum; The fourth logic module is used to perform an exponential shift operation based on the leading zero statistical result to obtain the target exponent; and to perform a rounding operation on the second mantissa to obtain the target mantissa.
9. The floating-point arithmetic unit according to claim 8, characterized in that, The first logic module includes an exponent subtraction unit, a first control unit, a mantissa shifting unit, a first selection unit, a second selection unit, and a third selection unit, wherein: The exponent subtraction unit is used to perform a subtraction operation on the first initial exponent and the second initial exponent to determine the exponent difference, and transmit the exponent difference to the first control unit. The first control unit is configured to receive the exponential difference and generate, based on the exponential difference, a first control signal to be transmitted to the first selection unit, a second control signal to be transmitted to the second selection unit, and a third control signal to be transmitted to the third selection unit. The first selection unit is configured to receive the first control signal and select a larger exponent from the first initial exponent bit and the second initial exponent bit according to the first control signal; The second selection unit is configured to receive the second control signal and select the mantissa corresponding to the smaller exponent from the first initial mantissa and the second initial mantissa according to the second control signal; The third selection unit is used to receive the third control signal and select the mantissa corresponding to the larger exponent from the first initial mantissa and the second initial mantissa according to the third control signal; The mantissa shifting unit is used to shift the mantissa corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa and the second initial mantissa.
10. The floating-point arithmetic unit according to claim 9, characterized in that, The second logic module includes an XOR unit, a second control unit, and a mantissa addition unit, wherein: The XOR unit is used to perform a sign XOR operation on the first initial sign bit and the second initial sign bit to obtain the sign XOR result, and transmit the sign XOR result to the second control unit; The second control unit is configured to receive the sign XOR result and generate a fourth control signal to be transmitted to the mantissa addition unit based on the sign XOR result. The mantissa addition unit is used to perform addition operations on the aligned first initial mantissa bits and second initial mantissa bits according to the fourth control signal to obtain the first mantissa sum and the target sign bit.
11. The floating-point arithmetic unit according to claim 10, characterized in that, The third logic module includes a leading zero processing unit, wherein: The leading zero processing unit is used to predict leading zeros for the first mantissa sum, determine the leading zero statistical result, and shift the first mantissa sum according to the leading zero statistical result to obtain the second mantissa sum.
12. The floating-point arithmetic unit according to claim 11, characterized in that, The fourth logic module includes an exponential shift unit and a rounding operation unit, wherein: The exponent shifting unit is used to receive the larger exponent sent by the first selection unit and the leading zero statistical result sent by the leading zero processing unit, and to perform an exponent shifting operation on the larger exponent according to the leading zero statistical result to obtain the target exponent bit; The rounding operation unit is used to perform a rounding operation on the second mantissa to obtain the target mantissa.
13. The floating-point arithmetic unit according to claim 12, characterized in that, The rounding operation unit is further configured to, when the calculation result is a non-standard number, perform a rounding operation on the second mantissa sum to obtain a third mantissa sum; when the third mantissa sum exceeds a preset range, perform normalization processing on the third mantissa sum to obtain the target mantissa digit, and add 1 to the target exponent digit to adjust the calculation result from a non-standard number to a standard number.
14. The floating-point arithmetic unit according to any one of claims 2 to 13, characterized in that, The input pipeline register is used to merge and store the data to be stamped in the initial sign bit, initial exponent bit and initial mantissa bit corresponding to the operand to be processed within the same clock cycle; The output pipeline register is used to merge and store the data to be stamped in the target sign bit, target exponent bit and target mantissa bit corresponding to the calculation result within the same clock cycle; The intermediate pipeline register is used to merge and store the intermediate results of logical operations performed on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed within the same clock cycle.
15. A floating-point processing device, characterized in that, The floating-point processing device includes a configuration module and a plurality of floating-point arithmetic units as described in any one of claims 1 to 14, wherein: The configuration module is used to call a target number of floating-point arithmetic units from among the multiple floating-point arithmetic units according to the precision of the operand to be processed; The floating-point arithmetic unit is used to perform floating-point operations on the operands to be processed to obtain the calculation result; wherein, the target quantity is related to the precision of the operands to be processed.
16. The floating-point processing apparatus according to claim 15, characterized in that, The configuration module is further configured to truncate the operand to be processed according to the precision of the operand to be processed, obtain a target number of data segments, and input the target number of data segments into the target number of floating-point arithmetic units.
17. The floating-point processing apparatus according to claim 15, characterized in that, When the precision of the operand to be processed is double precision, the target quantity is set to m; When the precision of the operand to be processed is single precision, the target quantity is set to n; When the precision of the operand to be processed is half precision, the target quantity is set to k; Where m, n, and k are all positive integers, and m:n:k satisfies the preset ratio requirement.
18. The floating-point processing apparatus according to claim 15, characterized in that, The floating-point processing device further includes a mode selection unit, wherein multiple input terminals of the mode selection unit are respectively connected to the output terminals of floating-point arithmetic units called at different precisions, and the control terminal of the mode selection unit is used to receive operation mode signals, wherein: The mode selection unit is used to output the operation results corresponding to the target number of floating-point arithmetic units called based on the precision of the operand to be processed, according to the operation mode signal.
19. The floating-point processing apparatus according to any one of claims 15 to 18, characterized in that, The plurality of floating-point arithmetic units include a plurality of multiplication units and a plurality of addition units, wherein: The configuration module is used to call a target number of the multiplication units and a target number of the addition units among the plurality of addition units, based on the precision of the operands to be processed. The multiplication unit is used to perform multiplication operations on the first operand and the second operand in the operands to be processed to obtain the multiplication result; The addition unit is used to perform an addition operation on the multiplication result and the third operand in the operand to be processed to obtain the operation result.
20. The floating-point processing apparatus according to claim 19, characterized in that, The floating-point processing device further includes a first trigger unit and a second trigger unit, wherein: The first triggering unit is used to input the multiplication result to a target number of the addition units during a preset clock cycle; The second triggering unit is used to input the flag information of the multiplication result to the target number of the addition units during a preset clock cycle.
21. A floating-point processing device, characterized in that, The floating-point processing device includes registers and a plurality of floating-point arithmetic units as described in any one of claims 1 to 14, wherein: The register is used to store at least two floating-point numbers in the same clock cycle; The floating-point arithmetic unit is used to obtain the operand to be processed from the register.
22. The floating-point processing apparatus according to claim 21, characterized in that, The floating-point arithmetic unit includes a multiplication unit, and the register includes a first register, wherein: The first register is used to store a first operand and a second operand in the same clock cycle, and to input the operand output by the first register into at least one multiplication unit; The at least one multiplication unit is used to perform multiplication operations on the operands output by the first register to determine at least one multiplication result.
23. The floating-point processing apparatus according to claim 21, characterized in that, The floating-point arithmetic unit includes an addition unit, and the register includes a second register, wherein: The second register is used to store the third and fourth operands in the same clock cycle, and to input the operands output by the second register into at least one addition unit; The at least one addition unit is used to perform addition operations on the operands output by the second register to determine at least one addition result.
24. The floating-point processing apparatus according to claim 21, characterized in that, The floating-point arithmetic unit includes a multiplication unit and an addition unit, and the register includes a first register and a second register, wherein: The first register is used to store a first operand and a second operand in the same clock cycle, and to input the operand output by the first register into at least one multiplication unit; The at least one multiplication unit is used to perform multiplication operations on the operands output by the first register to determine at least one multiplication result; The second register is used to store the third operand and the at least one multiplication result in the same clock cycle, and to input the operand output by the second register into at least one addition unit; The at least one addition unit is used to perform addition operations on the operands output by the second register to determine at least one addition result.
25. The floating-point processing apparatus according to claim 22 or 24, characterized in that, The floating-point processing device further includes a reallocation unit, wherein: The reallocation unit is used to reallocate inputs to the at least one multiplication unit when the first register outputs, and to input the reallocated operands into the at least one multiplication unit.
26. A floating-point processing method, characterized in that, The method includes: The logic processing module performs logical operations on the input operands to obtain the results. The delay control module controls the delay between the input of the operand to be processed into the logic processing module and the input floating-point arithmetic unit; and / or controls the delay between the output of the calculation result by the logic processing module and the output of the calculation result by the floating-point arithmetic unit; The delay control module includes an input pipeline register and an output pipeline register, and the method further includes: The input pipeline register controls whether to clock at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, according to the first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit meets the first delay requirement; and, The output pipeline register controls whether to clock at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the calculation result according to the second configuration parameter, so that the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit meets the second delay requirement; The first configuration parameter and the second configuration parameter are designed according to the overall requirements of the floating-point arithmetic unit so that the floating-point arithmetic unit can satisfy the delay balance when outputting the arithmetic result.
27. The method according to claim 26, characterized in that, At least one intermediate pipeline register is provided between the input pipeline register and the output pipeline register, and the method further includes: The at least one intermediate pipeline register performs logical operations on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, and determines the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
28. A floating-point processing method, characterized in that, The method includes: The configuration module calls a target number of floating-point arithmetic units from among a plurality of floating-point arithmetic units as described in any one of claims 1 to 14, depending on the precision of the operand to be processed; The floating-point arithmetic unit performs floating-point operations on the operands to be processed to obtain the calculation result; wherein, the target quantity is related to the precision of the operands to be processed.
29. The method according to claim 28, characterized in that, The method further includes: The configuration module truncates the operands to be processed according to their precision to obtain a target number of data segments, and inputs the target number of data segments into the target number of floating-point arithmetic units.
30. The method according to claim 28, characterized in that, The method further includes: The configuration module calls a target number of the multiplication units and a target number of the addition units in multiple addition units, based on the precision of the operands to be processed. The multiplication unit performs a multiplication operation on the first operand and the second operand in the operands to be processed to obtain the multiplication result; The addition unit performs an addition operation on the multiplication result and the third operand in the operands to be processed to obtain the operation result.
31. A floating-point processing method, characterized in that, The method includes: Registers store at least two floating-point numbers in the same clock cycle; The floating-point arithmetic unit obtains the operand to be processed from the register so that the floating-point arithmetic unit performs the floating-point processing method as described in claim 26 or 27.
32. The method according to claim 31, characterized in that, The floating-point arithmetic unit includes a multiplication unit and an addition unit, the register includes a first register and a second register, and the method further includes: The first register stores the first operand and the second operand in the same clock cycle, and inputs the operand output by the first register into at least one multiplication unit; The at least one multiplication unit performs a multiplication operation on the operands output by the first register to determine at least one multiplication result; The second register stores the third operand and the at least one multiplication result in the same clock cycle, and inputs the operand output by the second register into at least one addition unit; The at least one addition unit performs an addition operation on the operands output by the second register to determine at least one addition result.
33. The method according to claim 32, characterized in that, The method further includes: When the first register outputs, the reallocation unit reallocates the input to the at least one multiplication unit and inputs the reallocated operands into the at least one multiplication unit.
34. A chip, characterized in that, The chip includes: a floating-point arithmetic unit as described in any one of claims 1 to 14, or a floating-point processing device as described in any one of claims 15 to 20, or a floating-point processing device as described in any one of claims 21 to 25.
35. An electronic device, characterized in that, The electronic device includes a processor, wherein the processor includes: a floating-point arithmetic unit as claimed in any one of claims 1 to 14; or a floating-point processing device as claimed in any one of claims 15 to 20; or a floating-point processing device as claimed in any one of claims 21 to 25.
Citation Information
Patent Citations
Double-precision floating-point operation
CN109196465A
Fixed-floating-point SIMD multiply-add instruction fusion processing device and method and processor
CN117251132A
SRT operational circuit
CN118312129A