Floating point arithmetic unit, floating point processing device and method, chip and electronic equipment
Through the design of floating-point operation unit with time-controllable delay and mixed precision operation, the problem of floating-point operation unit being difficult to adjust in different application scenarios is solved, and the performance improvement of floating-point operation with high efficiency and low power consumption is achieved.
Patent Information
- Application Number
- CN202510719274.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-05-30
AI Technical Summary
Existing floating-point computing units are difficult to flexibly adjust in different application scenarios, resulting in increased accuracy loss and hardware complexity, limiting overall performance improvement.
The logic processing module and the delay control module are adopted to control the input and output delays to achieve a floating-point operation unit design, and the target number of floating-point operation units is called in multiple floating-point operation units for mixed precision design, combined with customized register optimization register design.
Effectively balance floating-point operation delay, improve computing efficiency, reduce power consumption, save circuit area, and improve overall performance.
Smart Images

Figure CN120233979A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of floating-point operation technologies, and particularly to a floating-point operation unit, a floating-point processing device, a method, a chip, and an electronic device. Background Art
[0002] With the rapid development of scientific computing, machine learning training, etc., multiplication-addition units capable of supporting floating-point data processing have emerged as the times require, such as floating-point multipliers, floating-point adders, floating-point multiply-adders, etc. They are widely used in fields such as scientific computing, digital signal processing, image processing, and machine learning.
[0003] In the related art, floating-point operation units usually adopt a fixed delay structure and are difficult to be flexibly adjusted in different application scenarios. Although some delay-controllable design schemes have been proposed currently, these design schemes still have some defects, such as precision loss, high hardware complexity, etc., thus limiting the improvement of the overall performance of floating-point operations. Summary of the Invention
[0004] This application proposes a floating-point operation unit, a floating-point processing device, a method, a chip, and an electronic device, which can effectively balance the delay of floating-point operations, improve the operation efficiency, and further enhance the overall performance of floating-point operations.
[0005] To achieve the above object, the technical solution of this application is implemented as follows: In a first aspect, an embodiment of this application provides a floating-point operation unit, which includes: A logic processing module, configured to perform a logic operation on the operand to be processed input to the floating-point operation unit to obtain an operation result; A delay control module, configured to control the delay between the input of the operand to be processed to the logic processing module and the input of the floating-point operation unit; and / or, control the delay between the output of the logic processing module of the operation result and the output of the floating-point operation unit of the operation result.
[0006] In a second aspect, an embodiment of this application provides a floating-point processing device, which includes a plurality of floating-point operation units as in the first aspect, wherein: The floating-point processing device is configured to call a target number of floating-point operation units among the plurality of floating-point operation units to perform floating-point operations on the operand to be processed according to the precision of the operand to be processed to obtain an operation result; wherein, the target number is related to the precision of the operand to be processed.
[0007] In a third aspect, an embodiment of this application provides a floating-point processing device, which includes a register and a plurality of floating-point operation units as in the first aspect, wherein: The register is configured to store at least two floating-point numbers in the same clock cycle; A floating-point arithmetic unit for obtaining operands to be processed from registers.
[0008] In a fourth aspect, an embodiment of the present application provides a floating-point processing method, which includes: A logic processing module performs a logic operation on the input operands to be processed to obtain an operation result; A delay control module controls the delay between the input of the operands to be processed to the logic processing module and the input to the floating-point arithmetic unit; and / or controls the delay between the output of the operation result by the logic processing module and the output of the operation result by the floating-point arithmetic unit.
[0009] In a fifth aspect, an embodiment of the present application provides a floating-point processing method, which includes: According to the precision of the operands to be processed, a target number of floating-point arithmetic units as described in the first aspect are called among multiple floating-point arithmetic units to perform floating-point arithmetic on the operands to be processed to obtain an operation result; wherein the target number is related to the precision of the operands to be processed.
[0010] In a sixth aspect, an embodiment of the present application provides a floating-point processing method, which includes: A register stores at least two floating-point numbers in the same clock cycle; A floating-point arithmetic unit obtains the operands to be processed from the register so that the floating-point arithmetic unit executes the floating-point processing method as described in the fourth aspect.
[0011] In a seventh aspect, an embodiment of the present application provides a chip, which includes: the floating-point arithmetic unit as described in the first aspect, or the floating-point processing device as described in the second aspect, or the floating-point processing device as described in the third aspect.
[0012] In an eighth aspect, an embodiment of the present application provides an electronic device, which includes a processor, wherein the processor includes: the floating-point arithmetic unit as described in the first aspect, or the floating-point processing device as described in the second aspect, or the floating-point processing device as described in the third aspect.
[0013] A floating-point arithmetic unit, a floating-point processing device, a method, a chip, and an electronic device provided by an embodiment of the present application. In the floating-point arithmetic unit, a logic processing module is configured to perform a logic operation on an operand to be processed input to the floating-point arithmetic unit to obtain an operation result; a delay control module is configured to control the delay between the input of the operand to be processed to the logic processing module and the input to the floating-point arithmetic unit; and / or, control the delay between the output of the logic processing module of the operation result and the output of the floating-point arithmetic unit of the operation result. In this way, for a floating-point arithmetic unit (such as a multiplication unit or an addition unit), the delay of floating-point arithmetic can be effectively balanced, and the delay of input and / or output during the entire processing process can be controlled. Based on this delay controllable strategy, waste of hardware resources can also be avoided, the operation efficiency can be improved, and the power consumption can be reduced. In addition, according to the precision of the operand to be processed, a target number of floating-point arithmetic units can be called among multiple floating-point arithmetic units to perform floating-point arithmetic on the operand to be processed to obtain an operation result, and the target number is related to the precision of the operand to be processed. In this way, a mixed-precision design can also be realized, so as to significantly improve the processing ability of the hardware while ensuring the precision requirements; and based on the reuse of basic modules such as multiplication units and addition units, the power consumption and resource consumption are further reduced, and the circuit area is saved. In addition, the design of registers can be optimized, and the standard registers in the related technology are replaced with customized devices, which can not only process data more efficiently, but also solve the power consumption waste caused by excessive unnecessary timing operations in the related technology, thereby improving the overall performance of floating-point arithmetic. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 Schematic diagram of the composition structure of a floating-point arithmetic unit provided by an embodiment of the present application Figure 1 ; Figure 2 Schematic diagram of the composition structure of a floating-point arithmetic unit provided by an embodiment of the present application Figure 2 ; Figure 3 Schematic diagram of the composition structure of a multiplication unit provided by an embodiment of the present application; Figure 4 Schematic diagram of the composition structure of an addition unit provided by an embodiment of the present application; Figure 5 Schematic diagram of the composition structure of a floating-point processing device provided by an embodiment of the present application Figure 1 ; Figure 6A Schematic diagram of the parameter configuration interface with multiple precisions provided by an embodiment of the present application Figure 1 ; Figure 6B Schematic diagram of the parameter configuration interface with multiple precisions provided by an embodiment of the present application Figure 2 ; Figure 6CSchematic diagram of a multi-precision parameter configuration interface provided by an embodiment of the present application Figure 3 ; Figure 6D Schematic diagram of a multi-precision parameter configuration interface provided by an embodiment of the present application Figure 4 ; Figure 7 Schematic diagram of the composition structure of a floating-point processing device provided by an embodiment of the present application Figure 2 ; Figure 8 Schematic diagram of the application framework of a floating-point processing device provided by an embodiment of the present application Figure 1 ; Figure 9 Schematic diagram of the application framework of a floating-point processing device provided by an embodiment of the present application Figure 2 ; Figure 10 Schematic diagram of the composition structure of a floating-point processing device provided by an embodiment of the present application Figure 3 ; Figure 11 Schematic diagram of the composition structure of a floating-point processing device provided by an embodiment of the present application Figure 4 ; Figure 12 Schematic diagram of the composition structure of a floating-point processing device provided by an embodiment of the present application Figure 5 ; Figure 13 Schematic diagram VI of the composition structure of a floating-point processing device provided by an embodiment of the present application; Figure 14 Schematic diagram of the process of a floating-point processing method provided by an embodiment of the present application Figure 1 ; Figure 15 Schematic diagram of the process of a floating-point processing method provided by an embodiment of the present application Figure 2 ; Figure 16 Schematic diagram of the process of a floating-point processing method provided by an embodiment of the present application Figure 3 . Detailed implementation manners
[0015] In order to understand the features and technical content of the embodiments of the present application in more detail, the implementation of the embodiments of the present application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are only for reference and illustration purposes and are not used to limit the embodiments of the present application.
[0016] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0017] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0018] It should also be noted that the terms "first / second / third" involved in the embodiments of the present application are only used to distinguish similar objects and do not represent a specific order for the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0019] To facilitate the understanding of the technical solutions of the embodiments of the present application, the following describes the relevant terms and related technologies of the embodiments of the present application. The following related technologies can be arbitrarily combined with the technical solutions of the embodiments of the present application as optional solutions, and they all fall within the protection scope of the embodiments of the present application.
[0020] A floating-point processing unit (FPU) is a structure that performs floating-point operations. The FPU is mainly responsible for performing mathematical operations involving floating-point numbers, such as basic arithmetic operations like addition, subtraction, multiplication, and division. Compared with the integer arithmetic unit, the FPU directly supports numerical calculations with decimal points through hardware circuits, significantly improving the accuracy and speed in scenarios such as scientific computing and graphics rendering. In addition, the FPU can also handle transcendental function operations (such as trigonometric functions and logarithmic functions), further expanding its application scope.
[0021] Multiply Accumulate (MAC) is a special operation in a digital signal processor or some microprocessors. The hardware circuit unit that implements this operation is called a "multiplier accumulator". The operation of this operation is to add the product result of the multiplication and the value in accumulator A, and then store it in the accumulator.
[0022] Not A Number (NAN) is a class of values in the numerical data type in computer science, representing undefined or non-representable values. It is often used in floating-point arithmetic.
[0023] A Leading Zero Counter (LZC) is a counter used to count the number of leading zeros in a binary number. In the leading zero counter, when the most significant bit of the input binary number is 0, the counter increments by 1; when the most significant bit is 1, it does not count.
[0024] It should be understood that a floating-point number mainly consists of three parts: a sign bit (Sign), an exponent bit (Exponent, exp), and a mantissa bit. Its encoding format is shown in Table 1.
[0025] Table 1
[0026] Taking single-precision floating-point numbers as an example, single-precision floating-point numbers occupy 32 bits and are divided into the following three parts: The sign bit (1 bit) is used to represent the positive or negative of the floating-point number; 0 represents a positive number, and 1 represents a negative number.
[0027] The exponent bit can also be called the "exponent code" (8 bits) and is used to represent the exponent part of the floating-point number; using the offset representation method, the actual exponent value is the stored value minus 127 (offset).
[0028] The mantissa bit (23 bits) represents the fractional part of the floating-point number.
[0029] In the embodiments of the present application, floating-point numbers include multiple formats, namely: normal numbers, subnormal numbers, and special numbers. Special numbers can include plus / minus zero, plus / minus infinity, and NaN. Among them, the exponent bit and mantissa bit of plus / minus zero are all 0, the exponent bit of plus / minus infinity is all 1 and the mantissa bit is all 0, the exponent bit of NaN is all 1 and the mantissa bit is not 0, the exponent bit of subnormal numbers is all 0 and the mantissa bit is not 0, and the rest of the cases are represented as normal numbers.
[0030] It should also be understood that in floating-point operations, for a unit that has both multiplication and addition functions, the multiply-accumulate unit (or called "floating-point multiply-accumulate unit") is the core component for processing real-number operations in modern computer systems. They are widely used in fields such as scientific computing, digital signal processing, and image processing. The design and implementation of these units need to follow the IEEE754 standard, which defines the representation and operation rules of floating-point numbers. With the development of scientific computing and machine learning, the demand for multi-precision floating-point operations is increasing day by day. Traditional fixed-point multipliers have a fixed input bit width and are difficult to meet the requirements of multi-precision calculations. Therefore, methods for supporting multi-precision floating-point multiplication operations have emerged, such as reconfigurable floating-point multiply-accumulate units. These units can dynamically adjust the precision according to needs, improve hardware utilization, and reduce bit redundancy. In terms of hardware implementation, the design of floating-point arithmetic units needs to balance speed, area, and power consumption. That is to say, in the process of implementing a high-speed and high-precision floating-point multiplier, the following main challenges need to be solved currently: 1. Timing Optimization and Pipeline Design: The greater the pipeline depth, the less computational effort required for each stage. However, pipeline stalls (bubbles) and data dependencies can also lead to performance bottlenecks. Balancing the pipeline depth and throughput is a key issue in the design.
[0031] 2. Power Consumption and Area Control: High-performance multipliers often employ complex parallel adder structures such as Wallace trees, but this incurs a large silicon area and power consumption. In mobile terminals and embedded systems, low-power design is particularly important, so optimized logic circuits and clock gating techniques are required.
[0032] 3. Error Control and Rounding Strategy: Floating-point operations require the result to meet the precision requirements of the IEEE 754 standard. During mantissa multiplication and normalization, rounding errors and improper rounding mode selection may occur. To ensure the correctness of the numerical value, a multi-level rounding verification mechanism is usually adopted, and a fine-grained control circuit is implemented in hardware.
[0033] It should also be understood that in the field of modern high-performance computing, the design of floating-point multipliers and adders faces strict requirements for latency, power consumption, and area. Traditional floating-point arithmetic units usually adopt a fixed latency structure, making it difficult to flexibly adjust in different application scenarios. In response to this challenge, in recent years, researchers have proposed various design schemes for floating-point multipliers and adders with controllable latency.
[0034] For example, a multi-channel floating-point multiply-add structure divides the processing process into four data paths by analyzing the relationship between operands and the type of operation to adapt to different situations and avoid unnecessary processing steps, thereby improving the operation speed and reducing power consumption. However, this design may lead to an increase in hardware complexity, which in turn affects the area and power consumption. Another example is the delay optimization of floating-point multipliers. Researchers have proposed a high-precision and low-power approximate floating-point multiplier design based on the probability analysis of partial products. This design proposes an approximate 4-2 compressor and a low-order OR gate compression method by analyzing the probability of partial products being 1, thus effectively reducing the consumption of hardware resources, power consumption, and delay. However, this design may have a problem of error accumulation in some applications with high-precision requirements.
[0035] That is to say, although there are already some design schemes for floating-point multiply-add units with controllable latency in related technologies, these design schemes still have some defects, such as precision loss and high hardware complexity, which limit the improvement of the overall performance of floating-point operations.
[0036] Based on this, embodiments of the present application provide a floating-point arithmetic unit, a floating-point processing device, a method, a chip, and an electronic device. For a floating-point arithmetic unit (such as a multiplication unit or an addition unit), it can effectively balance the latency of floating-point operations, and can control the latency of input and / or output during the entire processing process. Based on this latency controllable strategy, it can also avoid waste of hardware resources, improve the operation efficiency, and reduce power consumption. In addition, according to the precision of the operand to be processed, the target number of floating-point arithmetic units can be called from multiple floating-point arithmetic units to perform floating-point operations on the operand to be processed, so that a mixed-precision design can be realized, thereby significantly improving the processing ability of the hardware while ensuring the precision requirements; and based on the reuse of basic modules such as multiplication units and addition units, the power consumption and resource consumption are further reduced, and the circuit area is saved. In addition, the design of the register can be optimized, and the standard register in the related technology is replaced by a customized device, which can not only process data more efficiently, but also solve the power consumption waste caused by excessive unnecessary timing operations in the related technology, thereby improving the overall performance of floating-point operations.
[0037] The following will describe each embodiment of the present application in detail with reference to the accompanying drawings.
[0038] In an embodiment of the present application, Figure 1 is a schematic structural diagram of a floating-point arithmetic unit provided by an embodiment of the present application Figure 1 . As Figure 1 shown, the floating-point arithmetic unit 10 may include a logic processing module 101 and a latency control module 102, where: The logic processing module 101 is configured to perform a logic operation on the operand to be processed input to the floating-point arithmetic unit 10 to obtain an operation result; The latency control module 102 is configured to control the latency between the operand to be processed input to the logic processing module 101 and the input to the floating-point arithmetic unit 10; and / or, control the latency between the operation result output by the logic processing module 101 and the operation result output by the floating-point arithmetic unit 10.
[0039] In the embodiment of the present application, the logic processing module 101 is configured to perform a logic operation on the operand to be processed input to the floating-point arithmetic unit 10, such as a multiplication operation or an addition operation, to obtain a corresponding operation result. The latency control module 102 can achieve controllable latency before the operand to be processed is input to the logic processing module 101, and / or, achieve controllable latency before the operand to be processed is output from the floating-point arithmetic unit 10; thereby effectively balancing the latency of the logic operation, improving the operation efficiency of the floating-point arithmetic unit 10, and ensuring accurate result output.
[0040] In the embodiments of the present application, the floating-point arithmetic unit 10 may adopt a pipeline structure, dividing the entire floating-point processing process into several pipeline registers, which can effectively avoid the performance bottleneck caused by excessive pipeline delay. Exemplarily, the entire floating-point processing process may be divided into: an input pipeline register, at least one intermediate pipeline register, and an output pipeline register, thereby further improving the pipeline working efficiency and saving time.
[0041] In some embodiments, referring to Figure 2 , the delay control module 102 may include an input pipeline register 1021 and an output pipeline register 1022, where: The input pipeline register 1021 is used to control the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed; and / or, The output pipeline register 1022 is used to control the output path delay of at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result.
[0042] It should be noted that in the pipeline structure, the input pipeline register 1021 and the output pipeline register 1022 are controllable. That is to say, whether the input pipeline register 1021 and the output pipeline register 1022 are enabled can be controlled according to the corresponding configuration parameters, so as to achieve the purpose of delay balance.
[0043] In a possible implementation manner, for the input pipeline register 1021, the input pipeline register 1021 is used to control whether to beat at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit meets the first delay requirement.
[0044] In another possible implementation manner, for the output pipeline register 1022, the output pipeline register 1022 is used to control whether to beat at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result according to the second configuration parameter, so that the output path delay of at least one of the target sign bit, the target exponent bit, and the target mantissa bit meets the second delay requirement.
[0045] In the embodiments of the present application, the first delay requirement and the second delay requirement are preset and used to measure whether the path delay of each segment meets the requirements. In addition, the first configuration parameter and the second configuration parameter can be designed according to the overall requirements of the floating-point arithmetic unit 10 to control whether the corresponding pipeline register is enabled.
[0046] Exemplarily, the input pipeline register 1021 can be controlled to be enabled or not according to the first configuration parameter, that is, according to the first configuration parameter, at least one of the initial sign bit, initial exponent bit, and initial mantissa bit of the input is controlled to be latched; the output pipeline register 1022 can be controlled to be enabled or not according to the second configuration parameter, that is, according to the second configuration parameter, at least one of the target sign bit, target exponent bit, and target mantissa bit of the output is controlled to be latched.
[0047] It should also be noted that in the embodiments of the present application, for the controllable input pipeline register 1021 and output pipeline register 1022, these two pipeline registers can be all enabled, or can be all disabled, or only one of the pipeline registers can be enabled (for example, enable the input pipeline register 1021 and disable the output pipeline register 1022; or enable the output pipeline register 1022 and disable the input pipeline register 1021), and no limitation is made here.
[0048] That is to say, in the embodiments of the present application, the input pipeline register and output pipeline register of the floating-point arithmetic unit 10 can be controllably designed according to the corresponding configuration parameters. Here, the input pipeline register can control whether to latch at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, and the output pipeline register can control whether to latch at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result according to the second configuration parameter. In this way, when the delay of a certain path is large, by latching other paths, the purpose of delay balance can be achieved.
[0049] In some embodiments, continue to refer to Figure 2 , at least one intermediate pipeline register 1011 is provided between the input pipeline register 1021 and the output pipeline register 1022, where: At least one intermediate pipeline register 1011 is used to perform logical operations on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed, and determine the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0050] In the embodiments of the present application, the operand to be processed includes a first operand and a second operand. Here, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit, and the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit. Among them, the target sign bit can be obtained by performing a logical operation on the first initial sign bit and the second initial sign bit, the target exponent bit can be obtained by performing a logical operation on the first initial exponent bit and the second initial exponent bit, and the target mantissa bit can be obtained by performing a logical operation on the first initial mantissa bit and the second initial mantissa bit.
[0051] Thus, in the embodiments of the present application, the logic processing process can be divided into at least one intermediate pipeline register, and computing tasks are allocated between each intermediate pipeline register, so as to optimize the execution time of floating-point processing, avoid performance bottlenecks caused by too long pipeline delays, and improve the computing efficiency.
[0052] It can be understood that, in the embodiments of the present application, the floating-point operation unit 10 can be a multiplication unit (or referred to as "floating-point multiplication unit", "multiplier", etc.), and / or the floating-point operation unit 10 can also be an addition unit (or referred to as "floating-point addition unit", "adder", etc.). Hereinafter, the floating-point operation unit 10 will be described in detail by taking the multiplication unit and the addition unit as examples respectively.
[0053] In a possible implementation manner, the floating-point operation unit 10 can be a multiplication unit. In this case, at least one intermediate pipeline register can include one pipeline register (i.e., the first intermediate pipeline register). That is to say, the multiplication unit adopts a three-stage pipeline design, for example, an input pipeline register (the first stage of the pipeline), a first intermediate pipeline register (the second stage of the pipeline), and an output pipeline register (the third stage of the pipeline), and the input pipeline register and the output pipeline register can be controlled whether to insert a cycle according to the corresponding configuration parameters. Therefore, the multiplication unit can be called a "controllable three-stage pipeline multiplier".
[0054] That is to say, in the multiplication unit, there is an intermediate pipeline register between the input pipeline register and the output pipeline register, and the input pipeline register and the output pipeline register are controllable. That is, the input pipeline register can control whether to insert a cycle for at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, and the output pipeline register can control whether to insert a cycle for at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result according to the second configuration parameter.
[0055] In addition, in the embodiments of the present application, for the first operand and the second operand included in the operand to be processed, where the first operand can be the multiplicand and the second operand can be the multiplier; or the first operand can be the multiplier and the second operand can be the multiplicand; there is no limitation here. Exemplarily, assume that the first operand is the multiplicand x and the second operand is the multiplier y, then the operation result C = x × y.
[0056] Exemplarily, Figure 3 is a schematic diagram of the composition structure of a multiplication unit provided by the embodiments of the present application. As Figure 3As shown, from the perspective of the pipeline, this may include: an input pipeline register (first-stage pipeline) A1, a first intermediate pipeline register (second-stage pipeline) A2, and an output pipeline register (third-stage pipeline) A3.
[0057] In some embodiments, referring further to Figure 3 , the logic processing module 101 may include a first logic module (such as Figure 3 301-1 or 301-2 in ) and a second logic module 302, where: The first logic module is used to perform logical operations on the first initial sign bit and the second initial sign bit, and the first initial exponent bit and the second initial exponent bit respectively to determine the target sign bit and the target exponent bit;
[0058] In the embodiments of the present application, the second logic module 302 is located in the first intermediate pipeline register (i.e., the second-stage pipeline). For the first logic module (such as Figure 3 301-1 or 301-2 in
[0059] ), when the input pipeline register (i.e., the first-stage pipeline) is enabled, the first logic module 301-1 is located in the input pipeline register (i.e., the first-stage pipeline); or when the input pipeline register (i.e., the first-stage pipeline) is disabled, the first logic module 301-2 is located in the first intermediate pipeline register (i.e., the second-stage pipeline). That is to say, the position of the first logic module may be related to whether the input pipeline register (i.e., the first-stage pipeline) is enabled. Figure 3 In some embodiments, taking the first logic module 301-1 as an example, referring further to
[0060] , the first logic module 301-1 may include a sign bit processing unit 3011 and an exponent bit processing unit 3012. Among them, the sign bit processing unit 3011 is used to perform a sign exclusive OR operation on the first initial sign bit and the second initial sign bit to obtain the target sign bit; the exponent bit processing unit 3012 is used to perform an exponent addition operation on the first initial exponent bit and the second initial exponent bit to obtain the target exponent bit.
[0061] In the embodiments of the present application, for a target exponent bit, it can be obtained by performing an exponent addition operation on the initial exponent bits corresponding to two operands (for example, a first operand and a second operand). Here, the exponent bit is shifted based on a bias. Exemplarily, bias = 2 (k-1) -1, where k is the number of bits corresponding to the exponent bit. Thus, if the operand is a single-precision floating-point number, at this time k = 8, then bias is 127; if the operand is a half-precision floating-point number, at this time k = 5, then bias is 15; if the operand is a double-precision floating-point number, at this time k = 11, then bias is 1023.
[0062] That is to say, in the embodiments of the present application, for the exponent bit of each operand, if the operand is a single-precision floating-point number, then the exponent bit is used for 8-bit storage, and its storage format is the sum of the exponent value and 127; if the operand is a half-precision floating-point number, then the exponent bit is used for 5-bit storage, and its storage format is the sum of the exponent value and 15; otherwise, if the operand is a double-precision floating-point number, then the exponent bit is used for 11-bit storage, and its storage format is the sum of the exponent value and 1023.
[0063] Exemplarily, assuming that the initial exponent bits corresponding to the two operand formats are E1 and E2 respectively (i.e., the actual values of the binary exponent fields), then the exponent addition operation during multiplication is: actual exponent = (E1 - bias) + (E2 - bias) = E1 + E2 - 2×bias. Finally, it is necessary to add the single offset to this calculation result to obtain the target stored exponent value (i.e., the target exponent bit), specifically it can be E1 + E2 - bias. It should be noted here that in the multiplication operation, in fact, two actual exponents are added, rather than directly adding the stored exponents.
[0064] It should also be noted that for each operand, in addition to the initial sign bit and the initial exponent bit, a corresponding flag can be introduced to indicate whether the operand is a special number, such as zero, NAN, positive or negative infinity, or overflow, etc. Exemplarily, the first operand can also include a first initial flag bit for indicating whether the first operand is a special number; the second operand can also include a second initial flag bit for indicating whether the second operand is a special number.
[0065] In some embodiments, still taking the first logic module 301-1 as an example, continue to refer to Figure 3, the first logic module 301-1 may further include a flag bit processing unit 3013. The flag bit processing unit 3013 is configured to perform a logical operation based on a first initial flag bit and a second initial flag bit to determine a target flag bit corresponding to the operation result, and the target flag bit is used to indicate whether the operation result is a special number.
[0066] That is to say, in the embodiments of the present application, for each operand, in addition to the initial sign bit, initial exponent bit, and initial mantissa bit, there is also an initial flag bit flag(zero, nan, overflow) for determining whether the corresponding operand is a special number.
[0067] Exemplarily, when at least one of the first operand and the second operand is a special number, the following several cases are described: If the first operand is NAN and the second operand is any value, the operation result C is NAN, that is, the operation result of any value and NAN is NAN.
[0068] If the first operand is zero and the second operand is a non-zero finite number, the operation result C is zero, that is, the result of multiplying zero by a finite number is still 0, and the target sign bit is determined by performing an exclusive OR operation on the sign bits of the two operands.
[0069] If the first operand is zero and the second operand is positive or negative infinity, the operation result C is NAN, that is, multiplying zero by infinity is an undefined operation and returns NAN.
[0070] If the first operand is positive or negative infinity and the second operand is a non-zero finite number, the operation result C is positive or negative infinity, that is, multiplying infinity by a non-zero finite number results in infinity, and the target sign bit is determined by performing an exclusive OR operation on the sign bits of the two operands.
[0071] If the first operand is positive or negative infinity and the second operand is positive or negative infinity, the operation result C is positive or negative infinity, that is, multiplying infinity with the same sign results in positive infinity, and multiplying infinity with different signs results in negative infinity.
[0072] It can also be understood that in the second logic module 302, the target mantissa bit can be obtained by performing a mantissa operation on the mantissa bits corresponding to the multiplicand and the multiplier respectively. In some embodiments, the second logic module 302 may include a mantissa product unit 3021 and a rounding operation unit 3022. The mantissa product unit 3021 is configured to perform a mantissa multiplication operation on the first initial mantissa bit and the second initial mantissa bit to obtain a first mantissa product; the rounding operation unit 3022 is configured to perform a rounding operation on the first mantissa product to determine the target mantissa bit.
[0073] In the embodiments of the present application, for the mantissa multiplication operation, first, the implicit bits of the mantissa bits corresponding to each operand need to be restored. Exemplarily, if the operand is a normalized number, then the corresponding mantissa bits are added with an implicit leading 1 to form the mantissa 1.M actually participating in the operation (for example, if the stored mantissa bits M = 101, then the actual mantissa value is 1.101); if the operand is a denormalized number, then the corresponding mantissa bits are added with an implicit leading 0 to form the mantissa 0.M actually participating in the operation. Then, the actual mantissas of the two operands (for example, 1.M1 and 1.M2) are multiplied in unsigned binary, and the result is twice the length of the mantissa bits; for example, if it is single-precision (23-bit mantissa), then after the mantissas are multiplied, the first mantissa product is 46 bits; if it is double-precision (52-bit mantissa), then after the mantissas are multiplied, the first mantissa product is 104 bits.
[0074] Exemplarily, assume that the mantissa of the single-precision multiplicand is 1.0100..., and the mantissa of the single-precision multiplier is 1.1000..., then the first mantissa product is 1.0100...×1.1000... = 1.1110...; at this time, it is necessary to further perform a rounding operation on the first mantissa product.
[0075] In the embodiments of the present application, the rounding operation may include truncating rounding, rounding up, rounding down, or rounding to the nearest. Among them, truncating rounding is also called "rounding towards 0", and its rule is to directly truncate the redundant bits regardless of the value of the redundant bits and does not change the numerical sign. Rounding up is also called "rounding towards positive infinity", and its rule is that if the redundant bits of a positive number are non-zero, it carries, and for a negative number, it is directly truncated. Rounding down is also called "rounding towards negative infinity", and its rule is that if the redundant bits of a negative number are non-zero, it adjusts to a smaller direction, and for a positive number, it is directly truncated. Rounding to the nearest is also called "rounding towards an even number", and its rule is to preferentially select the rounding value closest to it. If it is a middle value (such as the redundant bits are 1000...), it rounds to the nearest even number. Exemplarily, for the floating-point number 2.5 (binary is 10.1), the value of the rounding operation is 2; for the floating-point number 1.375 (binary is 1.011, retaining 2 mantissa bits), the value of the rounding operation is 1.4.
[0076] In some embodiments, the rounding operation unit 3022 is further configured to perform a rounding operation on the first mantissa product to obtain a second mantissa product when the operation result is a denormalized number; and when the second mantissa product exceeds a preset range, perform a normalization process on the second mantissa product to obtain the target mantissa bits and perform a carry operation on the target exponent bits to transform the multiplication result from a denormalized number to a normalized number.
[0077] It should be noted that in the embodiments of the present application, whether the operation result is a denormal number can be determined according to whether the target exponent bits are all zeros, and whether the second mantissa product exceeds the preset range can refer to whether the second mantissa product overflows. Among them, if the second mantissa product does not exceed the preset range, then the exponent does not need to be adjusted, that is, there is no need to perform a carry operation on the target exponent bits; if the second mantissa product exceeds the preset range, then the exponent needs to be adjusted, that is, a carry operation is performed on the target exponent bits. That is to say, when the multiplication result of two initial mantissa bits undergoes a rounding operation, the following two situations may occur: Situation 1: The second mantissa product after the rounding operation does not overflow (for example, 1.111...111 → 1.000...000), and at this time, the exponent does not need to be adjusted.
[0078] Situation 2: The second mantissa product after the rounding operation overflows (for example, 1.111...111 + 1 → 10.000...000), and at this time, it needs to be shifted right by 1 bit, that is, the exponent is incremented by 1 (carry).
[0079] It should also be noted that in the embodiments of the present application, the rounding operation can be performed according to the IEEE-754 rule. For the first mantissa product, the processing of normalized numbers and denormalized numbers are two different branches and cannot be combined. Exemplarily, in one possible implementation, if the operation result is a normalized number, when the guard bit (G), round bit (R), and sticky bit (S) respectively meet the preset combination conditions, a carry is made forward (specifically, a carry is made to the target exponent bits, that is, the target exponent bits are adjusted by adding 1). In another possible implementation, if the operation result is a denormal number, after the rounding operation on the first mantissa product, if a carry is made to the target exponent bits, that is, the target exponent bits are adjusted by adding 1, then the target exponent bits will not be all zeros, resulting in the operation result changing from a denormal number to a normalized number. Thus, after performing the mantissa multiplication operation on the initial mantissa bits corresponding to the two operands respectively, when dealing with underflow numbers, this situation needs to be taken into account, and the case of denormal number carry is separately processed when assigning values to the final target exponent bits.
[0080] It should also be noted that in the embodiments of the present application, the multiplication unit may adopt a controllable three-stage pipelined structure. Among them, for the input pipeline register (i.e., the first-stage pipeline), it can be determined whether to be enabled according to the first configuration parameter, that is, to control whether to insert a pipeline stage for the input according to the first configuration parameter; for the output pipeline register (i.e., the third-stage pipeline), it can be determined whether to be enabled according to the second configuration parameter, that is, to control whether to insert a pipeline stage for the output according to the second configuration parameter. In a possible implementation, if both the input pipeline register and the output pipeline register are closed, then the multiplication unit at this time can be regarded as a single-stage pipelined structure.
[0081] It should also be noted that in the embodiments of the present application, the multiplication unit adopts a controllable three-stage pipelined structure. This is because the result of the mantissa multiplication needs one cycle to be obtained. Therefore, a first intermediate pipeline register is inserted between the input pipeline register and the output pipeline register, so that two initial mantissa bits are directly multiplied. At this time, the result of the mantissa multiplication arrives in the same cycle as the target flag bit, the target sign bit, and the target exponent bit. Here, it is considered that the multiplication operation involves complex calculation logic, and compared with logical operations such as addition and exclusive OR gates, it requires multiple logic gates and a longer calculation path, so its delay is relatively large. Here, other paths can be pipelined to achieve the purpose of delay balance.
[0082] In another possible implementation, the floating-point operation unit 10 may be an addition unit. In this case, at least one intermediate pipeline register may include three pipeline registers (i.e., the first intermediate pipeline register, the second intermediate pipeline register, and the third intermediate pipeline register). That is to say, the addition unit adopts a five-stage pipeline design, such as an input pipeline register (the first stage), a first intermediate pipeline register (the second stage), a second intermediate pipeline register (the third stage), a third intermediate pipeline register (the fourth stage), and an output pipeline register (the fifth stage). Moreover, the input pipeline register and the output pipeline register can be controlled whether to insert a pipeline stage according to the corresponding configuration parameters. Therefore, the addition unit can be called a "controllable five-stage pipelined adder".
[0083] That is to say, in the addition unit, three intermediate pipeline registers are arranged between the input pipeline register and the output pipeline register, and the input pipeline register and the output pipeline register are also controllable. In the embodiment of the present application, the input pipeline register and the output pipeline register are the same as the input pipeline register and the output pipeline register in the multiplication unit, and whether they are enabled can be determined by corresponding configuration parameters. Exemplarily, the input pipeline register can control whether to clock at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, and the output pipeline register can control whether to clock at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result according to the second configuration parameter.
[0084] In addition, in the embodiment of the present application, for the first operand and the second operand included in the operand to be processed, where the first operand can be the addend and the second operand can be the augend; or, the first operand can be the augend and the second operand can be the addend; however, which of the operands is used as the addend can be determined by the magnitude of the exponent of the operand (i.e., the "initial exponent bit"). Exemplarily, assume that the first operand is the addend x and the second operand is the augend y, then the operation result C = x + y.
[0085] Exemplarily, Figure 4 is a schematic diagram of the composition structure of an addition unit provided by an embodiment of the present application. As Figure 4 shown, from the perspective of the pipeline, this may include: an input pipeline register (first-stage pipeline) B1, a first intermediate pipeline register (second-stage pipeline) B2, a second intermediate pipeline register (third-stage pipeline) B3, a third intermediate pipeline register (fourth-stage pipeline) B4, and an output pipeline register (fifth-stage pipeline) B5.
[0086] In some embodiments, continue to refer to Figure 4 , the logic processing module 101 may include a first logic module 401, a second logic module 402, a third logic module 403, and a fourth logic module 404, where: The first logic module 401 is used to determine the exponent difference between the first initial exponent bit and the second initial exponent bit and the smaller exponent of the two, and perform a shift operation on the mantissa bit corresponding to the smaller exponent according to the exponent difference to align the first initial mantissa bit and the second initial mantissa bit; The second logic module 402 is used to determine the sign exclusive OR result of the first initial sign bit and the second initial sign bit, and perform an addition operation on the aligned first initial mantissa bit and the second initial mantissa bit according to the sign exclusive OR result to obtain the first mantissa sum and the target sign bit; The third logic module 403 is configured to perform leading zero processing on the first mantissa sum, determine the leading zero statistical result and the shifted second mantissa sum; The fourth logic module 404 is configured to perform an exponent shift operation according to the leading zero statistical result to obtain a target exponent bit; and perform a rounding operation on the second mantissa sum to obtain a target mantissa bit.
[0087] It should be noted that for the first operand and the second operand, in the addition operation, it is first necessary to determine the addend and the augend according to the exponents of the operands. Exemplarily, if the first initial exponent bit is greater than the second initial exponent bit, that is, the exponent of the first operand is larger, then the first operand can be determined as the addend, and at this time the second operand is the augend; otherwise, if the second initial exponent bit is greater than the first initial exponent bit, that is, the exponent of the second operand is larger, then the second operand can be determined as the addend, and at this time the first operand is the augend. In addition, the embodiment of the present application can also set a register flag bit to indicate whether the addend and the augend need to be swapped.
[0088] It should also be noted that the slight difference in parsing the two operands between the addition unit and the multiplication unit is that it is necessary to wait until the positions of the operands are determined in the previous step before the operands can be parsed, but the parts of the two operands are not operated here. That is to say, the addition unit needs to parse after determining the positions of the operands by comparing the exponent bits, while the multiplication unit does not need to parse after comparing the exponent bits.
[0089] It should also be noted that the first logic module 401 is located in the input pipeline register (i.e., the first-level pipeline), the second logic module 402 is located in the first intermediate pipeline register (i.e., the second-level pipeline), the second logic module 402 is located in the first intermediate pipeline register (i.e., the second-level pipeline), the third logic module 403 is located in the second intermediate pipeline register (i.e., the third-level pipeline), and the fourth logic module 404 is located in the third intermediate pipeline register (i.e., the fourth-level pipeline). Among them, if the input pipeline register (i.e., the first-level pipeline) is closed, then the first logic module 401 is located in the first intermediate pipeline register (i.e., the second-level pipeline). That is to say, the position of the first logic module 401 may be related to whether the input pipeline register (i.e., the first-level pipeline) is enabled.
[0090] In some embodiments, continue to refer to Figure 4 , the first logic module 401 may include an exponent subtraction unit a1, a first control unit a2, a first selection unit (MUX1) a3, a second selection unit (MUX2) a4, a third selection unit (MUX3) a5, and a mantissa shift unit a6, where: The exponent subtraction unit a1 is configured to perform a subtraction operation on the first initial exponent bit and the second initial exponent bit to determine the exponent difference, and transmit the exponent difference to the first control unit a2; The first control unit a2 is configured to receive the exponent difference, and generate a first control signal transmitted to the first selection unit a3, a second control signal transmitted to the second selection unit a4, and a third control signal transmitted to the third selection unit a5 according to the exponent difference; The first selection unit a3 is configured to receive the first control signal, and select the larger exponent from the first initial exponent bit and the second initial exponent bit according to the first control signal; The second selection unit a4 is configured to receive the second control signal, and select the mantissa bit corresponding to the smaller exponent from the first initial mantissa bit and the second initial mantissa bit according to the second control signal; The third selection unit a5 is configured to receive the third control signal, and select the mantissa bit corresponding to the larger exponent from the first initial mantissa bit and the second initial mantissa bit according to the third control signal; The mantissa shift unit a6 is configured to perform a shift operation on the mantissa bit corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa bit and the second initial mantissa bit.
[0091] In the embodiment of the present application, the first logic module 401 is configured to align the decimal points of the corresponding initial mantissa bits of two operands according to the exponent difference between the two operands. Among them, the exponent subtraction unit a1 is configured to determine the exponent difference (i.e., "order difference") between the first initial exponent bit and the second initial exponent bit. Then, the first control unit a2 can generate a first control signal, a second control signal, and a third control signal according to the order difference. The first control signal is used to select the larger exponent from the first initial exponent bit and the second initial exponent bit, the second control signal is used to select the mantissa bit corresponding to the smaller exponent from the first initial mantissa bit and the second initial mantissa bit, and the third control signal is used to select the mantissa bit corresponding to the larger exponent from the first initial mantissa bit and the second initial mantissa bit. In this way, according to the order difference, the mantissa shift unit a6 performs a right shift operation on the mantissa bit corresponding to the smaller exponent, so as to be able to align the first initial mantissa bit and the second initial mantissa bit.
[0092] That is to say, in the addition unit, it is first necessary to determine the smaller exponent, and perform a right shift operation on the mantissa bit corresponding to the smaller exponent to align the two initial mantissa bits. During this process, it is also necessary to adjust the smaller exponent in the two operands to the larger one, that is, the two exponent bits are equal.
[0093] Exemplarily, assume that the exponent of the first operand is 3 and the mantissa is 1.101 (the actual value is equal to 1.101×2 3 , and the converted decimal is 13); the exponent of the second operand is 1 and the mantissa is 1.110 (the actual value is equal to 1.110×2 1, when converted to decimal it is 3.5). In this way, since the exponent of the second operand is smaller and the difference between the exponents of the two is 3 - 1 = 2, not only does the exponent of the second operand need to be incremented by 2 to align with the larger exponent 3, but also the mantissa of the second operand needs to be shifted right by 2 bits, that is, the mantissa of the second operand becomes 0.0111, to achieve the alignment of the decimal points of the two mantissas.
[0094] In some embodiments, the second logic module 402 may include an exclusive - or unit b1, a second control unit b2, and a mantissa addition unit b3, where: The exclusive - or unit b1 is configured to perform an exclusive - or operation on the first initial sign bit and the second initial sign bit to obtain an exclusive - or result of signs, and transmit the exclusive - or result of signs to the second control unit b2; The second control unit b2 is configured to receive the exclusive - or result of signs and generate a fourth control signal transmitted to the mantissa addition unit b3 according to the exclusive - or result of signs; The mantissa addition unit b3 is configured to perform an addition operation on the aligned first initial mantissa bits and the second initial mantissa bits according to the fourth control signal to obtain a first mantissa sum and a target sign bit.
[0095] In the embodiments of the present application, for the second logic module 402, it mainly calculates the first mantissa sum between the aligned first initial mantissa bits and the second initial mantissa bits. Among them, the exclusive - or unit b1 can perform an exclusive - or operation on the initial sign bits corresponding to the two operands to determine the exclusive - or result of signs, then transmit the exclusive - or result of signs to the second control unit b2, and transmit the fourth control signal generated by the second control unit b2 to the mantissa addition unit b3. The fourth control signal is used to determine whether to perform an addition operation or a subtraction operation on the aligned first initial mantissa bits and the second initial mantissa bits, so as to obtain the first mantissa sum and determine the target sign bit corresponding to the operation result according to the first mantissa sum.
[0096] It should also be noted that for two operands (such as the first operand and the second operand), if the first initial sign bit and the second initial sign bit are the same, then it can be determined that the exclusive - or result of signs is false (0); otherwise, if these two sign bits are different, then it can be determined that the exclusive - or result of signs is true (1).
[0097] In the embodiments of the present application, if the symbol exclusive OR result is false (0), then the fourth control signal is used to indicate an addition operation on the aligned first initial tail digit and the second initial tail digit, and the target symbol digit remains unchanged, that is, the target symbol digit is the same as the symbol digit of the first operand or the second operand. If the symbol exclusive OR result is true (1), then the fourth control signal is used to indicate a subtraction operation on the aligned first initial tail digit and the second initial tail digit, and the target symbol digit is determined by the symbol digit with the larger absolute value among the aligned first initial tail digit and the second initial tail digit. Exemplarily, if the first initial symbol digit is positive and the second initial symbol digit is negative, and the absolute value of the first initial tail digit is larger among the aligned first initial tail digit and the second initial tail digit, then the target symbol digit can be determined to be positive; otherwise, if the absolute value of the second initial tail digit is larger among the aligned first initial tail digit and the second initial tail digit, then the target symbol digit can be determined to be negative.
[0098] That is to say, in the embodiments of the present application, if the symbol exclusive OR result is false (0), then the aligned first initial tail digit and the second initial tail digit can be subjected to an addition operation to obtain a first mantissa sum, and the target symbol digit remains unchanged; if the symbol exclusive OR result is true (1), then the aligned first initial tail digit and the second initial tail digit can be subjected to a subtraction operation, specifically, the initial tail digit with the larger absolute value subtracts the initial tail digit with the smaller absolute value, and the target symbol digit is determined by the symbol digit with the larger absolute value among the aligned first initial tail digit and the second initial tail digit.
[0099] In some embodiments, the third logic module 403 may include a leading zero processing unit c1. The leading zero processing unit c1 is configured to perform leading zero prediction on the first mantissa sum to determine a leading zero statistical result; and perform a shift processing on the first mantissa sum according to the leading zero statistical result to obtain a second mantissa sum.
[0100] In the embodiments of the present application, for the leading zero processing unit c1, a strategy of gradually shifting and judging zero is adopted here to determine the situation where all high bits are zero at the corresponding stage to obtain the leading zero statistical result; then, the first mantissa sum is shifted according to the leading zero statistical result to obtain a second mantissa sum, and the highest bit of the mantissa of the second mantissa sum is 1.
[0101] That is to say, if there are many zeros in front of the first mantissa sum after addition, then LZC needs to be introduced and the first mantissa sum is shifted left. Exemplarily, for the case of adding denormalized numbers, or when the sign bits of the two operands are different, resulting in a very small first mantissa sum. For example, the exponent of the first operand is 2 and the mantissa is 1.1; the exponent of the second operand is 1 and the mantissa is 1.1. Then after alignment, the exponent of the second operand is adjusted to 2 and the mantissa is shifted right by one bit, that is, the mantissa of the second operand becomes 0.11. At this time, when adding the two mantissas, 1.1 plus -0.11 in binary can get 0.11, that is, 0.75 in decimal. At this time, the first mantissa sum is 0.11 and needs to be normalized, that is, shifted left by one bit to become 1.1, and the exponent is decreased by 1, that is, from 2 to 1.
[0102] In this process, the number of leading zeros (i.e., the leading zero count result) determines the number of bits to be shifted left. Exemplarily, if the first mantissa sum is 0.00101, then the leading zero count result is 3. At this time, the first mantissa sum needs to be shifted left by three bits to get 1.01 (the second mantissa sum), and the exponent is decreased by 3.
[0103] In some embodiments, the fourth logic module 404 may include an exponent shift unit d1 and a rounding operation unit d2. Among them, the exponent shift unit d1 is configured to receive the larger exponent sent by the first selection unit and the leading zero count result sent by the leading zero processing unit c1, and perform an exponent shift operation on the larger exponent according to the leading zero count result to obtain the target exponent bit; the rounding operation unit d2 is configured to perform a rounding operation on the second mantissa sum to obtain the target mantissa bit.
[0104] In the embodiments of the present application, the exponent shift unit d1 mainly performs exponent shifting according to the leading zero count result sent by the leading zero processing unit c1. Here, "exponent shifting" refers to performing a subtraction operation on the exponent size. Exemplarily, if the first mantissa sum is 0.00101, then the leading zero count result is 3. At this time, not only does the first mantissa sum need to be shifted left by three bits to get 1.01 (the second mantissa sum), but also a subtraction operation needs to be performed on the larger exponent, that is, the larger exponent is decreased by 3 to obtain the target exponent bit.
[0105] In the embodiments of the present application, the rounding operation may include truncating rounding, rounding up, rounding down, or rounding to the nearest. Among them, the rounding operation is similar to the rounding operation in the multiplication unit and will not be elaborated here.
[0106] In a possible implementation, the rounding operation unit d2 is further configured to perform a rounding operation on the second mantissa sum to obtain a third mantissa sum when the operation result is a denormal number; when the third mantissa sum exceeds a preset range, perform a normalization process on the third mantissa sum to obtain a target mantissa digit, and adjust the target exponent digit by adding 1, so that the operation result is transformed from a denormal number to a normal number.
[0107] It should be noted that, in the embodiments of the present application, whether the operation result is a denormal number can be determined according to whether the target exponent digit is all zeros, and whether the third mantissa sum exceeds the preset range can refer to whether the third mantissa sum overflows. Among them, if the third mantissa sum does not exceed the preset range, then there is no need to adjust the exponent, that is, there is no need to carry to the target exponent digit; if the third mantissa sum exceeds the preset range, then the exponent needs to be adjusted, that is, carry to the target exponent digit. That is to say, after performing the rounding operation on the second mantissa sum, there may be the following two cases: Case 1: The third mantissa sum after the rounding operation does not overflow (for example, 1.111...111 → 1.000...000), and at this time, there is no need to adjust the exponent.
[0108] Case 2: The third mantissa sum after the rounding operation overflows (for example, 1.111...111 + 1 → 10.000...000), and at this time, it needs to be shifted one bit to the right, that is, the exponent is incremented by 1 (carry).
[0109] It should also be noted that, in the embodiments of the present application, the processing of normal numbers and denormal numbers for the operation result obtained by the addition operation is also two different branches and cannot be combined. Exemplarily, in a possible implementation, if the operation result is a normal number, when the guard bit (G), round bit (R), and sticky bit (S) respectively meet the preset combination conditions, carry forward (specifically, carry to the target exponent digit, that is, adjust the target exponent digit by adding 1). In another possible implementation, if the operation result is a denormal number, after performing the rounding operation on the second mantissa sum, if carrying to the target exponent digit, that is, adjusting the target exponent digit by adding 1, then the target exponent digit will not be all zeros, thereby making the operation result transformed from a denormal number to a normal number.
[0110] It should also be noted that for Figure 4The addition unit shown has the output of the exponent subtraction unit a1 connected to the first control unit a2. The first output of the first control unit a2 is connected to the control terminal of the first selection unit a3. The second output of the first control unit a2 is connected to the control terminal of the second selection unit a4. The third output of the first control unit a2 is connected to the control terminal of the third selection unit a5. The fourth output of the first control unit a2 is connected to the control terminal of the mantissa shift unit a6. The output of the second selection unit a4 and the output of the mantissa shift unit a6 are respectively connected to two input terminals of the mantissa addition unit b3. The output of the exclusive OR unit b1 is connected to the input terminal of the second control unit b2. The output of the second control unit b2 is connected to the control terminal of the mantissa addition unit b3. The output of the mantissa addition unit b3 is connected to the input terminal of the leading zero processing unit c1. The output of the leading zero processing unit c1 is respectively connected to the exponent shift unit d1 and the rounding operation unit d2 to obtain the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0111] Understandably, in the embodiments of the present application, it is also possible to use custom devices to improve performance, which is a very effective strategy, especially having significant advantages in power consumption optimization. Specifically, by adopting custom registers, the design can be precisely adjusted according to actual requirements, thereby reducing the power consumption of the system.
[0112] In some embodiments, the input pipeline register is used to merge and store the data to be pipelined among the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed within the same clock cycle; the output pipeline register is used to merge and store the data to be pipelined among the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result within the same clock cycle; the intermediate pipeline register is used to merge and store the intermediate result of the logical operation of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed within the same clock cycle.
[0113] That is to say, whether it is a multiplication unit or an addition unit, the registers in these modules (such as input pipeline registers, intermediate pipeline registers, output pipeline registers, etc.) can be replaced with custom devices, which is an exact optimization for timing bottlenecks. Among them, the customized registers can not only process data more efficiently, but also merge multiple data that need to be pipelined within the same clock cycle, making the input of the registers more flexible. When the registers output, they are precisely reallocated according to the requirements of different modules, avoiding the power consumption waste caused by excessive unnecessary timing operations in the related art.
[0114] A floating-point arithmetic unit provided by an embodiment of the present application can effectively balance the latency of floating-point operations for this floating-point arithmetic unit (such as a multiplication unit or an addition unit), and can control the latency of input and / or output during the entire processing process. Based on this latency controllable strategy, it can also avoid waste of hardware resources, improve the operation efficiency, and reduce power consumption; moreover, it can also optimize the design of registers, replacing standard registers in the related art with custom devices, which can not only process data more efficiently, but also solve the power consumption waste caused by excessive unnecessary timing operations in the related art, thereby improving the overall performance of floating-point operations.
[0115] In another embodiment of the present application, Figure 5 is a schematic structural diagram of a floating-point processing device provided by an embodiment of the present application Figure 1 . As Figure 5 shown, the floating-point processing device 50 may include a configuration module 501 and a plurality of floating-point arithmetic units 10 as described in any one of the foregoing embodiments, where: The configuration module 501 is configured to call a target number of floating-point arithmetic units 10 from a plurality of floating-point arithmetic units 10 according to the precision of the operand to be processed; The floating-point arithmetic unit 10 is configured to perform floating-point operations on the operand to be processed to obtain an operation result; wherein, the target number is related to the precision of the operand to be processed.
[0116] In the embodiment of the present application, the precision of the operand to be processed may include half precision, single precision, and double precision. Exemplarily, double precision can be represented by FP64, which includes 1 sign bit, 11 exponent bits, and 52 mantissa bits; single precision can be represented by FP32, which includes 1 sign bit, 8 exponent bits, and 23 mantissa bits; half precision can be represented by BF16, which includes 1 sign bit, 8 exponent bits, and 7 mantissa bits. BF16 is similar to FP32, but the mantissa bits are shortened to 7 bits.
[0117] Here, for the floating-point multiply-accumulate unit, the MAC design is a common optimization strategy aimed at balancing the performance, power consumption, and precision requirements of the hardware design. In high-performance computing, especially in the design of deep learning accelerators, how to improve the computing efficiency without sacrificing precision is a very critical issue that urgently needs to be solved currently. By adopting mixed-precision operations, the processing ability of the hardware can be significantly improved while ensuring the precision requirements. Specifically, through parametric design and module reuse, it is possible to construct multiply-accumulate units with multiple precisions and different ratios to meet the requirements in different application scenarios.
[0118] In some embodiments, the configuration module 501 is further configured to intercept the operand to be processed according to the precision of the operand to be processed, obtain a target number of data segments, and input the target number of data segments into the target number of floating-point arithmetic units correspondingly.
[0119] In the embodiments of the present application, the configuration module 501 can be configured according to Figures 6A - 6D the configuration interface, so as to intercept the operand according to the precision of the operand to be processed, and the purpose of switching the precision can be achieved. As Figures 6A - 6D shown, four parameter positions for intercepting the operand are schematically provided here, for example Figure 6A corresponding to the configuration interface with parameter position = 0, Figure 6B corresponding to the configuration interface with parameter position = 1, Figure 6C corresponding to the configuration interface with parameter position = 2, Figure 6D corresponding to the configuration interface with parameter position = 3.
[0120] In a possible implementation manner, if it is a single-precision operand, then 2 data segments can be intercepted, such as data segments b[31:0] and b[63:32]. Exemplarily, for parameter positions = 0, 1, the intercepted data segment is b[31:0]; for parameter positions = 2, 3, the intercepted data segment is b[63:32].
[0121] In another possible implementation manner, if it is a half-precision operand, then 4 data segments can be intercepted, such as data segments b[15:0], b[31:16], b[47:32], and b[63:48]. Exemplarily, for parameter position = 0, the intercepted data segment is b[15:0]; for parameter position = 1, the intercepted data segment is b[31:16]; for parameter position = 2, the intercepted data segment is b[47:32]; for parameter position = 3, the intercepted data segment is b[63:48].
[0122] In this way, after inputting the target number of data segments into the target number of floating-point arithmetic units, for example, if the floating-point arithmetic unit is a multiplication unit, the outputs of each multiplication unit can be placed at the top layer, merged and clocked, and then input into the adder unit in the MAC. In this way, not only can signal integration and unified management be achieved, but also by summarizing the outputs of each multiplier at the top layer, the signal flow can be unifiedly managed and controlled, simplifying the design complexity. The signal transmission path can be reduced, the latency can be lowered, and the overall performance of the system can be improved; moreover, the timing and latency balance can be optimized. In a multi-channel design, the calculations of each multiplication unit may have different latencies; by merging the signals at the top layer, the latencies of each channel can be balanced to ensure that the input signals received by the adder unit are consistent in timing, avoiding timing errors caused by latency mismatches; and modularity and maintainability can also be enhanced, which is conducive to modular design and makes the functions of each part more independent. The maintainability and scalability of the system are improved, facilitating subsequent debugging and optimization; at the same time, power consumption and resource consumption can also be reduced, for example, by reducing unnecessary intermediate storage and transmission, power consumption can be reduced. This helps to achieve efficient computing in resource-constrained environments.
[0123] In some embodiments, when the precision of the operand to be processed is double precision, the target number is set to m; when the precision of the operand to be processed is single precision, the target number is set to n; when the precision of the operand to be processed is half precision, the target number is set to k.
[0124] In the embodiments of the present application, m, n, and k are all positive integers, and m:n:k satisfies a preset ratio requirement. Exemplarily, the preset ratio requirement may be 1:2:4, but it is not specifically limited.
[0125] That is to say, the mixed-precision MAC unit can support calculations with different precision ratios through design. For example, a typical FP64:FP32:BF16 = 1:2:4 design. In this case, the precision can be flexibly adjusted according to different operation requirements to adapt to a wider range of computing tasks. Among them, FP64-precision operations are mainly used in application scenarios that require extremely high precision, while FP32 and BF16 are used in applications with relatively low precision requirements but extremely large amounts of calculations. Especially in machine learning and deep learning tasks, the FP32 and BF16 data formats can significantly improve the computing efficiency while ensuring the effectiveness of model training and inference.
[0126] In some embodiments, the floating-point processing device 50 may further include a mode selection unit. The multiple input ends of the mode selection unit are respectively connected to the output ends of the floating-point arithmetic units called under different precisions, and the control end of the mode selection unit is used to receive an operation mode signal, where: A mode selection unit, configured to output operation results corresponding to a target number of floating-point operation units called based on the precision of an operand to be processed according to an operation mode signal.
[0127] In an embodiment of the present application, the operation mode signal may be represented by op_mode. According to this operation mode signal, the operation results corresponding to the corresponding precision can be switched according to requirements, so as to achieve more efficient calculation.
[0128] In some embodiments, the multiple floating-point operation units may include multiple multiplication units and multiple addition units. Among them, a configuration module 501 is configured to call a target number of multiplication units in the multiple multiplication units and a target number of addition units in the multiple addition units according to the precision of the operand to be processed; the multiplication unit is configured to perform a multiplication operation on a first operand and a second operand in the operand to be processed to obtain a multiplication result; the addition unit is configured to perform an addition operation on the multiplication result and a third operand in the operand to be processed to obtain an operation result.
[0129] In an embodiment of the present application, if there are both a multiplication unit and an addition unit at the same time, then the number of mode selection units is two, such as a first mode selection unit and a second mode selection unit. In this way, the first mode selection unit is configured to selectively output the multiplication result corresponding to the precision of the operand to be processed in the multiplication operation according to the operation mode signal; the second mode selection unit is configured to selectively output the addition result corresponding to the precision of the operand to be processed in the addition operation according to the operation mode signal.
[0130] It should also be noted that the number of called multiplication units and the number of addition units are both related to the precision of the operand to be processed. Exemplarily, if it is a half-precision operand, then 4 multiplication units and 4 addition units can be called; if it is a single-precision operand, then 2 multiplication units and 2 addition units can be called, but no specific limitation is made here.
[0131] In a possible implementation manner, taking the multiple floating-point devices including multiple multiplication units and multiple addition units as an example, as Figure 7 shown, the floating-point processing device 50 may include multiple multiplication units 701, multiple addition units 702, a first mode selection unit (MUX_md_1) 703, a first trigger unit 704, a second trigger unit 705, and a second mode selection unit (MUX_md_2) 706.
[0132] It should be noted that the first trigger unit 704 is configured to input the multiplication result to a target number of addition units in a preset clock cycle; the second trigger unit 705 is configured to input the flag bit information of the multiplication result to a target number of addition units in a preset clock cycle.
[0133] It should also be noted that the operands to be processed here include the first operand \(d_{i\_x}\), the second operand \(d_{i\_y}\), and the third operand \(d_{i\_z}\), and each operand also includes flag bit information \(flags\). First, determine the precision of the operands to be processed, and according to the precision of the operands to be processed, call a target number of multiplication units among multiple multiplication units 701. The multiplication results corresponding to the target number of multiplication units are selected and output to the input end of the first trigger unit 704 through the first mode selection unit 703. The flag bit information of the multiplication results is input to the input end of the second trigger unit 705. The first trigger unit 704 is controlled by the clock signal \(clk\) to output the obtained multiplication results, and the multiplication results and the third operand \(d_{i\_z}\) are correspondingly input to the target number of addition units; finally, the addition results corresponding to the target number of addition units (i.e., the operation result \(d_o\)) are selected and output through the second mode selection unit 706. In this process, the second trigger unit 705 can also be controlled by the clock signal \(clk\) to output the flag bit information corresponding to the multiplication results, so that in the addition operation, it can be directly determined whether the multiplication results are special numbers according to the flag bit information.
[0134] That is to say, in the embodiments of the present application, for different precisions, multiple multiplication units and multiple addition units are reused, thereby saving circuit area. In addition, the first trigger unit 704 is a data trigger, which is used to correspondingly input the obtained multiplication results to the target number of addition units in a preset clock cycle; the second trigger unit 705 is a flag bit trigger, which is used to correspondingly input the flag bit information of the obtained multiplication results to the target number of addition units in a preset clock cycle, so that the addition units no longer need to judge whether the received multiplication results are special numbers, further saving circuit area.
[0135] Exemplarily, Figure 8 is a schematic application framework of a floating-point processing device provided by an embodiment of the present application Figure 1 , Figure 9 is a schematic application framework of a floating-point processing device provided by an embodiment of the present application Figure 2 . As Figure 8 and Figure 9 shown, the floating-point processing device 50 can be compatible with two precisions of FP32 and BF16. In Figure 7Based on this, it may further include a third trigger unit 801, a fourth trigger unit 802, a fifth trigger unit 803, a sixth trigger unit 804, a fourth selection unit 805, a fifth selection unit 806, and a seventh trigger unit 807. Among them, the third trigger unit 801, the fourth trigger unit 802, and the fifth trigger unit 803 are used to sample and output the first operand d_i_x, the second operand d_i_y, and the third operand d_i_z in a preset clock cycle, and then the fourth selection unit 805 selects data from the first operand d_i_x and 0 and inputs it into the corresponding multiplication unit, and the fifth selection unit 806 selects data from the second operand d_i_y and 0 and inputs it into the corresponding multiplication unit; the sixth trigger unit 804 is used to output the operation mode signal op_mode to the control end of the first mode selection unit 703 or the control end of the second mode selection unit 706 in a preset clock cycle, so as to select the operation result with FP32 precision or BF16 precision; finally, the seventh trigger unit 807 samples and outputs the operation result d_o in a preset clock cycle.
[0136] It should be noted that in Figure 8 or Figure 9 , multiple multiplication units 701 and multiple addition units 702 therein can be multiplexed at different precisions. Exemplarily, if the ratio of FP32 to BF16 is 2:4, then in Figure 9 , multiple multiplication units 701 include 6 multiplication units, and multiple addition units 702 include 6 addition units. At this time, only two multiplication units and two addition units are needed at FP32 precision, and only four multiplication units and four addition units are needed at BF16 precision. To further save circuit area, as Figure 8 shows, multiple multiplication units 701 include 4 multiplication units, and multiple addition units 702 include 4 addition units. At FP32 precision, only two multiplication units and two addition units are needed, and at BF16 precision, only four multiplication units and four addition units are needed. Additionally, if the ratio of FP32 to BF16 is adjusted to 2:5, then in Figure 9 , at this time, only two multiplication units and two addition units are needed at FP32 precision, and only five multiplication units and five addition units are needed at BF16 precision, and no specific limitation is made here.
[0137] That is to say, in the design of a mixed-precision MAC unit, the selection of data format is particularly important. Suppose the data format to be processed has a relatively low numerical range but requires high precision, such as TF32 (Tensor Float 32) or BF16 (BFloat16). For these data formats, the embodiments of the present application can also adopt the method of padding zeros at the lower bits to pass the input data into the module. Specifically, when the bit width of the input data is relatively low (for example, 16 bits), but its numerical range requirement is not high, zeros can be padded at the lower bits to convert the data into 32 bits for operation. In this way, not only can the calculation precision be effectively improved, but also the operation speed and power consumption efficiency can be increased by reducing the use of hardware resources.
[0138] For the addition part, especially the addition operation in BF16 precision, a similar optimization strategy also applies. Before performing the addition operation, the BF16 data can be extended to 32 bits by padding zeros at the lower bits (such as the BF16_to_FP32 module in Figure 9 ), and then the addition calculation is performed. In this way, not only is the precision of the addition improved, but also it can be ensured that there will be no loss of numerical precision due to insufficient bit width during the entire operation process; thus, the final operation result output will be more accurate, meeting the application scenarios with higher precision requirements, while maintaining a relatively low calculation cost and power consumption.
[0139] Briefly speaking, at the hardware implementation level, the design of a mixed-precision MAC unit is somewhat challenging. It is necessary to find a balance among hardware resources, timing, power consumption, and calculation precision. Therefore, the efficiency of multipliers and adders must be fully considered in the above process, especially their performance when processing data with different precisions. The introduction of mixed precision requires that during implementation, both efficient data transmission and processing should be ensured, and unnecessary calculation redundancy should be avoided. In the multiply-accumulate unit, multiple data with different precisions need to be precisely scheduled and optimized to ensure the efficiency of the data stream and the overall stability of the system.
[0140] In another embodiment of the present application, the embodiments of the present application also provide a floating-point processing device. Refer to Figure 10 , the floating-point processing device 50 may include a register 1001 and a plurality of floating-point operation units 10 as described in any one of the foregoing embodiments, where: The register 1001 is used to store at least two floating-point numbers in the same clock cycle; The floating-point operation unit 10 is used to obtain the operand to be processed from the register.
[0141] It should be noted that the standard registers in the related art cannot merge multiple data that need to be pipelined within the same clock cycle, resulting in unnecessary power consumption waste. In the embodiments of the present application, the register 1001 here can be a customized device, which is precisely optimized for sequential logic; it can not only process data more efficiently, but also merge multiple data that need to be pipelined within the same clock cycle, making the input of the register more flexible.
[0142] In some embodiments, referring to Figure 11 , the floating-point arithmetic unit 10 can be a multiplication unit, and the register 1001 can include a first register r1, where: The first register r1 is used to store a first operand and a second operand in the same clock cycle, and input the operand output by the first register r1 into at least one multiplication unit; At least one multiplication unit 1002 is used to perform a multiplication operation on the operand output by the first register r1 to determine at least one multiplication result.
[0143] In the embodiments of the present application, if the floating-point processing device 50 is only used to implement multiplication operations, then there can be at least one multiplication unit 1002 here. After merging and storing the first operand x and the second operand y to be pipelined input into the first register r1, the operand output by the first register r1 can be correspondingly input into at least one multiplication unit 1002, so that a multiplication result can be obtained.
[0144] In some embodiments, referring to Figure 12 , the floating-point arithmetic unit 10 can be an addition unit, and the register 1001 can include a second register r2, where: The second register r2 is used to store a third operand and a fourth operand in the same clock cycle, and input the operand output by the second register r2 into at least one addition unit; At least one addition unit 1003 is used to perform an addition operation on the operand output by the second register r2 to determine at least one addition result.
[0145] In the embodiments of the present application, if the floating-point processing device 50 is only used to implement addition operations, then there can be at least one addition unit 1003 here. After merging and storing the third operand z and the fourth operand u to be pipelined input into the second register r2, the operand output by the second register r2 can be correspondingly input into at least one addition unit 1003, so that an addition result can be obtained.
[0146] In some embodiments, referring to Figure 13, the floating-point arithmetic unit 10 may include a multiplication unit and an addition unit, and the register 1001 may include a first register r1 and a second register r2, where: The first register r1 is used to store a first operand and a second operand in the same clock cycle, and input the operand output by the first register r1 into at least one multiplication unit; At least one multiplication unit 1002 is used to perform a multiplication operation on the operand output by the first register r1 to determine at least one multiplication result; The second register r2 is used to store a third operand and at least one multiplication result in the same clock cycle, and input the operand output by the second register r2 into at least one addition unit; At least one addition unit 1003 is used to perform an addition operation on the operand output by the second register r2 to determine at least one addition result.
[0147] In some embodiments, continue to refer to Figure 13 , the floating-point processing device 50 may further include a reallocation unit 1004. Among them, the reallocation unit 1004 is used to reallocate the input for at least one multiplication unit when the first register r1 outputs, and input the reallocated operand into at least one multiplication unit 1002.
[0148] It should be noted that in the embodiments of the present application, the first register r1 and the second register r2 may also be customized devices. In this way, not only can data be processed more efficiently, but also multiple data that need to be pipelined can be combined within the same clock cycle, making the input of the register more flexible. When the register outputs, precise reallocation is performed according to the requirements of different modules, avoiding power consumption waste caused by excessive unnecessary timing operations in the related art.
[0149] It should also be noted that in the embodiments of the present application, the multiplication unit and the addition unit here may be located in the MAC unit. As Figure 13 shown, the three input operands also include flag bits flag (zero, nan, overflow) for indicating whether the corresponding input operand is a special number.
[0150] It should also be noted that taking Figure 13 as an example, the input third operand z arrives one clock cycle later than the input first operand x and second operand y according to the design requirements. Here, the input can be processed according to different design requirements. The reallocation unit allocates the input for each MAC (multiplication). The second register r2 simultaneously pipelines at least one multiplication result and the input third operand z, and then inputs them into each MAC for addition operation to obtain the final operation result = x×y + z.
[0151] An embodiment of the present application provides a floating-point processing device, which includes a plurality of floating-point arithmetic units. According to the precision of the operand to be processed, a target number of floating-point arithmetic units can be called among the plurality of floating-point arithmetic units to perform floating-point arithmetic on the operand to be processed, and an arithmetic result is obtained. In this way, a mixed-precision design can also be realized, so as to significantly improve the processing ability of the hardware while ensuring the precision requirements; and based on the reuse of basic modules such as multiplication units and addition units, the power consumption and resource consumption are further reduced, and the circuit area is saved. In addition, the design of the register can be optimized, and the standard register in the related art is replaced with a customized device, which can not only process data more efficiently, but also solve the power consumption waste caused by too many unnecessary timing operations in the related art, thereby improving the overall performance of floating-point arithmetic.
[0152] In another embodiment of the present application, Figure 14 is a flowchart of a floating-point processing method provided by an embodiment of the present application Figure 1 . As Figure 14 shown, the method may include: S1401, the logic processing module performs a logic operation on the input operand to be processed to obtain an arithmetic result.
[0153] S1402, the delay control module controls the delay between the input of the operand to be processed to the logic processing module and the input of the floating-point arithmetic unit; and / or, controls the delay between the output of the arithmetic result of the logic processing module and the output of the arithmetic result of the floating-point arithmetic unit.
[0154] In an embodiment of the present application, the floating-point processing method is applied to the floating-point arithmetic unit described in the foregoing embodiment. Thereby, the delay of the logic operation can be effectively balanced, the operation efficiency of the floating-point arithmetic unit 10 can be improved, and accurate result output can be ensured.
[0155] In an embodiment of the present application, the floating-point arithmetic unit may adopt a pipeline structure, and the entire floating-point processing process is divided into several pipeline registers, which can effectively avoid the performance bottleneck caused by too long pipeline delay. Exemplarily, the entire floating-point processing process may be divided into: an input pipeline register, at least one intermediate pipeline register, and an output pipeline register, thereby improving the pipeline working efficiency and saving time.
[0156] In some embodiments, the delay control module includes an input pipeline register and an output pipeline register, and the method may further include: The input pipeline register controls the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed; and / or, Output the path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result of the pipeline register.
[0157] In the embodiment of the present application, in the pipeline structure, the input pipeline register 1021 and the output pipeline register 1022 are controllable. That is to say, whether the input pipeline register 1021 and the output pipeline register 1022 are enabled can be controlled according to the corresponding configuration parameters to achieve the purpose of delay balance.
[0158] In some embodiments, the method may further include: the input pipeline register can control whether to latch at least one of the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed according to the first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, initial exponent bit, and initial mantissa bit meets the first delay requirement; and / or, the output pipeline register can control whether to latch at least one of the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result according to the second configuration parameter, so that the output path delay of at least one of the target sign bit, target exponent bit, and target mantissa bit meets the second delay requirement.
[0159] In the embodiment of the present application, the first delay requirement and the second delay requirement are preset and used to measure whether the path delay of each segment meets the requirement. In addition, the first configuration parameter and the second configuration parameter can be designed according to the overall requirements of the floating-point arithmetic unit to control whether the corresponding pipeline register is enabled.
[0160] In some embodiments, at least one intermediate pipeline register is provided between the input pipeline register and the output pipeline register, and the method may further include: At least one intermediate pipeline register performs a logical operation on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed to determine the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result.
[0161] In the embodiment of the present application, the operand to be processed includes a first operand and a second operand. Here, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit, and the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit. Among them, the target sign bit can be obtained by performing a logical operation on the first initial sign bit and the second initial sign bit, the target exponent bit can be obtained by performing a logical operation on the first initial exponent bit and the second initial exponent bit, and the target mantissa bit can be obtained by performing a logical operation on the first initial mantissa bit and the second initial mantissa bit.
[0162] In this way, the logical processing process can be divided into at least one intermediate pipeline register, and computing tasks are allocated between each intermediate pipeline register, so as to optimize the execution time of floating-point processing, avoid performance bottlenecks caused by too long pipeline latency, and improve the computing efficiency.
[0163] Those skilled in the art should understand that the description of the floating-point processing method in the embodiments of this application can be understood with reference to the relevant description of the floating-point arithmetic unit in the foregoing embodiments, and will not be elaborated here.
[0164] In still another embodiment of this application, Figure 15 is a flowchart of a floating-point processing method provided by an embodiment of this application Figure 2 . As Figure 15 shown, the method may include: S1501. The configuration module calls a target number of floating-point arithmetic units among multiple floating-point arithmetic units according to the precision of the operand to be processed.
[0165] S1502. The floating-point arithmetic unit performs floating-point arithmetic on the operand to be processed to obtain an operation result; wherein, the target number is related to the precision of the operand to be processed.
[0166] In the embodiments of this application, the floating-point processing method is applied to the floating-point processing device described in the foregoing embodiments. Among them, the precision of the operand to be processed may include half precision, single precision, and double precision. Here, by adopting mixed-precision arithmetic, while ensuring the precision requirements, the processing ability of the hardware can be significantly improved. Specifically, through parametric design and module reuse, it is possible to construct multiply-accumulate units with multiple precisions and different ratios to meet the requirements in different application scenarios.
[0167] In some embodiments, the method may further include: the configuration module intercepts the operand to be processed according to the precision of the operand to be processed to obtain a target number of data segments, and inputs the target number of data segments into the target number of floating-point arithmetic units correspondingly.
[0168] In the embodiments of this application, parameter configuration can be performed according to the configuration interface of the configuration module in Figures 6A - 6D , so as to intercept the operand according to the precision of the operand to be processed, and the purpose of switching precision can be achieved.
[0169] In some embodiments, the method may also include: the configuration module calls a target number of multiplication units in multiple multiplication units and a target number of addition units in multiple addition units according to the precision of the operand to be processed; the multiplication unit performs a multiplication operation on the first operand and the second operand in the operand to be processed to obtain a multiplication result; the addition unit performs an addition operation on the multiplication result and the third operand in the operand to be processed to obtain an operation result.
[0170] In the embodiment of the present application, the number of multiplication units and the number of addition units called are related to the precision of the operand to be processed. For example, if the operand is half-precision, then 4 multiplication units and 4 addition units can be called; if the operand is single-precision, then 2 multiplication units and 2 addition units can be called, but this is not specifically limited here.
[0171] That is to say, in the embodiment of the present application, after the target number of data segments are input to the target number of floating-point arithmetic units, for example, the floating-point arithmetic units are multiplication units, then the outputs of each multiplication unit can be placed in the top-level merged beat and input into the addition unit in the MAC. In this way, not only can signal integration and unified management be achieved, but also the signal flow can be uniformly managed and controlled by aggregating the outputs of each multiplier at the top level, simplifying the design complexity. Reduce the signal transmission path, reduce delay, and improve the overall performance of the system; it can also optimize the timing and delay balance. In a multi-channel design, the calculations of each multiplication unit may have different delays; by merging the signal at the top level, the delays of each channel can be balanced to ensure that the input signals received by the addition unit are consistent in timing, avoiding timing errors caused by delay mismatch; and it can also enhance modularity and maintainability, which is helpful for modular design and makes the functions of each part more independent. Improve the maintainability and scalability of the system, facilitate subsequent debugging and optimization; at the same time, it can also reduce power consumption and resource consumption, such as reducing unnecessary intermediate storage and transmission, and reducing power consumption. It helps to achieve efficient computing in a resource-constrained environment.
[0172] In addition, those skilled in the art should understand that the description of the floating-point processing method in the embodiment of the present application can be understood by referring to the relevant description of the floating-point processing device described in the aforementioned embodiment, and will not be described in detail here.
[0173] In yet another embodiment of the present application, Figure 16 A schematic diagram of a floating point processing method provided in an embodiment of the present application Figure 3 .like Figure 16 As shown, the method may include: S1601: The register stores at least two floating-point numbers in the same clock cycle.
[0174] S1602, the floating-point operation unit obtains the operand to be processed from the register, so that the floating-point operation unit executes: the logic processing module performs logic operation on the operand to be processed to obtain the operation result; the delay control module controls the delay between the input of the operand to be processed into the logic processing module and the input of the floating-point operation unit; and / or controls the delay between the output of the operation result by the logic processing module and the output of the operation result by the floating-point operation unit.
[0175] It should be noted that, for step S1602, after the floating-point operation unit obtains the operand to be processed from the register, the floating-point operation unit executes Figure 14 Floating point processing method shown.
[0176] It should also be noted that the standard register in the related art cannot merge multiple data that need to be beat in the same clock cycle, resulting in unnecessary power consumption. In the embodiment of the present application, the register here can be a customized device, which is precisely optimized for sequential logic; it can not only process data more efficiently, but also merge multiple data that need to be beat in the same clock cycle, making the input of the register more flexible.
[0177] In some embodiments, the floating-point operation unit includes a multiplication unit and an addition unit, and the register includes a first register and a second register. Accordingly, the method further includes: the first register stores the first operand and the second operand in the same clock cycle, and inputs the operand output by the first register to at least one multiplication unit; at least one multiplication unit performs a multiplication operation on the operand output by the first register to determine at least one multiplication result; the second register stores the third operand and at least one multiplication result in the same clock cycle, and inputs the operand output by the second register to at least one addition unit; at least one addition unit performs an addition operation on the operand output by the second register to determine at least one addition result.
[0178] In some embodiments, the method further includes: a reallocation unit reallocating inputs to at least one multiplication unit when the first register outputs, and inputting the reallocated operands to at least one multiplication unit.
[0179] It should be noted that in the embodiment of the present application, the first register and the second register can be customized devices. In this way, not only can data be processed more efficiently, but also multiple data that need to be beat can be merged in the same clock cycle, making the input of the register more flexible. When the register is output, it is accurately redistributed according to the needs of different modules, avoiding the waste of power consumption caused by too many unnecessary timing operations in the related art.
[0180] For example, Figure 13For example, the input third operand z is one beat later than the input first operand x and second operand y due to design requirements. Here, the input can be processed according to different design requirements. The reallocation unit allocates inputs for each MAC (multiplication). The second register r2 simultaneously beats at least one multiplication result with the input third operand z, and then inputs it to each MAC for addition operation to obtain the final operation result = x×y+z.
[0181] In addition, those skilled in the art should understand that the description of the floating-point processing method in the embodiment of the present application can also be understood by referring to the relevant description of the floating-point processing device described in the aforementioned embodiment, which will not be described in detail here.
[0182] In another embodiment of the present application, a delay-balanced, high-performance floating-point operation unit is proposed, which can be applied to the field of high-performance computing, such as large model training, artificial intelligence (AI), intelligent computing, etc. The floating-point operation unit here may include a multiplication unit and / or an addition unit, and the two are described in detail below.
[0183] In a possible implementation, a delay-balanced controllable three-stage pipeline multiplication unit design is provided. In the floating-point multiplication unit, delay balance and pipeline tuning design are adopted. The input will introduce a flag bit to indicate whether the operand is a special number, such as zero, NAN, overflow, etc. Figure 3 As shown, the first-level pipeline can choose whether to pass the input register data after the beat according to the configured parameters. After inserting the second-level pipeline, the mantissa is directly multiplied.
[0184] In this implementation, since the multiplication result needs one cycle to be obtained, the multiplication result arrives at the same time as the flag bit, the sign bit, and the exponent bit. Considering that the multiplication operation involves complex calculation logic, it requires multiple levels of logic gates and a longer calculation path compared to addition and XOR gates, so its delay is relatively large. The embodiment of the present application can beat other paths to achieve the purpose of delay balance.
[0185] The third level pipeline also controls whether the output is beat by configuring parameters.
[0186] It should be noted that the rounding operation can be performed according to the IEEE-754 rules. For the multiplication result, after the feedback of the test results, it was found that the processing of non-standard numbers and standard numbers is two different branches, which cannot be combined. Specifically, for standard numbers, according to the rounding rules of floating-point numbers, it can be summarized as when the reserved bit (G), approximate bit (R), and sticky bit (S) meet specific combination conditions, carry forward (normal carry). For non-standard numbers, there is a situation: after the mantissa is rounded, it carries to the exponent bit, resulting in the exponent bit not being all zero, which means that the non-standard number becomes a standard number after the carry. In this way, when processing underflow numbers, this situation needs to be taken into account, and the situation of non-standard number carry is handled separately when the last exponent bit is assigned.
[0187] In another possible implementation, a delay-balanced controllable five-stage pipeline addition unit design is provided. In high-performance computing, the addition unit is one of the basic computing units, and the performance of its design directly affects the efficiency of the overall system. In order to optimize the performance of the addition unit, especially when it comes to floating-point calculations, designing a delay-balanced controllable pipeline addition unit has become a very important technical solution. In this way, the delay of the addition operation can be effectively balanced, the computing efficiency of the addition unit can be improved, and accurate result output can be guaranteed.
[0188] In this implementation, the two-stage pipeline of the input and output of the addition unit is the same as that of the multiplication unit, and can be determined by configuration parameters. The addition operation first needs to select the larger one according to the exponent of the operand, and use this operand as the addend. Here, the register flag can be set to determine whether the addend and the addend need to be interchanged. The slight difference from the parsing of the operand by the multiplication unit is that the operand parsing can only be performed after the position of the operand is determined in the previous step, but the parts of the two operands are not operated here. Specifically, the exponent operation can align the decimal point of the mantissa according to the difference in the exponent of the two operands. Here, the shift alignment operation is performed according to the exponent difference, and the pipeline register is inserted after the shift. After obtaining the sum of the mantissa, it is determined whether the result needs to be inverted according to the operand sign bit. Then it needs to be shifted, the exponent size is adjusted appropriately, and it is normalized. After the normalization operation is completed, it is inserted into the pipeline register. LZC is also introduced in this process. Here, the strategy of step-by-step shifting and judging zero is adopted. The register records the situation of all zeros in the high bits of the corresponding stage. The number of leading zeros of the mantissa is judged and shifted according to the value of the final register. At the same time, it can also ensure that the highest bit of the final mantissa is 1. For details, see Figure 4 In addition, the rounding rules are similar to those of the aforementioned multiplication unit and will not be described in detail here.
[0189] Finally, the core advantage of the entire five-stage pipeline addition unit design lies in its controllable delay balance. Through the five-stage pipeline structure, the addition unit can distribute computing tasks between each stage, thereby optimizing the execution time of the addition operation. The delay of each stage is carefully designed to ensure that the final addition result will not violate the rules and avoid performance bottlenecks caused by excessive pipeline delays. In addition, the delay balance strategy adopted here can also effectively avoid the waste of hardware resources, improve computing efficiency, and reduce power consumption.
[0190] In another possible implementation, a mixed-precision multiplication-addition unit design is provided. Among them, MAC unit design is a common optimization strategy that aims to balance the performance, power consumption and accuracy requirements of hardware design. In high-performance computing, especially in the design of deep learning accelerators, how to improve computing efficiency without sacrificing accuracy is a very critical issue that needs to be solved urgently. By adopting mixed-precision operations, the processing power of the hardware can be significantly improved while ensuring the accuracy requirements. Based on basic modules such as multiplication units and addition units, through parameterized design and module reuse, we can build multiplication-addition units with multiple precisions and different ratios to meet the needs of different application scenarios.
[0191] For example, Figures 6A - 6D As shown, at the top level of MAC, the operands can be intercepted according to the configured parameters to achieve the purpose of switching accuracy.
[0192] In addition, placing the output of each multiplication unit at the top level for merging and beating and then inputting it back into the addition unit of the MAC unit has the following important significance: (1) Signal integration and unified management. By summarizing the output of each multiplier at the top level, the signal flow can be managed and controlled in a unified manner, simplifying the design complexity. It reduces the signal transmission path, reduces the delay, and improves the overall performance of the system. (2) Optimize timing and delay balance. In a multi-channel design, the calculation of each multiplier may have different delays. By merging the signals at the top level, the delays of each channel can be balanced to ensure that the input signals received by the adder are consistent in timing and avoid timing errors caused by delay mismatch. (3) Enhance modularity and maintainability. It helps modular design and makes the functions of each part more independent. It improves the maintainability and scalability of the system and facilitates subsequent debugging and optimization. (4) Reduce power consumption and resource consumption. It reduces unnecessary intermediate storage and transmission and reduces power consumption. It helps to achieve efficient computing in a resource-constrained environment.
[0193] Specifically, the mixed-precision MAC unit is designed to support calculations with different precision ratios, such as the typical FP64:FP32:BF16 = 1:2:4 design. This design can flexibly adjust the precision according to different computing requirements to adapt to a wider range of computing tasks. In this design mode, FP64 precision calculations are mainly used in application scenarios that require extremely high precision, while FP32 and BF16 are used for applications with relatively low precision requirements but extremely large computational workloads. Especially in machine learning and deep learning tasks, the data formats of FP32 and BF16 can greatly improve computing efficiency while ensuring the effectiveness of model training and reasoning.
[0194] In addition, the embodiments of the present application can also select a suitable precision ratio according to specific needs. For example, in some applications, the ratio of FP32 to BF16 can be adjusted to 2:4 to optimize resource usage and further improve the computing speed. This flexible precision selection enables the mixed-precision MAC unit to perform well in different tasks, meeting the precision requirements, improving computing efficiency, and reducing hardware power consumption and resource consumption. Figures 7 - 9 An illustrative mixed-precision application is provided.
[0195] In the design of mixed-precision MAC units, the choice of data format is particularly important. Among them, it is assumed that the data format that needs to be processed is a data format with a lower numerical range but higher precision requirements, such as TF32 or BF16. For these data formats, the embodiment of the present application can use the method of padding the low bits with zeros to pass the input data into the module. Specifically, when the input data bit width is low (for example, 16 bits), but its numerical range requirement is not high, you can pad the low bits with zeros and convert the data to 32 bits for calculation. This method can not only effectively improve the calculation accuracy, but also improve the calculation speed and power efficiency by reducing the use of hardware resources.
[0196] For the addition part, especially the addition operation under BF16 precision, similar optimization strategies are also applicable. Before the addition operation, the BF16 data can be expanded to 32 bits by padding the low bits with zeros (such as Figure 9 This method not only improves the accuracy of addition, but also ensures that there will be no loss of numerical precision due to insufficient bit width during the entire operation. Through this technical means, the final output result will be more accurate, meeting application scenarios with higher precision requirements, while maintaining low computing costs and power consumption.
[0197] Understandably, at the hardware implementation level, the design of a mixed-precision MAC unit poses certain challenges. It is necessary to find a balance among hardware resources, timing, power consumption, and computational precision. To this end, the efficiency of the multiplication unit and the addition unit must be fully considered during the design process, especially their performance when processing data with different precisions. The introduction of mixed precision requires that during implementation, both efficient data transmission and processing be ensured, while unnecessary computational redundancy be avoided. In the multiply-accumulate unit, multiple data with different precisions need to be precisely scheduled and optimized to ensure the efficiency of the data flow and the overall stability of the system.
[0198] In addition, the design of the hardware accelerator also needs to fully consider flexibility. The design of the mixed-precision MAC unit is not just an optimization for a certain fixed task, but rather to dynamically adjust the precision and resource allocation according to the requirements of different tasks. Exemplarily, in deep learning training, the proportion of precision used may be adjusted according to different levels and computational loads of the model to maximize the performance of the hardware. Designers also need to provide sufficient control mechanisms, such as the operation mode signal (op_mode), so as to switch the precision and operation mode according to the task requirements during runtime, thereby achieving more efficient computing.
[0199] In yet another possible implementation, a method of using custom devices to improve performance is provided here, and specifically, it may be necessary to negotiate with the backend. Using custom devices to improve performance is a very effective strategy, especially having significant advantages in power consumption optimization. Specifically speaking, by adopting custom registers, precise adjustment can be made according to actual needs, thereby reducing the power consumption of the system. The implementation motivation of this technical solution stems from the problem of a high proportion of sequential logic in the comprehensive test report. Due to the increase in the complexity of sequential logic, the bottlenecks of power consumption and performance gradually emerge, especially in the data transmission and storage of registers. Therefore, by optimizing the design of registers, unnecessary power consumption can be significantly reduced, thereby achieving remarkable improvement in power consumption.
[0200] In the embodiments of this application, the standard registers in the module can be replaced with custom devices, which is an exact optimization for the timing bottleneck. Custom registers can not only process data more efficiently but also merge multiple data that need to be pipelined within the same clock cycle, making the input of the register more flexible. When the register outputs, precise reallocation is performed according to the requirements of different modules, avoiding the power consumption waste caused by excessive unnecessary sequential operations in the related art. In this way, not only can the performance be improved, but also obvious optimization in the overall power consumption can be achieved. Especially in complex systems, the overall efficiency of the system can be effectively enhanced.
[0201] Exemplarily, taking the top layer as an example, in Figure 13Among them, when inputting the operands, flags (zero, nan, overflow) are also included to indicate whether the operand is a special number. Among them, the input of the third operand z arrives one clock cycle later than the input of the second operand x and the third operand y according to the design requirements. At this time, the input can be processed according to different design requirements. The reallocation unit allocates inputs to each MAC. The second register r2 simultaneously stalls the multiplication results of each MAC and the input third operand z, and then inputs them into each MAC for addition operations to obtain the final operation result.
[0202] It should also be noted that the multiply-accumulate unit (MAC) with latency balance provided in the embodiments of the present application plays a key role in digital signal processing. Its efficient multiplication and addition operations have been widely used in many fields, such as intelligent computing, high-performance computing and other fields.
[0203] Exemplarily, in the field of Digital Signal Processing (DSP), the MAC unit is used to perform filtering, convolution and other signal processing tasks. Its efficient multiply-accumulate operations have been widely used in audio, video and communication systems. In the field of High Performance Computing (HPC), the MAC unit is used for matrix multiplication and linear algebra operations. Its efficient computing power has been widely used in scientific computing, engineering simulation and data analysis. In the fields of AI and Machine Learning (ML), the MAC unit is used for the forward and backward propagation of neural networks. Its efficient computing power has been widely used in the training and inference of deep learning models. In the field of Graphics Processing Unit (GPU), the MAC unit is used for graphics rendering and image processing. Its efficient computing power has been widely used in games, virtual reality and computer vision. In the field of communication systems, the MAC unit is used for modulation and demodulation, encoding and decoding, and signal processing. Its efficient computing power has been widely used in wireless communication, satellite communication and fiber optic communication. In the field of Internet of Things (IoT), the MAC unit is used for sensor data processing and edge computing. Its efficient computing power has been widely used in smart homes, smart cities and industrial automation. In the field of medical devices, the MAC unit is used for image processing and signal analysis. Its efficient computing power has been widely used in medical imaging, diagnosis and monitoring.
[0204] In yet another embodiment of the present application, an embodiment of the present application provides a chip, which may include: the floating-point arithmetic unit 10 described in any one of the foregoing embodiments, or the floating-point processing device 50 described in any one of the foregoing embodiments.
[0205] In yet another embodiment of the present application, an embodiment of the present application provides an electronic device, which includes a processor. Among them, the processor includes: the floating-point arithmetic unit 10 described in any one of the foregoing embodiments, or the floating-point processing device 50 described in any one of the foregoing embodiments.
[0206] In summary, through the above embodiments, the specific implementation of the foregoing embodiments is elaborated in detail. It can be seen that through the technical solutions of the foregoing embodiments, for a floating-point arithmetic unit (such as a multiplication unit or an addition unit), the delay of floating-point arithmetic can be effectively balanced, and the delay of input and / or output during the entire processing process can be controlled. Based on this delay control strategy, waste of hardware resources can also be avoided, the operation efficiency can be improved, and the power consumption can be reduced. In addition, through the mixed-precision design, while ensuring the accuracy requirements, the processing ability of the hardware can be significantly improved; and based on the reuse of basic modules such as multiplication units and addition units, the power consumption and resource consumption can be further reduced, and the circuit area can be saved. In addition, the design of registers can be optimized, and the standard registers in the related art can be replaced with customized devices, which can not only process data more efficiently, but also solve the power consumption waste caused by excessive unnecessary timing operations in the related art, thereby improving the overall performance of floating-point arithmetic.
[0207] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0208] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0209] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0210] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0211] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0212] It should be noted that in the present application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitations, the element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0213] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments.
[0214] In the method embodiments disclosed in several method embodiments provided in the present application, they can be arbitrarily combined without conflict to obtain new method embodiments.
[0215] The features disclosed in several product embodiments provided in the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0216] The features disclosed in several method or device embodiments provided in the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0217] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. A floating-point arithmetic unit, characterized in that, The floating-point arithmetic unit includes: A logic processing module for performing a logic operation on the operand to be processed input to the floating-point arithmetic unit to obtain an operation result; A delay control module for controlling the delay between the input of the operand to be processed to the logic processing module and the input to the floating-point arithmetic unit; and / or, controlling the delay between the output of the operation result by the logic processing module and the output of the operation result by the floating-point arithmetic unit; The delay control module includes an input pipeline register and an output pipeline register, where: The input pipeline register is used to control whether to clock at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed according to a first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit meets a first delay requirement; and / or, The output pipeline register is used to control whether to clock at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result according to a second configuration parameter, so that the output path delay of at least one of the target sign bit, the target exponent bit, and the target mantissa bit meets a second delay requirement.
2. The floating-point arithmetic unit according to claim 1, wherein At least one intermediate pipeline register is provided between the input pipeline register and the output pipeline register, where: The at least one intermediate pipeline register is used to perform a logic operation on the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed to determine the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result; where the operand to be processed includes a first operand and a second operand, the first operand includes a first initial sign bit, a first initial exponent bit, and a first initial mantissa bit, and the second operand includes a second initial sign bit, a second initial exponent bit, and a second initial mantissa bit.
3. The floating-point arithmetic unit according to claim 2, wherein When the floating-point arithmetic unit is a multiplication unit, the at least one intermediate pipeline register includes one pipeline register; The logic processing module includes a first logic module and a second logic module, where: The first logic module is used to perform a logic operation on the first initial sign bit and the second initial sign bit and the first initial exponent bit and the second initial exponent bit respectively to determine the target sign bit and the target exponent bit; The second logic module is used to perform a multiplication operation on the first initial mantissa bit and the second initial mantissa bit to determine the target mantissa bit.
4. The floating-point arithmetic unit according to claim 3, wherein The second logic module includes a mantissa product unit and a rounding operation unit, where: The mantissa product unit is used to perform a mantissa multiplication operation on the first initial mantissa bit and the second initial mantissa bit to obtain a first mantissa product; The rounding operation unit is used to perform a rounding operation on the first mantissa product to determine the target mantissa bit.
5. The floating-point arithmetic unit according to claim 4, wherein The rounding operation unit is further configured to perform a rounding operation on the first mantissa product when the operation result is a denormal number, to obtain a second mantissa product; when the second mantissa product exceeds a preset range, perform a normalization process on the second mantissa product to obtain the target mantissa bit, and carry out a carry operation to the target exponent bit, so that the operation result is transformed from a denormal number to a normalized number.
6. The floating-point arithmetic unit according to claim 3, wherein The first logic module includes a sign bit processing unit and an exponent bit processing unit, where: The sign bit processing unit is configured to perform an exclusive OR operation on the first initial sign bit and the second initial sign bit to obtain the target sign bit; The exponent bit processing unit is configured to perform an exponent addition operation on the first initial exponent bit and the second initial exponent bit to obtain the target exponent bit.
7. The floating-point arithmetic unit according to claim 6, wherein The first operand includes a first initial flag bit, and the second operand includes a second initial flag bit; the first logic module further includes a flag bit processing unit, where: The flag bit processing unit is configured to perform a logic operation according to the first initial flag bit and the second initial flag bit to determine the target flag bit corresponding to the operation result, and the target flag bit is used to indicate whether the operation result is a special number.
8. The floating-point arithmetic unit according to claim 2, wherein When the floating-point operation unit is an addition unit, the at least one intermediate pipeline register includes three pipeline registers; The logic processing module includes a first logic module, a second logic module, a third logic module, and a fourth logic module, where: The first logic module is configured to determine the exponent difference between the first initial exponent bit and the second initial exponent bit and the smaller exponent of the two, and perform a shift operation on the mantissa bit corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa bit and the second initial mantissa bit; The second logic module is configured to determine the exclusive OR result of the signs of the first initial sign bit and the second initial sign bit, and perform an addition operation on the aligned first initial mantissa bit and the second initial mantissa bit according to the exclusive OR result to obtain a first mantissa sum and the target sign bit; The third logic module is configured to perform a leading zero processing on the first mantissa sum to determine the leading zero statistical result and the shifted second mantissa sum; The fourth logic module is configured to perform an exponent shift operation according to the leading zero statistical result to obtain the target exponent bit; and perform a rounding operation on the second mantissa sum to obtain the target mantissa bit.
9. The floating-point arithmetic unit according to claim 8, wherein The first logic module includes an exponent subtraction unit, a first control unit, a mantissa shift unit, a first selection unit, a second selection unit, and a third selection unit, where: The exponent subtraction unit is configured to perform a subtraction operation on the first initial exponent bit and the second initial exponent bit to determine the exponent difference, and transmit the exponent difference to the first control unit; The first control unit is configured to receive the exponent difference, and generate a first control signal transmitted to the first selection unit, a second control signal transmitted to the second selection unit, and a third control signal transmitted to the third selection unit according to the exponent difference; The first selection unit is configured to receive the first control signal, and select the larger exponent from the first initial exponent bit and the second initial exponent bit according to the first control signal; The second selection unit is configured to receive the second control signal, and select the mantissa bit corresponding to the smaller exponent from the first initial mantissa bit and the second initial mantissa bit according to the second control signal; The third selection unit is configured to receive the third control signal, and select the mantissa bit corresponding to the larger exponent from the first initial mantissa bit and the second initial mantissa bit according to the third control signal; The mantissa shift unit is configured to perform a shift operation on the mantissa bit corresponding to the smaller exponent according to the exponent difference, so as to align the first initial mantissa bit and the second initial mantissa bit.
10. The floating-point arithmetic unit according to claim 9, characterized in that, The second logic module includes an exclusive-or unit, a second control unit, and a mantissa addition unit, wherein: The exclusive-or unit is configured to perform a sign exclusive-or operation on the first initial sign bit and the second initial sign bit to obtain the sign exclusive-or result, and transmit the sign exclusive-or result to the second control unit; The second control unit is configured to receive the sign exclusive-or result, and generate a fourth control signal transmitted to the mantissa addition unit according to the sign exclusive-or result; The mantissa addition unit is configured to perform an addition operation on the aligned first initial mantissa bit and second initial mantissa bit according to the fourth control signal to obtain the first mantissa sum and the target sign bit.
11. The floating-point arithmetic unit according to claim 10, characterized in that, The third logic module includes a leading zero processing unit, wherein: The leading zero processing unit is configured to perform a leading zero prediction on the first mantissa sum to determine the leading zero statistical result; and perform a shift processing on the first mantissa sum according to the leading zero statistical result to obtain the second mantissa sum.
12. The floating-point arithmetic unit according to claim 11, wherein The fourth logic module includes an exponent shift unit and a rounding operation unit, wherein: The exponent shift unit is configured to receive the larger exponent sent by the first selection unit and the leading zero statistical result sent by the leading zero processing unit, and perform an exponent shift operation on the larger exponent according to the leading zero statistical result to obtain the target exponent bit; The rounding operation unit is configured to perform a rounding operation on the second mantissa sum to obtain the target mantissa bit.
13. The floating-point operation unit according to claim 12, wherein The rounding operation unit is further configured to, in the case that the operation result is a denormal number, perform a rounding operation on the second mantissa sum to obtain a third mantissa sum; when the third mantissa sum exceeds a preset range, perform a normalization process on the third mantissa sum to obtain the target mantissa bit, and perform a plus 1 adjustment on the target exponent bit, so that the operation result is transformed from a denormal number to a normal number.
14. The floating-point arithmetic unit according to any one of claims 2 to 13, wherein: The input pipeline register is configured to merge and store the data to be pipelined among the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed within the same clock cycle; The output pipeline register is configured to merge and store the data to be pipelined among the target sign bit, target exponent bit, and target mantissa bit corresponding to the operation result within the same clock cycle; The intermediate pipeline register is configured to merge and store the intermediate result of the logical operation on the initial sign bit, initial exponent bit, and initial mantissa bit corresponding to the operand to be processed within the same clock cycle.
15. A floating-point processing device, characterized in that, The floating-point processing device includes a configuration module and a plurality of floating-point arithmetic units according to any one of claims 1 to 14, wherein: The configuration module is configured to call a target number of the floating-point arithmetic units from among the plurality of floating-point arithmetic units according to the precision of the operand to be processed; The floating-point arithmetic unit is configured to perform floating-point arithmetic on the operand to be processed to obtain an operation result; wherein, the target number is related to the precision of the operand to be processed.
16. The floating-point processing device according to claim 15, wherein: The configuration module is further configured to truncate the operand to be processed according to the precision of the operand to be processed to obtain a target number of data segments, and input the target number of data segments to the target number of the floating-point arithmetic units correspondingly.
17. The floating-point processing device according to claim 15, wherein: When the precision of the operand to be processed is double precision, the target number is set to m; When the precision of the operand to be processed is single precision, the target number is set to n; When the precision of the operand to be processed is half precision, the target number is set to k; Wherein, m, n, and k are all positive integers, and m:n:k satisfies a preset ratio requirement.
18. The floating-point processing device according to claim 15, wherein, The floating-point processing device further includes a mode selection unit, and a plurality of input ends of the mode selection unit are respectively connected to the output ends of the floating-point arithmetic units called under different precisions correspondingly, and a control end of the mode selection unit is configured to receive an operation mode signal, wherein: The mode selection unit is configured to output the operation results corresponding to the target number of the floating-point arithmetic units called based on the precision of the operand to be processed according to the operation mode signal.
19. The floating-point processing device according to any one of claims 15 to 18, characterized in that, The plurality of floating-point arithmetic units include a plurality of multiplication units and a plurality of addition units, wherein: The configuration module is configured to call a target number of the multiplication units from among the plurality of multiplication units and call a target number of the addition units from among the plurality of addition units according to the precision of the operand to be processed; The multiplication unit is configured to perform a multiplication operation on a first operand and a second operand in the operand to be processed to obtain a multiplication result; The addition unit is configured to perform an addition operation on the multiplication result and a third operand in the operand to be processed to obtain the operation result.
20. The floating-point processing device according to claim 19, wherein The floating-point processing device further includes a first trigger unit and a second trigger unit, wherein: The first trigger unit is configured to input the multiplication result to a target number of the addition units in a preset clock cycle; The second trigger unit is configured to input the flag bit information of the multiplication result to a target number of the addition units in a preset clock cycle.
21. A floating-point processing device, characterized in that, The floating-point processing device includes a register and a plurality of floating-point operation units as described in any one of claims 1 to 14, wherein: The register is configured to store at least two floating-point numbers in the same clock cycle; The floating-point operation unit is configured to obtain the operand to be processed from the register.
22. The floating-point processing device according to claim 21, wherein The floating-point operation unit includes a multiplication unit, and the register includes a first register, wherein: The first register is configured to store a first operand and a second operand in the same clock cycle, and input the operand output by the first register to at least one multiplication unit; The at least one multiplication unit is configured to perform a multiplication operation on the operand output by the first register to determine at least one multiplication result.
23. The floating-point processing device according to claim 21, wherein The floating-point operation unit includes an addition unit, and the register includes a second register, wherein: The second register is configured to store a third operand and a fourth operand in the same clock cycle, and input the operand output by the second register to at least one addition unit; The at least one addition unit is configured to perform an addition operation on the operand output by the second register to determine at least one addition result.
24. The floating-point processing device according to claim 21, wherein The floating-point operation unit includes a multiplication unit and an addition unit, and the register includes a first register and a second register, wherein: The first register is configured to store a first operand and a second operand in the same clock cycle, and input the operand output by the first register to at least one multiplication unit; The at least one multiplication unit is configured to perform a multiplication operation on the operand output by the first register to determine at least one multiplication result; The second register is configured to store a third operand and the at least one multiplication result in the same clock cycle, and input the operand output by the second register to at least one addition unit; The at least one addition unit is configured to perform an addition operation on the operand output by the second register to determine at least one addition result.
25. The floating-point processing device according to claim 22 or 24, characterized in that, The floating-point processing device further includes a reallocation unit, wherein: The reallocation unit is configured to reallocate the input for the at least one multiplication unit when the first register outputs, and input the reallocated operand to the at least one multiplication unit.
26. A floating-point processing method, characterized in that, The method includes: A logic processing module performs a logic operation on the input operand to be processed to obtain an operation result; A delay control module controls the delay between the input of the operand to be processed to the logic processing module and the input to the floating-point operation unit; and / or, controls the delay between the output of the operation result by the logic processing module and the output of the operation result by the floating-point operation unit; Wherein, the delay control module includes an input pipeline register and an output pipeline register, and the method further includes: The input pipeline register controls whether to pipeline at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed according to a first configuration parameter, so that the calculation path delay of at least one of the initial sign bit, the initial exponent bit, and the initial mantissa bit meets a first delay requirement; and / or, The output pipeline register controls whether to pipeline at least one of the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result according to a second configuration parameter, so that the output path delay of at least one of the target sign bit, the target exponent bit, and the target mantissa bit meets a second delay requirement.
27. The method according to claim 26, wherein At least one intermediate pipeline register is arranged between the input pipeline register and the output pipeline register, and the method further includes: The at least one intermediate pipeline register performs a logical operation on the initial sign bit, the initial exponent bit, and the initial mantissa bit corresponding to the operand to be processed, and determines the target sign bit, the target exponent bit, and the target mantissa bit corresponding to the operation result.
28. A floating-point processing method, characterized in that, The method includes: The configuration module calls a target number of the floating-point operation units from a plurality of floating-point operation units as described in any one of claims 1 to 14 according to the precision of the operand to be processed; The floating-point operation unit performs a floating-point operation on the operand to be processed to obtain an operation result; wherein, the target number is related to the precision of the operand to be processed.
29. The method according to claim 28, wherein The method further includes: The configuration module truncates the operand to be processed according to the precision of the operand to be processed to obtain a target number of data segments, and inputs the target number of data segments to the target number of the floating-point operation units correspondingly.
30. The method according to claim 28, characterized in that, The method further includes: The configuration module calls a target number of the multiplication units from a plurality of multiplication units and calls a target number of the addition units from a plurality of addition units according to the precision of the operand to be processed; The multiplication unit performs a multiplication operation on a first operand and a second operand in the operand to be processed to obtain a multiplication result; The addition unit performs an addition operation on the multiplication result and a third operand in the operand to be processed to obtain the operation result.
31. A floating-point processing method, characterized in that, The method includes: The register stores at least two floating-point numbers in the same clock cycle; The floating-point operation unit obtains the operand to be processed from the register, so that the floating-point operation unit executes the floating-point processing method as described in claim 26 or 27.
32. The method according to claim 31, wherein The floating-point operation unit includes a multiplication unit and an addition unit, the register includes a first register and a second register, and the method further includes: The first register stores a first operand and a second operand in the same clock cycle, and inputs the operand output by the first register to at least one multiplication unit; The at least one multiplication unit performs a multiplication operation on the operand output by the first register to determine at least one multiplication result; The second register stores a third operand and the at least one multiplication result in the same clock cycle, and inputs the operand output by the second register to at least one addition unit; The at least one adder unit performs an addition operation on the operands output by the second register to determine at least one addition result.
33. The method according to claim 32, wherein The method further includes: The reallocation unit reallocates inputs for the at least one multiplication unit when the first register outputs, and inputs the reallocated operands to the at least one multiplication unit.
34. A chip, characterized in that, The chip includes: a floating-point arithmetic unit according to any one of claims 1 to 14, or a floating-point processing device according to any one of claims 15 to 20, or a floating-point processing device according to any one of claims 21 to 25.
35. An electronic device, characterized in that, The electronic device includes a processor, wherein the processor includes: a floating-point arithmetic unit according to any one of claims 1 to 14; or a floating-point processing device according to any one of claims 15 to 20; or a floating-point processing device according to any one of claims 21 to 25.
Citation Information
Patent Citations
Double-precision floating-point operation
CN109196465A
Supporting 8-bit floating point format operands in computing architecture
CN115129370A
Fixed-floating-point SIMD multiply-add instruction fusion processing device and method and processor
CN117251132A
SRT operational circuit
CN118312129A
Multiply-Accumulate Engine with Non-Normalized Floating-Point Accumulator
US20240103808A1