Multi-precision floating point number addition and subtraction operation device and floating point operation unit

By designing a multi-precision floating-point number addition and subtraction operation device for multi-precision floating-point numbers in the RISC-V architecture, using separation, order and operation sorting circuits, the problems of low resource utilization, high power consumption and complex design in the prior art are solved, and efficient addition and subtraction operation for floating-point numbers with different precisions are achieved.

CN120045158APending Publication Date: 2025-05-27SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510167607.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the RISC-V architecture, the addition and subtraction operations of half-precision floating point numbers, single-precision floating point numbers and double-precision floating point numbers require the design of independent circuit structures, resulting in low resource utilization, high power consumption, and high design and verification complexity.

Method used

A multi-precision floating-point number addition and subtraction operation device is designed, including a separation operation circuit, a staging operation circuit and an operation sorting circuit. The accuracy is distinguished through the encoding segments in the target instruction, and the separation, staging, completion, shift and format output operations are performed to realize the addition and subtraction operation of floating-point number with different precisions.

Benefits of technology

It realizes the addition and subtraction of double-precision, single-precision and half-precision floating-point numbers using a type of hardware circuit, which reduces circuit complexity and power consumption and improves resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045158A_ABST
    Figure CN120045158A_ABST
Patent Text Reader

Abstract

The invention provides a multi-precision floating-point number addition and subtraction operation device and a floating-point operation unit, and belongs to the technical field of computers.The device comprises a separation operation circuit used for obtaining two floating-point numbers from a floating-point register set of an instruction set architecture processor according to a target instruction and conducting separation operation on the two floating-point numbers respectively, two separation numbers are obtained; the target instruction comprises an operation instruction for performing operation processing on double-precision, single-precision and half-precision floating-point numbers; the order matching operation circuit is used for carrying out order code order matching, mantissa complementing and shifting operation on the two separated numbers to obtain a first processing number and a second processing number; and the operation sorting circuit is used for performing addition and subtraction operation on the first processing number and the second processing number and performing formatting output on an operation result. Through the circuit architecture, not only can the circuit complexity and power consumption during addition and subtraction operation of the floating-point number be reduced, but also the resource utilization rate of the circuit can be remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technologies, and in particular, to an addition and subtraction operation device for multi-precision floating-point numbers and a floating-point operation unit. Background Art

[0002] The operation of floating-point numbers is a fundamental and crucial task in a computer, which involves the accurate representation and calculation of real numbers. Among them, floating-point numbers include data formats such as double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Double-precision floating-point numbers are often used in scenarios such as scientific computing, high-precision numerical simulation, and financial computing that require high precision and a large range of numerical representation; single-precision floating-point numbers are widely used in fields such as graphics processing, machine learning, and game development that have high requirements for performance and storage; half-precision floating-point numbers are suitable for application scenarios such as neural network inference, mobile devices, and embedded systems that have high requirements for computing speed and storage efficiency but relatively low requirements for precision. In many actual application scenarios, the floating-point operations of these several precisions may need to be performed alternately.

[0003] Currently, in the RISC-V (Reduced Instruction Set Computer - Version 5, instruction set) architecture, the addition and subtraction operations of half-precision floating-point data, single-precision floating-point data, and double-precision floating-point data are usually designed as independent circuit units, each occupying different hardware resources. Although this design improves the operation speed of floating-point data in specific cases, it also brings problems such as low resource utilization rate, high power consumption, and increased design and verification complexity. Currently, there is no relatively effective solution to this technical problem. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide an addition and subtraction operation device for multi-precision floating-point numbers and a floating-point operation unit, so as to solve the technical problems in the related art that when performing addition and subtraction operations on half-precision floating-point numbers, single-precision floating-point numbers, and double-precision floating-point numbers, it is necessary to design independent circuit structures, each occupying different hardware resources, resulting in low resource utilization rate, high power consumption, and high design and verification complexity.

[0005] To solve the above technical problems, the present invention provides an addition and subtraction operation device for multi-precision floating-point numbers, including:

[0006] A separation operation circuit is configured to obtain a first floating-point number and a second floating-point number from a floating-point register bank of an instruction set architecture processor according to a target instruction, and perform separation operations on the first floating-point number and the second floating-point number respectively to obtain a first separated number and a second separated number; the target instruction includes: arithmetic instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers; when the target instruction operates on the double-precision floating-point number, the single-precision floating-point number, and the half-precision floating-point number, it is distinguished by an encoding segment for specifying the floating-point precision in the arithmetic instruction;

[0007] An exponent alignment operation circuit is configured to perform exponent alignment, mantissa complementation, and shifting operations on the first separated number and the second separated number to obtain a first processed number and a second processed number;

[0008] An arithmetic and formatting circuit is configured to perform addition and subtraction operations on the first processed number and the second processed number, and perform formatted output on the operation result.

[0009] In a specific embodiment of the present application, if the target instruction is an arithmetic instruction for operating on a double-precision floating-point number, the encoding segment for specifying the floating-point precision in the target instruction is 2'b01; if the target instruction is an arithmetic instruction for operating on a single-precision floating-point number, the encoding segment for specifying the floating-point precision in the target instruction is 2'b00; if the target instruction is an arithmetic instruction for operating on a half-precision floating-point number, the encoding segment for specifying the floating-point precision in the target instruction is 2'b10 or 2'b11.

[0010] In a specific embodiment of the present application, the separation operation circuit includes:

[0011] A first separation circuit is configured to separate the first floating-point number into a sign bit, an exponent, and a mantissa to obtain the first separated number;

[0012] A second separation circuit is configured to separate the second floating-point number into a sign bit, an exponent, and a mantissa to obtain the second separated number; the first separation circuit and the second separation circuit have the same setting structure.

[0013] In a specific embodiment of the present application, the first floating-point number is stored in a first general floating-point register of the floating-point register bank, and the number of bits of the first general floating-point register is 64 bits;

[0014] If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 63 of the first general floating-point register;

[0015] If the first floating-point number is a single-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 31 of the first general-purpose floating-point register, and bits 32 to 63 of the first general-purpose floating-point register are all filled with zeros;

[0016] If the first floating-point number is a half-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 15 of the first general-purpose floating-point register, and bits 16 to 63 of the first general-purpose floating-point register are all filled with zeros.

[0017] In a specific embodiment of the present application, the first separation circuit includes: a first multiplexer, a second multiplexer, and a third multiplexer;

[0018] Among them, the signal selection ports of the first multiplexer, the second multiplexer, and the third multiplexer are all used to receive the encoding segment in the target instruction for specifying the floating-point precision;

[0019] The first input terminal, the second input terminal, and the third input terminal of the first multiplexer are respectively used to receive the data stored in bits 63, 31, and 15 of the first general-purpose floating-point register, and the output terminal of the first multiplexer is used to output the sign bit of the first separated number;

[0020] The first input terminal of the second multiplexer is used to receive the data stored in bits 62 to 52 of the first general-purpose floating-point register, the second input terminal of the second multiplexer is used to receive the data stored in bits 30 to 23 of the first general-purpose floating-point register, the third input terminal of the second multiplexer is used to receive the data stored in bits 14 to 10 of the first general-purpose floating-point register, and the output terminal of the second multiplexer is used to output the exponent of the first separated number;

[0021] The first input terminal of the third multiplexer is used to receive the data stored in bits 51 to 0 of the first general-purpose floating-point register, the second input terminal of the third multiplexer is used to receive the data stored in bits 22 to 0 of the first general-purpose floating-point register, the third input terminal of the third multiplexer is used to receive the data stored in bits 9 to 0 of the first general-purpose floating-point register, and the output terminal of the third multiplexer is used to output the mantissa of the first separated number.

[0022] In a specific embodiment of the present application, the exponent alignment operation circuit includes:

[0023] An exponent alignment unit, configured to perform exponent alignment on the exponent of the first separated number and the exponent of the second separated number to obtain a target alignment result;

[0024] A mantissa complementing unit, configured to respectively complement the mantissas of the first separated number and the second separated number according to the target alignment result to obtain a mantissa complementing result of the first separated number and a mantissa complementing result of the second separated number;

[0025] A shift operation unit, configured to perform a shift operation on the mantissa complementing result of the first separated number and the mantissa complementing result of the second separated number to obtain the first processed number and the second processed number.

[0026] In a specific embodiment of the present application, when both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the bit width of the alignment result obtained by the exponent alignment unit performing exponent alignment on the exponent of the first separated number and the exponent of the second separated number is 12 bits;

[0027] When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the alignment result obtained by the exponent alignment unit performing exponent alignment on the exponent of the first separated number and the exponent of the second separated number is 9 bits;

[0028] When both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the bit width of the alignment result obtained by the exponent alignment unit performing exponent alignment on the exponent of the first separated number and the exponent of the second separated number is 6 bits.

[0029] In a specific embodiment of the present application, when both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the bit width of the complementing result obtained by the mantissa complementing unit complementing the mantissas of the first separated number and the second separated number is 53 bits;

[0030] When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the complementing result obtained by the mantissa complementing unit complementing the mantissas of the first separated number and the second separated number is 24 bits;

[0031] When both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the bit width of the complementing result obtained by the mantissa complementing unit complementing the mantissas of the first separated number and the second separated number is 11 bits.

[0032] In a specific embodiment of the present application, the arithmetic arrangement circuit includes:

[0033] An operator unit is configured to splice the first processed number by using a guard bit, a rounding bit, and a sticky bit to obtain first spliced data, splice the second processed number by using the guard bit, the rounding bit, and the sticky bit to obtain second spliced data, and perform an addition and subtraction operation on the first spliced data and the second spliced data to obtain target operation data;

[0034] An arrangement unit is configured to perform truncation processing on the target operation data according to the floating-point type of the first floating-point number and / or the second floating-point number, and perform formatted output on the truncation processing result.

[0035] To solve the above technical problems, the present invention further provides a floating-point operation unit, including an addition and subtraction operation device for a multi-precision floating-point number as disclosed above.

[0036] Beneficial effects: In the present invention, the encoding segments for specifying the floating-point precision in the operation instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers in the instruction set architecture processor are distinguished in advance, so that the operation instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers in the instruction set architecture processor have the same structural form.

[0037] When performing an addition and subtraction operation on floating-point numbers, the separation operation circuit first obtains a first floating-point number and a second floating-point number from the floating-point register of the instruction set architecture processor according to the target instruction, and separately performs a separation operation on the first floating-point number and the second floating-point number to obtain a first separated number and a second separated number; then, a mantissa alignment, mantissa complementation, and shift operation are performed on the first separated number and the second separated number by using the alignment operation circuit to obtain a first processed number and a second processed number; finally, an addition and subtraction operation is performed on the first processed number and the second processed number by using the operation and arrangement circuit, and formatted output is performed on the operation result. In this setup architecture, it is equivalent to using only one type of hardware circuit to perform addition and subtraction operation processing on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Compared with the current situation where three different types of hardware circuits are required to separately perform addition and subtraction operation processing on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers, through this circuit architecture, not only can the complexity and power consumption of the circuit be reduced, but also the resource utilization rate of the circuit can be significantly improved.

[0038] Correspondingly, a floating-point operation unit provided by the present invention also has the above beneficial effects. Description of the Drawings

[0039] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0040] Figure 1 It is a structural diagram of an addition and subtraction operation device for multi-precision floating-point numbers provided by an embodiment of the present invention;

[0041] Figure 2 It is a structural schematic diagram of a target instruction;

[0042] Figure 3 It is a schematic diagram when a RISC-V processor processes the target instruction using a four-stage pipeline technology;

[0043] Figure 4 It is a structural diagram of a first separation circuit provided by an embodiment of the present invention;

[0044] Figure 5 It is a schematic flow diagram when performing addition and subtraction operations on multi-precision floating-point numbers provided by an embodiment of the present invention. Specific Embodiments

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0046] The terms "including" and "having" in the specification of the present invention and the above drawings, and any variations related to "including" and "having", are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units not listed.

[0047] To enable those skilled in the art to better understand the solution of the present invention, the following will further elaborate on the present invention with reference to the drawings and specific embodiments.

[0048] Currently, when performing operations on floating-point numbers, the most widely used is IEEE745 (IEEE Binary Floating-Point Arithmetic Standard). Please refer to Table 1, and Table 1 is the IEEE754 data type.

[0049] Table 1

[0050] Data type Sign bit Exponent Mantissa Normalized number 0 / 1 Any value Any value Denormalized number 0 / 1 All 0s Not all 0s Zero 0 / 1 All 0s All 0s Infinity 0 / 1 All 1s All 0s Not a number 0 / 1 All 1s Not all 0s

[0051] In different instruction sets, the implementation methods for floating-point operations are different. For example: In the x86 architecture, the SSE (Streaming SIMD Extensions) instruction set is provided for floating-point operations. This instruction set includes basic operation instructions such as addition, subtraction, multiplication, and division of floating-point numbers, as well as auxiliary instructions for converting, comparing, and rounding floating-point numbers. The design of this instruction set aims to improve the efficiency and performance of floating-point operations. In the ARM (Advanced RISC Machine) architecture, the NEON (New Engine for Next Generation Objects) instruction set is provided for floating-point operations. This instruction set includes vectorized operation instructions for floating-point numbers. This instruction set can process multiple floating-point data simultaneously, and its design purpose is to utilize hardware parallelism and vectorized computing to improve the operation parallelism and efficiency of floating-point numbers.

[0052] Although the floating-point operation technology has different implementation methods in different instruction sets, their ultimate goals are to improve the operation efficiency and performance of floating-point numbers, or to control the power consumption of the processor, so as to meet the requirements of different application scenarios.

[0053] The instruction set architecture processor (RISC-V architecture processor, Reduced Instruction Set Computer - Version 5 processor) originated from a research project at the University of California, Berkeley, aiming to provide an open, modular, and extensible instruction set design. The RISC-V processor (that is, the instruction set architecture processor) adopts a simple and flexible design principle. The instruction set in this architecture processor includes both a basic integer instruction set and an optional extended instruction set. The F extension instruction set and the D extension instruction set in the extended instruction set are used to process single-precision floating-point operations and double-precision floating-point operations respectively. The F extension instruction set and the D extension instruction set not only cover basic arithmetic operation instructions such as addition, subtraction, multiplication, and division of single-precision floating-point numbers and double-precision floating-point numbers, but also cover auxiliary instructions such as conversion, comparison, and rounding of single-precision floating-point numbers and double-precision floating-point numbers.

[0054] Since the F extension instruction and D extension instruction set in the RISC-V processor provide a series of instructions for floating-point operations, and their implementation technologies involve multiple aspects such as the hardware design of the processor and the optimization of floating-point operations, the RISC-V processor needs to have hardware support for floating-point operations. For example, the RISC-V processor needs to set up a floating-point register file, a floating-point arithmetic unit, a floating-point status register, etc. to implement the execution of floating-point operation instructions and the precision control of data.

[0055] To improve the instruction execution efficiency of the RISC-V processor, the RISC-V processor usually adopts pipeline technology to execute floating-point operation instructions in parallel. In terms of hardware implementation, the RISC-V processor also needs to execute efficient floating-point operation algorithms, including not only basic arithmetic operations such as addition, subtraction, multiplication, and division of floating-point numbers, but also auxiliary operations such as conversion, comparison, and rounding of floating-point numbers, so as to ensure the accuracy and reliability of the operation results. In addition, the RISC-V processor also needs to design corresponding exception handling mechanisms and status management mechanisms to handle possible exceptions during floating-point operations.

[0056] In the prior art, when performing floating-point operations in the RISC-V processor, the addition and subtraction operations of half-precision floating-point data, single-precision floating-point data, and double-precision floating-point data are usually designed as independent circuit units. Although this design improves the operation speed in specific cases, it also brings problems such as low resource utilization, high power consumption, and increased design and verification complexity.

[0057] To solve the above technical problems, the present invention provides an addition and subtraction operation device for multi-precision floating-point numbers. Under this operation device, only one type of hardware circuit can be used to perform addition and subtraction operation processing on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers, which can not only reduce the complexity and power consumption of the circuit, but also significantly improve the resource utilization rate of the circuit.

[0058] Please refer to Figure 1 , Figure 1 which is the structural diagram of an addition and subtraction operation device for multi-precision floating-point numbers provided by an embodiment of the present invention. The device includes:

[0059] A separation operation circuit 11 is configured to obtain a first floating-point number and a second floating-point number from the floating-point register bank of an instruction set architecture processor according to a target instruction, and perform separation operations on the first floating-point number and the second floating-point number respectively to obtain a first separated number and a second separated number; the target instruction includes: an arithmetic instruction for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers; when the target instruction operates on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers, it is distinguished by an encoding segment for specifying the floating-point precision in the arithmetic instruction.

[0060] An exponent alignment operation circuit 12 is configured to perform exponent alignment, mantissa complementation, and shifting operations on the first separated number and the second separated number to obtain a first processed number and a second processed number.

[0061] An arithmetic and formatting circuit 13 is configured to perform addition and subtraction operations on the first processed number and the second processed number, and perform formatted output on the operation result.

[0062] In this embodiment, in order to enable those skilled in the art to clearly understand the implementation principle of the present invention, the target instruction and the application scenario of the multi-precision floating-point number addition and subtraction operation device of the present invention are specifically described first.

[0063] Among them, the target instruction includes: an arithmetic instruction for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Currently, in the RISC-V processor, the arithmetic instruction for operating on double-precision floating-point numbers is usually called a D extension instruction, and the arithmetic instruction for operating on single-precision floating-point numbers is usually called an F extension instruction. In this application, we can call the arithmetic instruction for operating on half-precision floating-point numbers an H extension instruction. In other words, in the present invention, the target instruction can be either a D extension instruction, an F extension instruction, or an H extension instruction.

[0064] Please refer to Figure 2 , Figure 2 for the structural schematic diagram of the target instruction. In Figure 2Among them, bits 0 to 6 of the target instruction are the opcode encoding segment of the instruction, which is used to determine the type of the instruction. The encoding of the floating-point operation instruction in this segment is 7’b1010011; bits 27 to 31 of the target instruction are the funct5 encoding segment of the instruction, which is used to further distinguish the instruction operations. For example, the encoding of the floating-point addition in this segment is 0x00, the encoding of the floating-point subtraction in this segment is 0x01, and the encoding of the floating-point multiplication in this segment is 0x02; bits 25 to 26 of the target instruction are the fmt encoding segment of the instruction, which is used to specify the precision format of the floating-point number. For single-precision floating-point instructions, this encoding segment is generally set to 2’b00, while for double-precision floating-point instructions, this encoding segment is generally set to 2’b01; bits 12 to 14 of the target instruction are the rm encoding segment of the instruction, which is used to specify the rounding mode of the floating-point operation. For details, please refer to Table 2, which shows the rounding modes of the floating-point operation; bits 20 to 24, bits 15 to 19, and bits 7 to 11 of the target instruction represent the register address of the floating-point operand 2, the register address of the floating-point operand 1, and the destination register address, respectively.

[0065] From Figure 2 it can be seen that since the fmt encoding segment in the target instruction is used to specify the precision format of the floating-point number, we can distinguish the floating-point operation type corresponding to the target instruction through the fmt encoding segment (which is used to specify the floating-point precision).

[0066] Table 2

[0067] Rounding encoding format Abbreviation Explanation 000 RNE Round to nearest even 001 RTZ Round to zero 010 RDN Round down (to negative infinity) 011 RUP Round up (to positive infinity) 100 RMM Round to nearest maximum magnitude 101 -- Invalid 110 -- Invalid 111 -- Dynamic rounding mode

[0068] In addition, in the RISC-V processor, a four-stage pipeline technology is usually adopted to process the target instruction. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the RISC-V processor using the four-stage pipeline technology to process the target instruction. From Figure 3 it can be seen that when the RISC-V processor uses the four-stage pipeline technology to process the target instruction, it is usually divided into four stages: instruction fetch, decoding, execution, and write-back.

[0069] Among them, the execution of floating-point operation instructions is implemented by a floating-point unit (FPU, Floating Point Unit). In the instruction fetch stage, the RISC-V processor fetches instructions from memory; in the decoding stage, the RISC-V processor compiles the fetched instructions to obtain information such as the operand register address, destination register address, and instruction type required by the instructions. After decoding is completed, it enters the execution stage. In the execution stage, the RISC-V processor fetches the operands according to the register address and executes the corresponding operations. For floating-point operation instructions, the floating-point numbers and information related to the instruction type obtained after decoding are passed to the floating-point unit pipeline, and the final result is written back to the floating-point register file through the write-back stage.

[0070] In the present invention, in order to reduce the circuit complexity when performing addition and subtraction operations on multi-precision floating-point numbers, a separation operation circuit 11, an exponent alignment operation circuit 12, and an arithmetic arrangement circuit 13 are provided in the floating-point unit. Among them, the separation operation circuit 11 obtains the first floating-point number and the second floating-point number from the floating-point register file of the RISC-V processor according to the target instruction, and performs separation operations on the first floating-point number and the second floating-point number respectively to obtain the first separated number and the second separated number.

[0071] It should be noted that since the 20th to 24th bits of the target instruction represent the register address of the floating-point operand 1, and the 15th to 19th bits of the target instruction represent the register address of the floating-point operand 2, therefore, according to the data encoding on the 20th to 24th bits of the target instruction and the register addresses corresponding to the data encoding on the 15th to 19th bits, the corresponding first floating-point number and second floating-point number can be found from the floating-point register.

[0072] After that, the exponent alignment operation circuit 12 is used to perform exponent alignment, mantissa complementation, and shift operations on the first separated number and the second separated number to obtain the first processed number and the second processed number, so as to facilitate the subsequent floating-point calculation process; finally, the arithmetic arrangement circuit 13 is used to perform addition and subtraction operations on the first processed number and the second processed number, and format and output the operation result.

[0073] Obviously, in this setup architecture, it is equivalent to using only one type of hardware circuit to perform addition and subtraction operation processing on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Compared with the current situation where three different types of hardware circuits are required to perform addition and subtraction operation processing on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers respectively, through this circuit architecture, not only can the structural complexity and power consumption of the circuit be reduced, but also the resource utilization rate of the circuit can be significantly improved.

[0074] Based on the above embodiments, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation, if the target instruction is an arithmetic instruction for operating on double-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2'b01; if the target instruction is an arithmetic instruction for operating on single-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2'b00; if the target instruction is an arithmetic instruction for operating on half-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2'b10 or 2'b11.

[0075] According to Figure 2 it can be seen that in the target instruction, the encoding segment for specifying the floating-point precision is the fmt encoding segment. Therefore, we can distinguish the arithmetic instructions for double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers according to the encoding data on the fmt encoding segment in the target instruction.

[0076] For the sake of convenient description, we call the arithmetic instruction for operating on double-precision floating-point numbers the D extension instruction set, the arithmetic instruction for operating on single-precision floating-point numbers the F extension instruction set, and the arithmetic instruction for operating on half-precision floating-point numbers the H extension instruction set.

[0077] In the official documentation of the RISC-V processor, when the encoding data on the fmt encoding segment are 2'b01 and 2'b00 respectively, they are used to represent the arithmetic instructions for processing double-precision floating-point numbers and single-precision floating-point numbers. Therefore, when the fmt encoding segment in the target instruction is 2'b01, it indicates that the target instruction is an arithmetic instruction for operating on double-precision floating-point numbers; when the fmt encoding segment in the target instruction is 2'b00, it indicates that the target instruction is an arithmetic instruction for operating on single-precision floating-point numbers.

[0078] In the official documentation of the RISC-V processor, since 2'b10 and 2'b11 on the fmt encoding segment are not defined, we can use 2'b10 or 2'b11 on the fmt encoding segment to define the arithmetic instruction for operating on half-precision floating-point numbers, so that the arithmetic instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers are compatible with each other and can be distinguished at the same time.

[0079] Here, we assume that when the fmt encoding segment is 2'b10, it indicates that the target instruction is an arithmetic instruction for operating on half-precision floating-point numbers. Then the following are some specific examples of target instructions for different precision floating-point operation instructions:

[0080] When the target instruction is 00000 01 00010 00001 000 00011 1010011, it indicates that the target instruction is an instruction for performing an addition operation on double-precision floating-point numbers;

[0081] When the target instruction is 00000 10 00010 00001 000 00011 1010011, it indicates that the target instruction is an instruction for performing an addition operation on half-precision floating-point numbers;

[0082] When the target instruction is 00001 00 00010 00001 000 00011 1010011, it indicates that the target instruction is an instruction for performing a subtraction operation on single-precision floating-point numbers;

[0083] When the target instruction is 00001 10 00010 00001 000 00011 1010011, it indicates that the target instruction is an instruction for performing a subtraction operation on half-precision floating-point numbers.

[0084] Obviously, through the technical solution provided in this embodiment, it is possible to accurately distinguish the arithmetic instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers.

[0085] Based on the above embodiment, this embodiment further describes and optimizes the technical solution. As a preferred implementation manner, the separation operation circuit includes:

[0086] A first separation circuit for separating the first floating-point number into a sign bit, an exponent, and a mantissa to obtain a first separated number;

[0087] A second separation circuit for separating the second floating-point number into a sign bit, an exponent, and a mantissa to obtain a second separated number; the first separation circuit and the second separation circuit have the same setup structure.

[0088] Since the separation operation circuit needs to perform separation operations on the first floating-point number and the second floating-point number simultaneously to obtain the first separated number and the second separated number, therefore, in this embodiment, two first separation circuits and second separation circuits with the same structure are provided in the separation operation circuit, and the first separation circuit is used to separate the first floating-point number into a sign bit, an exponent, and a mantissa to obtain the first separated number, and the second separation circuit is used to separate the second floating-point number into a sign bit, an exponent, and a mantissa to obtain the second separated number.

[0089] In practical applications, logic devices such as multiplexers, signal selection circuits, or tri-state gates can be used to build the first separation circuit and the second separation circuit, and the built first separation circuit and second separation circuit are respectively used to perform separation operations on the first floating-point number and the second floating-point number, so as to separate the first floating-point number and the second floating-point number into a sign bit, an exponent, and a mantissa to correspondingly obtain a first separated number and a second separated number.

[0090] Obviously, through the technical solution provided by this embodiment, the first separation circuit and the second separation circuit can be used to perform separation operations on the first floating-point number and the second floating-point number respectively, and the first separated number and the second separated number can be correspondingly obtained.

[0091] As a preferred implementation manner, the first floating-point number is stored in the first general floating-point register of the floating-point register bank, and the number of bits of the first general floating-point register is 64 bits;

[0092] If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 63 of the first general floating-point register;

[0093] If the first floating-point number is a single-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 31 of the first general floating-point register, and bits 32 to 63 of the first general floating-point register are all filled with zeros;

[0094] If the first floating-point number is a half-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 15 of the first general floating-point register, and bits 16 to 63 of the first general floating-point register are all filled with zeros.

[0095] Since the RISC-V architecture stipulates that if the RISC-V processor needs to support single-precision floating-point instructions or double-precision floating-point instructions, a separate set of floating-point register banks must be added inside it, and there are 32 general floating-point registers in this floating-point register bank, labeled f0 to f31.

[0096] If the RISC-V processor only needs to support instructions for operating on double-precision floating-point numbers (D extension instructions), the width of each general floating-point register in the floating-point register bank is 64 bits; if the RISC-V processor only needs to support instructions for operating on single-precision floating-point numbers (F extension instructions), the width of each general floating-point register in the floating-point register bank is 32 bits.

[0097] In order to enable the addition and subtraction operation device described in the present invention to process double-precision floating-point instructions, single-precision floating-point instructions, and half-precision floating-point instructions simultaneously, it is necessary to Figure 3The 32 general-purpose floating-point registers in the internal floating-point register bank are all set to 64-bit registers.

[0098] If the general-purpose floating-point register used to store the first floating-point number in the floating-point register bank is the first general-purpose floating-point register. Then, if the first floating-point number is a double-precision floating-point number, the first floating-point number will be stored in the entire first general-purpose floating-point register, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 63 of the first general-purpose floating-point register; if the first floating-point number is a single-precision floating-point number, the first floating-point number will be stored in the lower 32 bits of the first general-purpose floating-point register, and the upper 32 bits of the first general-purpose floating-point register will be filled with zeros, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 31 of the first general-purpose floating-point register, and bits 32 to 63 of the first general-purpose floating-point register will all be filled with zeros; if the first floating-point number is a half-precision floating-point number, the first floating-point number will be stored in the lower 16 bits of the first general-purpose floating-point register, and the upper 48 bits of the first general-purpose floating-point register will be filled with zeros, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 15 of the first general-purpose floating-point register, and bits 16 to 63 of the first general-purpose floating-point register will all be filled with zeros.

[0099] Since the storage principle of the second floating-point number in the floating-point register bank is the same as that of the first floating-point number in the floating-point register bank, therefore, the storage method of the second floating-point number in the floating-point register bank will not be specifically described here.

[0100] Obviously, through the technical solution provided by this embodiment, the first floating-point number can be accurately and reliably stored in the general-purpose floating-point register of the floating-point register bank.

[0101] Please refer to Figure 4 , Figure 4 FIG. is a structural diagram of a first separation circuit provided by an embodiment of the present invention. As a preferred embodiment, the first separation circuit includes: a first multiplexer MUX1, a second multiplexer MUX2, and a third multiplexer MUX3;

[0102] Among them, the signal selection ports of the first multiplexer MUX1, the second multiplexer MUX2, and the third multiplexer MUX3 are all used to receive the encoding segment fmt for specifying the floating-point precision in the target instruction;

[0103] The first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are respectively used to receive the data stored in bits 63, 31, and 15 of the first general-purpose floating-point register, and the output terminal of the first multiplexer MUX1 is used to output the sign bit of the first separated number;

[0104] The first input terminal of the second multiplexer MUX2 is used to receive the data stored in bits 62 to 52 of the first general floating-point register, the second input terminal of the second multiplexer MUX2 is used to receive the data stored in bits 30 to 23 of the first general floating-point register, the third input terminal of the second multiplexer MUX2 is used to receive the data stored in bits 14 to 10 of the first general floating-point register, and the output terminal of the second multiplexer MUX2 is used to output the exponent of the first separated number;

[0105] The first input terminal of the third multiplexer MUX3 is used to receive the data stored in bits 51 to 0 of the first general floating-point register, the second input terminal of the third multiplexer MUX3 is used to receive the data stored in bits 22 to 0 of the first general floating-point register, the third input terminal of the third multiplexer MUX3 is used to receive the data stored in bits 9 to 0 of the first general floating-point register, and the output terminal of the third multiplexer MUX3 is used to output the mantissa of the first separated number.

[0106] In this embodiment, the structure of the first separation circuit is specifically described. Since the multiplexer has a simple structure, small footprint and low cost, in this embodiment, in order to reduce the structural complexity and design cost of the first separation circuit, three multiplexers are used to build the first separation circuit. Among them, the output signal of the first separation circuit is determined by the signals received by the signal selection ports of the first multiplexer MUX1, the second multiplexer MUX2 and the third multiplexer MUX3.

[0107] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b01, it indicates that the first floating-point number is a double-precision floating-point number, and the first separation circuit needs to separate the first floating-point number in double-precision floating-point format into a sign bit sign, an exponent exp and a mantissa fract to obtain a first separated number.

[0108] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the conducting, off, and off states respectively. At this time, the first multiplexer MUX1 receives the data stored in the 63rd bit of the first general-purpose floating-point register and outputs the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the conducting, off, and off states respectively. At this time, the second multiplexer MUX2 receives the data stored in the 62nd to 52nd bits of the first general-purpose floating-point register and outputs the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the conducting, off, and off states respectively. At this time, the third multiplexer MUX3 receives the data stored in the 51st to 0th bits of the first general-purpose floating-point register and outputs the mantissa of the first separated number.

[0109] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b00, it indicates that the first floating-point number is a single-precision floating-point number. The first separation circuit needs to separate the first floating-point number in the single-precision floating-point format into a sign bit, an exponent, and a mantissa to obtain the first separated number.

[0110] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the off, conducting, and off states respectively. At this time, the first multiplexer MUX1 receives the data stored in the 31st bit of the first general-purpose floating-point register and outputs the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the off, conducting, and off states respectively. At this time, the second multiplexer MUX2 receives the data stored in the 30th to 23rd bits of the first general-purpose floating-point register and outputs the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the off, conducting, and off states respectively. At this time, the third multiplexer MUX3 receives the data stored in the 22nd to 0th bits of the first general-purpose floating-point register and outputs the mantissa of the first separated number.

[0111] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b10, it indicates that the first floating-point number is a half-precision floating-point number. The first separation circuit needs to separate the first floating-point number in the half-precision floating-point format into a sign bit, an exponent, and a mantissa to obtain the first separated number.

[0112] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the off state, the off state, and the on state respectively. At this time, the first multiplexer MUX1 will receive the data stored in the 15th bit of the first general-purpose floating-point register and output the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the off state, the off state, and the on state respectively. At this time, the second multiplexer MUX2 will receive the data stored in the 14th bit to the 10th bit of the first general-purpose floating-point register and output the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the off state, the off state, and the on state respectively. At this time, the third multiplexer MUX3 will receive the data stored in the 9th bit to the 0th bit of the first general-purpose floating-point register and output the mantissa of the first separated number.

[0113] Please refer to Table 3. Table 3 shows the corresponding relationship of the separation operation of the first floating-point number under different floating-point precisions.

[0114] Table 3

[0115] fmt is 2’b00 fmt is 2’b01 fmt is 2’b10 Sign bit sign operand

[31] operand

[63] operand

[15] Exponent exp operand[30:23] operand[62:52] operand[14:10] Mantissa fract operand[22:0] operand[51:0] operand[9:0]

[0116] In Table 3, fmt represents the fmt encoding segment in the target instruction, the sign bit sign represents the output signal of the first multiplexer MUX1, the exponent exp represents the output signal of the second multiplexer MUX2, the fraction fract represents the output signal of the third multiplexer MUX3, and operand represents the operand.

[0117] It should be noted that considering compatibility issues, the separation results of the first floating-point number are all saved using the general-purpose floating-point registers with the corresponding bit widths of double-precision floating-point numbers. That is, when the first floating-point number is a double-precision floating-point number, except that the sign bit in the first separated number is still saved using the 1-bit general-purpose floating-point register sign, the exponent of the first separated number will be saved using the 11-bit general-purpose floating-point register exp[10:0], and the mantissa of the first separated number will be saved using the 52-bit general-purpose floating-point register fract[51:0]. When the first floating-point number is a single-precision floating-point number and a half-precision floating-point number, the separation results corresponding to the first floating-point number will have 0 filled in the highest bit of their corresponding general-purpose floating-point registers.

[0118] Obviously, through the technical solution provided in this embodiment, the first floating-point number can be separated using the first separation circuit with a simple structure and low cost.

[0119] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation, the exponent operation circuit includes:

[0120] An exponent alignment unit, configured to perform exponent alignment on the exponent of the first separated number and the exponent of the second separated number to obtain a target alignment result;

[0121] A mantissa complementing unit, configured to respectively perform mantissa complementing on the mantissa of the first separated number and the mantissa of the second separated number according to the target alignment result to obtain a mantissa complementing result of the first separated number and a mantissa complementing result of the second separated number;

[0122] A shift operation unit, configured to perform a shift operation on the mantissa complementing result of the first separated number and the mantissa complementing result of the second separated number to obtain a first processed number and a second processed number.

[0123] In this embodiment, the alignment operation circuit is essentially composed of an exponent alignment unit, a mantissa complementing unit, and a shift operation unit. Among them, the exponent alignment unit is configured to perform exponent alignment on the exponent of the first separated number and the exponent of the second separated number to obtain a target alignment result. When the exponent alignment unit performs the exponent alignment operation on the exponent of the first separated number and the exponent of the second separated number, it actually calculates the difference between the exponent of the first separated number and the exponent of the second separated number, and can compare the magnitudes of the two exponents while obtaining the difference. The mantissa complementing unit is configured to perform mantissa complementing on the mantissa of the first separated number and the mantissa of the second separated number according to the target alignment result, that is, the mantissa complementing unit will perform the operation of adding implicit bits to the mantissa and generate the complete mantissas of the first separated number and the second separated number; the shift operation unit is configured to perform a shift operation on the mantissa complementing result of the first separated number and the mantissa complementing result of the second separated number.

[0124] It should be noted that the exponent alignment unit, the mantissa complementing unit, and the shift operation unit can all be built by an adder circuit. In actual operation, since the exponent alignment unit, the mantissa complementing unit, and the shift operation unit are all well-known to those skilled in the art, the circuit structures of the exponent alignment unit, the mantissa complementing unit, and the shift operation unit will not be elaborated here.

[0125] Obviously, through the technical solution provided in this embodiment, the alignment operation circuit can complete the exponent alignment, mantissa complementing, and shift operation of the first separated number and the second separated number, and facilitate the continuation of the subsequent process.

[0126] As a preferred implementation manner, when both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the bit width of the alignment result obtained by the exponent alignment unit performing exponent alignment on the exponent of the first separated number and the exponent of the second separated number is 12 bits;

[0127] When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the alignment result obtained by the exponent alignment unit performing exponent alignment on the exponent of the first separated number and the exponent of the second separated number is 9 bits;

[0128] When both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the bit width of the alignment result of the exponent alignment unit for aligning the exponents of the first separated number and the second separated number is 6 bits.

[0129] According to Table 3, it can be seen that during the process of the exponent alignment unit aligning the exponents of the first separated number and the second separated number, if both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the exponent alignment unit actually subtracts two exponents with a bit width of 11 bits. Then, after the exponent alignment unit subtracts two exponents with a bit width of 11 bits, the alignment result will become an exponent with a bit width of 12 bits. If both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the exponent alignment unit actually subtracts two exponents with a bit width of 8 bits. Then, after the exponent alignment unit subtracts two exponents with a bit width of 8 bits, the alignment result will become an exponent with a bit width of 9 bits. If both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the exponent alignment unit actually subtracts two exponents with a bit width of 5 bits. Then, after the exponent alignment unit subtracts two exponents with a bit width of 5 bits, the alignment result will become an exponent with a bit width of 6 bits.

[0130] It should be noted that in practical applications, the alignment result after the exponent alignment unit aligns the exponents of the first separated number and the second separated number is still to select the calculation result with the corresponding precision through the encoding segment fmt in the target instruction.

[0131] In addition, it is worth noting that in actual operation, in order to further reduce the structural complexity of the exponent alignment unit, the exponent alignment unit can be set as an adder circuit with a bit width of 12 bits. In this setting method, regardless of whether the bit widths of the exponents of the first separated number and the second separated number are 11 bits, 8 bits, or 5 bits, the 12-bit adder circuit can be reused to perform the exponent alignment operation on the exponents of the first separated number and the second separated number. This can not only save the cumbersome process of setting 3 adder circuits with different bit widths in the exponent alignment unit, but also relatively reduce the occupied space volume of the exponent alignment unit.

[0132] As a preferred implementation manner, when both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the bit width of the complement result of the mantissa complement unit for complementing the mantissas of the first separated number and the second separated number is 53 bits;

[0133] When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the complement result of the mantissa complement unit for complementing the mantissas of the first separated number and the second separated number is 24 bits;

[0134] When both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the bit width of the complementation result of the mantissa complementation unit for the mantissa of the first separated number and the mantissa of the second separated number is 11 bits.

[0135] According to Table 3, when the mantissa complementation unit performs mantissa complementation on the mantissa of the first separated number and the mantissa of the second separated number, if both the first floating-point number and the second floating-point number are double-precision floating-point numbers, then the mantissa complementation unit needs to complement the mantissa with a bit width of 52 bits. At this time, the bit width of the complementation result of the mantissa of the first separated number and the mantissa of the second separated number is 53 bits. If both the first floating-point number and the second floating-point number are single-precision floating-point numbers, then the mantissa complementation unit needs to complement the mantissa with a bit width of 23 bits. At this time, the bit width of the complementation result of the mantissa of the first separated number and the mantissa of the second separated number is 24 bits. If both the first floating-point number and the second floating-point number are half-precision floating-point numbers, then the mantissa complementation unit needs to complement the 10-bit mantissa. At this time, the complementation result of the mantissa of the first separated number and the mantissa of the second separated number is 11 bits.

[0136] Table 4

[0137] fmt is 2’b00 fmt is 2’b01 fmt is 2’b10 Result of exponent alignment exp_sub[8:0] exp_sub[11:0] exp_sub[5:0] Result of mantissa complementation mant[8:0] mant[52:0] mant[10:0]

[0138] Please refer to Table 4. Table 4 shows the alignment results and mantissa complementation results corresponding to different precision floating-point numbers. In Table 4, when fmt is 2'b00, it means that the fmt encoding segment of the target instruction is 2'b00, indicating that both the first floating-point number and the second floating-point number are single-precision floating-point numbers. At this time, the alignment result of the exponent of the first separated number and the exponent of the second separated number is exp_sub[8:0] (occupying 9 bits), and the mantissa complementation result of the mantissa of the first separated number and the mantissa of the second separated number is mant[23:0] (occupying 24 bits); when fmt is 2'b01, it means that the fmt encoding segment of the target instruction is 2'b01, indicating that both the first floating-point number and the second floating-point number are double-precision floating-point numbers. At this time, the alignment result of the exponent of the first separated number and the exponent of the second separated number is exp_sub[11:0] (occupying 12 bits), and the mantissa complementation result of the mantissa of the first separated number and the mantissa of the second separated number is mant[52:0] (occupying 53 bits); when fmt is 2'b10, it means that the fmt encoding segment of the target instruction is 2'b10, indicating that both the first floating-point number and the second floating-point number are half-precision floating-point numbers. At this time, the alignment result of the exponent of the first separated number and the exponent of the second separated number is exp_sub[5:0] (occupying 6 bits), and the mantissa complementation result of the mantissa of the first separated number and the mantissa of the second separated number is mant[10:0] (occupying 11 bits).

[0139] Obviously, through the technical solution provided in this embodiment, the exponent alignment of the exponent of the first separated number and the exponent of the second separated number can be accurately and reliably performed, and the mantissa complement of the mantissa of the first separated number and the mantissa of the second separated number can be performed.

[0140] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation manner, the operation and arrangement circuit includes:

[0141] An operation sub-unit, configured to splice the first processed number by using a guard bit, a round bit, and a sticky bit to obtain a first spliced data, splice the second processed number by using a guard bit, a round bit, and a sticky bit to obtain a second spliced data, and perform an addition and subtraction operation on the first spliced data and the second spliced data to obtain a target operation data;

[0142] An arrangement sub-unit, configured to perform an interception process on the target operation data according to the floating-point type of the first floating-point number and / or the second floating-point number, and perform a formatted output on the result of the interception process.

[0143] In this embodiment, an operation sub-unit and an arrangement sub-unit are provided in the operation and arrangement circuit. Among them, the operation sub-unit is configured to splice the first processed number and the second processed number respectively by using three auxiliary bits, namely, a guard bit (Guard), a round bit (Round), and a sticky bit (Sticky), to obtain a first spliced data and a second spliced data, and perform an addition and subtraction operation on the first spliced data and the second spliced data to obtain a target operation data. And the arrangement sub-unit is configured to perform an interception process on the target operation data according to the floating-point type of the first floating-point number and the second floating-point number, and perform a formatted output on the result of the interception process.

[0144] Specifically, if both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the first processed number and the second processed number participating in the addition and subtraction operation are essentially 53-bit complete mantissas. Because of the rounding requirement, the operation sub-unit can first splice the first processed number and the second processed number simultaneously by using the three auxiliary bits, namely, the guard bit, the round bit, and the sticky bit, to obtain two 56-bit data; then, perform an addition and subtraction operation on these two spliced data, and a 57-bit target operation data will be generated, which can be marked as mant_value[0:56] here. On this basis, continue to perform judgments such as rounding and overflow, perform an interception process on mant_value[0:56] according to the double-precision floating-point type, and perform a formatted output on the result of the interception process.

[0145] When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the first operand and the second operand participating in the addition or subtraction operation are essentially 24-bit complete mantissas. Because of the rounding requirement, the arithmetic sub-unit can first use three auxiliary bits, namely the guard bit, the round bit, and the sticky bit, to splice the first operand and the second operand simultaneously, obtaining two 27-bit data; then, perform the addition or subtraction operation on these two spliced data, generating a 28-bit target operation data, which can be marked as mant_value[0:27] here. On this basis, continue to perform judgments such as rounding and overflow, intercept mant_value[0:27] according to the single-precision floating-point type, and format and output the intercepted result.

[0146] Similarly, when both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the first operand and the second operand participating in the addition or subtraction operation are essentially 11-bit complete mantissas. Because of the rounding requirement, the arithmetic sub-unit can first use three auxiliary bits, namely the guard bit, the round bit, and the sticky bit, to splice the first operand and the second operand simultaneously, obtaining two 14-bit data; then, perform the addition or subtraction operation on these two spliced data, generating a 15-bit target operation data, which can be marked as mant_value[0:14] here. On this basis, continue to perform judgments such as rounding and overflow, intercept mant_value[0:14] according to the single-precision floating-point type, and format and output the intercepted result.

[0147] Obviously, through the technical solution provided in this embodiment, the addition and subtraction operations on floating-point data with different precisions can be completed.

[0148] To enable those skilled in the art to more clearly understand the implementation principle of the present invention, the operation process when performing addition and subtraction operations on multi-precision floating-point numbers is specifically described here. Please refer to Figure 5 , Figure 5 which is a schematic flowchart of the process of performing addition and subtraction operations on multi-precision floating-point numbers provided by the embodiment of the present invention.

[0149] When the RSIC-V processor determines according to the target instruction that the addition or subtraction operation needs to be performed on the first floating-point number operand1 and the second floating-point number operand2, it will first use the first separation circuit and the second separation circuit in the separation operation circuit to perform separation operations on the first floating-point number and the second floating-point number respectively, and separate the first floating-point number into three parts: the sign bit sign_op1, the exponent exp_op1, and the mantissa fract_op1, and at the same time separate the second floating-point number into three parts: the sign bit sign_op2, the exponent exp_op2, and the mantissa fract_op2.

[0150] It should be noted that the first floating-point number operand1 and the second floating-point number operand2 can be double-precision floating-point numbers, single-precision floating-point numbers, or half-precision floating-point numbers. Among them, the floating-point number types of the first floating-point number operand1 and the second floating-point number operand2 are determined by the fmt encoding in the target instruction.

[0151] Then, the exponent alignment unit in the exponent alignment operation circuit aligns the exponents exp_op1 and exp_op2, and the mantissa complementing unit and the shift operation unit in the exponent alignment operation circuit complement and shift the mantissas fract_op1 and fract_op2 according to the exponent difference exp_sub between exp_op1 and exp_op2, thereby obtaining mant_op1 and mant_op2.

[0152] Finally, the arithmetic and arrangement circuit performs an addition operation or a subtraction operation on mant_op1 and mant_op2, intercepts the operation results of mant_op1 and mant_op2 according to the floating-point number types of the first floating-point number operand1 and / or the second floating-point number operand2, and combines the sign bits sign_op1 and sign_op2, and the exponents exp_op1 and exp_op2 to obtain the final operation result when performing addition and subtraction operations on the first floating-point number and the second floating-point number.

[0153] Obviously, in this setting architecture, it is equivalent to using only one type of hardware circuit to perform addition and subtraction operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Compared with the prior art that requires using three different types of hardware circuits to perform addition and subtraction operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers respectively, this circuit architecture can not only reduce the complexity and power consumption of the circuit, but also significantly improve the resource utilization rate of the circuit.

[0154] Correspondingly, an embodiment of the present invention further provides a floating-point operation unit, including an addition and subtraction operation device for multi-precision floating-point numbers as disclosed above.

[0155] The floating-point operation unit provided by the embodiment of the present invention has the beneficial effects of an addition and subtraction operation device for multi-precision floating-point numbers as disclosed above.

[0156] The above has introduced in detail a device for addition and subtraction operations of multi-precision floating-point numbers and a floating-point operation unit provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A multi-precision floating point addition and subtraction device, characterized in that: include: A separation operation circuit is used to obtain a first floating point number and a second floating point number from a floating point register group of an instruction set architecture processor according to a target instruction, and perform separation operations on the first floating point number and the second floating point number respectively to obtain a first separated number and a second separated number; the target instruction includes: an operation instruction for operating a double-precision floating point number, a single-precision floating point number and a half-precision floating point number; when the target instruction operates the double-precision floating point number, the single-precision floating point number and the half-precision floating point number, the target instruction distinguishes the double-precision floating point number, the single-precision floating point number and the half-precision floating point number through a coding segment for specifying the precision of the floating point number in the operation instruction; an order-matching circuit, used for performing order-matching, mantissa-complementing and shifting operations on the first separated number and the second separated number to obtain a first processed number and a second processed number; The operation arrangement circuit is used to perform addition and subtraction operations on the first processing number and the second processing number, and format the operation results and output them.

2. The multi-precision floating point addition and subtraction operation device according to claim 1, characterized in that: If the target instruction is an operation instruction for operating double-precision floating-point numbers, the coding segment in the target instruction for specifying the precision of the floating-point number is 2'b01; if the target instruction is an operation instruction for operating single-precision floating-point numbers, the coding segment in the target instruction for specifying the precision of the floating-point number is 2'b00; if the target instruction is an operation instruction for operating half-precision floating-point numbers, the coding segment in the target instruction for specifying the precision of the floating-point number is 2'b10 or 2'b11.

3. The multi-precision floating point addition and subtraction operation device according to claim 1, characterized in that: The separation operation circuit comprises: A first separation circuit is used to separate the first floating point number into a sign bit, an exponent and a mantissa to obtain the first separated number; The second separation circuit is used to separate the second floating point number into a sign bit, an exponent and a mantissa to obtain the second separated number; the first separation circuit and the second separation circuit have the same setting structure.

4. The multi-precision floating point addition and subtraction operation device according to claim 3, characterized in that: The first floating-point number is stored in a first general floating-point register of the floating-point register group, and the number of bits of the first general floating-point register is 64 bits; If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number are sequentially stored in bits 0 to 63 of the first general floating-point register; If the first floating-point number is a single-precision floating-point number, data in the first floating-point number are sequentially stored at bits 0 to 31 of the first general floating-point register, and bits 32 to 63 of the first general floating-point register are all filled with zeros; If the first floating point number is a half-precision floating point number, data in the first floating point number are sequentially stored in bits 0 to 15 of the first general floating point register, and bits 16 to 63 of the first general floating point register are all filled with zero.

5. The multi-precision floating point addition and subtraction operation device according to claim 4, characterized in that: The first separation circuit includes: a first multiplexer, a second multiplexer and a third multiplexer; The signal gating port of the first multiplexer, the signal gating port of the second multiplexer and the signal gating port of the third multiplexer are all used to receive the encoding segment for specifying the floating point precision in the target instruction; The first input terminal, the second input terminal and the third input terminal of the first multiplexer are respectively used to receive the data stored on the 63rd bit, the 31st bit and the 15th bit of the first general floating-point register, and the output terminal of the first multiplexer is used to output the sign bit of the first separated number; The first input end of the second multiplexer is used to receive the data stored in the 62nd to 52nd bits of the first general floating-point register, the second input end of the second multiplexer is used to receive the data stored in the 30th to 23rd bits of the first general floating-point register, the third input end of the second multiplexer is used to receive the data stored in the 14th to 10th bits of the first general floating-point register, and the output end of the second multiplexer is used to output the order code of the first separation number; The first input end of the third multiplexer is used to receive data stored at the 51st to 0th bits in the first general floating-point register, the second input end of the third multiplexer is used to receive data stored at the 22nd to 0th bits in the first general floating-point register, the third input end of the third multiplexer is used to receive data stored at the 9th to 0th bits in the first general floating-point register, and the output end of the third multiplexer is used to output the mantissa of the first separated number.

6. The multi-precision floating point addition and subtraction operation device according to claim 1, characterized in that: The pair-order operation circuit comprises: an rank matching unit, used for matching the rank of the first separation number with the rank of the second separation number to obtain a target matching result; A mantissa completion unit, configured to complete the mantissa of the first separated number and the mantissa of the second separated number according to the target order result, to obtain a mantissa completion result of the first separated number and a mantissa completion result of the second separated number; A shift operation unit is used to perform a shift operation on the mantissa completion result of the first separated number and the mantissa completion result of the second separated number to obtain the first processing number and the second processing number.

7. The multi-precision floating point addition and subtraction operation device according to claim 6, characterized in that: When both the first floating point number and the second floating point number are double-precision floating point numbers, the bit width of the result of the exponent alignment performed by the exponent alignment unit on the exponent of the first separated number and the exponent of the second separated number is 12 bits; When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the result of the exponent alignment performed by the exponent alignment unit on the exponent of the first separated number and the exponent of the second separated number is 9 bits; When both the first floating point number and the second floating point number are half-precision floating point numbers, the bit width of the result of the exponent alignment performed by the exponent alignment unit on the exponent of the first separated number and the exponent of the second separated number is 6 bits.

8. The multi-precision floating point addition and subtraction operation device according to claim 6, characterized in that: When both the first floating-point number and the second floating-point number are double-precision floating-point numbers, the bit width of the completion result of the mantissa completion performed by the mantissa completion unit on the mantissa of the first separated number and the mantissa of the second separated number is 53 bits; When both the first floating-point number and the second floating-point number are single-precision floating-point numbers, the bit width of the completion result of the mantissa completion performed by the mantissa completion unit on the mantissa of the first separated number and the mantissa of the second separated number is 24 bits; When both the first floating-point number and the second floating-point number are half-precision floating-point numbers, the bit width of the completion result of the mantissa completion performed by the mantissa completion unit on the mantissa of the first separated number and the mantissa of the second separated number is 11 bits.

9. The multi-precision floating point addition and subtraction operation device according to claim 1, characterized in that: The operation arrangement circuit comprises: an operation subunit, configured to splice the first processing number using a guard bit, a rounding bit, and a sticky bit to obtain first spliced ​​data, splice the second processing number using the guard bit, the rounding bit, and the sticky bit to obtain second spliced ​​data, and perform addition and subtraction operations on the first spliced ​​data and the second spliced ​​data to obtain target operation data; The sorting subunit is used to intercept the target operation data according to the floating point type of the first floating point number and / or the second floating point number, and format and output the interception processing result.

10. A floating point unit, characterized in that: It comprises a multi-precision floating point addition and subtraction operation device as described in any one of claims 1 to 9.