Multi-precision floating-point number multiplication device and floating-point arithmetic unit
By designing a multi-precision floating-point number multiplication operation device, the separation, order code, compression and sorting circuits are used to solve the problems of low resource utilization and high power consumption caused by the independence of floating-point number multiplication circuits of different precision floating-point number in the RISC-V architecture, and a more efficient resource utilization and simplified design process is achieved.
Patent Information
- Application Number
- CN202510167651.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-27
AI Technical Summary
In the RISC-V architecture, multiplication operations of half-precision floating point numbers, single-precision floating point numbers and double-precision floating point numbers require the design of independent circuit units, resulting in low resource utilization, high power consumption and increased design complexity.
A multiplication operation device for multi-precision floating-point numbers is designed, including separation circuits, order code circuits, compression circuits and finishing circuits. Through these circuits, multiplication operations of floating-point numbers of different precisions are realized, and the same type of hardware circuit is used to perform operations.
The circuit complexity and power consumption during floating point multiplication operations of different precisions are reduced, resource utilization is improved, and design and verification process is simplified.
Smart Images

Figure CN120045159A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a multiplication operation device for multi-precision floating-point numbers and a floating-point operation unit. Background Art
[0002] The operation of floating-point numbers is a fundamental and crucial task in a computer, which involves the accurate representation and calculation of real numbers. Among them, floating-point numbers include data formats such as double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Double-precision floating-point numbers are commonly used in scenarios such as scientific computing, high-precision numerical simulation, and financial computing that require high precision and a large range of numerical representation; single-precision floating-point numbers are widely used in fields such as graphics processing, machine learning, and game development that have high requirements for performance and storage; half-precision floating-point numbers are suitable for application scenarios such as neural network inference, mobile devices, and embedded systems that have high requirements for computing speed and storage efficiency but relatively low requirements for precision. In many actual application scenarios, the floating-point operations of these several precisions may need to be performed alternately.
[0003] Currently, in the RISC-V (Reduced Instruction Set Computer - Version 5, instruction set) architecture, the multiplication operations of half-precision floating-point data, single-precision floating-point data, and double-precision floating-point data are usually designed as independent circuit units, and each occupies different hardware resources. Although this design improves the operation speed of floating-point data in specific cases, it also brings problems such as low resource utilization, high power consumption, and increased design and verification complexity. Currently, there is no relatively effective solution to this technical problem. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a multiplication operation device for multi-precision floating-point numbers and a floating-point operation unit, so as to solve the technical problem in the related art that when performing multiplication operations on half-precision floating-point numbers, single-precision floating-point numbers, and double-precision floating-point numbers, it is necessary to design independent circuit structures, and each occupies different hardware resources, resulting in low resource utilization, high power consumption, and high design and verification complexity.
[0005] To solve the above technical problem, the present invention provides a multiplication operation device for multi-precision floating-point numbers, including:
[0006] A separation circuit is used to obtain two floating-point numbers from the floating-point register file of an instruction set architecture processor according to a target instruction, and perform separation operations on the two floating-point numbers respectively to obtain two separated numbers; the target instruction includes: an arithmetic instruction for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers; when the target instruction operates on the double-precision floating-point numbers, the single-precision floating-point numbers, and the half-precision floating-point numbers, it is distinguished by an encoding segment for specifying the floating-point precision in the arithmetic instruction.
[0007] An exponent circuit is used to obtain the sum value of the exponents of the two separated numbers to obtain a target sum value.
[0008] A compression circuit is used to perform compression processing on the product result of the mantissas of the two separated numbers to obtain a mantissa processed number.
[0009] An arrangement circuit is used to output the product of two floating-point numbers according to the sign bits of the two separated numbers, the target sum value, and the mantissa processed number.
[0010] In a specific embodiment of the present application, the exponent circuit is specifically an adder with an 11-bit input.
[0011] In a specific embodiment of the present application, when the two floating-point numbers are a first floating-point number and a second floating-point number respectively, and the two separated numbers are a first separated number and a second separated number respectively, the separation circuit includes:
[0012] A first separation sub-circuit is used to separate the first floating-point number into a sign bit, an exponent, and a mantissa to obtain the first separated number.
[0013] A second separation sub-circuit is used to separate the second floating-point number into a sign bit, an exponent, and a mantissa to obtain the second separated number; the first separation sub-circuit and the second separation sub-circuit have the same setting structure.
[0014] In a specific embodiment of the present application, the first floating-point number is stored in a first general floating-point register of the floating-point register file, and the number of bits of the first general floating-point register is 64 bits;
[0015] If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 63 of the first general floating-point register;
[0016] If the first floating-point number is a single-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 31 of the first general floating-point register, and bits 32 to 63 of the first general floating-point register are all filled with zeros;
[0017] If the first floating-point number is a half-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 15 of the first general-purpose floating-point register, and bits 16 to 63 of the first general-purpose floating-point register are all filled with zeros.
[0018] In a specific embodiment of the present application, the first separation sub-circuit includes: a first multiplexer, a second multiplexer, and a third multiplexer;
[0019] Among them, the signal selection ports of the first multiplexer, the signal selection port of the second multiplexer, and the signal selection port of the third multiplexer are all used to receive the encoding segment for specifying the floating-point precision in the target instruction;
[0020] The first input terminal, the second input terminal, and the third input terminal of the first multiplexer are respectively used to receive the data stored in bit 63, bit 31, and bit 15 of the first general-purpose floating-point register, and the output terminal of the first multiplexer is used to output the sign bit of the first separated number;
[0021] The first input terminal of the second multiplexer is used to receive the data stored in bits 62 to 52 of the first general-purpose floating-point register, the second input terminal of the second multiplexer is used to receive the data stored in bits 30 to 23 of the first general-purpose floating-point register, the third input terminal of the second multiplexer is used to receive the data stored in bits 14 to 10 of the first general-purpose floating-point register, and the output terminal of the second multiplexer is used to output the exponent of the first separated number;
[0022] The first input terminal of the third multiplexer is used to receive the data stored in bits 51 to 0 of the first general-purpose floating-point register, the second input terminal of the third multiplexer is used to receive the data stored in bits 22 to 0 of the first general-purpose floating-point register, the third input terminal of the third multiplexer is used to receive the data stored in bits 9 to 0 of the first general-purpose floating-point register, and the output terminal of the third multiplexer is used to output the mantissa of the first separated number.
[0023] In a specific embodiment of the present application, the compression circuit is specifically a circuit built based on a Wallace tree.
[0024] In a specific embodiment of the present application, when both floating-point numbers are double-precision floating-point numbers, the bit width of the mantissas of the two separated numbers is 53 bits; when both floating-point numbers are single-precision floating-point numbers, starting from the lowest bit of the mantissas of the two separated numbers, the 24th to 53rd bits of the mantissas of the two separated numbers are filled with zeros; when both floating-point numbers are half-precision floating-point numbers, starting from the lowest bit of the mantissas of the two separated numbers, the 11th to 53rd bits of the mantissas of the two separated numbers are filled with zeros.
[0025] In a specific embodiment of the present application, the compression circuit includes: 22 4-2 compressors, 11 3-2 compressors, a first adder, a second adder, a third adder, a first clock gater, a second clock gater, and a third clock gater;
[0026] Among them, the 52 signal input ends corresponding to the first 4-2 compressor, the second 4-2 compressor, the third 4-2 compressor, the fourth 4-2 compressor, the fifth 4-2 compressor, the sixth 4-2 compressor, the seventh 4-2 compressor, the eighth 4-2 compressor, the ninth 4-2 compressor, the tenth 4-2 compressor, the eleventh 4-2 compressor, the twelfth 4-2 compressor, and the thirteenth 4-2 compressor are sequentially used to receive the first partial product to the 52nd partial product; the i-th partial product is: the result corresponding to the multiplication of the i-th number in the mantissa of the second separated number and the mantissa of the first separated number; the i-th number in the mantissa of the second separated number is: starting from the lowest bit in the mantissa of the second separated number, the number corresponding to the i-th position in the mantissa of the second separated number; 1 ≤ i ≤ 53;
[0027] The 27 signal input ends corresponding to the first 3-2 compressor, the second 3-2 compressor, the third 3-2 compressor, the fourth 3-2 compressor, the fifth 3-2 compressor, the sixth 3-2 compressor, the seventh 3-2 compressor, the eighth 3-2 compressor, and the ninth 3-2 compressor are sequentially used to receive the 26 signals output by the first 4-2 compressor, the second 4-2 compressor, the third 4-2 compressor, the fourth 4-2 compressor, the fifth 4-2 compressor, the sixth 4-2 compressor, the seventh 4-2 compressor, the eighth 4-2 compressor, the ninth 4-2 compressor, the tenth 4-2 compressor, the eleventh 4-2 compressor, the twelfth 4-2 compressor, and the thirteenth 4-2 compressor, and the 53rd partial product;
[0028] The 20 signal input terminals corresponding to the fourteenth 4-2 compressor, fifteenth 4-2 compressor, sixteenth 4-2 compressor, seventeenth 4-2 compressor, and eighteenth 4-2 compressor are successively used to receive the 18 signals output by the first 3-2 compressor, the second 3-2 compressor, the third 3-2 compressor, the fourth 3-2 compressor, the fifth 3-2 compressor, the sixth 3-2 compressor, the seventh 3-2 compressor, the eighth 3-2 compressor, and the ninth 3-2 compressor, as well as two zero-level signals;
[0029] The 12 signal input terminals corresponding to the nineteenth 4-2 compressor, twentieth 4-2 compressor, and twenty-first 4-2 compressor are successively used to receive the 10 signals output by the fourteenth 4-2 compressor, fifteenth 4-2 compressor, sixteenth 4-2 compressor, seventeenth 4-2 compressor, and eighteenth 4-2 compressor, as well as two zero-level signals;
[0030] The 6 signal input terminals corresponding to the tenth 3-2 compressor and the eleventh 3-2 compressor are successively used to receive the 6 signals output by the nineteenth 4-2 compressor, twentieth 4-2 compressor, and twenty-first 4-2 compressor;
[0031] The four input terminals of the twenty-second 4-2 compressor are respectively used to receive the two output signals of the tenth 3-2 compressor and the two output signals of the eleventh 3-2 compressor;
[0032] The two input terminals of the first adder are respectively used to receive the two output signals of the fourteenth 4-2 compressor, and the output terminal of the first adder is the first output terminal of the compression circuit; the two input terminals of the second adder are respectively used to receive the two output signals of the nineteenth 4-2 compressor, and the output terminal of the second adder is the second output terminal of the compression circuit; the two input terminals of the third adder are respectively used to receive the two output signals of the twenty-second 4-2 compressor, and the output terminal of the third adder is the third output terminal of the compression circuit;
[0033] The first clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path from the first partial product to the twelfth partial product; the second clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path from the first partial product to the twenty-fourth partial product; the third clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path from the first partial product to the fifty-third partial product.
[0034] In a specific embodiment of the present application, when both floating-point numbers are half-precision floating-point numbers, and the first clock gating controller activates the 4-2 compressors and 3-2 compressors on the operation path from the 1st partial product to the 12th partial product, the output end of the first adder is used to output the mantissa processing number;
[0035] When both floating-point numbers are single-precision floating-point numbers, and the second clock gating controller activates the 4-2 compressors and 3-2 compressors on the operation path from the 1st partial product to the 24th partial product, the output end of the second adder is used to output the mantissa processing number;
[0036] When both floating-point numbers are double-precision floating-point numbers, and the third clock gating controller activates the 4-2 compressors and 3-2 compressors on the operation path from the 1st partial product to the 53rd partial product, the output end of the third adder is used to output the mantissa processing number.
[0037] To solve the above technical problems, the present invention also provides a floating-point operation unit, including a multiplication operation device for multi-precision floating-point numbers as disclosed above.
[0038] Beneficial effects: In the present invention, the encoding segments for specifying the floating-point number precision in the operation instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers in the instruction set architecture processor are distinguished in advance, so that the operation instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers in the instruction set architecture processor have the same structural form.
[0039] When performing a multiplication operation on two floating-point numbers, the separation circuit first obtains the two floating-point numbers from the floating-point register bank of the instruction set architecture processor according to the target instruction, and separately performs a separation operation on the two floating-point numbers to obtain two separated numbers; then, the exponent circuit calculates the sum of the exponents of the two separated numbers to obtain the target sum value, and the compression circuit performs a compression process on the product result of the mantissas of the two separated numbers to obtain the mantissa processing number; finally, the sorting circuit outputs the product of the two floating-point numbers according to the sign bits, the target sum value, and the mantissa processing number of the two separated numbers. In this setup architecture, it is equivalent to using only one type of hardware circuit to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Compared with the current situation where three different types of hardware circuits are required to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers respectively, through this circuit architecture, not only can the circuit complexity and power consumption be reduced when performing multiplication operations on floating-point numbers with different precisions, but also the resource utilization rate of the circuit can be significantly improved.
[0040] Correspondingly, a floating-point operation unit provided by the present invention also has the above beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 Structural diagram of a multiplication operation device for multi-precision floating-point numbers provided by an embodiment of the present invention;
[0043] Figure 2 Structural schematic diagram of a target instruction;
[0044] Figure 3 Schematic diagram when a RISC-V processor processes a target instruction using a four-stage pipeline technology;
[0045] Figure 4 Structural diagram of a first sub-circuit provided by an embodiment of the present invention;
[0046] Figure 5 Schematic diagram when filling the mantissas of the separated numbers corresponding to different-precision floating-point numbers;
[0047] Figure 6 Schematic diagram of the partial product generated when multiplying the mantissas of the separated numbers corresponding to two single-precision floating-point numbers;
[0048] Figure 7 Structural diagram of a compression circuit provided by an embodiment of the present invention;
[0049] Figure 8 Principle schematic diagram when performing multiplication operation on multi-precision floating-point numbers provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0051] As used in the description of the present invention and the above-mentioned drawings, the terms "comprising" and "having", and any variations related to "comprising" and "having", are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may include steps or units not listed.
[0052] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0053] Currently, when performing floating-point operations, the most widely used is IEEE745 (IEEE Binary Floating-Point Arithmetic Standard). Please refer to Table 1, which shows the IEEE754 data types.
[0054] Table 1
[0055] Data type Sign bit Exponent Mantissa Normalized number 0 / 1 Any value Any value Denormalized number 0 / 1 All 0s Not all 0s Zero 0 / 1 All 0s All 0s Infinity 0 / 1 All 1s All 0s Not a number 0 / 1 All 1s Not all 0s
[0056] In different instruction sets, the implementation methods for floating-point operations are different. For example: In the x86 architecture, the SSE (Streaming SIMD Extensions) instruction set is provided for floating-point operations. This instruction set includes basic operation instructions such as addition, subtraction, multiplication, and division of floating-point numbers, as well as auxiliary instructions for converting, comparing, and rounding floating-point numbers. The design of this instruction set aims to improve the efficiency and performance of floating-point operations. In the ARM (Advanced RISC Machine) architecture, the NEON (New Engine for Next Generation Objects) instruction set is provided for floating-point operations. This instruction set includes vectorized operation instructions for floating-point numbers. This instruction set can process multiple floating-point data simultaneously, and its design purpose is to utilize hardware parallelism and vectorized computing to improve the operation parallelism and efficiency of floating-point numbers.
[0057] Although the floating-point operation technology has different implementation methods in different instruction sets, their ultimate goal is to improve the operation efficiency and performance of floating-point numbers, or to control the power consumption of the processor, so as to meet the requirements of different application scenarios.
[0058] The Instruction Set Architecture Processor (RISC-V architecture processor, Reduced Instruction Set Computer - Version 5 processor) originated from a research project at the University of California, Berkeley, aiming to provide an open, modular, and scalable instruction set design. The RISC-V processor (i.e., the Instruction Set Architecture Processor) adopts simple and flexible design principles. The instruction set in this architecture processor includes both a basic integer instruction set and optional extension instruction sets. The F extension instruction set and D extension instruction set in the extension instruction sets are used to handle single-precision floating-point operations and double-precision floating-point operations respectively. The F extension instruction set and D extension instruction set not only cover basic arithmetic operation instructions such as addition, subtraction, multiplication, and division of single-precision and double-precision floating-point numbers, but also cover auxiliary instructions such as conversion, comparison, and rounding of single-precision and double-precision floating-point numbers.
[0059] Since the F extension instructions and D extension instruction set in the RISC-V processor provide a series of instructions for floating-point operations, and its implementation technology involves multiple aspects such as the hardware design of the processor and floating-point operation optimization, the RISC-V processor needs to have hardware support for floating-point operations. For example: A floating-point register file, a floating-point arithmetic unit, and a floating-point status register, etc., need to be set in the RISC-V processor to implement the execution of floating-point operation instructions and the precision control of data.
[0060] To improve the instruction execution efficiency of the RISC-V processor, the RISC-V processor usually adopts pipeline technology to execute floating-point operation instructions in parallel. In terms of hardware implementation, the RISC-V processor also needs to execute efficient floating-point operation algorithms, including not only basic arithmetic operations such as addition, subtraction, multiplication, and division of floating-point numbers, but also auxiliary operations such as conversion, comparison, and rounding of floating-point numbers, and to ensure the accuracy and reliability of the operation results. In addition, the RISC-V processor also needs to design corresponding exception handling mechanisms and status management mechanisms to handle possible exceptions during floating-point operations.
[0061] In the existing technology, when performing floating-point operations in the RISC-V processor, the multiplication operations of half-precision floating-point data, single-precision floating-point data, and double-precision floating-point data are usually designed as independent circuit units. Although this design improves the operation speed in specific cases, it also brings problems such as low resource utilization, high power consumption, and increased design and verification complexity.
[0062] To solve the above technical problems, the present invention provides a multiplication operation device for multi-precision floating-point numbers. With this operation device, only one type of hardware circuit can be used to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. This can not only reduce the circuit complexity and power consumption during the multiplication operation of multi-precision floating-point numbers, but also significantly improve the resource utilization rate of the circuit.
[0063] Please refer to Figure 1 , Figure 1 which is a structural diagram of a multiplication operation device for multi-precision floating-point numbers provided by an embodiment of the present invention. The device includes:
[0064] A separation circuit 11, configured to obtain two floating-point numbers from the floating-point register bank of the instruction set architecture processor according to a target instruction, and perform separation operations on the two floating-point numbers respectively to obtain two separated numbers; the target instruction includes: an operation instruction for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers; when the target instruction operates on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers, it is distinguished by an encoding segment for specifying the floating-point precision in the operation instruction;
[0065] An exponent circuit 12, configured to obtain the sum value of the exponents of the two separated numbers to obtain a target sum value;
[0066] A compression circuit 13, configured to perform compression processing on the product result of the mantissas of the two separated numbers to obtain a mantissa processed number;
[0067] An arrangement circuit 14, configured to output the product of the two floating-point numbers according to the sign bits of the two separated numbers, the target sum value, and the mantissa processed number.
[0068] In this embodiment, in order to enable those skilled in the art to clearly understand the implementation principle of the present invention, first, a specific description is given of the target instruction and the application scenario of the multiplication operation device for multi-precision floating-point numbers of the present invention.
[0069] Among them, the target instruction includes: an operation instruction for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Currently, in the RISC-V processor, the operation instruction for operating on double-precision floating-point numbers is usually called a D extension instruction, and the operation instruction for operating on single-precision floating-point numbers is usually called an F extension instruction. In this application, we can call the operation instruction for operating on half-precision floating-point numbers an H extension instruction. In other words, in the present invention, the target instruction can be either a D extension instruction, an F extension instruction, or an H extension instruction.
[0070] Please refer to Figure 2 , Figure 2It is a structural schematic diagram of the target instruction. In Figure 2 the 0th to 6th bits of the target instruction are the opcode encoding segment of the instruction, which is used to determine the type of the instruction. The encoding of the floating-point operation instruction in this segment is 7’b1010011; the 27th to 31st bits of the target instruction are the funct5 encoding segment of the instruction, which is used to further distinguish the instruction operations. For example, the encoding of the floating-point addition in this segment is 0x00, the encoding of the floating-point subtraction in this segment is 0x01, and the encoding of the floating-point multiplication in this segment is 0x02; the 25th to 26th bits of the target instruction are the fmt encoding segment of the instruction, which is used to specify the precision format of the floating-point number. For single-precision floating-point instructions, this encoding segment is generally set to 2’b00, while for double-precision floating-point instructions, this encoding segment is generally set to 2’b01; the 12th to 14th bits of the target instruction are the rm encoding segment of the instruction, which is used to specify the rounding mode of the floating-point operation. For details, please refer to Table 2, which is the rounding mode of the floating-point operation; the 20th to 24th bits, the 15th to 19th bits, and the 7th to 11th bits of the target instruction represent the register address of the floating-point operand 2, the register address of the floating-point operand 1, and the destination register address, respectively.
[0071] Table 2
[0072] Rounding encoding format Abbreviation Description 000 RNE Round to nearest even 001 RTZ Round to zero 010 RDN Round down (to negative infinity) 011 RUP Round up (to positive infinity) 100 RMM Round to nearest maximum magnitude 101 -- Invalid 110 -- Invalid 111 -- Dynamic rounding mode
[0073] As can be seen from Figure 2 since the fmt encoding segment in the target instruction is used to specify the precision format of the floating-point number, we can distinguish the floating-point operation type corresponding to the target instruction through the fmt encoding segment (which is used to specify the floating-point precision).
[0074] In addition, in the RISC-V processor, a four-stage pipeline technology is usually adopted to process the target instruction. Please refer to Figure 3 , Figure 3 which is a schematic diagram when the RISC-V processor uses the four-stage pipeline technology to process the target instruction. As can be seen from Figure 3 when the RISC-V processor uses the four-stage pipeline technology to process the target instruction, it is usually divided into four stages: instruction fetch, decoding, execution, and write-back.
[0075] Among them, the execution of floating-point operation instructions is achieved by relying on a floating-point unit (FPU, Floating Point Unit). In the instruction fetch stage, the RISC-V processor fetches instructions from memory; in the decoding stage, the RISC-V processor compiles the fetched instructions to obtain information such as the operand register address, destination register address, and instruction type required by the instructions. After decoding is completed, it enters the execution stage. In the execution stage, the RISC-V processor fetches the operands according to the register address and performs the corresponding operations. For floating-point operation instructions, the floating-point numbers and information related to the instruction type obtained after decoding are passed to the floating-point unit pipeline, and the final result is written back to the floating-point register file through the write-back stage.
[0076] In the present invention, in order to reduce the circuit complexity when performing multiplication operations on multi-precision floating-point numbers, a separation circuit 11, an exponent circuit 12, a compression circuit 13, and an arrangement circuit 14 are provided in the floating-point unit. Among them, the separation circuit 11 obtains two floating-point numbers from the floating-point register file of the RISC-V processor according to the target instruction, and respectively performs separation operations on the two floating-point numbers to obtain two separated numbers. Among them, when the separation circuit 11 performs separation operations on the two floating-point numbers, it respectively separates the two floating-point numbers into three parts: the sign bit, the exponent, and the mantissa.
[0077] It should be noted that since the 20th to 24th bits of the target instruction represent the register address of the floating-point operand 1, and the 15th to 19th bits of the target instruction represent the register address of the floating-point operand 2, therefore, according to the data encoding on the 20th to 24th bits of the target instruction and the register addresses corresponding to the data encoding on the 15th to 19th bits, the corresponding two floating-point numbers can be found from the floating-point register file.
[0078] After that, the exponent circuit 12 is used to obtain the sum value of the exponents of the two separated numbers to obtain the target sum value. Since the mantissa of a floating-point number usually has many bits, if the mantissas of the two separated numbers are multiplied, a data with a particularly large number of bits will be obtained. If the result of directly multiplying the mantissas of the two separated numbers is allowed to participate in subsequent operations, it will greatly increase the computational difficulty of the floating-point product operation. Therefore, in this embodiment, after using the exponent circuit 12 to obtain the sum value of the exponents of the two separated numbers, it is also necessary to use the compression circuit 13 to perform compression processing on the mantissa processing results of the two separated numbers to obtain the mantissa processing number, so as to reduce the computational complexity of the floating-point product operation.
[0079] After determining the target sum value and the mantissa processing number, it is equivalent to obtaining the sum of the exponents of the two separated numbers and the product result of the mantissas of the two separated numbers. At this time, only need to use the sorting circuit 14 to splice the sign bits, the target sum value and the mantissa processing number of the two separated numbers, then the product result corresponding to the two floating-point numbers can be determined.
[0080] Obviously, in this setting architecture, it is equivalent to only using one type of hardware circuit to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers. Compared with the current situation where three different types of hardware circuits are required to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers respectively, through this circuit architecture, not only can the circuit complexity and power consumption be reduced when performing multiplication operations on floating-point numbers with different precisions, but also the resource utilization rate of the circuit can be significantly improved.
[0081] As a preferred implementation manner, if the target instruction is an arithmetic instruction for operating on double-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2’b01; if the target instruction is an arithmetic instruction for operating on single-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2’b00; if the target instruction is an arithmetic instruction for operating on half-precision floating-point numbers, the encoding segment for specifying the floating-point precision in the target instruction is 2’b10 or 2’b11.
[0082] According to Figure 2 it can be known that in the target instruction, the encoding segment for specifying the floating-point precision is the fmt encoding segment. Therefore, we can distinguish the arithmetic instructions for double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers according to the encoding data on the fmt encoding segment in the target instruction.
[0083] For the convenience of description, we call the arithmetic instruction for operating on double-precision floating-point numbers the D extension instruction set, the arithmetic instruction for operating on single-precision floating-point numbers the F extension instruction set, and the arithmetic instruction for operating on half-precision floating-point numbers the H extension instruction set.
[0084] In the official document of the RISC-V processor, when the encoding data on the fmt encoding segment are 2’b01 and 2’b00 respectively, they are used to represent the arithmetic instructions for processing double-precision floating-point numbers and single-precision floating-point numbers. Therefore, when the fmt encoding segment in the target instruction is 2’b01, it indicates that the target instruction is an arithmetic instruction for operating on double-precision floating-point numbers, and when the fmt encoding segment in the target instruction is 2’b00, it indicates that the target instruction is an arithmetic instruction for operating on single-precision floating-point numbers.
[0085] In the official documentation of the RISC-V processor, since 2'b10 and 2'b11 on the fmt encoding segment are not defined, we can use 2'b10 or 2'b11 on the fmt encoding segment to define the arithmetic instructions for operating on half-precision floating-point numbers. This can make the arithmetic instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers compatible with each other while still being distinguishable.
[0086] Here, we assume that when the fmt encoding segment is 2'b10, it indicates that the target instruction is an arithmetic instruction for operating on half-precision floating-point numbers. Then, the following are some specific examples of target instructions for different-precision floating-point operations:
[0087] When the target instruction is 00010 01 00010 00001 000 00011 1010011, it means that the target instruction is an instruction for adding double-precision floating-point numbers;
[0088] When the target instruction is 00010 00 00010 00001 000 00011 1010011, it means that the target instruction is an instruction for subtracting single-precision floating-point numbers;
[0089] When the target instruction is 00010 10 00010 00001 000 00011 1010011, it means that the target instruction is an instruction for subtracting half-precision floating-point numbers.
[0090] Obviously, through the technical solution provided in this embodiment, the arithmetic instructions for operating on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers can be accurately distinguished.
[0091] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation manner, when the two floating-point numbers are the first floating-point number and the second floating-point number respectively, and the two separated numbers are the first separated number and the second separated number respectively, the separation circuit includes:
[0092] The first separation sub-circuit is used to separate the first floating-point number into a sign bit, an exponent, and a mantissa to obtain the first separated number;
[0093] The second separation sub-circuit is used to separate the second floating-point number into a sign bit, an exponent, and a mantissa to obtain the second separated number; the setting structures of the first separation sub-circuit and the second separation sub-circuit are the same.
[0094] Since the separation circuit needs to perform separation operations on the first floating-point number and the second floating-point number simultaneously to obtain the first separated number and the second separated number, in this embodiment, two first sub-separation circuits and second sub-separation circuits with the same structure are provided in the separation circuit. The first sub-separation circuit is used to separate the first floating-point number into a sign bit, an exponent, and a mantissa to obtain the first separated number, and the second sub-separation circuit is used to separate the second floating-point number into a sign bit, an exponent, and a mantissa to obtain the second separated number.
[0095] In practical applications, logic devices such as a multiplexer, a signal selection circuit, or a tri-state gate can be used to build the first sub-separation circuit and the second sub-separation circuit, and the built first sub-separation circuit and second sub-separation circuit are respectively used to perform separation operations on the first floating-point number and the second floating-point number, so as to separate the first floating-point number and the second floating-point number into three parts: a sign bit, an exponent, and a mantissa, corresponding to obtaining the first separated number and the second separated number.
[0096] Obviously, through the technical solution provided in this embodiment, the first sub-separation circuit and the second sub-separation circuit can be used to perform separation operations on the first floating-point number and the second floating-point number respectively, and the first separated number and the second separated number can be obtained correspondingly.
[0097] As a preferred implementation manner, the first floating-point number is stored in the first general-purpose floating-point register of the floating-point register bank, and the number of bits of the first general-purpose floating-point register is 64 bits;
[0098] If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 63 of the first general-purpose floating-point register;
[0099] If the first floating-point number is a single-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 31 of the first general-purpose floating-point register, and bits 32 to 63 of the first general-purpose floating-point register are all filled with zeros;
[0100] If the first floating-point number is a half-precision floating-point number, the data in the first floating-point number is sequentially stored in bits 0 to 15 of the first general-purpose floating-point register, and bits 16 to 63 of the first general-purpose floating-point register are all filled with zeros.
[0101] Since the RISC-V architecture stipulates that if the RISC-V processor needs to support single-precision floating-point instructions or double-precision floating-point instructions, a separate set of floating-point register banks must be added inside it. There are 32 general-purpose floating-point registers in this floating-point register bank, labeled f0 to f31.
[0102] If the RISC-V processor only needs to support instructions for operating on double-precision floating-point numbers (D extension instructions), the width of each general-purpose floating-point register in the floating-point register file is 64 bits; if the RISC-V processor only needs to support instructions for operating on single-precision floating-point numbers (F extension instructions), the width of each general-purpose floating-point register in the floating-point register file is 32 bits.
[0103] In order to enable the multiplication operation device according to the present invention to process double-precision floating-point instructions, single-precision floating-point instructions, and half-precision floating-point instructions simultaneously, it is necessary to Figure 3 set all 32 general-purpose floating-point registers in the internal floating-point register file to 64-bit registers.
[0104] If the general-purpose floating-point register for storing the first floating-point number in the floating-point register file is the first general-purpose floating-point register. Then, if the first floating-point number is a double-precision floating-point number, the first floating-point number will be stored in the entire first general-purpose floating-point register, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 63 of the first general-purpose floating-point register; if the first floating-point number is a single-precision floating-point number, the first floating-point number will be stored in the lower 32 bits of the first general-purpose floating-point register, and the upper 32 bits of the first general-purpose floating-point register will be filled with zeros, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 31 of the first general-purpose floating-point register, and bits 32 to 63 of the first general-purpose floating-point register will all be filled with zeros; if the first floating-point number is a half-precision floating-point number, the first floating-point number will be stored in the lower 16 bits of the first general-purpose floating-point register, and the upper 48 bits of the first general-purpose floating-point register will be filled with zeros, that is, the data in the first floating-point number will be sequentially stored in bits 0 to 15 of the first general-purpose floating-point register, and bits 16 to 63 of the first general-purpose floating-point register will all be filled with zeros.
[0105] Since the storage principle of the second floating-point number in the floating-point register file is the same as that of the first floating-point number in the floating-point register file, therefore, the storage method of the second floating-point number in the floating-point register file will not be specifically described here.
[0106] Obviously, through the technical solution provided by this embodiment, the first floating-point number can be accurately and reliably stored in the general-purpose floating-point register of the floating-point register file.
[0107] Please refer to Figure 4 , Figure 4 which is a structural diagram of a first sub-circuit provided by an embodiment of the present invention. As a preferred embodiment, the first sub-circuit includes: a first multiplexer MUX1, a second multiplexer MUX2, and a third multiplexer MUX3;
[0108] Among them, the signal selection ports of the first multiplexer MUX1, the second multiplexer MUX2, and the third multiplexer MUX3 are all used to receive the encoding segment fmt in the target instruction for specifying the floating-point precision;
[0109] The first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are respectively used to receive the data stored in the 63rd bit, the 31st bit, and the 15th bit of the first general-purpose floating-point register. The output terminal of the first multiplexer MUX1 is used to output the sign bit of the first separated number;
[0110] The first input terminal of the second multiplexer MUX2 is used to receive the data stored in the 62nd bit to the 52nd bit of the first general-purpose floating-point register. The second input terminal of the second multiplexer MUX2 is used to receive the data stored in the 30th bit to the 23rd bit of the first general-purpose floating-point register. The third input terminal of the second multiplexer MUX2 is used to receive the data stored in the 14th bit to the 10th bit of the first general-purpose floating-point register. The output terminal of the second multiplexer MUX2 is used to output the exponent of the first separated number;
[0111] The first input terminal of the third multiplexer MUX3 is used to receive the data stored in the 51st bit to the 0th bit of the first general-purpose floating-point register. The second input terminal of the third multiplexer MUX3 is used to receive the data stored in the 22nd bit to the 0th bit of the first general-purpose floating-point register. The third input terminal of the third multiplexer MUX3 is used to receive the data stored in the 9th bit to the 0th bit of the first general-purpose floating-point register. The output terminal of the third multiplexer MUX3 is used to output the mantissa of the first separated number.
[0112] In this embodiment, the structure of the first sub-circuit is specifically described. Since the multiplexer has a simple structure, a small footprint, and a low price, in this embodiment, in order to reduce the structural complexity and design cost of the first sub-circuit, three multiplexers are used to build the first sub-circuit. Among them, the output signal of the first sub-circuit is determined by the signals received by the signal selection ports of the first multiplexer MUX1, the second multiplexer MUX2, and the third multiplexer MUX3.
[0113] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b01, it indicates that the first floating-point number is a double-precision floating-point number. The first sub-circuit needs to separate the first floating-point number in double-precision floating-point format into a sign bit sign, an exponent exp, and a mantissa fract to obtain the first separated number.
[0114] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the conducting, off, and off states respectively. At this time, the first multiplexer MUX1 will receive the data stored in the 63rd bit of the first general-purpose floating-point register and output the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the conducting, off, and off states respectively. At this time, the second multiplexer MUX2 will receive the data stored in the 62nd to 52nd bits of the first general-purpose floating-point register and output the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the conducting, off, and off states respectively. At this time, the third multiplexer MUX3 will receive the data stored in the 51st to 0th bits of the first general-purpose floating-point register and output the mantissa of the first separated number.
[0115] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b00, it indicates that the first floating-point number is a single-precision floating-point number. The first separation sub-circuit needs to separate the first floating-point number in the single-precision floating-point format into a sign bit, an exponent, and a mantissa to obtain the first separated number.
[0116] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the off, conducting, and off states respectively. At this time, the first multiplexer MUX1 will receive the data stored in the 31st bit of the first general-purpose floating-point register and output the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the off, conducting, and off states respectively. At this time, the second multiplexer MUX2 will receive the data stored in the 30th to 23rd bits of the first general-purpose floating-point register and output the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the off, conducting, and off states respectively. At this time, the third multiplexer MUX3 will receive the data stored in the 22nd to 0th bits of the first general-purpose floating-point register and output the mantissa of the first separated number.
[0117] When the encoding segment fmt for specifying the floating-point precision in the target instruction is 2'b10, it indicates that the first floating-point number is a half-precision floating-point number. The first separation sub-circuit needs to separate the first floating-point number in the half-precision floating-point format into a sign bit, an exponent, and a mantissa to obtain the first separated number.
[0118] Specifically, the first input terminal, the second input terminal, and the third input terminal of the first multiplexer MUX1 are in the off state, the off state, and the on state respectively. At this time, the first multiplexer MUX1 will receive the data stored in the 15th bit of the first general-purpose floating-point register and will output the sign bit of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the second multiplexer MUX2 are in the off state, the off state, and the on state respectively. At this time, the second multiplexer MUX2 will receive the data stored in the 14th bit to the 10th bit of the first general-purpose floating-point register and will output the exponent of the first separated number. The first input terminal, the second input terminal, and the third input terminal of the third multiplexer MUX3 are in the off state, the off state, and the on state respectively. At this time, the third multiplexer MUX3 will receive the data stored in the 9th bit to the 0th bit of the first general-purpose floating-point register and will output the mantissa of the first separated number.
[0119] Please refer to Table 3. Table 3 shows the corresponding relationship of the separation operation of the first floating-point number under different floating-point precisions. In Table 3, fmt represents the fmt encoding segment in the target instruction, the sign bit sign represents the output signal of the first multiplexer MUX1, the exponent exp represents the output signal of the second multiplexer MUX2, the fraction fract represents the output signal of the third multiplexer MUX3, and the operand represents the operand.
[0120] Table 3
[0121] fmt is 2’b00 fmt is 2’b01 fmt is 2’b10 Sign bit sign operand
[31] operand
[63] operand
[15] Exponent exp operand[30:23] operand[62:52] operand[14:10] Mantissa fract operand[22:0] operand[51:0] operand[9:0]
[0122] It should be noted that considering compatibility issues, the separation results of the first floating-point number are all saved using general-purpose floating-point registers with the corresponding bit widths of double-precision floating-point numbers. That is, when the first floating-point number is a double-precision floating-point number, except that the sign bit in the first separated number is still saved using a 1-bit general-purpose floating-point register sign, the exponent of the first separated number will be saved using an 11-bit general-purpose floating-point register exp[10:0], and the mantissa of the first separated number will be saved using a 52-bit general-purpose floating-point register fract[51:0]. When the first floating-point number is a single-precision floating-point number and a half-precision floating-point number, the separation results corresponding to the first floating-point number will have 0s filled in the highest bits of their corresponding general-purpose floating-point registers.
[0123] Obviously, through the technical solution provided by this embodiment, the first floating-point number can be separated using a first separation sub-circuit with a simple structure and low cost.
[0124] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation manner, the exponent circuit is specifically an adder with an 11-bit input.
[0125] As can be seen from Table 3, when both floating-point numbers are double-precision floating-point numbers, the exponent bit width of the two separated numbers is 11 bits; when both floating-point numbers are single-precision floating-point numbers, the exponent bit width of the two separated numbers is 8 bits; when both floating-point numbers are half-precision floating-point numbers, the exponent bit width of the two separated numbers is 5 bits. Since the essence of the exponent circuit is to sum the exponents of the two separated numbers, in practical applications, the exponent circuit can be set as an adder, and the adder is used to sum the exponents of the two separated numbers. Moreover, in this embodiment, because the highest exponent bit width of the two separated numbers is 11 bits, the input of the adder needs to be set to 11 bits. In this way, when the two floating-point numbers are single-precision floating-point numbers or half-precision floating-point numbers, the adder with an 11-bit input can also be reused to sum the exponents of the separated numbers corresponding to the single-precision floating-point numbers and the half-precision floating-point numbers.
[0126] It should be noted that when both floating-point numbers are single-precision floating-point numbers, the high three bits of the exponents of the two separated numbers need to be filled with 0 so that the exponents of the separated numbers corresponding to the single-precision floating-point numbers can reuse the adder with an 11-bit input for exponent summation operations. Similarly, when both floating-point numbers are half-precision floating-point numbers, the high six bits of the exponents of the two separated numbers need to be filled with 0 so that the exponents of the separated numbers corresponding to the half-precision floating-point numbers can reuse the adder with an 11-bit input for exponent summation operations.
[0127] Obviously, through the technical solution provided in this embodiment, the same adder can be reused to sum the exponents of the separated numbers corresponding to floating-point numbers with different precisions.
[0128] Based on the above embodiment, this embodiment further illustrates and optimizes the technical solution. As a preferred implementation manner, when both floating-point numbers are double-precision floating-point numbers, the mantissa bit width of the two separated numbers is 53 bits; when both floating-point numbers are single-precision floating-point numbers, starting from the lowest bit of the mantissas of the two separated numbers, the 24th to 53rd bits of the mantissas of the two separated numbers are filled with zeros; when both floating-point numbers are half-precision floating-point numbers, starting from the lowest bit of the mantissas of the two separated numbers, the 11th to 53rd bits of the mantissas of the two separated numbers are filled with zeros.
[0129] In practical applications, when calculating the sum of the exponents of the two separated numbers, it is also necessary to generate the product of the mantissa parts of the two separated numbers. As can be seen from Table 3, when both floating-point numbers are double-precision floating-point numbers, the mantissa bit width of these two floating-point numbers is 52 bits, and adding the implicit bit on the mantissa will become 53 bits. Therefore, when both floating-point numbers are double-precision floating-point numbers, the mantissa bit width of the two separated numbers is 53 bits. Obviously, when the mantissas of the two corresponding separated numbers are multiplied when both floating-point numbers are double-precision floating-point numbers, 53 53-bit partial products will be generated according to the bit-by-bit multiplication.
[0130] In order to enable the multiplication operation device described in this application to perform product operations on double - precision floating - point numbers, single - precision floating - point numbers, and half - precision floating - point numbers simultaneously, it is necessary to unify the formats of the mantissas in the separated numbers corresponding to various types of floating - point numbers.
[0131] According to Table 3, when both floating - point numbers are single - precision floating - point numbers, the mantissa bit widths in the separated numbers corresponding to these two floating - point numbers are 24 bits (23 - bit mantissa plus 1 - bit implicit bit). At this time, in order to unify the mantissa format of the separated number corresponding to the single - precision floating - point number with the mantissa format of the separated number corresponding to the double - precision floating - point number, it is necessary to fill the 24th to 53rd bits of the mantissas of the two separated numbers with zeros starting from the lowest bits of the mantissas of the two separated numbers.
[0132] Similarly, according to Table 3, when both floating - point numbers are half - precision floating - point numbers, the mantissa bit widths in the separated numbers corresponding to these two floating - point numbers are 10 bits (9 - bit mantissa plus 1 - bit implicit bit). At this time, in order to unify the mantissa format of the separated number corresponding to the half - precision floating - point number with the mantissa format of the separated number corresponding to the double - precision floating - point number, it is necessary to fill the 11th to 53rd bits of the mantissas of the two separated numbers with zeros starting from the lowest bits of the mantissas of the two separated numbers. For details, please refer to Figure 5 , Figure 5 the schematic diagram when filling the mantissas of the separated numbers corresponding to floating - point numbers with different precisions.
[0133] Obviously, through the technical solution provided in this embodiment, the mantissa formats in the separated numbers corresponding to floating - point numbers with different precisions can be unified, which is convenient for the operation and processing of subsequent processes.
[0134] Based on the above - mentioned embodiment, this embodiment further explains and optimizes the technical solution. As a preferred implementation manner, the compression circuit is specifically a circuit built based on the Wallace tree.
[0135] In this embodiment, the compression circuit is built based on the Wallace tree. That is, the compression circuit compresses the product result of the two separated numbers based on the working principle of the Wallace tree to obtain the mantissa processing number.
[0136] Since the Wallace tree can, compared with other types of multiplication operation circuits, parallel - process the gradual compression process of multiple partial products through a tree - shaped structure, this can reduce the length of the critical path and significantly reduce the time required for performing product operations on two floating - point numbers. Therefore, in this embodiment, the compression circuit is built according to the Wallace tree.
[0137] Obviously, the technical solution provided by this embodiment can significantly reduce the efficiency and speed of the compression circuit when compressing the product result of the mantissas of two separated numbers.
[0138] Before describing the structure of the compression circuit, the result of multiplying the mantissas of two separated numbers corresponding to two floating-point numbers will be briefly described.
[0139] For double-precision floating-point numbers, the bit width of the mantissa of the two corresponding separated numbers is 53 bits (52-bit mantissa plus 1 implicit bit), so when the mantissas of the two separated numbers are multiplied, 53 53-bit partial products will be generated. For single-precision floating-point numbers, the bit width of the mantissa of the two corresponding separated numbers is 24 bits (23-bit mantissa plus 1 implicit bit), so when the mantissas of the two separated numbers are multiplied, 24 24-bit partial products will be generated. For half-precision floating-point numbers, the bit width of the mantissa of the two corresponding separated numbers is 11 bits (10-bit mantissa plus 1 implicit bit), so when the mantissas of the two separated numbers are multiplied, 11 11-bit partial products will be generated.
[0140] See also Figure 6 , Figure 6 The following is a diagram of the partial product generated when the mantissas of the corresponding separated numbers of two single-precision floating-point numbers are multiplied. When two single-precision floating-point numbers are multiplied, the partial product of the mantissas of their corresponding separated numbers will appear in a step-like shape. Assuming that the two single-precision floating-point numbers are A1 and A2, and their corresponding separated numbers are B1 and B2, then Figure 6 Among them, pp00, pp01, pp02, ... pp23 are the corresponding products when the first, second, ... 24th numbers starting from the lowest bit in the mantissa of B2 are multiplied by the mantissa of B1 respectively.
[0141] From this we can conclude that the i-th partial product is: the result corresponding to the multiplication of the i-th number in the mantissa of the second separate number and the mantissa of the first separate number; the i-th number in the mantissa of the second separate number is: the number corresponding to the i-th position in the mantissa of the second separate number with the lowest bit in the mantissa of the second separate number as the starting point; 1≤i≤53.
[0142] See also Figure 7 , Figure 7 A structural diagram of a compression circuit provided by an embodiment of the present invention. As a preferred implementation, the compression circuit includes: 22 4-2 compressors, 11 3-2 compressors, a first adder ADD1, a second adder ADD2, a third adder ADD3, a first clock gate controller, a second clock gate controller, and a third clock gate controller;
[0143] Among them, the 52 signal input terminals corresponding to the first 4:2 compressor 4:2 CSA1, the second 4:2 compressor 4:2 CSA2, the third 4:2 compressor 4:2 CSA3, the fourth 4:2 compressor 4:2 CSA4, the fifth 4:2 compressor 4:2 CSA5, the sixth 4:2 compressor 4:2 CSA6, the seventh 4:2 compressor 4:2 CSA7, the eighth 4:2 compressor 4:2 CSA8, the ninth 4:2 compressor 4:2 CSA9, the tenth 4:2 compressor 4:2 CSA10, the eleventh 4:2 compressor 4:2 CSA11, the twelfth 4:2 compressor 4:2 CSA12, and the thirteenth 4:2 compressor 4:2 CSA13 are successively used to receive the 1st partial product pp00 to the 52nd partial product pp51; the i-th partial product is: the result corresponding to the multiplication of the i-th number in the mantissa of the second separated number by the mantissa of the first separated number; the i-th number in the mantissa of the second separated number is: the number corresponding to the i-th position in the mantissa of the second separated number starting from the lowest bit in the mantissa of the second separated number; 1 ≤ i ≤ 53;
[0144] The 27 signal input terminals corresponding to the first 3:2 compressor 3:2 CSA1, the second 3:2 compressor 3:2 CSA2, the third 3:2 compressor 3:2 CSA3, the fourth 3:2 compressor 3:2 CSA4, the fifth 3:2 compressor 3:2 CSA5, the sixth 3:2 compressor 3:2 CSA6, the seventh 3:2 compressor 3:2 CSA7, the eighth 3:2 compressor 3:2 CSA8, and the ninth 3:2 compressor 3:2 CSA9 are successively used to receive the 26 signals output by the first 4:2 compressor 4:2 CSA1, the second 4:2 compressor 4:2 CSA2, the third 4:2 compressor 4:2 CSA3, the fourth 4:2 compressor 4:2 CSA4, the fifth 4:2 compressor 4:2 CSA5, the sixth 4:2 compressor 4:2 CSA6, the seventh 4:2 compressor 4:2 CSA7, the eighth 4:2 compressor 4:2 CSA8, the ninth 4:2 compressor 4:2 CSA9, the tenth 4:2 compressor 4:2 CSA10, the eleventh 4:2 compressor 4:2 CSA11, the twelfth 4:2 compressor 4:2 CSA12, and the thirteenth 4:2 compressor 4:2 CSA13, and the 53rd partial product pp52;
[0145] The 20 signal input terminals corresponding to the fourteenth 4:2 compressor 4:2 CSA14, fifteenth 4:2 compressor 4:2 CSA15, sixteenth 4:2 compressor 4:2 CSA16, seventeenth 4:2 compressor 4:2 CSA17, and eighteenth 4:2 compressor 4:2 CSA18 are successively used to receive the 18 signals output by the first 3:2 compressor 3:2 CSA1, second 3:2 compressor 3:2 CSA2, third 3:2 compressor 3:2 CSA3, fourth 3:2 compressor 3:2 CSA4, fifth 3:2 compressor 3:2 CSA5, sixth 3:2 compressor 3:2 CSA6, seventh 3:2 compressor 3:2 CSA7, eighth 3:2 compressor 3:2 CSA8, and ninth 3:2 compressor 3:2 CSA9, as well as two zero-level signals;
[0146] The 12 signal input terminals corresponding to the nineteenth 4:2 compressor 4:2 CSA19, twentieth 4:2 compressor 4:2 CSA20, and twenty-first 4:2 compressor 4:2 CSA21 are successively used to receive the 10 signals output by the fourteenth 4:2 compressor 4:2 CSA14, fifteenth 4:2 compressor 4:2 CSA15, sixteenth 4:2 compressor 4:2 CSA16, seventeenth 4:2 compressor 4:2 CSA17, and eighteenth 4:2 compressor 4:2 CSA18, as well as two zero-level signals;
[0147] The 6 signal input terminals corresponding to the tenth 3:2 compressor 3:2 CSA10 and eleventh 3:2 compressor 3:2 CSA11 are successively used to receive the 6 signals output by the nineteenth 4:2 compressor 4:2 CSA19, twentieth 4:2 compressor 4:2 CSA20, and twenty-first 4:2 compressor 4:2 CSA21;
[0148] The four input terminals of the twenty-second 4:2 compressor 4:2 CSA22 are respectively used to receive the two output signals of the tenth 3:2 compressor 3:2 CSA10 and the two output signals of the eleventh 3:2 compressor 3:2 CSA11;
[0149] The two input terminals of the first adder ADD1 are respectively used to receive the two output signals of the fourteenth 4:2 compressor 4:2 CSA14, and the output terminal of the first adder ADD1 is the first output terminal of the compression circuit; the two input terminals of the second adder ADD2 are respectively used to receive the two output signals of the nineteenth 4:2 compressor 4:2 CSA19, and the output terminal of the second adder ADD2 is the second output terminal of the compression circuit; the two input terminals of the third adder ADD3 are respectively used to receive the two output signals of the twenty-second 4:2 compressor 4:2 CSA22, and the output terminal of the third adder ADD3 is the third output terminal of the compression circuit;
[0150] The first clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation paths where the first partial product pp00 to the twelfth partial product pp11 are located; the second clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation paths where the first partial product pp00 to the twenty-fourth partial product pp23 are located; the third clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation paths where the first partial product pp00 to the fifty-third partial product pp52 are located.
[0151] In this embodiment, the structure of the compression circuit is specifically described. In Figure 7 pp00, pp01, pp02... pp52 are the partial products corresponding to the multiplication of the mantissas of two separate numbers respectively, where pp00 is the first partial product, pp01 is the second partial product, pp02 is the third partial product,... pp52 is the fifty-third partial product.
[0152] In Figure 7 In the compression circuit described above, the Wallace tree has a total of six levels. Among them, the first-level circuit of the Wallace tree is:
[0153] The four input terminals of the first 4-2 compressor 4:2 CSA1 are respectively used to receive the 1st partial product pp00 to the 4th partial product pp03; the four input terminals of the second 4-2 compressor 4:2 CSA2 are respectively used to receive the 5th partial product pp04 to the 8th partial product pp07; the four input terminals of the third 4-2 compressor 4:2 CSA3 are respectively used to receive the 9th partial product pp08 to the 12th partial product pp11; the four input terminals of the fourth 4-2 compressor 4:2 CSA4 are respectively used to receive the 13th partial product pp12 to the 16th partial product pp15; the four input terminals of the fifth 4-2 compressor 4:2 CSA5 are respectively used to receive the 17th partial product pp16 to the 20th partial product pp19; the four input terminals of the sixth 4-2 compressor 4:2 CSA6 are respectively used to receive the 21st partial product pp20 to the 24th partial product pp23; the four input terminals of the seventh 4-2 compressor 4:2 CSA7 are respectively used to receive the 25th partial product pp24 to the 28th partial product pp27; the four input terminals of the eighth 4-2 compressor 4:2 CSA8 are respectively used to receive the 29th partial product pp28 to the 32nd partial product pp31; the four input terminals of the ninth 4-2 compressor 4:2 CSA9 are respectively used to receive the 33rd partial product pp32 to the 36th partial product pp35; the four input terminals of the tenth 4-2 compressor 4:2 CSA10 are respectively used to receive the 37th partial product pp36 to the 40th partial product pp39; the four input terminals of the eleventh 4-2 compressor 4:2 CSA11 are respectively used to receive the 41st partial product pp40 to the 44th partial product pp43; the four input terminals of the twelfth 4-2 compressor 4:2 CSA12 are respectively used to receive the 45th partial product pp44 to the 48th partial product pp47; the four input terminals of the thirteenth 4-2 compressor 4:2 CSA13 are respectively used to receive the 49th partial product pp48 to the 52nd partial product pp51;
[0154] The second-stage circuit of the Wallace tree is:
[0155] The three input terminals of the first 3-2 compressor 3:2 CSA1 are respectively used to receive two output signals of the first 4-2 compressor 4:2 CSA1 and the first output signal of the second 4-2 compressor 4:2 CSA2; the three input terminals of the second 3-2 compressor 3:2 CSA2 are respectively used to receive the second output signal of the second 4-2 compressor 4:2 CSA2 and two output signals of the third 4-2 compressor 4:2 CSA3; the three input terminals of the third 3-2 compressor 3:2 CSA3 are respectively used to receive two output signals of the fourth 4-2 compressor 4:2 CSA4 and the first output signal of the fifth 4-2 compressor 4:2 CSA5; the three input terminals of the fourth 3-2 compressor 3:2 CSA4 are respectively used to receive the second output signal of the fifth 4-2 compressor 4:2 CSA5 and two output signals of the sixth 4-2 compressor 4:2 CSA6; the three input terminals of the fifth 3-2 compressor 3:2 CSA5 are respectively used to receive two output signals of the seventh 4-2 compressor 4:2 CSA7 and the first output signal of the eighth 4-2 compressor 4:2 CSA8; the three input terminals of the sixth 3-2 compressor 3:2 CSA6 are respectively used to receive the second output signal of the eighth 4-2 compressor 4:2 CSA8 and two output signals of the ninth 4-2 compressor 4:2 CSA9; the three input terminals of the seventh 3-2 compressor 3:2 CSA7 are respectively used to receive two output signals of the tenth 4-2 compressor 4:2 CSA10 and the first output signal of the eleventh 4-2 compressor 4:2 CSA11; the three input terminals of the eighth 3-2 compressor 3:2 CSA8 are respectively used to receive the second output signal of the eleventh 4-2 compressor 4:2 CSA11 and two output signals of the twelfth 4-2 compressor 4:2 CSA12; the three input terminals of the ninth 3-2 compressor 3:2 CSA9 are respectively used to receive two output signals of the thirteenth 4-2 compressor 4:2 CSA13 and the 53rd partial product pp52;
[0156] The third stage circuit of the Wallace tree is:
[0157] The four input terminals of the fourteenth 4-2 compressor 4:2 CSA14 are respectively used to receive two output signals of the first 3-2 compressor 3:2 CSA1 and two output signals of the second 3-2 compressor 3:2 CSA2; the four input terminals of the fifteenth 4-2 compressor 4:2 CSA15 are respectively used to receive two output signals of the third 3-2 compressor 3:2 CSA3 and two output signals of the fourth 3-2 compressor 3:2 CSA4; the four input terminals of the sixteenth 4-2 compressor 4:2 CSA16 are respectively used to receive two output signals of the fifth 3-2 compressor 3:2 CSA5 and two output signals of the sixth 3-2 compressor 3:2 CSA6; the four input terminals of the seventeenth 4-2 compressor 4:2 CSA17 are respectively used to receive two output signals of the seventh 3-2 compressor 3:2 CSA7 and two output signals of the eighth 3-2 compressor 3:2 CSA8; the four input terminals of the eighteenth 4-2 compressor 4:2 CSA18 are respectively used to receive two output signals of the ninth 3-2 compressor 3:2 CSA9 and two zero levels.
[0158] The fourth-stage circuit of the Wallace tree is:
[0159] The four input terminals of the nineteenth 4-2 compressor 4:2 CSA19 are respectively used to receive two output signals of the fourteenth 4-2 compressor 4:2 CSA14 and two output signals of the fifteenth 4-2 compressor 4:2 CSA15; the four input terminals of the twentieth 4-2 compressor 4:2 CSA20 are respectively used to receive two output signals of the sixteenth 4-2 compressor 4:2 CSA16 and two output signals of the seventeenth 4-2 compressor 4:2 CSA17; the four input terminals of the twenty-first 4-2 compressor 4:2 CSA21 are respectively used to receive two output signals of the eighteenth 4-2 compressor 4:2 CSA18 and two zero levels.
[0160] The fifth-stage circuit of the Wallace tree is:
[0161] The three input terminals of the tenth 3-2 compressor 3:2 CSA10 are respectively used to receive two output signals of the nineteenth 4-2 compressor 4:2 CSA19 and the first output signal of the twentieth 4-2 compressor 4:2 CSA20, and the three input terminals of the eleventh 3-2 compressor 3:2 CSA11 are respectively used to receive the second output signal of the twentieth 4-2 compressor 4:2 CSA20 and two output signals of the twenty-first 4-2 compressor 4:2 CSA21;
[0162] The sixth-stage circuit of the Wallace tree is:
[0163] The four input terminals of the twenty-second 4-2 compressor 4:2 CSA22 are respectively used to receive the two output signals of the tenth 3-2 compressor 3:2 CSA10 and the two output signals of the eleventh 3-2 compressor 3:2 CSA11.
[0164] In this compression circuit, the Wallace tree has a total of six levels. In the first level of the Wallace tree, 9 4-2 compressors are used to process the first partial product pp00 to the 52nd partial product pp51, and the 53rd partial product pp52 enters the second level of the Wallace tree; in the second level of the Wallace tree, 9 3-2 compressors are used to reduce the number of partial products to 18; in the third level of the Wallace tree, 5 4-2 compressors are used to reduce the number of partial products to 10. Among them, the ninth 3-2 compressor 3:2 CSA9 has two input signals at zero level, which is set for balancing the delay; in the fourth level of the Wallace tree, 3 4-2 compressors are used to reduce the number of partial products to 6. The eighteenth 4-2 compressor 4:2 CSA18 has two input signals at zero level, which is also set for balancing the delay; in the fifth level of the Wallace tree, 2 3-2 compressors are used to reduce the number of partial products to 4; in the sixth level of the Wallace tree, 1 4-2 compressor is used to finally reduce the number of partial products to 2.
[0165] In addition, in this embodiment, the 4-2 compressors and 3-2 compressors in the compression circuit are also grouped, and the clock signal clk_h generated by the first clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path where the first partial product pp00 to the 12th partial product pp11 are located. The clock signal clk_s generated by the second clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path where the first partial product pp00 to the 24th partial product pp23 are located. At the same time, the clock signal clk_d generated by the third clock gater is used to control the 4-2 compressors and 3-2 compressors on the operation path where the first partial product pp00 to the 53rd partial product pp52 are located.
[0166] It should be noted that in practical applications, other types of compressors can also be used to build the compression circuit. However, the higher the number of input bits of the compressor, the higher the corresponding setting cost. Therefore, in this embodiment, considering the actual application scenario and design cost, 4-2 compressors and 3-2 compressors are selected to build the compression circuit.
[0167] Obviously, through the technical solution provided by this embodiment, the cost investment required for building the compression circuit can be relatively reduced.
[0168] As a preferred embodiment, when both floating-point numbers are half-precision floating-point numbers, and the first clock gater activates the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 12th partial product pp11 are located, the output terminal of the first adder ADD1 is used to output the mantissa processing number;
[0169] When both floating-point numbers are single-precision floating-point numbers, and the second clock gater activates the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 24th partial product pp23 are located, the output terminal of the second adder ADD2 is used to output the mantissa processing number;
[0170] When both floating-point numbers are double-precision floating-point numbers, and the third clock gater activates the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 53rd partial product pp52 are located, the output terminal of the third adder ADD3 is used to output the mantissa processing number.
[0171] In this embodiment, when both floating-point numbers are half-precision floating-point numbers, it is only necessary to use the clock signal clk_h generated by the first clock gater to activate the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 12th partial product pp11 are located. At this time, the output signal of the fourteenth 4-2 compressor 4:2 CSA14 in the third-level circuit of the Wallace tree will be connected to the input terminal of the first adder ADD1, and the first adder ADD1 will output the product calculation result of the mantissas of the separated numbers corresponding to the two half-precision floating-point numbers.
[0172] When both floating-point numbers are single-precision floating-point numbers, it is only necessary to use the clock signal clk_s generated by the second clock gater to activate the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 24th partial product pp23 are located. At this time, the output signal of the nineteenth 4-2 compressor 4:2 CSA19 in the fourth-level circuit of the Wallace tree will be connected to the input terminal of the second adder ADD2, and the second adder ADD2 will output the product calculation result of the mantissas of the separated numbers corresponding to the two single-precision floating-point numbers.
[0173] When both floating-point numbers are double-precision floating-point numbers, it is only necessary to use the clock signal clk_d generated by the third clock gater to activate the 4-2 compressors and 3-2 compressors on the operation path where the 1st partial product pp00 to the 53rd partial product pp52 are located. At this time, the output signal of the twenty-second 4-2 compressor 4:2 CSA22 in the sixth-level circuit of the Wallace tree will be connected to the input terminal of the third adder ADD3, and the third adder ADD3 will output the product calculation result of the mantissas in the separated numbers corresponding to the two double-precision floating-point numbers.
[0174] In addition, in this embodiment, by grouping different compressors in the compression circuit and triggering the compressors in different groups to work through a clock signal, the mantissa of the separated numbers corresponding to different-precision floating-point numbers can be compressed. Moreover, by grouping different compressors in the compression circuit, the compressors activated by the clock signal are in a normal working state, and the compressors not activated by the clock signal remain stationary without flipping, thereby further reducing the power consumption required by the compression circuit.
[0175] Obviously, through the technical solution provided by this embodiment, the same compression circuit can be reused to compress the product results of the mantissas of the separated numbers corresponding to different-precision floating-point numbers.
[0176] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the principle when performing multiplication operations on multi-precision floating-point numbers according to an embodiment of the present invention. When performing multiplication operations on two floating-point numbers, first, two sub-circuits in the multiplication operation device for multi-precision floating-point numbers ( Figure 4 the circuit shown) are used to separate the two floating-point numbers to obtain two separated numbers; among them, the two floating-point numbers can be double-precision floating-point numbers, single-precision floating-point numbers, or half-precision floating-point numbers; then, the exponent circuits are used to add the exponents of the two separated numbers to obtain a target sum value and generate the partial product of the mantissa of the two separated numbers; after that, the clock signal is used to activate Figure 7 the compressors on the paths where different partial products are located in the compression circuit shown to separately compress the product results of the mantissas of the separated numbers corresponding to double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers to obtain the mantissa processed numbers; specifically, when both floating-point numbers are double-precision floating-point numbers, the 4-2 compressors and 3-2 compressors on the operation paths from the first partial product pp00 to the 53rd partial product pp52 in Figure 7 are all activated by the clock signal clk_d, and at this time, the third adder ADD3 can output the mantissa processed number corresponding to the double-precision floating-point number; when both floating-point numbers are single-precision floating-point numbers, the 4-2 compressors and 3-2 compressors on the operation paths from the first partial product pp00 to the 24th partial product pp23 in Figure 7 are activated by the clock signal clk_s, and at this time, the second adder ADD2 can output the mantissa processed number corresponding to the single-precision floating-point number; when both floating-point numbers are half-precision floating-point numbers, the clock signal clk_h is used to Figure 7The 4-2 compressors and 3-2 compressors on the operation path where the first partial product pp00 to the 12th partial product pp11 are located in [the context] are activated. At this time, the first adder ADD1 can output the mantissa processing number corresponding to the half-precision floating-point number. Finally, according to the sign bits of the two separated numbers, the target sum value, and the mantissa processing number, the product calculation result of the two floating-point numbers can be output.
[0177] Obviously, through the multiplication operation device for multi-precision floating-point numbers provided by this application, the same circuit structure can be used to perform product operations on floating-point numbers of different precisions, avoiding the need to use three different types of hardware circuits to perform multiplication operations on double-precision floating-point numbers, single-precision floating-point numbers, and half-precision floating-point numbers respectively. Through this circuit structure, not only can the circuit structure complexity and power consumption be reduced when performing multiplication operations on floating-point numbers of different precisions, but also the resource utilization rate of the circuit can be significantly improved.
[0178] In addition, through this circuit structure, the occupied space volume and the required design cost when performing multiplication operations on multi-precision floating-point numbers can also be reduced. At the same time, the development cycle of the circuit chip can be shortened.
[0179] Correspondingly, the embodiment of the present invention also provides a floating-point operation unit, including a multiplication operation device for multi-precision floating-point numbers as disclosed above.
[0180] The floating-point operation unit provided by the embodiment of the present invention has the beneficial effects of a multiplication operation device for multi-precision floating-point numbers as disclosed above.
[0181] The above has introduced in detail a multiplication operation device for multi-precision floating-point numbers and a floating-point operation unit provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.
Claims
1. A multi-precision floating point multiplication device, characterized in that: include: A separation circuit is used to obtain two floating-point numbers from a floating-point register group of an instruction set architecture processor according to a target instruction, and to perform separation operations on the two floating-point numbers respectively to obtain two separated numbers; the target instruction includes: an operation instruction for operating and processing a double-precision floating-point number, a single-precision floating-point number, and a half-precision floating-point number; when the target instruction operates and processes the double-precision floating-point number, the single-precision floating-point number, and the half-precision floating-point number, the target instruction distinguishes the double-precision floating-point number, the single-precision floating-point number, and the half-precision floating-point number through a coding segment in the operation instruction for specifying the precision of the floating-point number; The order code circuit is used to obtain the sum of the orders of two separated numbers to obtain the target sum value; A compression circuit is used for compressing the product result of the mantissas of two separated numbers to obtain a mantissa processed number; The sorting circuit is used to output the product of two floating point numbers according to the sign bits of the two separated numbers, the target sum value and the mantissa processing number.
2. A multi-precision floating-point number multiplication device according to claim 1, characterized in that: The order code circuit is specifically an adder with an 11-bit input.
3. The multiplication device of multiple precision floating point numbers according to claim 1, characterized in that: When the two floating-point numbers are respectively the first floating-point number and the second floating-point number, and the two separated numbers are respectively the first separated number and the second separated number, the separation circuit comprises: A first separation subcircuit is used to separate the first floating point number into a sign bit, an exponent and a mantissa to obtain the first separated number; The second separation subcircuit is used to separate the second floating point number into a sign bit, an exponent and a mantissa to obtain the second separated number; the first separation subcircuit and the second separation subcircuit have the same setting structure.
4. A multi-precision floating point multiplication device according to claim 3, characterized in that: The first floating-point number is stored in a first general floating-point register of the floating-point register group, and the number of bits of the first general floating-point register is 64 bits; If the first floating-point number is a double-precision floating-point number, the data in the first floating-point number are sequentially stored in bits 0 to 63 of the first general floating-point register; If the first floating-point number is a single-precision floating-point number, data in the first floating-point number are sequentially stored at bits 0 to 31 of the first general floating-point register, and bits 32 to 63 of the first general floating-point register are all filled with zeros; If the first floating point number is a half-precision floating point number, data in the first floating point number are sequentially stored in bits 0 to 15 of the first general floating point register, and bits 16 to 63 of the first general floating point register are all filled with zero.
5. A multi-precision floating-point multiplication device according to claim 4, characterized in that: The first separation subcircuit includes: a first multiplexer, a second multiplexer and a third multiplexer; The signal gating port of the first multiplexer, the signal gating port of the second multiplexer and the signal gating port of the third multiplexer are all used to receive the encoding segment for specifying the floating point precision in the target instruction; The first input terminal, the second input terminal and the third input terminal of the first multiplexer are respectively used to receive the data stored on the 63rd bit, the 31st bit and the 15th bit of the first general floating-point register, and the output terminal of the first multiplexer is used to output the sign bit of the first separated number; The first input end of the second multiplexer is used to receive the data stored in the 62nd to 52nd bits of the first general floating-point register, the second input end of the second multiplexer is used to receive the data stored in the 30th to 23rd bits of the first general floating-point register, the third input end of the second multiplexer is used to receive the data stored in the 14th to 10th bits of the first general floating-point register, and the output end of the second multiplexer is used to output the order code of the first separation number; The first input end of the third multiplexer is used to receive data stored at the 51st to 0th bits in the first general floating-point register, the second input end of the third multiplexer is used to receive data stored at the 22nd to 0th bits in the first general floating-point register, the third input end of the third multiplexer is used to receive data stored at the 9th to 0th bits in the first general floating-point register, and the output end of the third multiplexer is used to output the mantissa of the first separated number.
6. A multi-precision floating point multiplication device according to claim 1, characterized in that: The compression circuit is specifically a circuit built based on the Wallace tree.
7. A multi-precision floating point multiplication device according to claim 6, characterized in that: When both floating-point numbers are double-precision floating-point numbers, the bit width of the mantissas of the two separated numbers is 53 bits; when both floating-point numbers are single-precision floating-point numbers, the 24th to 53rd bits of the mantissas of the two separated numbers are filled with zeros starting from the lowest bit of the mantissas of the two separated numbers; when both floating-point numbers are half-precision floating-point numbers, the 11th to 53rd bits of the mantissas of the two separated numbers are filled with zeros starting from the lowest bit of the mantissas of the two separated numbers.
8. The multi-precision floating point multiplication device according to claim 7, characterized in that: The compression circuit includes: 22 4-2 compressors, 11 3-2 compressors, a first adder, a second adder, a third adder, a first clock gate controller, a second clock gate controller and a third clock gate controller; wherein the 52 signal input terminals corresponding to the first 4-2 compressor, the second 4-2 compressor, the third 4-2 compressor, the fourth 4-2 compressor, the fifth 4-2 compressor, the sixth 4-2 compressor, the seventh 4-2 compressor, the eighth 4-2 compressor, the ninth 4-2 compressor, the tenth 4-2 compressor, the eleventh 4-2 compressor, the twelfth 4-2 compressor and the thirteenth 4-2 compressor are used to receive the first partial product to the fifty-second partial product in sequence; the i-th partial product is: the result corresponding to the multiplication of the i-th number in the mantissa of the second separated number and the mantissa of the first separated number; the i-th number in the mantissa of the second separated number is: the number corresponding to the i-th position in the mantissa of the second separated number with the lowest bit in the mantissa of the second separated number as the starting end; 1≤i≤53; 27 signal input terminals corresponding to the first 3-2 compressor, the second 3-2 compressor, the third 3-2 compressor, the fourth 3-2 compressor, the fifth 3-2 compressor, the sixth 3-2 compressor, the seventh 3-2 compressor, the eighth 3-2 compressor and the ninth 3-2 compressor are used to receive 26 signals output by the first 4-2 compressor, the second 4-2 compressor, the third 4-2 compressor, the fourth 4-2 compressor, the fifth 4-2 compressor, the sixth 4-2 compressor, the seventh 4-2 compressor, the eighth 4-2 compressor, the ninth 4-2 compressor, the tenth 4-2 compressor, the eleventh 4-2 compressor, the twelfth 4-2 compressor and the thirteenth 4-2 compressor and the 53rd partial product in sequence; The 20 signal input terminals corresponding to the fourteenth 4-2 compressor, the fifteenth 4-2 compressor, the sixteenth 4-2 compressor, the seventeenth 4-2 compressor and the eighteenth 4-2 compressor are used to receive 18 signals output by the first 3-2 compressor, the second 3-2 compressor, the third 3-2 compressor, the fourth 3-2 compressor, the fifth 3-2 compressor, the sixth 3-2 compressor, the seventh 3-2 compressor, the eighth 3-2 compressor and the ninth 3-2 compressor and two zero-level signals in sequence; The 12 signal input terminals corresponding to the nineteenth 4-2 compressor, the twentieth 4-2 compressor and the twenty-first 4-2 compressor are used to receive the ten signals output by the fourteenth 4-2 compressor, the fifteenth 4-2 compressor, the sixteenth 4-2 compressor, the seventeenth 4-2 compressor and the eighteenth 4-2 compressor and the two zero-level signals in sequence; The six signal input terminals corresponding to the tenth 3-2 compressor and the eleventh 3-2 compressor are used to receive the six signals output by the nineteenth 4-2 compressor, the twentieth 4-2 compressor and the twenty-first 4-2 compressor in sequence; The four input ends of the twenty-second 4-2 compressor are respectively used to receive the two output signals of the tenth 3-2 compressor and the two output signals of the eleventh 3-2 compressor; The two input ends of the first adder are respectively used to receive the two output signals of the fourteenth 4-2 compressor, and the output end of the first adder is the first output end of the compression circuit; the two input ends of the second adder are respectively used to receive the two output signals of the nineteenth 4-2 compressor, and the output end of the second adder is the second output end of the compression circuit; the two input ends of the third adder are respectively used to receive the two output signals of the twenty-second 4-2 compressor, and the output end of the third adder is the third output end of the compression circuit; The first clock gate controller is used to control the 4-2 compressor and the 3-2 compressor on the computing path from the 1st partial product to the 12th partial product; the second clock gate controller is used to control the 4-2 compressor and the 3-2 compressor on the computing path from the 1st partial product to the 24th partial product; the third clock gate controller is used to control the 4-2 compressor and the 3-2 compressor on the computing path from the 1st partial product to the 53rd partial product.
9. A multi-precision floating point multiplication device according to claim 8, characterized in that: When both floating-point numbers are half-precision floating-point numbers, and the first clock gating device activates the 4-2 compressor and the 3-2 compressor on the operation path from the first partial product to the twelfth partial product, the output end of the first adder is used to output the mantissa processing number; When both floating-point numbers are single-precision floating-point numbers, and the second clock gating device activates the 4-2 compressor and the 3-2 compressor on the operation path from the first partial product to the 24th partial product, the output end of the second adder is used to output the mantissa processing number; When both floating-point numbers are double-precision floating-point numbers and the third clock gating device activates the 4-2 compressor and the 3-2 compressor on the operation path from the 1st partial product to the 53rd partial product, the output end of the third adder is used to output the mantissa processing number.
10. A floating point unit, characterized in that: A multi-precision floating-point number multiplication device comprising the multi-precision floating-point number multiplication device as claimed in any one of claims 1 to 9.