Floating point calculation device and method for processor, electronic equipment and storage medium

By using the mantissa alignment module in the floating-point computing device to intercept and shift the mantissa of the floating-point number, it can achieve exponential alignment before addition, solving the problem of long calculation paths of floating-point addition and improving the calculation speed.

CN120196367AActive Publication Date: 2025-06-24MOORE THREADS TECH CO LTD

Patent Information

Application Number
CN202510677232.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

During the floating point addition calculation process, multiple operation processing of selectors, shifters and adders is required, resulting in a long calculation path and affecting the calculation speed.

Method used

By introducing a mantissa alignment module in the floating-point calculation device, the mantissa of the floating-point number is processed by using the intercepted bit number and shifted bit number, so that the two mantissa achieve exponential alignment before addition, thereby reducing the selector processing process before shifting.

Benefits of technology

The path of floating point addition calculation is shortened, the calculation speed is improved, and the complexity and cost of hardware are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196367A_ABST
    Figure CN120196367A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a floating point calculation device and method for a processor, electronic equipment and a storage medium, and the device comprises a mantissa alignment module which is used for carrying out the interception processing of a first mantissa of a first floating point number based on an interception digit to obtain a first addend, and carrying out the first addend based on a shift digit to obtain a second addend; the second mantissa of the second floating-point number is shifted to obtain a second addend, and the first addend and the second addend are two mantissas with aligned indexes; the interception bit number and the shift bit number are determined based on an index difference value between a first index of the first floating-point number and a second index of the second floating-point number; the adder is used for adding the first addend and the second addend to obtain a first mantissa sum; and the post-processing module is used for performing normalization processing and rounding processing on the first mantissa sum to obtain a target mantissa. Therefore, the selector processing process before the mantissa is shifted in the floating point addition calculation process can be reduced, the calculation path is shortened, and the calculation speed is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chip technology, and particularly to a floating-point calculation device and method for a processor, an electronic device, and a storage medium. Background Art

[0002] Floating-point calculation is an important type of data operation, which has the characteristics of high calculation accuracy and wide calculation range, and is widely used in multiple fields. For example, floating-point calculation can be applied to scenarios such as scientific calculation, graphics processing, artificial intelligence deep learning, and multimedia data processing. Among them, floating-point addition calculation is a common floating-point calculation. During the floating-point addition calculation process, it is usually necessary to first select the one with the smaller corresponding exponent from the mantissa of the first floating-point number and the mantissa of the third floating-point number, and then shift the selected mantissa with the smaller corresponding exponent to align the two mantissas, and then add the two aligned mantissas. That is to say, during the addition process, it is necessary to perform arithmetic processing such as a selector, a shifter, and an adder in sequence; thus, the addition calculation process path is long, which affects the calculation speed. Summary of the Invention

[0003] The embodiments of this application provide a floating-point calculation device and method for a processor, an electronic device, and a storage medium, which can reduce the selector processing process before shifting the mantissa during the floating-point addition calculation process, shorten the calculation path, and thus improve the calculation speed.

[0004] The technical solution of this application is implemented as follows: The embodiments of this application provide a floating-point calculation device for a processor, including: A mantissa alignment module, configured to perform truncation processing on the first mantissa of the first floating-point number based on the truncation bits to obtain a first addend, and perform shifting processing on the second mantissa of the second floating-point number based on the shift bits to obtain a second addend, where the first addend and the second addend are two mantissas with aligned exponents; the truncation bits and the shift bits are determined based on the exponent difference between the first exponent of the first floating-point number and the second exponent of the second floating-point number; An adder, configured to add the first addend and the second addend to obtain a first mantissa sum; A post-processing module, configured to perform normalization processing and rounding processing on the first mantissa sum to obtain a target mantissa.

[0005] The embodiments of this application provide a floating-point calculation method for a processor, including: Based on the number of bits to be truncated, truncate the first mantissa of the first floating-point number to obtain a first addend, and based on the number of bits to be shifted, shift the second mantissa of the second floating-point number to obtain a second addend. The first addend and the second addend are two mantissas with aligned exponents. The number of bits to be truncated and the number of bits to be shifted are determined based on the exponent difference between the first exponent of the first floating-point number and the second exponent of the second floating-point number. Add the first addend and the second addend to obtain a sum of the first mantissas. Perform normalization processing and rounding processing on the sum of the first mantissas to obtain a target mantissa.

[0006] An embodiment of the present application provides an electronic device, including: A memory for storing a computer program; A processor for executing the steps of the above floating-point calculation method when running the computer program.

[0007] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above floating-point calculation method are implemented.

[0008] An embodiment of the present application provides a floating-point calculation device and method for a processor, an electronic device, and a computer-readable storage medium. Since, in the process of calculating the sum of the mantissas of the first floating-point number and the second floating-point number, the number of bits to be shifted and the number of bits to be truncated can be determined in advance according to the exponent difference between the exponent of the first floating-point number and the exponent of the second floating-point number, it is possible to, while truncating the mantissa of the first floating-point number to obtain a first addend, shift the mantissa of the second floating-point number to obtain a second addend, so that the first addend and the second addend have aligned exponents. In this way, the mantissa of the second floating-point number can be directly shifted to the right, without first selecting the mantissa with the smaller corresponding exponent from the mantissas of the first floating-point number and the second floating-point number and then shifting the mantissa with the smaller corresponding exponent to the right, reducing the selector processing process before shifting, shortening the calculation path, and improving the calculation speed. Description of the Drawings

[0009] Figure 1 It is a schematic diagram of the data structure of a floating-point format provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of the implementation of a floating-point multiplication and addition calculation method in the related art provided by an embodiment of the present application; Figure 3 It is a schematic diagram of the structure of a floating-point calculation device in the related art provided by an embodiment of the present application; Figure 4 It is a schematic diagram of the composition structure of a floating-point calculation device provided by an embodiment of the present application Figure 1 ; Figure 5 Schematic diagram of the composition structure of a floating-point calculation device provided by an embodiment of the present application Figure 2 ; Figure 6 Schematic diagram of the composition structure of a floating-point calculation device provided by an embodiment of the present application Figure 3 ; Figure 7 Schematic diagram of the data structure of a first addend and a second addend provided by an embodiment of the present application Figure 1 ; Figure 8 Schematic diagram of the composition structure of a floating-point calculation device provided by an embodiment of the present application Figure 4 ; Figure 9 Schematic diagram of the data structure of a first addend and a second addend provided by an embodiment of the present application Figure 2 ; Figure 10 Schematic diagram of the data structure of a first addend and a second addend provided by an embodiment of the present application Figure 3 ; Figure 11 Schematic diagram of the data structure of a first addend and a second addend provided by an embodiment of the present application Figure 4 ; Figure 12 Schematic diagram of the composition structure of a floating-point calculation device provided by an embodiment of the present application Figure 5 ; Figure 13 Schematic diagram of the implementation process of a floating-point calculation method provided by an embodiment of the present application; Figure 14 Schematic diagram of the composition structure of an electronic device provided by an embodiment of the present application Detailed implementation manners

[0010] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0011] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0012] In the following description, the terms "first", "second", and "third" only distinguish similar objects and do not represent a specific order for the objects. Understandably, "first", "second", and "third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0013] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0014] To facilitate the understanding of this solution, before describing the embodiments of this application, the application background in the embodiments of this application will be described.

[0015] Floating-point numbers have high precision and a wide calculation range. In the related art, the commonly used floating-point number format is the Institute of Electrical and Electronics Engineers (IEEE) 754 standard, that is, a floating-point number consists of three parts: a sign bit (S), an exponent bit (Exp), and a mantissa bit (Mant). Figure 1 The structure of a single-precision floating-point number FP32 is shown, as Figure 1 shown, a single-precision floating-point number FP32 occupies a total of 32 bits. Among them, the sign bit occupies 1 bit, the exponent bit occupies 8 bits, and the mantissa bit occupies 23 bits. It can be understood that in the embodiments of this application, the structure of the floating-point number may include but is not limited to FP32, FP64, or FP16, etc., and the embodiments of this application do not make any limitations in this regard.

[0016] The floating-point multiply-add instruction is one of the common floating-point calculation instructions. Among them, the floating-point addition calculation is an important calculation process in the floating-point multiply-add instruction. Therefore, improving the floating-point addition calculation speed is very important for improving the floating-point operation performance.

[0017] Taking the calculation process of the floating-point multiply-add instruction as an example, the floating-point addition calculation in the related art will be introduced below.

[0018] Figure 2 A floating-point number multiply-add calculation method in the related art is shown, including: S21. Multiply the mantissa of the floating-point number A1 by the mantissa of the floating-point number B1 to obtain a multiplication number.

[0019] Among them, the mantissa of the floating-point number is a decimal number composed of the hidden integer and the mantissa bit of the floating-point number. The hidden integer is the hidden integer bit, usually the integer bit 1.

[0020] S22. Obtain an exponent difference by subtracting the exponent of the multiplication number from the exponent of the floating-point number C1.

[0021] S23. According to the positive or negative nature of the exponent difference, select the one with the larger corresponding exponent as the addend 1 from the mantissas of the multiplication number and the floating-point number C1, and select the one with the smaller corresponding exponent and shift it to the right to obtain the addend 2.

[0022] S24. Add the addend 1 and the addend 2 to obtain an addition sum.

[0023] S25. Normalize the addition sum to obtain a normalized sum; S26. Perform a rounding process on the normalized sum to obtain a sum mantissa.

[0024] Among them, the sum mantissa is the mantissa of the product of the floating-point number A1 and the floating-point number B1, and then the sum of the floating-point number C1.

[0025] The addition calculation process therein includes S22 - S24, that is, the process of adding the floating-point decimals of the multiplication number and the floating-point number C1 to obtain the addition sum. The hardware device structure adopted in this process is as Figure 3 shown, including: a selector 31, a selector 32, a right shifter 33, and an adder 34. Among them, the selector 31 and the selector 32 can receive selection signals, and the selection signals are determined according to the exponent difference. When the exponent difference is positive, the selection signal is at a low level. At this time, the selector 31 selects the multiplication number for output, and the selector 32 selects the mantissa of the floating-point number C1 for output to the right shifter 33; when the exponent difference is negative, the selection signal is at a high level. At this time, the selector 31 selects the mantissa of the floating-point number C1 for output, and the selector 32 selects the multiplication number for output to the right shifter 33. The adder 34 is used to add the data output by the selector 31 and the data output by the right shifter 33 to obtain the addition sum. It can be seen that to implement the addition calculation process in S22 - S24, two selectors, a right shifter, and an adder need to be used for operation processing in sequence, and the path is relatively long, which affects the calculation speed.

[0026] The embodiment of the present application provides a floating-point calculation device, method, electronic device, and computer-readable storage medium for a processor, which can reduce the selector processing process before shifting the mantissa in the floating-point addition calculation process, shorten the calculation path, and thus improve the calculation speed. This device can be applied to electronic devices that require floating-point addition calculations. In some embodiments, the electronic device includes devices that need to perform addition operations through a Graphics Processing Unit (GPU). For example, the electronic device can be a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable game device), etc.

[0027] Figure 4 Schematic diagram of the composition structure of an optional floating-point calculation device provided by an embodiment of the present application Figure 1 , such as Figure 4 shown, the floating-point calculation device may include: A mantissa alignment module 41, configured to perform truncation processing on the first mantissa of the first floating-point number based on the truncation bits to obtain a first addend, and perform shift processing on the second mantissa of the second floating-point number based on the shift bits to obtain a second addend. The first addend and the second addend are two mantissas with aligned exponents; the truncation bits and the shift bits are determined based on the exponent difference between the first exponent of the first floating-point number and the second exponent of the second floating-point number; An adder 42, configured to add the first addend and the second addend to obtain a sum of the first mantissas; A post-processing module 43, configured to perform normalization processing and rounding processing on the sum of the first mantissas to obtain a target mantissa.

[0028] Among them, the mantissa of a floating-point number includes a mantissa bit and a hidden integer bit, and the hidden integer bit may be 0 or 1. The first floating-point number has a first mantissa and a first exponent, and the second floating-point number has a second mantissa and a second exponent. When adding the first floating-point number and the second floating-point number, it is necessary to align their exponents before adding.

[0029] It can be understood that performing truncation processing on the first mantissa of the first floating-point number is equivalent to shifting the first mantissa. On this basis, performing shift processing on the second mantissa of the second floating-point number can align the exponents of the first floating-point number and the second floating-point number to obtain two mantissas with aligned exponents, that is, the first addend and the second addend. Since the exponent difference between the first exponent and the second exponent can represent the difference in exponents between the first mantissa and the second mantissa, therefore, according to this exponent difference, the truncation bits used for performing truncation processing on the first mantissa and the shift bits used for performing shift processing on the second mantissa can be determined to obtain the first addend and the second addend with aligned exponents. Among them, the exponent difference between the first exponent and the second exponent may be a positive number, may be 0, or may be a negative number.

[0030] In some embodiments, based on the number of bits to be intercepted, data bit segments can be intercepted starting from the highest bit of the first mantissa to obtain the first addend. The data bit width of the first addend is the number of bits intercepted. That is, the data bit segment starting from the highest bit and with a bit width of the number of bits intercepted in the first mantissa can be intercepted as the first addend. For example, when the number of bits to be intercepted is 0, the bit width of the intercepted first addend is 0, that is, 0 is determined as the first addend. Another example is when the number of bits to be intercepted is greater than 0 and not greater than the data bit width of the first mantissa, the data bit segment starting from the highest bit and with a bit width of the number of bits intercepted in the first mantissa can be intercepted as the first addend. In this way, the effect of shifting the first mantissa to the right can be achieved, and the sum of the number of bits shifted to the right and the number of bits intercepted is the data bit width of the first mantissa. Still another example is when the number of bits to be intercepted is greater than the data bit width of the first mantissa, zeros can be filled at the end of the first mantissa to make the data bit width of the first mantissa greater than or equal to the number of bits to be intercepted, and based on the number of bits to be intercepted, data bit segments can be intercepted starting from the highest bit of the first mantissa after filling with zeros to obtain the first addend. In this way, the effect of shifting the first mantissa to the left can be achieved, and the number of bits shifted to the left is the same as the number of zeros filled at the end.

[0031] In some embodiments, the shifting process for the second mantissa may include a left-shift process or a right-shift process. For example, when the exponent difference is positive, the exponent of the first floating-point number is greater than the exponent of the second floating-point number. At this time, the second mantissa of the second floating-point number is shifted to the right, while the first mantissa of the first floating-point number remains unchanged, that is, exponent alignment can be achieved. Another example is when the exponent difference is negative, the exponent of the first floating-point number is less than the exponent of the second floating-point number. At this time, if there are multiple leading 0s in the second mantissa of the second floating-point number, the second mantissa of the second floating-point number can be shifted to the left, and the number of bits shifted to the left does not exceed the number of leading 0s. The first mantissa of the first floating-point number can remain unchanged or be shifted to the right through an interception process to align with the exponent of the left-shifted second mantissa.

[0032] It should be noted that those skilled in the art can adopt any suitable interception method to intercept the first mantissa and any suitable shifting method to shift the second mantissa according to the actual situation, as long as the exponents of the first addend and the second addend are aligned. The embodiments of the present application do not limit this.

[0033] In some embodiments, a data bit segment starting from the highest bit and having a bit width equal to the truncation number can be truncated from the first mantissa as the first addend; the second mantissa is right-shifted by the shift number of bits to obtain the second addend. In this way, exponent alignment can be achieved by right-shifting the mantissa of the second floating-point number, and the second floating-point number can be an addend or an augend in the addition operation. For example, when the exponent difference is positive, the exponent of the first floating-point number is greater than the exponent of the second floating-point number. At this time, the second mantissa of the second floating-point number is right-shifted, while the first mantissa of the first floating-point number remains unchanged, so that exponent alignment can be achieved. Therefore, the truncation number used for truncating the first mantissa can be the bit width of the first mantissa, and the shift number of bits used for right-shifting the second mantissa can be the exponent difference. Another example is when the exponent difference is negative, the exponent of the first floating-point number is less than the exponent of the second floating-point number. At this time, if the second mantissa of the second floating-point number is right-shifted and the first mantissa of the first floating-point number remains unchanged, the exponent difference between the two will be even greater. Therefore, in this case, the first mantissa of the first floating-point number also needs to be right-shifted, and the number of bits shifted is greater than the number of bits by which the second mantissa of the second floating-point number is right-shifted, in order to achieve exponent alignment between the first mantissa and the second mantissa. That is, the truncation number used for truncating the first mantissa is less than or equal to the difference between the bit width of the first mantissa and the absolute value of the exponent difference, and the sum of the shift number of bits used for right-shifting the second mantissa, the truncation number, and the absolute value of the exponent difference is the bit width of the first mantissa. Another example is when the exponent difference is 0, the first floating-point number and the second floating-point number are exponent-aligned. Then, the first mantissa of the first floating-point number does not need to be shifted, and the second mantissa of the second floating-point number does not need to be shifted either. That is, the truncation number used for truncating the first mantissa can be the bit width of the first mantissa, and the shift number of bits used for right-shifting the second mantissa can be 0. Or, the first mantissa of the first floating-point number and the second mantissa of the second floating-point number can be right-shifted by the same number of bits. That is, the truncation number used for truncating the first mantissa can be less than the bit width of the first mantissa, and the sum of the shift number of bits used for right-shifting the second mantissa and the truncation number is the bit width of the first mantissa.

[0034] Exemplarily, the most significant bit of the bit width of the adder 42 is reserved for the sign bit. Taking the FP32 floating-point format as an example, the first mantissa of the first floating-point number and the second mantissa of the second floating-point number are both 24 bits, expressed as "1.23". When the exponent difference is 2, if the second mantissa "1.23" of the second floating-point number is shifted right by 2 bits, the first mantissa "1.23" of the first floating-point number does not need to be shifted right, that is, the number of bits to be truncated is 24; if the second mantissa "1.23" of the second floating-point number is shifted right by 3 bits, the first mantissa "1.23" of the first floating-point number needs to be shifted right by 1 bit, that is, the number of bits to be truncated is 23. When the exponent difference is -5, if the second mantissa "1.23" of the second floating-point number is shifted right by 3 bits, the first mantissa "1.23" of the first floating-point number needs to be shifted right by 8 bits, that is, the number of bits to be truncated is 16; if the second mantissa "1.23" of the second floating-point number is shifted right by 4 bits, the first mantissa "1.23" of the first floating-point number needs to be shifted right by 9 bits, that is, the number of bits to be truncated is 15.

[0035] In some embodiments, the right shift of the first mantissa is implemented by truncating the high bits, and the number of truncated bits can be set in combination with the exponent difference and the bit width of the data bits of the adder 42. For example, when the exponent difference is -5, if the second mantissa of the second floating-point number is shifted right by 3 bits, in order to implement the right shift of the first mantissa by 8 bits, when the bit width of the data bits of the adder 42 is 24, the number of truncated bits is 16 bits; when the bit width of the data bits of the adder is 50, the number of truncated bits is 42 bits; when the bit width of the data bits of the adder is 53, the number of truncated bits is 45 bits.

[0036] In some embodiments, when the positive and negative properties of the exponent difference are different, the number of shift bits for the second mantissa of the second floating-point number to shift right can be different or the same. It should be noted that when the positive and negative properties of the exponent difference are different, even if the number of shift bits for the second mantissa of the second floating-point number to shift right is the same, the number of truncated bits is different. When the positive and negative properties of the exponent difference are determined, the number of shift bits and the number of truncated bits are corresponding, so that the exponents of the first addend and the second addend can be aligned.

[0037] The post-processing module 43 can perform normalization processing and rounding processing on the first mantissa using any suitable normalization method and rounding processing method, and the embodiments of the present application do not limit this.

[0038] It can be understood that the normalization process is used to make the representation of floating-point numbers conform to a preset floating-point format. For example, the normalization process can make the representation of floating-point numbers conform to the IEEE754 standard, ensuring that the highest bit is opposite to the sign bit. If not normalized, it may lead to precision loss and calculation errors. The normalization process can include left normalization and / or right normalization. Left normalization means that if the highest bit of the mantissa is not a valid value (for example, there are multiple leading 0s), the mantissa needs to be shifted left until the highest bit is a valid value. Right normalization means that if the mantissa overflows, the mantissa needs to be shifted right by one bit.

[0039] The rounding process can include, but is not limited to, at least one of rounding to the nearest, forced setting to 1, truncation method, rounding to the nearest even number (Round to Nearest, RNE), rounding toward zero (Round toward Zero, RTZ), rounding toward positive infinity (Round toward +Infinity, RPI), rounding toward negative infinity (Round toward +Infinity, RNI), etc.

[0040] In some embodiments, the post-processing module 43 can first perform a normalization process on the first mantissa sum to remove the leading 0s and obtain a second mantissa sum; then perform a rounding process on the second mantissa sum according to the rounding process method to obtain a target mantissa. The number of bits of the target mantissa is the same as the effective bit width of the mantissa of the floating-point number. Among them, the effective bit width of the mantissa of the floating-point number can be determined in advance according to the actual application scenario.

[0041] In some embodiments, the rounding process method adopted in the post-processing module 43 can be a preset rounding process method. In some embodiments, multiple rounding process methods can be set in the post-processing module 43, and in the case of receiving a rounding selection instruction, perform a rounding process according to the rounding process method indicated by the rounding selection instruction.

[0042] In some embodiments, the post-processing module 43 can intercept the high-order bits of the number of bits of the mantissa of the floating-point number from the second mantissa sum as the mantissa sum to be rounded (that is, the valid bits of the second mantissa sum), and determine the rounding-related bits from the second mantissa sum, and perform a rounding process on the mantissa sum to be rounded according to the rounding-related bits to obtain a target mantissa.

[0043] In some embodiments, the mantissa sum to be rounded can be rounded based on the rounding-related bits in the second mantissa sum, the low-order discarded bits stick1 in the first mantissa except for the intercepted data field, and the discarded bits stick2 after the rounding-related bits in the second mantissa sum, to obtain a target mantissa, such that the number of bits of the target mantissa is the same as the effective bit width of the mantissa of the floating-point number. Among them, the rounding-related bits can include the least significant bit of the mantissa sum to be rounded, or the rounding-related bits can include the least significant bit of the mantissa sum to be rounded and at least one bit after the least significant bit. Here, the number of bits of the rounding-related bits can be set according to actual needs, and the embodiments of the present application do not limit it.

[0044] In some embodiments, the rounding-related bits can include 3 bits, namely the least significant digit lsd, the rounding digit rnd, and the guard digit gard. Among them, the least significant digit lsd is the least significant bit of the mantissa sum to be rounded, the bit after the least significant digit lsd is the rounding digit rnd, and the bit after the rounding digit is the guard digit gard. The post-processing module 43 can round the mantissa sum to be rounded according to the least significant digit lsd, the rounding digit rnd, the guard digit gard, the low-order discarded bits stick1, and the discarded bits stick2, to obtain a target mantissa.

[0045] In some embodiments, the post-processing module 43 can determine a discard value according to the discarded bits stick1 and the discarded bits stick2, and round the mantissa sum to be rounded together with the least significant digit lsd, the rounding digit rnd, the guard digit gard, and the discard value, to obtain a target mantissa.

[0046] In some embodiments, when the rounding-related bits and the discard value satisfy the carry condition, add 1 to the mantissa sum to be rounded to obtain a target mantissa; or when the rounding-related bits and the discard value do not satisfy the carry condition, determine the mantissa sum to be rounded as the target mantissa. For example, the carry condition can include but is not limited to: both the rounding-related bits and the discard value are 1, or one of the rounding-related bits and the discard value is 1.

[0047] In the embodiments of the present application, since during the process of calculating the sum of the mantissa of the first floating-point number and the mantissa of the second floating-point number, the shift number of bits and the interception number of bits can be determined in advance according to the exponent difference between the exponent of the first floating-point number and the exponent of the second floating-point number, so that while intercepting the mantissa of the first floating-point number to obtain the first addend, the mantissa of the second floating-point number can be shifted to obtain the second addend, enabling the first addend and the second addend to achieve exponent alignment. In this way, the mantissa of the second floating-point number can be directly shifted to the right, instead of first selecting the mantissa with the smaller corresponding exponent from the mantissa of the first floating-point number and the mantissa of the second floating-point number, and then shifting the mantissa with the smaller corresponding exponent to the right, reducing the selector processing process before shifting, being able to shorten the calculation path and improve the calculation speed.

[0048] In some embodiments, as Figure 5 shown, the mantissa alignment module 41 may include: a determination module 411, configured to determine a reference exponent, and an interception bit width and a shift bit width corresponding to the reference exponent based on the exponent difference; an interception module 412, configured to intercept a data bit segment starting from the highest bit and having a bit width of the interception bit width in the first mantissa as the first addend; a shifter 413, configured to shift the second mantissa to the right by the shift bit width to obtain a second addend, and the first addend and the second addend are aligned based on the reference exponent.

[0049] Here, the exponent corresponding to the first addend and the exponent corresponding to the second addend are both the reference exponent, that is, the first addend and the second addend are aligned based on the reference exponent.

[0050] It should be noted that those skilled in the art can, according to the actual application scenario, adopt any suitable method to determine the reference exponent, the interception bit width, and the shift bit width based on the exponent difference, as long as the first addend obtained by intercepting the first mantissa with the interception bit width and the second addend obtained by shifting the second mantissa to the right by the shift bit width are aligned based on the reference exponent. The embodiments of the present application do not make any limitations in this regard.

[0051] In some embodiments, the reference exponent may be determined based on the exponent difference, the interception bit width may be determined based on the reference exponent and the first exponent of the first floating-point number, and the shift bit width may be determined based on the reference exponent and the second exponent of the second floating-point number. Here, since the first addend and the second addend are aligned based on the reference exponent, that is, the exponent corresponding to the first addend and the exponent corresponding to the second addend are both the reference exponent, therefore, the interception bit width can be determined based on the reference exponent and the first exponent of the first floating-point number, and the shift bit width can be determined based on the reference exponent and the second exponent of the second floating-point number. For example, when the exponent difference is greater than or equal to 0, the first exponent may be determined as the reference exponent. In this case, the first mantissa does not need to be shifted, and the exponent corresponding to the second mantissa after being shifted to the right (i.e., the second addend) is the first exponent. Thus, the interception bit width may be equal to the data bit width of the first mantissa, and the shift bit width used for shifting the second mantissa may be the exponent difference. Another example is when the exponent difference is less than 0, the second exponent may be determined as the reference exponent. In this case, the second mantissa does not need to be shifted, and the exponent corresponding to the first mantissa after being intercepted (i.e., the first addend) is the second exponent. Thus, the interception bit width may be equal to the difference between the data bit width of the first mantissa and the absolute value of the exponent difference, and the shift bit width used for shifting the second mantissa may be 0.

[0052] In some embodiments, the shifter 413 may include a right shifter.

[0053] In the above embodiments, the determination module determines a reference exponent, as well as an intercept bit number and a shift bit number corresponding to the reference exponent, based on an exponent difference; the intercept module intercepts a data bit segment starting from the highest bit and having a bit width equal to the intercept bit number in the first mantissa as the first addend; the shifter shifts the second mantissa to the right by the shift bit number to obtain a second addend, and the first addend and the second addend are aligned based on the reference exponent. Since the reference exponent is determined based on the exponent difference, in this way, the first mantissa and the second mantissa can be efficiently and accurately exponent-aligned to obtain the exponent-aligned first addend and second addend.

[0054] In some embodiments, as Figure 6 shown, the floating-point calculation device further includes: a multiplication module 44 for multiplying the third mantissa of the third floating-point number and the fourth mantissa of the fourth floating-point number to obtain a first mantissa; an exponent difference determination module 45 for summing the third exponent of the third floating-point number and the fourth exponent of the fourth floating-point number to obtain a first exponent, and subtracting the second exponent from the first exponent to obtain an exponent difference.

[0055] Here, the first floating-point number is the product of the third floating-point number and the fourth floating-point number. The third floating-point number has a third exponent and a third mantissa, and the fourth floating-point number has a fourth exponent and a fourth mantissa.

[0056] In the process of calculating the product of the third floating-point number and the fourth floating-point number, the exponent (i.e., the first exponent) and the mantissa (i.e., the first mantissa) of the product can be calculated separately. Among them, the third mantissa and the fourth mantissa can be input into the multiplication module 44 to obtain the first mantissa; and the first exponent can be obtained by adding the third exponent and the fourth exponent. After obtaining the first exponent, the exponent difference determination module 45 can also subtract the second exponent from the first exponent to obtain an exponent difference.

[0057] In some embodiments, the multiplication module 44 includes a multiplier. Exemplarily, the highest bits of the bit widths of the multiplier and the adder 42 are reserved for sign bits. Taking the FP32 floating-point format as an example, the mantissas of the first floating-point number and the second floating-point number are both 24 bits, expressed as "1.23". The bit width of the multiplier is 50 bits, where the first highest bit is the sign bit "a", the second highest bit is the reserved overflow bit "b", and the multiplicand is expressed as "b.48".

[0058] In some embodiments, the exponent difference determination module 45 may include: an addition module for summing the third exponent of the third floating-point number and the fourth exponent of the fourth floating-point number to obtain a first exponent; a subtraction module for subtracting the second exponent from the first exponent to obtain an exponent difference. Among them, the addition module may include any suitable adder, and the subtraction module may include any suitable subtractor, which are not limited in the embodiments of the present application.

[0059] In some embodiments, the most significant bit of the exponent difference output by the exponent difference determination module 45 is the sign bit, which can represent the positive or negative nature of the exponent difference.

[0060] In some embodiments, the exponent difference determination module 45 may send the exponent difference to the mantissa alignment module 41, and the mantissa alignment module 41 may determine the truncation bit number and the shift bit number according to the exponent difference. Among them, different exponent differences correspond to different truncation bit numbers and shift bit numbers, and the shift bit number corresponds to the truncation bit number to ensure the exponent alignment of the first addend and the second addend.

[0061] In some embodiments, the exponent difference determination module 45 may send the exponent difference to the determination module 411 in the mantissa alignment module 41, and the determination module 411 may determine the reference exponent according to the exponent difference, as well as the truncation bit number and the shift bit number corresponding to the reference exponent.

[0062] In the above embodiments, the multiplication module multiplies the third mantissa of the third floating-point number and the fourth mantissa of the fourth floating-point number to obtain the first mantissa; the exponent difference determination module sums the third exponent of the third floating-point number and the fourth exponent of the fourth floating-point number to obtain the first exponent, and subtracts the second exponent from the first exponent to obtain the exponent difference. In this way, the calculation path of the floating-point multiply-add operation can be shortened, the calculation speed can be improved, and the exponent difference can be determined by a simple hardware structure, reducing the cost.

[0063] In some embodiments of the present application, the bit width of the data bits of the adder 42 is twice the preset threshold; the preset threshold is the sum of the preset mantissa significant bit width and the preset rounding-related bit width minus 1.

[0064] The bit width of the adder 42 includes a sign bit and data bits; the bit width of the sign bit is 1 bit, and the bit width of the data bits is twice the preset threshold. The preset threshold can be set in advance according to the mantissa significant bit width and the rounding-related bit width, or can be calculated according to the preset mantissa significant bit width and the preset rounding-related bit width.

[0065] The mantissa significant bit width refers to the data bit width of the significant bits of the mantissa, and the rounding-related bit width refers to the data bit width of the rounding-related bits. The mantissa significant bit width and the rounding-related bit width can both be preset by those skilled in the art according to the actual application scenario, and the embodiments of the present application do not limit this. For example, the mantissa significant bit width can be 24 bits, or 25 bits, etc., and the rounding-related bits can be 1 bit, 2 bits, or 3 bits, etc.

[0066] It can be understood that the preset threshold can represent the number of consecutive bits from the first bit of the significant bits of the mantissa to the last bit of the bits related to rounding when rounding the mantissa with a preset significant bit width of the mantissa. Since the last bit of the significant bits of the mantissa is the starting bit of the bits related to rounding, the preset threshold is the sum of the preset significant bit width of the mantissa and the preset bit width of the bits related to rounding minus 1.

[0067] In addition, since during the shifting process, there may be a situation where the first addend is much smaller than the second addend, or the second addend is much smaller than the first addend. Thus, by setting the bit width of the data bits of the adder to twice the preset threshold, while retaining more significant bits of the mantissa to improve the accuracy of the multiply-add operation, the bit width of the adder 42 can be minimized to the greatest extent, reducing the hardware cost.

[0068] Exemplarily, taking FP32 as an example, the preset significant bit width of the mantissa of the floating-point number is 24, the bit width of the bits related to rounding is 3, then the preset threshold is 26, the bit width of the data bits of the adder is 52, and the bit width of the adder is the bit width of 52 data bits and the bit width of 1 sign bit, that is, 53.

[0069] In some embodiments, the multiplication module 44 includes a multiplier, and the bit width of the multiplier includes the bit width of the sign bit and the bit width of the multiplication data; the bit width of the multiplication data includes twice the preset significant bit width of the mantissa of the floating-point number plus 1, and the added 1 bit is the overflow bit. In this way, while ensuring the accuracy of the multiplication calculation result, the bit width of the multiplier can be minimized to the greatest extent, reducing the hardware cost.

[0070] In some embodiments, the determination module 411 is further configured to: when the exponent difference is greater than or equal to 0, determine the first exponent as the reference exponent, determine the bit width of the data bits of the adder 42 as the truncation number of bits, and determine the exponent difference as the shift number of bits.

[0071] In the embodiments of the present application, since the exponent difference is a non-negative number, the second mantissa of the second floating-point number is shifted to the right by the number of bits of the exponent difference, and then it can be exponent-aligned with the first mantissa of the first floating-point number. The shift number of bits is the exponent difference, and the truncation number of bits is the bit width of the data bits of the adder 42, that is, the bit width of the adder 42 minus 1. Here, the subtracted 1 is the highest sign bit. Thus, it is equivalent to shifting the second mantissa to the right by 0 bits, that is to say, the first addend is the second mantissa itself.

[0072] It should be noted that when the number of digits of the second mantissa is less than the bit width of the data bits of the adder 42, the adder 42 can fill in numbers at the end of the second mantissa until the number of digits of the second mantissa after filling is the same as the bit width of the data bits of the adder 42. Among them, the value filled at the end can be 0 or 1, which can be determined according to the positive or negative nature of the second mantissa. For example, if the second mantissa is positive, 0 can be filled at the end of the second mantissa; if the second mantissa is negative, 1 can be filled at the end of the second mantissa.

[0073] Exemplarily, Figure 7 shows a schematic data structure of a first addend and a second addend Figure 1 , such as Figure 7 shown, the highest bit of the adder is the sign bit S. The exponent difference is 3. The first mantissa is "b.48" and the second mantissa is "1.23". Then the first addend is "b.48" after the sign bit S, and the second addend is "1.23" shifted right by 3 bits.

[0074] It can be understood that when the exponent difference is greater than or equal to 0, the second mantissa is directly shifted right by the number of bits of the exponent difference, and the exponent alignment with the first mantissa can be directly achieved. In this way, the operations in the exponent alignment process can be reduced, and the efficiency of exponent alignment can be improved.

[0075] In some embodiments of the present application, such as Figure 8 shown, the post-processing module 43 includes a normalization module 431 and a rounding module 432; The shifter 413 is further configured to: when the exponent difference is greater than or equal to 0, transmit the data bit segment shifted out after shifting the second mantissa by the number of shifting bits as the first discarded bit to the rounding module 432; The normalization module 431 is configured to remove the leading 0s in the first mantissa and obtain the second mantissa sum; The rounding module 432 is configured to perform a rounding process on the second mantissa sum based on the first discarded bit to obtain the target mantissa.

[0076] It can be understood that when the shifter shifts the second mantissa by the number of shifting bits, the lowest bit data bit segment with the data bit width of the number of shifting bits in the second mantissa can be shifted out from the shifter. When the exponent difference is greater than or equal to 0, the shifted out lowest bit data bit segment can be transmitted as the first discarded bit to the rounding module 432.

[0077] For example, if the second mantissa is denoted as "1.23" and the exponent difference is 3, then the second addend is "1.23" shifted right by 3 bits, and the lowest 3 bits of the second mantissa are the lowest bit data bit segment shifted out from the shifter, which is also the first discarded bit.

[0078] The normalization module 431 can remove the leading 0s in the first mantissa sum to obtain a second mantissa sum, and output the second mantissa sum to the rounding module 432.

[0079] The rounding module 432 can perform a rounding process on the second mantissa sum based on the first discard bit using any suitable rounding method to obtain a target mantissa, which is not limited in the embodiments of the present application.

[0080] In some embodiments, when the first discard bit is greater than 0, the significant bits of the second mantissa sum can be carried (i.e., incremented by 1) to obtain the target mantissa bits; when the first discard bit is 0, the significant bits of the second mantissa sum are determined as the target mantissa bits.

[0081] In some embodiments, the second mantissa sum can be rounded based on the first discard bit and the rounding-related bits in the second mantissa sum to obtain a target mantissa. For example, when both the first discard bit and the rounding-related bits are greater than 0, the significant bits of the second mantissa sum are carried (i.e., incremented by 1) to obtain the target mantissa bits; when the first discard bit is 0 or the rounding-related bits are 0, the significant bits of the second mantissa sum are determined as the target mantissa bits. Another example is that when at least one of the first discard bit and the rounding-related bits is greater than 0, the significant bits of the second mantissa sum are carried (i.e., incremented by 1) to obtain the target mantissa bits; when both the first discard bit and the rounding-related bits are 0, the significant bits of the second mantissa sum are determined as the target mantissa bits.

[0082] In the above embodiments, when the exponent difference is greater than or equal to 0, the shifter shifts the second mantissa to the right by the shift number of bits, and the data bit segment shifted out is transmitted to the rounding module as the first discard bit; the normalization module removes the leading 0s in the first mantissa sum to obtain a second mantissa sum; the rounding module rounds the second mantissa sum based on the first discard bit to obtain a target mantissa. In this way, when the exponent difference is greater than or equal to 0, the normalization process and the rounding process can be efficiently and accurately implemented, improving the calculation accuracy and simplifying the rounding process logic.

[0083] In some embodiments, the determination module 411 is further configured to perform at least one of the following: When the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to a preset threshold, the second exponent is determined as the reference exponent, and it is determined that both the truncation number of bits and the shift number of bits are 0; When the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, the sum of the first exponent and the preset threshold is determined as the reference exponent, the preset threshold is determined as the truncation number of bits, and the difference between the preset threshold and the absolute value of the exponent difference is determined as the shift number of bits.

[0084] Here, when the exponent difference is less than 0, the determination module can determine the reference exponent, as well as the truncation bit number and shift bit number corresponding to the reference exponent, according to the magnitude relationship between the absolute value of the exponent difference and a preset threshold.

[0085] It can be understood that when the exponent difference is less than 0, if the absolute value of the exponent difference is greater than or equal to the preset threshold, it indicates that the difference between the second exponent and the first exponent is greater than or equal to the preset threshold. In this case, the second exponent can be used as the reference exponent. Therefore, the shift bit number used for right-shifting the second mantissa is 0, that is, the second mantissa is not right-shifted; moreover, the first mantissa will not affect the preset threshold number of data bits starting from the highest bit in the non-shifted second mantissa. Therefore, the first mantissa can be directly used as the low-order discarded bits. That is to say, the truncation bit number can be 0. In this way, when the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, the normalization process and rounding process can be efficiently and accurately implemented, improving the calculation accuracy and simplifying the rounding process logic.

[0086] Exemplarily, Figure 9 shows a schematic data structure of a first addend and a second addend Figure 2 , as Figure 9 shown. The bit width of the adder is 53, and the highest bit is the sign bit S; the preset threshold is 26. If the exponent difference is -27, both the truncation bit number and the shift bit number are 0. Thus, the first mantissa is represented as "b.48" and is used as the low-order discarded bit stick1 as a whole, and the first addend is 0; the second mantissa is represented as "1.23", and the second addend is the second mantissa itself.

[0087] When the exponent difference is less than 0, if the absolute value of the exponent difference is less than the preset threshold, it indicates that the difference between the second exponent and the first exponent is less than the preset threshold. In this case, the sum of the first exponent and the preset threshold can be used as the reference exponent. In this way, the truncation bit number can be the preset threshold, that is, the data bit segment with a width of the preset threshold starting from the highest bit in the first mantissa can be truncated as the low order of the first addend; equivalently, the number of bits by which the first mantissa is right-shifted is the bit width of the 42 data bits of the adder (i.e., twice the preset threshold) minus the preset threshold, which is the preset threshold. At this time, the shift bit number used for right-shifting the second mantissa is the preset threshold minus the absolute value of the exponent difference, which is equivalent to left-shifting the second mantissa by the absolute value of the exponent difference relative to the truncated first mantissa. In this way, the first addend and the second addend can be exponent-aligned. In this way, when the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, the normalization process and rounding process can be efficiently and accurately implemented, improving the calculation accuracy and simplifying the rounding process logic Exemplarily, Figure 10 shows a schematic data structure of a first addend and a second addendFigure 3 , as Figure 10 shown, the adder bit width is 53, and the highest bit is the sign bit S; the preset threshold is 26. If the exponent difference is -5, the first mantissa is represented as "b.48" and the second mantissa is represented as "1.23", then the truncation bit number is 26, that is, the lower bits of the first addend are "b.25", and except for the sign bit, the higher bits are filled with 0. At this time, the shift bit number for right-shifting the second mantissa is 21, that is, the second addend is obtained by right-shifting the second mantissa "1.23" by 21 bits. In this way, it is equivalent to the second mantissa being left-shifted by 5 bits relative to the lower 26 bits "b.25" of the first addend.

[0088] It can be understood that by truncating the first mantissa, the right-shifting of the first mantissa is realized, so that it is exponentially aligned with the right-shifted second mantissa, which can reduce the setting of the selector, shorten the calculation path, and improve the calculation speed.

[0089] In some embodiments, continue to refer to Figure 8 , the post-processing module 43 includes a normalization module 431 and a rounding module 432; The truncation module 412 is further configured to: in the case where the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, transmit the first mantissa as the first discarded bit to the rounding module 432; and / or, in the case where the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, transmit the remaining data bit segment after truncation in the first mantissa as the first discarded bit to the rounding module 432; The normalization module 431 is configured to remove the leading 0s in the first mantissa and, to obtain the second mantissa and; The rounding module 432 is configured to perform a rounding process on the second mantissa and based on the first discarded bit to obtain the target mantissa.

[0090] It can be understood that in the case where the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, the second exponent is used as the reference exponent, and both the truncation bit number and the shift bit number are 0. At this time, the first addend will not affect the preset threshold number of data bits starting from the highest bit in the unshifted second mantissa. Therefore, the first mantissa can be directly transmitted as the first discarded bit to the rounding module 432.

[0091] When the exponential difference is less than 0 and the absolute value of the exponential difference is less than the preset threshold, the sum of the first exponent and the preset threshold is used as the reference exponent, the preset threshold is used as the truncation bit width, and the difference between the preset threshold and the absolute value of the exponential difference is used as the shift bit width. At this time, the data bit segment starting from the highest bit and with a bit width of the preset threshold in the first mantissa can be intercepted as the low bits of the first addend, which is equivalent to shifting the first mantissa to the right by the number of bits of the preset threshold, and the second addend is obtained by shifting the second mantissa to the left by the absolute value of the exponential difference relative to the truncated first mantissa. Therefore, the remaining intercepted data bit segment in the first mantissa can be transmitted to the rounding module 432 as the first discarded bit, which is equivalent to taking the number of bits of the preset threshold shifted out from the right in the first mantissa as the first discarded bit.

[0092] In the above embodiments, when the exponential difference is less than 0 and the absolute value of the exponential difference is greater than or equal to the preset threshold, the intercepting module transmits the first mantissa to the rounding module as the first discarded bit; and / or, when the exponential difference is less than 0 and the absolute value of the exponential difference is less than the preset threshold, the remaining intercepted data bit segment in the first mantissa is transmitted to the rounding module as the first discarded bit; the normalization module removes the leading 0s in the sum of the first mantissas to obtain the sum of the second mantissas; the rounding module performs rounding processing on the sum of the second mantissas based on the first discarded bit to obtain the target mantissa. In this way, when the exponential difference is less than 0, the normalization processing and rounding processing can be efficiently and accurately implemented, improving the calculation accuracy and simplifying the rounding processing logic.

[0093] In some embodiments of the present application, the rounding module 432 is further configured to intercept the data bit segment starting from the highest bit and with a bit width of the effective bit width of the mantissa in the sum of the second mantissas as the mantissa sum to be rounded; starting from the lowest bit of the mantissa sum to be rounded, at least two consecutive rounding-related bits are determined from the sum of the second mantissas; according to the first discarded bit, the rounding-related bits, and the second discarded bit after the rounding-related bits, the mantissa sum to be rounded is rounded according to the target rounding condition to obtain the target mantissa.

[0094] Among them, the rounding-related bits can be at least two consecutive data bits starting from the lowest bit of the mantissa sum to be rounded in the sum of the second mantissas. The number of bits of the rounding-related bits can include the preset bit width of the rounding-related bits, which can be preset according to the actual application scenario. For example, the bit width of the rounding-related bits can be 2 bits or 3 bits, etc.

[0095] The second discarded bit can be at least one data bit starting from the next bit after the last rounding-related bit in the sum of the second mantissas.

[0096] In some implementation manners, the first discarded bit can be, for example but not limited to, the aforementioned low-order discarded bit stick1, and the second discarded bit can be, for example but not limited to, the aforementioned discarded bit stick2.

[0097] The target input value condition can be determined by those skilled in the art according to the rounding method adopted in the actual application scenario, and the embodiments of the present application do not limit this. For example, the target input value condition may include that at least one of the first discarded bit, the rounding-related bit, and the rounding-related bit is greater than 0. Another example is that the target input value condition may include that the first discarded bit, the rounding-related bit, and the rounding-related bit are all greater than 0.

[0098] In some embodiments, it is possible to determine whether the first discarded bit, the rounding-related bit, and the second discarded bit satisfy the target input value condition, and based on the determination result, perform rounding processing on the mantissa sum to be rounded to obtain the target mantissa. For example, when the first discarded bit, the rounding-related bit, and the second discarded bit satisfy the target input value condition, the mantissa sum to be rounded can be incremented by 1 and used as the target mantissa. Another example is that when the first discarded bit, the rounding-related bit, and the second discarded bit do not satisfy the target input value condition, the mantissa sum to be rounded can be used as the target mantissa.

[0099] In the above embodiments, by comprehensively considering the first discarded bit, the rounding-related bit, and the second discarded bit after the rounding-related bit, the rounding processing of the mantissa sum to be rounded can be more accurate, obtaining a more precise target mantissa and improving the calculation accuracy.

[0100] In some embodiments, the rounding module 432 is further configured to: Determine the target discard value according to the first discarded bit and the second discarded bit; When the target discard value and the rounding-related bit satisfy the target input value condition, determine the input value as 1; or, when the target discard value and the rounding-related bit do not satisfy the target input value condition, determine the input value as 0; Add the mantissa sum to be rounded and the input value to obtain the target mantissa.

[0101] In some embodiments, a rounding logic module, an adder, and a selector may be provided in the rounding module 432. The rounding logic module is configured to determine the target discard value according to the first discarded bit and the second discarded bit.

[0102] In some embodiments, the rounding logic module may determine that the target discard value is 0 when both the first discarded bit and the second discarded bit are 0; otherwise, determine that the target discard value is 1.

[0103] In some embodiments, the rounding logic module may determine that the first discard value is 1 when the first discarded bit contains at least one 1, otherwise determine that the first discard value is 0; and, determine that the second discard value is 1 when the second discarded bit contains at least one 1, otherwise determine that the second discard value is 0; when both the first discard value and the second discard value are 0, determine that the target discard value is 0, otherwise the target discard value is 1.

[0104] In some embodiments, when the target discard value and the rounding-related bits satisfy the target input value condition, the rounding logic module may generate a first input value selection signal, and in response to the first input value selection signal, the selector selects 1 as the input value and outputs it to the adder. Otherwise, the rounding logic module may generate a second input value selection signal, and in response to the second input value selection signal, the selector selects 0 as the input value and outputs it to the adder. The adder adds the input value from the selector to the mantissa to be rounded and the sum to obtain the target mantissa.

[0105] In some embodiments, the rounding module 432 may process the rounding process as a process of adding the mantissa sum to be rounded and 0, or a process of adding the mantissa sum to be rounded and 1. In this way, it is convenient for the rounding module 432 to implement the rounding process through hardware logic.

[0106] In some embodiments of the present application, the rounding-related bits include the least significant digit lsd, the rounding digit rnd, and the guard digit gard; the rounding processing method includes rounding to even; the target input value condition includes: the rounding digit rnd is 1, and at least one of the least significant digit lsd, the target discard value, and the guard digit is 1.

[0107] In some embodiments, when the rounding digit rnd is 1, and at least one of the least significant digit lsd, the target discard value, and the guard digit is 1, the rounding module 432 may determine that the input value is 1. The mantissa sum to be rounded is added to 1 to obtain the target mantissa.

[0108] In some embodiments, if the rounding digit is 0, the rounding module 432 may determine that the rounding-related bits and the target discard value do not satisfy the target input value condition, and thus determine that the input value is 0.

[0109] In some embodiments, if the rounding digit is 1, but the least significant digit lsd, the target discard value, and the guard digit are all 0, the rounding module 432 may determine that the rounding-related bits and the target discard value do not satisfy the target input value condition, and thus determine that the input value is 0.

[0110] It can be understood that by setting the target input value condition corresponding to the rounding processing method that the target discard value and the rounding-related bits need to satisfy, the rounding module 432 can perform a rounding operation on the mantissa sum to be rounded, achieve the rounding processing effect corresponding to the rounding processing method, and improve the accuracy of calculation.

[0111] Exemplarily, the following shows three FP32 floating-point numbers: operand A, operand B, and operand C. Among them, operand A is the third floating-point number, operand B is the fourth floating-point number, and operand C is the second floating-point number.

[0112] Operand A: 0 10000011 000_0010_0000_0011_0000_0000 (hexadecimal 32‘h41820300); Operand B: 0 10000100 000_0001_0100_0000_0101_0000 (hexadecimal 32‘h42014050); Operand C: 0 10001111 100_0000_0100_0000_0000_0000 (hexadecimal 32‘h47c04000); The bit width of the multiplier is 50, and the highest bit is the sign bit; the bit width of the adder is 53, and the highest bit is the sign bit. The preset threshold is 26.

[0113] Operand A is multiplied by Operand B to obtain the product AB (i.e., the first floating-point number). The mantissa of the product AB is the first mantissa, and the exponent of the product AB (i.e., the first exponent) is as shown in formula (1).

[0114]

[0115]

[0116] (1); Among them, the exponent part of Operand A , the exponent part of Operand B and the exponent part of Operand C all adopt the biased exponent. The bias value is 127 (binary representation is 0111_1111), that is, the real exponent in decimal needs to be added to be equal to the corresponding biased exponent.

[0117] The exponent difference is the exponent of the product AB minus the exponent of Operand C , see formula (2).

[0118]

[0119]

[0120] (2); Among them, The decimal number is -6. Since the exponent difference is negative and the absolute value of the exponent difference is less than 26, it can be determined that the truncation bit number is 26, and the difference between the preset threshold and 6, that is, 20, is the shift bit number. That is to say, the adder 42 needs to place the high 26 bits (0100 0001 1010 0100 0010 1100 01) of the 48-bit mantissa AB (i.e., the first mantissa) at the low 26 bits of the 53-bit adder, that is, truncate the high 25 bits of "b.48" as the low bits of the first addend, and the remaining low 22 bits "100000 1111 0000 0000 0000" are denoted as stick1. The shifter performs a right shift operation on the mantissa C, and the shift bit number is 20, to obtain the second addend, as Figure 11 shown. The adder 42 can obtain the sum of the first mantissas as "00000000000000000000110000010100011010010000 1011000". The rounding module 432 removes the leading 0s of the sum of the first mantissas, obtains the sum of the second mantissas as "110000010100011010010000 1011000", truncates 24 bits of the mantissa sum to be processed as "110000010100011010010000", lsd is "0", rnd is "1", gard is "0", and the remaining "11000" is denoted as stick2. It can be seen that both stick1 and stick2 contain "1", and the target discard value is 1. According to the rounding-to-even method, rnd is 1 in the rounding-related bits, and among lsd, gard, and the target discard value, the target discard value is 1. Therefore, it is determined that the rounding-in value is 1. The target mantissa is the sum of the mantissa to be rounded and 1, that is, "110000010100011010010001".

[0121] In the above embodiment, the rounding-related bits include the least significant digit lsd, the rounding digit rnd, and the guard digit gard; the rounding processing method includes rounding to even; the target rounding-in value condition includes: the rounding digit rnd is 1, and at least one of the least significant digit lsd, the target discard value, and the guard digit is 1. In this way, the rounding processing of the mantissa sum to be rounded can be accurately realized by using the rounding-to-even method.

[0122] In some embodiments, the target mantissa is the mantissa of the target floating-point number, as Figure 12 shown, and the floating-point calculation device further includes an exponent calculation module 46; A post-processing module 43, configured to perform normalization processing and rounding processing on the sum of the first mantissas to obtain a target mantissa and an exponent correction number; An exponent calculation module 46, configured to correct the reference exponent based on the exponent correction number to obtain the target exponent of the target floating-point number.

[0123] Here, during the process in which the post - processing module 43 normalizes and rounds the first mantissa sum to obtain the target mantissa, the shift processing performed will affect the exponent of the finally obtained target floating - point number. Among them, a right shift will move the decimal point of the mantissa to the left, and the exponent needs to be increased accordingly; a left shift will move the decimal point of the mantissa to the right, and the exponent needs to be decreased accordingly. For each 1 - bit movement of the mantissa, the exponent needs to change by 1 (in decimal). Thus, during the process of normalizing and rounding the first mantissa sum to obtain the target mantissa, the post - processing module 43 can also generate a corresponding exponent correction number according to the shift processing during the normalization and rounding processes. This exponent correction number can represent the exponent change amount corresponding to the target mantissa during the normalization and rounding processes.

[0124] In some embodiments, the exponent calculation module 46 can be used to add the exponent correction number to the reference exponent to obtain the target exponent of the target floating - point number.

[0125] In some embodiments, the exponent calculation module 46 is also used to obtain the first exponent of the first floating - point number based on the sum of the exponents of the third floating - point number and the fourth floating - point number.

[0126] In some embodiments, the positive or negative of the exponent correction number can represent the shift direction of the mantissa; thus, the exponent calculation module 46 can adjust the reference exponent according to the exponent correction number to obtain the target exponent, that is, the exponent of the sum of the product of the third floating - point number and the fourth floating - point number and the second floating - point number. According to the target exponent and the target mantissa, the target floating - point number can be obtained, that is, the floating - point format number of the multiply - add result of the sum of the product of the third floating - point number and the fourth floating - point number and the third floating - point number.

[0127] In the above - mentioned embodiments, the exponent calculation module corrects the reference exponent based on the exponent correction number output by the post - processing module, and a more accurate target exponent can be obtained, thereby improving the calculation accuracy of the target floating - point number.

[0128] It can be understood that through the floating - point calculation device provided by the embodiments of the present application, the multiply - add calculation between any three floating - point numbers can be completed. In this way, since the selector processing process before the shift is reduced during the calculation process, the calculation path can be shortened, thereby improving the calculation speed of the floating - point multiply - add calculation.

[0129] Based on the above - mentioned embodiments, the embodiments of the present application provide a floating - point calculation method for a processor. This method can be applied to the floating - point calculation device described in the above - mentioned embodiments. As Figure 13 shown, this method can include the following steps S101 to step S103.

[0130] S101. Based on the number of bits to be truncated, truncate the first mantissa of the first floating-point number to obtain the first addend, and based on the number of bits to be shifted, shift the second mantissa of the second floating-point number to obtain the second addend. The first addend and the second addend are two mantissas with aligned exponents; the number of bits to be truncated and the number of bits to be shifted are determined based on the exponent difference between the first exponent of the first floating-point number and the second exponent of the second floating-point number; S102. Add the first addend and the second addend to obtain the sum of the first mantissas; S103. Perform normalization processing and rounding processing on the sum of the first mantissas to obtain the target mantissa.

[0131] In some embodiments, the above step S101 is executed by the mantissa alignment module in the floating-point calculation device, the above step S102 is executed by the adder in the floating-point calculation device, and the above step S103 is executed by the post-processing module in the floating-point calculation device.

[0132] In some embodiments, the mantissa alignment module includes: a determination module, a truncation module, and a shifter; the above step S101 may include the following steps S111 to S113: Step S111. The determination module determines the reference exponent, and the number of bits to be truncated and the number of bits to be shifted corresponding to the reference exponent based on the exponent difference; Step S112. The truncation module truncates the data bit segment starting from the highest bit and with a bit width of the number of bits to be truncated in the first mantissa as the first addend; Step S113. The shifter shifts the second mantissa to the right by the number of bits to be shifted to obtain the second addend, and the first addend and the second addend are aligned based on the reference exponent.

[0133] In some embodiments, the floating-point calculation device further includes: a multiplication module and an exponent difference determination module; the above floating-point calculation method may further include the following steps S121 to S122: Step S121. The multiplication module multiplies the third mantissa of the third floating-point number and the fourth mantissa of the fourth floating-point number to obtain the first mantissa; Step S122. The exponent difference determination module sums the third exponent of the third floating-point number and the fourth exponent of the fourth floating-point number to obtain the first exponent, and subtracts the second exponent from the first exponent to obtain the exponent difference.

[0134] In some embodiments, the bit width of the data bits of the adder is twice the preset threshold; the preset threshold is the sum of the preset effective bit width of the mantissa and the preset bit width related to rounding minus 1.

[0135] In some embodiments, the above step S111 may include the following step S131: In step S131, when the exponent difference is greater than or equal to 0, the first exponent is determined as the reference exponent, the bit width of the data bits of the adder is determined as the truncation number of bits, and the exponent difference is determined as the shift number of bits.

[0136] In some embodiments, the post-processing module includes a normalization module and a rounding module; the above floating-point calculation method may further include the following step S141: When the exponent difference is greater than or equal to 0, the shifter shifts the second mantissa to the right by the shift number of bits, and the data bit segment shifted out is transmitted to the rounding module as the first discarded bit; The above step S103 may include the following steps S142 and S143: Step S142, the normalization module removes the leading 0s in the first mantissa sum to obtain the second mantissa sum; Step S143, the rounding module performs a rounding process on the second mantissa sum based on the first discarded bit to obtain the target mantissa.

[0137] In some embodiments, the above step S111 may include the following step S151 or step S152: Step S151, when the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to a preset threshold, the second exponent is determined as the reference exponent, and it is determined that both the truncation number of bits and the shift number of bits are 0; Step S152, when the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, the sum of the first exponent and the preset threshold is determined as the reference exponent, the preset threshold is determined as the truncation number of bits, and the difference between the preset threshold and the absolute value of the exponent difference is determined as the shift number of bits.

[0138] In some embodiments, the post-processing module includes a normalization module and a rounding module; the above floating-point calculation method may further include the following step S161 or step S162: Step S161, when the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, the truncation module transmits the first mantissa to the rounding module as the first discarded bit; Step S162, when the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, the truncation module transmits the remaining data bit segment intercepted from the first mantissa to the rounding module as the first discarded bit; The above step S103 may include the following steps S163 and S164: Step S163, the normalization module removes the leading 0s in the first mantissa sum to obtain the second mantissa sum; Step S164, the rounding module performs a rounding process on the second mantissa sum based on the first discarded bit to obtain the target mantissa.

[0139] In some embodiments, the above step S164 may include the following steps S171 to S173: Step S171, the rounding module intercepts the data bit segment starting from the highest bit and having a bit width equal to the effective bit width of the mantissa from the second mantissa sum as the mantissa sum to be rounded; Step S172, the rounding module determines at least two consecutive rounding-related bits from the second mantissa sum starting from the lowest bit of the mantissa sum to be rounded; Step S173, the rounding module performs a rounding process on the mantissa sum to be rounded according to the first discarded bit, the rounding-related bits, and the second discarded bit after the rounding-related bits according to the target rounding condition to obtain the target mantissa.

[0140] In some embodiments, the above step S173 may include the following steps S181 to S183: Step S181, the rounding module determines the target discard value according to the first discarded bit and the second discarded bit; Step S182, the rounding module determines that the rounding value is 1 when the target discard value and the rounding-related bits meet the target rounding condition; or determines that the rounding value is 0 when the target discard value and the rounding-related bits do not meet the target rounding condition; Step S183, the rounding module adds the mantissa sum to be rounded and the rounding value to obtain the target mantissa.

[0141] In some embodiments, the rounding-related bits include the least significant bit, the rounding bit, and the guard bit; the rounding process includes rounding to even; the target rounding condition includes: the rounding bit is 1, and at least one of the least significant bit, the target discard value, and the guard bit is 1.

[0142] In some embodiments, the target mantissa is the mantissa of the target floating-point number, and the floating-point calculation device further includes an exponent calculation module; the above floating-point calculation method further includes the following steps S191 to S192: Step S191, the post-processing module performs a normalization process and a rounding process on the first mantissa sum to obtain the target mantissa and the exponent correction number; Step S192, the exponent calculation module corrects the reference exponent based on the exponent correction number to obtain the target exponent of the target floating-point number.

[0143] Figure 14 The structure of an electronic device is shown, such as Figure 14 As shown, the electronic device 1700 includes a memory 1707, a processor 1708, and a computer program stored on the memory 1707 and executable on the processor 1708; wherein, when the processor 1708 is used to run the computer program, it executes the floating-point calculation method in the foregoing embodiments.

[0144] It can be understood that the electronic device 1700 further includes a bus system 1709; each component in the electronic device 1700 is coupled together through the bus system 1709. It can be understood that the bus system 1709 is used to realize the connection and communication between these components. In addition to including a data bus, the bus system 1709 further includes a power bus, a control bus, and a status signal bus.

[0145] It can be understood that the memory in the embodiments of the present application can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a Read-Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a Random Access Memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as a Static Random Access Memory (SRAM), a Synchronous Static Random Access Memory (SSRAM), a Dynamic Random Access Memory (DRAM), a Synchronous Dynamic Random Access Memory (SDRAM), a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), an Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), a SyncLink Dynamic Random Access Memory (SLDRAM), and a Direct Rambus Random Access Memory (DRRAM).The memories described in the embodiments of the present application are intended to include, but are not limited to, these and any other suitable types of memories.

[0146] The methods disclosed in the embodiments of the present application above can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, the steps of the above methods can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The above-mentioned processor may be a general-purpose processor, a DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the methods disclosed in the embodiments of the present application, it can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory. The processor reads the signals in the memory and combines its hardware to complete the steps of the foregoing methods.

[0147] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above methods are implemented.

[0148] The embodiments of the present application provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above methods are implemented. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0149] It should be noted here that: the above descriptions of the various embodiments tend to emphasize the differences between the various embodiments, and their similarities can be referred to each other. In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be electrical, mechanical, or other forms.

[0150] As described above, the above are only embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A floating-point calculation device for a processor, characterized in that Including: A mantissa alignment module, configured to perform truncation processing on a first mantissa of a first floating-point number based on a truncation bit number to obtain a first addend, and perform shifting processing on a second mantissa of a second floating-point number based on a shift bit number to obtain a second addend, where the first addend and the second addend are two mantissas with aligned exponents; the truncation bit number and the shift bit number are determined based on an exponent difference between a first exponent of the first floating-point number and a second exponent of the second floating-point number; An adder, configured to add the first addend and the second addend to obtain a sum of the first mantissas; A post-processing module, configured to perform normalization processing and rounding processing on the sum of the first mantissas to obtain a target mantissa.

2. The floating-point calculation device according to claim 1, characterized in that, The mantissa alignment module includes: A determination module, configured to determine a reference exponent, and a truncation bit number and a shift bit number corresponding to the reference exponent based on the exponent difference; A truncation module, configured to truncate a data bit segment starting from the highest bit and with a bit width of the truncation bit number in the first mantissa as the first addend; A shifter, configured to shift the second mantissa to the right by the shift bit number to obtain the second addend, and the first addend and the second addend are aligned based on the reference exponent.

3. The floating-point calculation device according to claim 2, wherein The floating-point calculation device further includes: A multiplication module, configured to multiply a third mantissa of a third floating-point number and a fourth mantissa of a fourth floating-point number to obtain the first mantissa; An exponent difference determination module, configured to sum a third exponent of the third floating-point number and a fourth exponent of the fourth floating-point number to obtain the first exponent, and subtract the second exponent from the first exponent to obtain the exponent difference.

4. The floating-point calculation device according to claim 3, characterized in that The bit width of the data bits of the adder is twice a preset threshold; the preset threshold is the sum of a preset effective bit width of the mantissa and a preset bit width related to rounding minus 1.

5. The floating-point calculation device according to claim 4, wherein: The determination module is further configured to: when the exponent difference is greater than or equal to 0, determine the first exponent as the reference exponent, determine the bit width of the data bits of the adder as the truncation bit number, and determine the exponent difference as the shift bit number.

6. The floating-point calculation device according to claim 5, wherein The post-processing module includes a normalization module and a rounding module; The shifter is further configured to: when the exponent difference is greater than or equal to 0, transmit a data bit segment shifted out after shifting the second mantissa to the right by the shift bit number as a first discarded bit to the rounding module; The normalization module, configured to remove leading 0s in the sum of the first mantissas to obtain a sum of the second mantissas; The rounding module, configured to perform rounding processing on the sum of the second mantissas based on the first discarded bit to obtain the target mantissa.

7. The floating-point calculation device according to claim 4, characterized in that, The determination module is further configured to at least one of the following: When the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, determine the second exponent as the reference exponent, and determine that both the truncation bit number and the shift bit number are 0; When the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, the sum of the first exponent and the preset threshold is determined as the reference exponent, the preset threshold is determined as the truncation bit number, and the difference between the preset threshold and the absolute value of the exponent difference is determined as the shift bit number.

8. The floating-point calculation device according to claim 7, characterized in that, The post-processing module includes a normalization module and a rounding module; The truncation module is further configured to: when the exponent difference is less than 0 and the absolute value of the exponent difference is greater than or equal to the preset threshold, transmit the first mantissa as the first discarded bit to the rounding module; And / or, when the exponent difference is less than 0 and the absolute value of the exponent difference is less than the preset threshold, transmit the remaining data bit segment intercepted from the first mantissa as the first discarded bit to the rounding module; The normalization module is configured to remove the leading 0s in the first mantissa sum to obtain a second mantissa sum; The rounding module is configured to perform a rounding process on the second mantissa sum based on the first discarded bit to obtain the target mantissa.

9. The floating-point calculation device according to claim 6 or 8, wherein The rounding module is further configured to: intercept the data bit segment starting from the highest bit and having a bit width equal to the effective bit width of the mantissa in the second mantissa sum as the mantissa sum to be rounded; starting from the lowest bit of the mantissa sum to be rounded, determine at least two consecutive rounding-related bits from the second mantissa sum; according to the first discarded bit, the rounding-related bits, and the second discarded bit after the rounding-related bits, perform a rounding process on the mantissa sum to be rounded according to the target rounding condition to obtain the target mantissa.

10. The floating-point calculation device according to claim 9, characterized in that, The rounding module is further configured to: Determine a target discarded value according to the first discarded bit and the second discarded bit; When the target discarded value and the rounding-related bits meet the target rounding condition, determine the rounding value as 1; or, when the target discarded value and the rounding-related bits do not meet the target rounding condition, determine the rounding value as 0; Add the mantissa sum to be rounded and the rounding value to obtain the target mantissa.

11. The floating-point calculation device according to claim 10, wherein The rounding-related bits include the least significant bit, the rounding bit, and the guard bit; the rounding process includes rounding to an even number; the target rounding condition includes: the rounding bit is 1, and at least one of the least significant bit, the target discarded value, and the guard bit is 1.

12. The floating-point calculation device according to any one of claims 2 to 8, 10, and 11, characterized in that, The target mantissa is the mantissa of the target floating-point number, and the floating-point calculation device further includes an exponent calculation module; The post-processing module is configured to perform a normalization process and a rounding process on the first mantissa sum to obtain a target mantissa and an exponent correction number; The exponent calculation module is configured to correct the reference exponent based on the exponent correction number to obtain the target exponent of the target floating-point number.

13. A floating-point calculation method for a processor, characterized in that, Including: Based on the truncation bit number, truncate the first mantissa of the first floating-point number to obtain the first addend, and based on the shift bit number, shift the second mantissa of the second floating-point number to obtain the second addend. The first addend and the second addend are two mantissas with aligned exponents; the truncation bit number and the shift bit number are determined based on the exponent difference between the first exponent of the first floating-point number and the second exponent of the second floating-point number; Add the first addend and the second addend to obtain the sum of the first mantissas; Perform normalization processing and rounding processing on the sum of the first mantissas to obtain the target mantissa.

14. An electronic device, characterized in that, It includes: A memory for storing a computer program; A processor for executing the floating-point calculation method as described in claim 13 when running the computer program.

15. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by the processor, the steps of the floating-point calculation method as described in claim 13 are implemented.

Citation Information

Patent Citations

  • Floating point addition operation device and method, electronic device and storage medium

    CN118519608A

  • Efficient Dual-path Floating-Point Arithmetic Operators

    US20220206747A1

  • Data processing method and apparatus, device, and storage medium

    US20250004711A1

  • Method for multiplying and accumulating operands, and device therefor

    WO2023231363A1

Cited By

  • Floating-point number multiplication method and floating-point number multiplication circuit

    CN120803394A

  • Data format conversion device and method, electronic equipment and computer storage medium

    CN120848840A

  • Processor, floating-point number processing method, storage medium and program product

    CN120929042A