An FPGA-based floating-point multiplier, calculation method and device

By combining the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm on the FPGA platform, the problem of difficulty in achieving high speed, low power consumption and high precision in traditional designs is solved, and efficient and low power consumption floating point multiplication operations are achieved.

CN119536684BActive Publication Date: 2025-06-17SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104500.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-06-17
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

Traditional floating point multiplier designs are difficult to meet the needs of high speed, low power consumption and high precision, especially in smart products and high performance computing applications.

Method used

The floating point multiplier based on FPGA is used, and the mantissa multiplication operation is performed by combining the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm, and floating point multiplication in IEEE-754 format is realized through the symbol calculation module, the exponential addition module and the result normalization module.

Benefits of technology

It significantly improves the speed of floating point multiplication operations, reduces hardware complexity and power consumption, and is suitable for high-speed application scenarios, while maintaining high precision and good compatibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119536684B_ABST
    Figure CN119536684B_ABST
Patent Text Reader

Abstract

This application relates to the field of floating-point multiplication, and specifically discloses a floating-point multiplier, a calculation method, and a device based on an FPGA. The floating-point multiplier includes: a sign calculation module for determining the sign of the target output floating-point number by means of an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number; an exponent addition module for adding the exponents of the first input floating-point number and the second input floating-point number and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number; a binary multiplier module that performs a multiplication operation based on the mantissa bit widths of the first input floating-point number and the second input floating-point number using the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to obtain the mantissa product of the target output floating-point number; and a result normalization module for performing a normalization operation based on the mantissa product. It can reduce the calculation delay and also reduce the percentage increase in the hardware area.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of floating-point multiplication, and particularly to a floating-point multiplier, a calculation method, and a device based on FPGA. Background Art

[0002] In recent years, with the development and iteration of intelligent products, electronic devices are currently being rapidly updated towards the direction of small size, low power consumption, and high speed. For electronic products, their speed depends on arithmetic operations. At the same time, in the current intelligent era, the applications of technologies such as multimedia, artificial intelligence, machine learning, deep learning, and the Internet of Things also involve huge basic arithmetic calculations. Floating-point multiplication is a key operation in high-performance computing applications. However, floating-point operations are not only complex, but also require more hardware area and power consumption compared to fixed-point multipliers. With the improvement of precision, the area, latency, and power consumption of floating-point multipliers will increase sharply. In the past few decades, people have been working hard to improve the performance of floating-point calculations.

[0003] Traditional floating-point multiplier design methods have certain limitations when facing the above problems and often cannot meet the requirements of high speed, low power consumption, and high precision. Although the IEEE 754 standard supports different floating-point formats, such as single-precision format, double-precision format, etc., the design of multipliers for each format faces similar challenges. Therefore, it is necessary to propose a new floating-point multiplier to overcome the deficiencies of the prior art. Summary of the Invention

[0004] To solve the above problems, this application proposes a floating-point multiplier, a calculation method, and a device based on FPGA. The floating-point multiplier is applied to the multiplication of floating-point numbers in the IEEE-754 format, including: a sign calculation module for determining the sign of the target output floating-point number by means of an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number; an exponent addition module for adding the exponents of the first input floating-point number and the second input floating-point number and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number; a binary multiplier module for performing a multiplication operation on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, using the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to obtain the mantissa product of the target output floating-point number; and a result normalization module for performing a normalization operation based on the mantissa product, where the normalization operation includes at least one of shifting the mantissa and adjusting the exponent value.

[0005] In one example, based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm are used to perform a multiplication operation on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number, which specifically includes: determining that the mantissa bit width meets a first preset condition, and using an improved Karatsuba algorithm to perform a divide-and-conquer multiplication operation on the first mantissa and the second mantissa until the mantissa product of the target output floating-point number is obtained, or the intermediate product result meets a second preset condition; when the intermediate product result meets the second preset condition, using the intermediate product result as the input of the improved Urdhva-Tiryagbhyam algorithm and continuing the multiplication operation until the mantissa product of the target output floating-point number is obtained.

[0006] In one example, using the improved Karatsuba algorithm to perform a divide-and-conquer multiplication operation on the first mantissa and the second mantissa specifically includes: determining the length ratio between the first mantissa and the second mantissa; determining the numerical distribution of the first mantissa and the second mantissa, where the numerical distribution is the number and digit positions of digits with a value of zero within adjacent preset digits; based on the length ratio and the numerical distribution, determining the splitting points of the first mantissa and the second mantissa, and splitting the first mantissa and the second mantissa according to the splitting points.

[0007] In one example, using the improved Karatsuba algorithm to perform a divide-and-conquer multiplication operation on the first mantissa and the second mantissa specifically includes: setting a parallel computing unit, a register, and a control logic module in the FPGA-based floating-point multiplier; the parallel computing unit is used to concurrently execute computing tasks within the same clock cycle; the computing tasks include multiplication calculation, addition calculation, and result merging; the register is used to temporarily store intermediate results and transfer data between different computing tasks; the control logic module is used to manage the start, completion, and error handling of each computing task.

[0008] In one example, before using the intermediate product result as the input of the improved Urdhva-Tiryagbhyam algorithm and continuing the multiplication operation, the binary multiplier module is further used to: when using the Urdhva-Tiryagbhyam algorithm, determine the number of calculation times corresponding to different preset computing tasks, where the preset computing tasks include pairs of the same values that appear at different time points or different positions and fixed digit combinations that appear in the multiplication task; store the input item groups of the preset computing tasks, the types of adders used, and the calculation results after hash mapping in a preset calculation table.

[0009] In one example, taking the intermediate product result as the input of the improved Urdhva-Tiryagbhyam algorithm and continuing with the multiplication operation specifically includes: determining the local correlation between the intermediate product result and each input item group in the pre-designed calculation table; if there is an input item group whose local correlation with the intermediate product result is higher than a first preset threshold and lower than a second preset threshold, then reading the adder model corresponding to the input item group in the pre-designed calculation table and using the adder model as the adder for the intermediate product result; if the local correlation is higher than the second preset threshold, then reading the corresponding calculation result in the pre-designed calculation table as the calculation result of a partial calculation task of the first mantissa and the second mantissa.

[0010] In one example, multiple types of adders are provided in the binary multiplier module, and the types of the adders include at least one of a carry-propagate adder, a carry-save adder, a carry-select adder, and a carry-lookahead adder.

[0011] In one example, the binary multiplier module is further configured to dynamically select the type of adder according to the bit width and data characteristics of the intermediate product result; the dynamically selecting the type of adder according to the bit width and data characteristics of the intermediate product result specifically includes: determining the carry probability of the intermediate product result based on a vector machine model to determine the first adder expectation of the intermediate product result; determining the second adder expectation of the intermediate product result based on the bit width of the intermediate product result; determining the time limit of the addition operation based on the delay requirement set by the user and determining the third adder expectation of the intermediate product result based on the time limit; determining the fourth adder expectation of the intermediate product result based on the quantity and type of available hardware resources; and determining the adder type corresponding to the intermediate product result based on the first adder expectation, the second adder expectation, the third adder expectation, and the fourth adder expectation.

[0012] The present application also provides a calculation method for a floating-point multiplier based on FPGA, including: determining the sign of the target output floating-point number through an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number; adding the exponents of the first input floating-point number and the second input floating-point number, and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number; based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, using the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to perform multiplication operations on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number to obtain the mantissa product of the target output floating-point number; performing a normalization operation based on the mantissa product, and the normalization operation includes at least one of shifting the mantissa and adjusting the exponent value.

[0013] The present application also provides a calculation device for a floating-point multiplier based on FPGA, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute: determining the sign of the target output floating-point number through an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number; adding the exponents of the first input floating-point number and the second input floating-point number, and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number; based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, using the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to perform multiplication operations on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number to obtain the mantissa product of the target output floating-point number; performing a normalization operation based on the mantissa product, and the normalization operation includes at least one of shifting the mantissa and adjusting the exponent value.

[0014] The method proposed by the present application can bring the following beneficial effects:

[0015] 1. By combining the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm, the time for mantissa multiplication operations can be significantly reduced, thereby improving the operation speed of the entire floating-point multiplier. Especially in high-speed application scenarios such as real-time signal processing and graphics rendering, this fast floating-point multiplication operation ability can greatly improve the performance and response speed of the system.

[0016] 2. Compared with the traditional floating-point multiplier design method, the algorithm combination adopted in this application reduces the number of multipliers and the hardware complexity. Especially when dealing with high-precision floating-point operations, a large number of complex hardware circuits are not required to implement the multiplication operation, thus reducing the chip area requirement. For hardware platforms such as FPGAs, this can make more effective use of limited resources and reduce the manufacturing cost.

[0017] 3. Due to the reduction of hardware complexity and the improvement of operation speed, the power consumption of the floating-point multiplier proposed in this application is also correspondingly reduced during operation. In power-sensitive applications such as mobile devices and embedded systems, this low-power feature can extend the battery life of the device and improve the practicality of the device.

[0018] 4. The floating-point multiplier design proposed in this application is based on the IEEE - 754 standard format and has good compatibility with existing computer systems and software. It can be easily integrated into various computer hardware platforms and application programs without large-scale modification of the existing system, thus reducing the application cost and the difficulty of popularization. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of this application and form a part of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0020] Figure 1 is a schematic structural diagram of a floating-point multiplier based on FPGA in an embodiment of this application;

[0021] Figure 2 is a schematic diagram of the calculation process inside a floating-point multiplier based on FPGA in an embodiment of this application;

[0022] Figure 3 is a schematic diagram of the calculation process inside a binary multiplier module in an embodiment of this application;

[0023] Figure 4 is a schematic diagram of the calculation process of a Karatsuba multiplier in an embodiment of this application;

[0024] Figure 5 is a schematic structural diagram of a 4x4 Urdhva-Tiryagbhyam multiplier hardware in an embodiment of this application;

[0025] Figure 6 is a schematic flowchart of a calculation method of a floating-point multiplier based on FPGA in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0027] The technical solutions provided by each embodiment of this application will be described in detail below in conjunction with the drawings.

[0028] As Figure 1 and Figure 2 shown, this application provides a floating-point multiplier based on FPGA, which is applied to the multiplication of floating-point numbers in IEEE-754 format. At this time, a floating-point number consists of a sign, an exponent, a mantissa, and an exponent base. For the multiplication operation of two floating-point numbers, the final result is obtained by separately calculating the products of the sign, exponent, and mantissa. The floating-point multiplier includes:

[0029] A sign calculation module for determining the sign of the target output floating-point number by means of an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number. Specifically, the sign calculation module determines the sign of the product according to the most significant bit (MSB) of the input floating-point number. For example, a simple XOR gate can be used as the sign calculator. If the sign bits of the two floating-point numbers are the same, the sign of the product is positive; if the sign bits are different, the sign of the product is negative. This sign calculation method based on an exclusive-OR gate is simple and efficient, does not require a complex circuit structure, can quickly determine the sign of the product, and saves time for the entire floating-point multiplication operation.

[0030] An exponent addition module for adding the exponents of the first input floating-point number and the second input floating-point number and subtracting the bias value of the corresponding format to obtain the exponent output of the target output floating-point number. Specifically, the exponent addition module is used to add the input exponents and subtract the bias value of the corresponding format to obtain the actual product exponent. In the IEEE-754 standard, the exponent of the single-precision format is 8 bits wide, and the bias value is 127; the exponent of the double-precision format is 11 bits wide, and the bias value is 1023. Since the calculation time of the mantissa multiplication operation is much longer than that of the exponent addition, a simple carry-ripple adder and a borrow-ripple subtractor can be used for the addition and subtraction operations of the exponents, which occupy less resources and can reduce the hardware complexity and cost.

[0031] The binary multiplier module, based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, uses the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to perform a multiplication operation on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number to obtain the mantissa product of the target output floating-point number. In floating-point multiplication, the most important and complex part is the mantissa multiplication. The multiplication operation takes more time than the addition operation. Moreover, as the number of bits increases, it consumes more area and time. To reduce the area of the multiplication operation and improve the efficiency of the multiplication operation, the present invention adopts a method combining the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm.

[0032] The result normalization module is used to perform a normalization operation based on the mantissa product. The normalization operation includes at least one of shifting the mantissa and adjusting the exponent value. In the IEEE-754 floating-point representation, there is a hidden bit in the mantissa, usually 1. This hidden bit has a special processing method during storage and operation. When the highest bit of the multiplication operation result is not the first bit to the left of the hidden bit, the result needs to be normalized to meet the requirements of the IEEE - 754 standard format. The operations of the result normalization module include shifting the mantissa and correspondingly adjusting the exponent value. If the highest bit of the multiplication operation result is not the first bit to the left of the hidden bit, the result is shifted to the left until the highest bit is 1 (in base 2, non-zero is 1). Each time a left shift operation is performed, the exponent value is increased by 1.

[0033] By combining the advantages of the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm, this application not only reduces the latency but also reduces the percentage increase in hardware area, having significant advantages over traditional methods.

[0034] Specifically, as Figure 3As shown, when the binary multiplier module performs floating-point multiplication calculations, since the Karatsuba algorithm is a divide-and-conquer algorithm suitable for handling high-precision multiplication operations, it is efficient when the input bit width is large. However, at lower bit widths, the Karatsuba algorithm is not efficient. To solve this problem, the present invention uses the Urdhva-Tiryagbhyam algorithm suitable for low bit widths at lower bit widths. Specifically, when it is determined that the tail bit width meets the first preset condition (such as when the bit widths of the operands are both greater than eight bits), the improved Karatsuba algorithm is used to perform divide-and-conquer multiplication on the first mantissa and the second mantissa until the mantissa product of the target output floating-point number is obtained, or the intermediate product result meets the second preset condition (such as when the bit widths of the operands are both eight bits). When the intermediate product result meets the second preset condition, the intermediate product result is used as the input of the improved Urdhva-Tiryagbhyam algorithm to continue the multiplication operation until the mantissa product of the target output floating-point number is obtained.

[0035] Here, the difference between the improved Karatsuba algorithm in this application and the existing Karatsuba algorithm is described:

[0036] The traditional Karatsuba algorithm uses a fixed splitting method to divide each of the two large integers into two parts, and then recursively calculates the results of each group of sub-problems. This method is suitable for most cases, but it may not be the optimal choice in some special cases. For example, when one operand is much larger than the other, directly applying Karatsuba may lead to unnecessary complexity.

[0037] Therefore, this application introduces a dynamic segmentation strategy to adaptively adjust the segmentation point according to the characteristics of the input data. For unbalanced operands, the larger number can be asymmetrically segmented so that the speed advantage of the addition operation can be maximally utilized in each recursive call. This approach reduces unnecessary recursive levels and decreases the overall latency. When performing asymmetric segmentation, the length ratio between the first mantissa (or the first operand in the calculation process) and the second mantissa (the second operand in the calculation process) can be determined, and the numerical distribution of the first mantissa and the second mantissa can be determined. The numerical distribution is the number and digit positions of the digits with a value of zero within adjacent preset digits. Then, based on the length ratio and the numerical distribution, the segmentation points of the first mantissa and the second mantissa are determined, and the first mantissa and the second mantissa are segmented according to the segmentation points. When determining the segmentation points, the corresponding segmentation expectations when the segmentation points are set at different positions can be calculated through preset weights, and appropriate segmentation points can be selected through the segmentation expectations. In hardware implementation, the segmentation points can be determined according to the performance indicators of the hardware (such as the latency of the adder, the latency of the multiplier, etc.). If the performance overheads of the multiplication and addition operations of operands with different lengths are known on the hardware, these information can be used to select the segmentation points to minimize the total calculation overhead. The asymmetric segmentation proposed in this application can reduce the recursive levels, thus simplifying the calculation process. For unbalanced operands, the asymmetric segmentation can segment the larger number into parts closer to the smaller number, thereby reducing the number of recursive calls. At the same time, it balances the addition load and improves the parallel processing ability. Through asymmetric segmentation, the addition tasks generated by each recursive call can be more evenly distributed, avoiding excessive or insufficient addition requirements in some stages. It also adapts to operands of different scales and avoids resource waste. For operands with a large length difference, symmetric segmentation may result in waste of some computing resources in processing redundant data. However, the asymmetric segmentation can be flexibly adjusted according to the actual input to ensure that each part can be processed most effectively.

[0038] Meanwhile, the traditional Karatsuba algorithm is executed serially, that is, it must wait for the previous step to complete before starting the next step. This will result in a long latency time in hardware implementation, especially when dealing with high-precision numerical values. Therefore, this application adopts a multi-stage pipeline structure, decomposes the tasks of each stage into smaller subtasks, and executes these subtasks in parallel. Specifically, multiple parallel units can be set in the hardware, each responsible for different sub-multiplication and addition operations. This can not only reduce the overall latency but also improve the throughput because multiple multiplications can be performed at the same time. Specifically, a parallel computing unit, a register, and a control logic module are set in the binary multiplier module of the floating-point multiplier provided in this application. The parallel computing unit here is used to concurrently execute computing tasks within the same clock cycle, and the computing tasks include multiplication calculation, addition calculation, and result merging. The register here is used to temporarily store intermediate results and transfer data between different computing tasks, and the control logic module is used to manage the start, completion, and error handling of each computing task.

[0039] In one embodiment, the existing Urdhva-Tiryagbhyam algorithm is essentially a method of vertical and cross multiplication, which quickly generates partial products through a series of rules. However, for some common patterns or recurring data combinations, recalculating each time will cause waste of resources.

[0040] Therefore, in the improved Urdhva-Tiryagbhyam algorithm in this application, when generating partial products, the local correlation of data can be used for pre-calculation and caching. At the same time, a pre-calculation table mechanism is added, and the partial multiplication results are pre-calculated and stored in a small lookup table. Whenever the same pattern of data is encountered, the result can be directly read from the lookup table to avoid repeated calculation. The lookup table can be customized according to the expected application scenario to cover the most frequently occurring data patterns. In addition, compression technology can be combined to reduce the size of the lookup table, thereby saving memory space. Specifically, when using the Urdhva-Tiryagbhyam algorithm, the number of calculations corresponding to different pre-designed computing tasks can be determined. The pre-designed computing tasks here include the same numerical pairs that appear at different time points or different positions and the fixed digit combinations that appear in the multiplication task; then the input item groups of the pre-designed computing tasks, the types of adders used, and the calculation results after hash mapping are stored in the pre-calculation table.

[0041] Further, when using the improved Urdhva-Tiryagbhyam algorithm for calculation, it is necessary to determine the local correlation between the intermediate product result and each input item group in the pre-designed calculation table. If there is an input item group whose local correlation with the intermediate product result is higher than the first preset threshold and lower than the second preset threshold, then read the adder model corresponding to the input item group in the pre-designed calculation table, and use the adder model as the adder for the intermediate product result. If the local correlation is higher than the second preset threshold, then read the corresponding calculation result in the pre-designed calculation table as the calculation result of the partial calculation task of the first mantissa and the second mantissa. Here, the calculation process of the local correlation is described: By traversing the two input operands, it can be determined whether there are repeated numerical values, the same bit-level patterns, or local digits with mathematical characteristics in the two operands. The bit-level pattern here refers to some bit combinations that frequently appear in multiplication operations. For example, the lower few bits or the higher few bits are always fixed. The mathematical characteristics here refer to the patterns formed based on mathematical laws, such as the input operand being the power or multiple relationship of another input operand or a historical operand.

[0042] In one embodiment, in order to accelerate the calculation process of the floating-point multiplier, in the floating-point multiplier provided by the present application, multiple types of adders are provided. When performing addition calculations, the adder can be selected based on the characteristics of the input operands. In one embodiment, the adder types provided by the present application include carry-propagation adder, carry-save adder, carry-select adder, and carry-lookahead adder.

[0043] And each adder corresponds to different characteristics of the input operands. Specifically, the carry-propagation adder is suitable for small-scale or low-latency-requirement addition operations. The carry-save adder is suitable for accumulating the results of multiple addition operations to reduce the carry-propagation delay. The carry-select adder is suitable for large-bit-width data and can quickly generate the final result. The carry-lookahead adder is suitable for latency-sensitive applications, such as high-speed operations. When selecting an adder, a set of selection criteria can be defined to determine the most suitable adder type for the current situation. Here, an example of the criteria is given: Dynamically select the type of adder according to the bit width and data characteristics of the intermediate product result, specifically including:

[0044] Determine the carry probability of addition for the intermediate product result based on a vector machine model to determine the first adder expectation of the intermediate product result; determine the second adder expectation of the intermediate product result based on the bit width of the intermediate product result; determine the time limit for the addition operation based on the delay requirement set by the user, and determine the third adder expectation of the intermediate product result based on the time limit; determine the fourth adder expectation of the intermediate product result based on the quantity and type of available hardware resources; finally, determine the adder type corresponding to the intermediate product result based on the first adder expectation, the second adder expectation, the third adder expectation, and the fourth adder expectation. Finally, the dynamic selection process of the adder type can be managed by using a decision tree or a finite state machine. At this time, in the decision tree or the finite state machine, each node or state represents a possible selection condition, and the edges or transitions represent the transition from one condition to another.

[0045] Figure 6 It is a schematic flowchart of a calculation method of a floating-point multiplier based on FPGA provided by one or more embodiments of this specification. This process can be executed by a computing device in the corresponding field, and some input parameters or intermediate results in the process allow manual intervention and adjustment to help improve accuracy.

[0046] The implementation of the analysis method involved in the embodiments of this application can be a terminal device or a server, and this application does not make special restrictions on this. For the convenience of understanding and description, the following embodiments will be described in detail taking the server as an example.

[0047] It should be noted that this server can be a single device or a system composed of multiple devices, that is, a distributed server, and this application does not make specific limitations on this.

[0048] As Figure 6 shown, the embodiments of this application provide a calculation method of a floating-point multiplier based on FPGA, including:

[0049] S601: Determine the sign of the target output floating-point number through an exclusive-OR gate and the sign bits of the first input floating-point number and the second input floating-point number.

[0050] S602: Add the exponents of the first input floating-point number and the second input floating-point number, and subtract the bias value of the corresponding format to obtain the exponent output of the target output floating-point number.

[0051] S603: Based on the mantissa bit widths of the first input floating-point number and the second input floating-point number, use the Karatsuba algorithm and the Urdhva-Tiryagbhyam algorithm to perform a multiplication operation on the first mantissa of the first input floating-point number and the second mantissa of the second input floating-point number to obtain the mantissa product of the target output floating-point number.

[0052] S604: Perform a normalization operation based on the product of the mantissas, where the normalization operation includes at least one of performing a shift operation on the mantissa and adjusting the exponent value.

[0053] Each embodiment in this application is described in a progressive manner. For the parts that are the same or similar among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.

[0054] The devices and media provided in the embodiments of this application correspond one by one to the methods. Therefore, the devices and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.

[0055] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0056] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0057] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the steps in the process Figure 1 a process or processes and / or blocks Figure 1 steps of the functions specified in a block or blocks.

[0059] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0060] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.

[0061] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0062] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an ……" does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0063] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A floating-point multiplier based on FPGA, characterized in that: Applies to floating-point multiplication in IEEE-754 format, including: A sign calculation module, used for determining the sign of the target output floating point number through an XOR gate and the sign bits of the first input floating point number and the second input floating point number; An exponential addition module, used for adding the exponents of the first input floating point number and the second input floating point number, and subtracting a deviation value of a corresponding format to obtain an exponential output of the target output floating point number; A binary multiplier module, based on the mantissa bit widths of the first input floating point number and the second input floating point number, uses a Karatsuba algorithm and an Urdhva-Tiryagbhyam algorithm to perform a multiplication operation on a first mantissa of the first input floating point number and a second mantissa of the second input floating point number to obtain a mantissa product of the target output floating point number; a result normalization module, configured to perform a normalization operation based on the mantissa product, wherein the normalization operation includes at least one of shifting the mantissa and adjusting the exponent value; The method of performing a multiplication operation on a first mantissa of the first input floating point number and a second mantissa of the second input floating point number by using a Karatsuba algorithm and an Urdhva-Tiryagbhyam algorithm based on the mantissa bit widths of the first input floating point number and the second input floating point number specifically includes: Determine that the mantissa bit width satisfies a first preset condition, and use an improved Karatsuba algorithm to perform a divide-and-conquer multiplication operation on the first mantissa and the second mantissa until the mantissa product of the target output floating-point number is obtained, or an intermediate product result satisfies a second preset condition; When the intermediate product result satisfies the second preset condition, the intermediate product result is used as an input of the improved Urdhva-Tiryagbhyam algorithm, and the multiplication operation is continued until the mantissa product of the target output floating-point number is obtained; The step of using the improved Karatsuba algorithm to perform a divide-and-conquer multiplication operation on the first mantissa and the second mantissa specifically includes: determining a length ratio between the first mantissa and the second mantissa; Determine the numerical distribution of the first mantissa and the second mantissa, wherein the numerical distribution is the number of digits with zero values ​​and the digit positions in adjacent preset digits; Determine a segmentation point between the first mantissa and the second mantissa based on the length ratio and the numerical distribution, and segment the first mantissa and the second mantissa according to the segmentation point; The intermediate product result is used as an input of the improved Urdhva-Tiryagbhyam algorithm, and before continuing the multiplication operation, the binary multiplier module is further used for: When using the Urdhva-Tiryagbhyam algorithm, determining the number of calculations corresponding to different preset calculation tasks, wherein the preset calculation tasks include the same value pairs appearing at different time points or different positions and the fixed digit combination appearing in the multiplication task; The input item group of the preset calculation task, the adder type used, and the calculation result after hash mapping are stored in a preset calculation table; The step of using the intermediate product result as an input of the improved Urdhva-Tiryagbhyam algorithm and continuing the multiplication operation specifically includes: Determine the local correlation between the intermediate product result and each input item group in the preset calculation table; If there is a local correlation between an input item group and the intermediate product result that is higher than a first preset threshold and lower than a second preset threshold, reading an adder model corresponding to the input item group in the preset calculation table, and using the adder model as an adder for the intermediate product result; If the local correlation is higher than a second preset threshold, reading a corresponding calculation result from the preset calculation table as a calculation result of a part of the calculation tasks of the first mantissa and the second mantissa; The binary multiplier module is provided with multiple types of adders, and the types of the adders include at least one of a carry propagation adder, a carry save adder, a carry select adder, and a carry look-ahead adder; the binary multiplier module is also used to dynamically select the type of the adder according to the bit width and data characteristics of the intermediate product result; The method of dynamically selecting the type of adder according to the bit width and data characteristics of the intermediate product result specifically includes: Determine the addition carry probability of the intermediate product result based on the vector machine model to determine the first adder expectation of the intermediate product result; determining a second adder expectation of the intermediate product result based on a bit width of the intermediate product result; determining a time limit for the addition operation based on a latency requirement set by a user, and determining a third adder expectation of the intermediate product result based on the time limit; determining a fourth adder expectation of the intermediate product result based on the amount and type of available hardware resources; Based on the first adder expectation, the second adder expectation, the third adder expectation, and the fourth adder expectation, the adder type corresponding to the intermediate product result is determined.

2. The FPGA-based floating-point multiplier according to claim 1, wherein: Using the improved Karatsuba algorithm, a divide-and-conquer multiplication operation is performed on the first mantissa and the second mantissa, specifically including: A parallel computing unit, a register, and a control logic module are provided in the FPGA-based floating-point multiplier; The parallel computing unit is used to concurrently execute computing tasks in the same clock cycle; the computing tasks include multiplication calculation, addition calculation and result merging; The registers are used to temporarily store intermediate results and transfer data between different computing tasks; The control logic module is used to manage the initiation, completion and error handling of each computing task.