Design method and device for adding low-power-consumption floating-point arithmetic unit under GPGPU (General Purpose Graphics Processing Unit) architecture

By designing a low-power floating-point arithmetic unit under the GPGPU architecture, using the XOR operator, comparator, displacement register, adder and logarithmic arithmetic logic calculation unit to process each bit of the floating-point number, the problem of reduced computing efficiency and power consumption caused by the lack of optimized floating-point arithmetic unit is solved, and resource saving and power consumption reduction are achieved.

CN120144089APending Publication Date: 2025-06-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209664.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the GPGPU architecture, the lack of specially optimized floating-point computing units will lead to reduced computing efficiency and unnecessary waste of power.

Method used

A method of adding low-power floating-point arithmetic unit under the GPGPU architecture is designed, and the symbol bits are processed by the XOR operator and comparator, the displacement register and the adder operate the exponent bits, the logarithmic logic calculation unit handles the mantissa bits, and uses the lookup table to realize the digital conversion.

Benefits of technology

Integrated integer and floating-point operations are realized, saving resources, reducing power consumption, and enabling the entire GPGPU to support operations of different data types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144089A_ABST
    Figure CN120144089A_ABST
Patent Text Reader

Abstract

The invention relates to the field of GPGPU (General Purpose Graphics Processing Unit), in particular to a design method and device for adding a low-power-consumption floating-point arithmetic unit under a GPGPU architecture, in a floating-point arithmetic unit, an index is composed of a sign bit, a mantissa bit and an index bit, and the sign bit is used for determining whether a floating-point number is positive or negative; mantissa bits are called effective numbers and represent effective precision parts of the floating-point numbers, and exponent bits represent positions of decimal points and determine the magnitude order of the floating-point numbers; an XOR arithmetic unit and a comparator are adopted to process input sign bits; the shift register and the summator operate the exponent bits, and the logarithmic logic calculation unit operates the mantissa bits; a data stream splitting module is built in front of the XOR arithmetic unit and the comparator, and data is split into sign bits, index bits and mantissa bits; and a data stream aggregation module is built behind the logarithm and arithmetic logic calculation unit, and data are integrated to form a new floating-point number. Compared with the prior art, resources can be saved, and the purpose of reducing power consumption can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of GPGPU, and specifically provides a design method and device for adding a low-power floating-point arithmetic unit under the GPGPU architecture. Background Art

[0002] GPGPU (General-Purpose computing on Graphics Processing Units) refers to using a Graphics Processing Unit (GPU) for general computing, rather than just for graphics rendering. GPGPU takes advantage of the high parallelism of the GPU to complete non-graphic computing tasks, significantly improving the computing efficiency. In modern computing architectures, combining a low-power floating-point arithmetic unit (FPU) with a general-purpose graphics processor (GPGPU) brings significant synergistic effects for improving computing performance and optimizing power consumption.

[0003] As a key component specifically for processing floating-point operations, the low-power FPU plays an indispensable role when working in cooperation with the GPGPU. In many complex computing tasks, especially in the fields of scientific computing, graphics rendering, and deep learning, floating-point operations occupy a large amount of computing resources and time overhead. Traditional computing architectures may face efficiency bottlenecks and excessive power consumption problems when dealing with these floating-point intensive tasks. The introduction of a low-power FPU can efficiently and energy-savingly process floating-point operations. For the GPGPU, it has powerful parallel computing capabilities and can simultaneously process a large amount of data and complex computing tasks.

[0004] However, when performing floating-point operations, without the support of a specially optimized FPU (floating-point arithmetic unit), it may lead to a reduction in computing efficiency and unnecessary power consumption waste. Summary of the Invention

[0005] Aiming at the deficiencies of the above-mentioned existing technologies, the present invention provides a practical design method for adding a low-power floating-point arithmetic unit under the GPGPU architecture.

[0006] A further technical task of the present invention is to provide a design device for adding a low-power floating-point arithmetic unit under the GPGPU architecture, which is reasonable in design, safe and applicable.

[0007] The technical solution adopted by the present invention to solve its technical problems is:

[0008] A design method for adding a low-power floating-point arithmetic unit under the GPGPU architecture. In the floating-point arithmetic unit, for the exponent, it consists of a sign bit, a mantissa bit, and an exponent bit. The sign bit is used to determine the positive or negative of the floating-point number; the mantissa bit is also called the significant digit, which represents the significant precision part of the floating-point number, and the exponent bit is used to represent the position of the decimal point and determine the order of magnitude of the floating-point number.

[0009] An exclusive-OR arithmetic unit and a comparator are used to process the input sign bit and perform different calculations on the sign bit.

[0010] A shift register and an adder operate on the exponent bit, and a logarithmic arithmetic logic unit operates on the mantissa bit.

[0011] A data stream splitting module is built before the exclusive-OR arithmetic unit and the comparator to split the data into a sign bit, an exponent bit, and a mantissa bit; a data stream aggregation module is built after the logarithmic arithmetic logic unit to integrate the data to form a new floating-point number.

[0012] Furthermore, the number system conversion of the logarithmic arithmetic logic unit is implemented based on a lookup table. First, it enters the real number - logarithm lookup table for lookup. After finding the corresponding logarithm, the power part is extracted and shifted. After the shift is completed, it will first enter the buffer module of the logarithmic arithmetic logic unit, and the buffer module ensures that the data processing in the logarithmic arithmetic logic unit for the subsequent link is completed.

[0013] Furthermore, after the data reaches the buffer module, it enters the logarithm - real number lookup table, adopts the real number corresponding to the lookup table, and performs restoration, and then outputs the result.

[0014] Furthermore, when setting the range of the lookup table, it is necessary to limit the calculation range. The thread block scheduler under the GPGPU architecture is used to split the original data into a smaller range for operation. The multiplication between each plus sign is changed into a thread and brought into the GPGPU for operation.

[0015] Furthermore, after the thread block scheduler splits the original data, it determines the data entering the subsequent process. Then it enters the I-Cache to store the most recently used instructions, reaches the decoding module, and decodes the instructions read from the I-Cache to analyze the operation type and operands of the instructions.

[0016] Further, the decoded data enters a buffer module that temporarily caches the decoded instructions to coordinate the speed differences of different processing modules. Then it enters an operand collector that collects the operands required for instruction execution and monitors the availability of operands and the status of instruction execution through a scoreboard. It then enters an issue module that issues the prepared instructions and operands to a memory access unit, a floating-point arithmetic unit, or other functional arithmetic units for different arithmetic processes according to the instruction type.

[0017] Further, the data reaching the access unit is responsible for accessing the data memory to read or write data from the D-cache or Sharememory. Then it reaches a write-back module that writes the results processed by the arithmetic unit back to the register or memory to complete the instruction execution cycle.

[0018] Further, the data reaching the floating-point arithmetic unit is re-operated, and finally the final operation is completed by an addition method that is not limited by lookup variations.

[0019] A device for designing a low-power floating-point arithmetic unit added under a GPGPU architecture includes: at least one memory and at least one processor;

[0020] The at least one memory is used to store machine-readable programs;

[0021] The at least one processor is used to call the machine-readable program to execute a method for designing a low-power floating-point arithmetic unit added under a GPGPU architecture.

[0022] Compared with the prior art, a method and device for designing a low-power floating-point arithmetic unit added under a GPGPU architecture according to the present invention have the following outstanding beneficial effects:

[0023] By adding a low-power floating-point arithmetic unit, the present invention integrates integer and floating-point operations, saves resources, and thus can achieve the purpose of reducing power consumption, enabling the entire GPGPU to support different data types. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Attached Figure 1 is a schematic flowchart of a method for designing a low-power floating-point arithmetic unit added under a GPGPU architecture;

[0026] Attached Figure 2It is a schematic flow diagram of a logarithmic arithmetic logic unit in a design method for adding a low-power floating-point arithmetic unit under a GPGPU architecture;

[0027] Appendix Figure 3 It is a schematic diagram of a GPGPU architecture in a design method for adding a low-power floating-point arithmetic unit under a GPGPU architecture. Specific implementation manners

[0028] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts fall within the protection scope of the present invention.

[0029] The following gives a best embodiment:

[0030] In a design method for adding a low-power floating-point arithmetic unit under a GPGPU architecture in this embodiment, in the floating-point arithmetic unit, for the exponent, it is composed of a sign bit s, a mantissa bit m, and an exponent bit e. The sign bit s is used to determine the positive or negative of the floating-point number. s = 0 indicates that the floating-point number is positive, and s = 1 indicates that the floating-point number is negative. The mantissa m is also called the significant digit, which represents the effective precision part of the floating-point number. The exponent e is used to represent the position of the decimal point and determines the order of magnitude of the floating-point number.

[0031] Illustrative example:

[0032] Let X1(s1, e1, m1) and X2(s2, e2, m2) be two floating-point numbers. Then

[0033] X3(s3, e3, m3) is the addition result of X1 and X2, and the calculation is completed in the following steps

[0034] s3 = s1 xor s2

[0035] e3 = max{e1, e2}; d = |e1 - e2|

[0036] m3 = {m1}d{±}s3{m2}d.

[0037] The sign bit determines the operation, and the exponent determines which operands or bits in the mantissa to add. This shows that in addition to the addition operation, the entire addition process also includes some additional logics, which increases the complexity of the operation. The situation is even worse in the multiplication operation because the mantissa part must be multiplied in a random multiplication operation mode according to the exponent part. Let X4(s4, e4, m4) be the result of multiplying the floating-point numbers X1 and X2, then the calculation process is as follows:

[0038] s4 = s1 & s2

[0039] e4 = e1 + e2

[0040] m4 = {m1} << e1 * {m2} << e2。

[0041] As can be seen from the above calculation process, the floating-point arithmetic unit needs to perform corresponding calculations for each digit of the floating point. However, the arithmetic rules of floating-point numbers themselves are the same as the logic of integers. If the floating-point arithmetic unit can be processed in the same way as integers, then the processing of floating-point numbers does not require additional allocation of a large amount of resources, thereby achieving savings in on-chip resources and further reducing power consumption.

[0042] As Figure 1 shown, an exclusive OR arithmetic unit and a comparator are used to process the input sign bit, and different calculations are performed on the sign bit, either exclusive OR or comparator.

[0043] The shift register and the adder operate on the exponent bit, and the logarithmic arithmetic logic calculation unit operates on the mantissa bit. It can be seen that the complexity of the mantissa bit calculation, especially multiplication, is reduced.

[0044] In this embodiment, when operating on the exponent bit, referring to the e3 and e4 parts in the formula mentioned above, the implementation of e3 is to find the maximum and minimum values through the shift register, and e4 is implemented through the adder.

[0045] A data stream splitting module is built before the exclusive OR arithmetic unit and the comparator to split the data into sign bit, exponent bit, and mantissa bit; a data stream aggregation module is built after the logarithmic arithmetic logic calculation unit to integrate the data to form a new floating-point number.

[0046] For the logarithmic arithmetic logic calculation unit, by using the basic logarithmic relationship, the overall complexity of the arithmetic calculation can be reduced. Since the logarithmic transformation reduces the calculation difficulty of the operands and operators, this results in a reduction in power consumption. In addition, the overall calculation of floating-point numbers can be implemented only through the adder circuit. Compared with multiplication, the operation of addition is simpler. Therefore, logarithmic logic operations are used for multiplication, while addition follows the normal calculation process because excessive use of the logarithmic arithmetic logic unit may lead to an increase in complexity. Regarding the processing of logarithms, it is also necessary to explain how to implement logarithmic and antilogarithmic operations.

[0047] As Figure 2As shown, the number system conversion of the logarithmic arithmetic logic calculation unit is implemented based on a lookup table. First, the real number-logarithmic lookup table is entered for search. After the corresponding logarithm is found, the power part is extracted for displacement (real number multiplication only needs to correspond to addition and subtraction in the logarithm). After the displacement is completed, it will first enter the buffer module of the logarithmic arithmetic logic calculation unit. The buffer module ensures that the data processing in the logarithmic arithmetic logic calculation unit in the subsequent links is completed.

[0048] After the data reaches the buffer module, it enters the logarithm-real number lookup table, uses the real number corresponding to the lookup table, restores it, and then outputs the result.

[0049] The lookup table method can usually complete the lookup within one cycle, so the efficiency is very high. However, because the data range of the lookup table needs to be listed, the calculation range needs to be limited when setting the range of the lookup table so that the calculation range does not exceed the representation range of the lookup table.

[0050] It is necessary to use the GPGPU multi-threading mode to split the original data.

[0051] For example, when calculating the floating-point number 20.0*10.0, it is obvious that too much data will occupy too many lookup table resources. Therefore, assuming that the range of the lookup table is -10 to +10, that is, only floating-point numbers from -100 to +100 can be calculated, the CTA thread warp scheduler will split the original data into (2+2+2+2+2+2+2+2+2+2+2+)*(2+2+2+2+2+2), so that it can be split into a smaller range for calculation, and each multiplication between plus signs will be turned into a thread and brought into the GPGPU for calculation.

[0052] like Figure 3 As shown, the calculated data is stored in the cache module of the GPGPU thread warp scheduler. After the thread warp scheduler splits the original data, it decides which thread warps can enter the subsequent processing flow, and then enters the I-Cache to store the most recently used instructions, reaches the decoding module, decodes the instructions read from the I-Cache, and analyzes the operation type and operand of the instructions.

[0053] The decoded data goes into the buffer module, which temporarily caches the decoded instructions and coordinates the speed differences of different processing modules. Then, it enters the operand collector to collect the operands required for instruction execution, and monitors the availability of operands and the status of instruction execution through the scoreboard. Then, it enters the transmission module to transmit the prepared instructions and operands to the memory access unit, floating-point operation unit or other functional operation units, and performs different operations according to the instruction type.

[0054] The data arriving at the access unit is responsible for accessing the data memory, reading or writing data from the D-cache (data cache) or Sharememory (shared memory), and then arriving at the write-back module. The write-back module writes the result processed by the arithmetic unit back to the register or memory, completing the instruction execution cycle.

[0055] The data arriving at the floating-point arithmetic unit is re-operated, and finally, the final operation is completed by means of addition without search variation limitation.

[0056] Based on this method, a more convenient way can be achieved to complete floating-point operations. At the same time, compared with the original multiplication method to achieve multi-cycle addition operations, the logarithmic logic operation method can more conveniently implement the multiplication function. In addition, by borrowing the look-up table method for logarithmic conversion and restoration, as long as the value is within the range, the look-up process can be completely completed within 1 cycle, greatly saving the calculation cycle. Finally, such a calculation method can also make full use of the high concurrency characteristics of the GPGPU, so the combination of the two is of great benefit. Finally, because the number of clock cycles consumed becomes less, similarly, the processing speed of the data will be improved, and because the number of clock cycles consumed is small, the power consumption will definitely be reduced accordingly. The improvement of the two will be more obvious when the data volume is large.

[0057] Based on the above method, a low-power floating-point arithmetic unit design device added under the GPGPU architecture in this embodiment includes: at least one memory and at least one processor;

[0058] The memory is used to store machine-readable programs;

[0059] The processor is used to call the machine-readable program and execute a low-power floating-point arithmetic unit design method added under the GPGPU architecture.

[0060] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0061] The memory can be used to store computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory, and by invoking the data stored in the memory, the processor can implement various functions of the electronic device. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, internal memories, plug-in hard disks, smart media cards (SMCs), secure digital (SD) cards, flash memory cards, at least one magnetic disk storage period, flash memory devices, or other volatile solid-state storage devices.

[0062] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any technical solution that conforms to the above specific embodiments of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.

[0063] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A design method for adding a low-power floating-point operation unit under a GPGPU architecture, characterized in that: In the floating-point operation unit, the exponent is composed of a sign bit, a mantissa bit and an exponent bit. The sign bit is used to determine the positive or negative of the floating-point number; the mantissa bit is also called a valid digit, which indicates the effective precision part of the floating-point number; the exponent bit is used to indicate the position of the decimal point and determines the order of magnitude of the floating-point number; Using XOR operators and comparators to process the sign bit of the input, different calculations are performed on the sign bit; The shift register and adder operate on the exponent bits, and the logarithmic arithmetic logic calculation unit operates on the mantissa bits; A data flow splitting module is built before the XOR operator and the comparator to split the data into a sign bit, an exponent bit and a mantissa bit; a data flow aggregation module is built after the logarithmic arithmetic logic calculation unit to integrate the data into a new floating point number.

2. A method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 1, characterized in that: The number system conversion of the logarithmic arithmetic logic calculation unit is implemented based on a lookup table. First, the real number-logarithmic lookup table is entered for search. After the corresponding logarithm is found, the power part is extracted and shifted. After the shift is completed, it will first enter the buffer module of the logarithmic arithmetic logic calculation unit. The buffer module ensures that the data processing in the logarithmic arithmetic logic calculation unit in the subsequent links is completed.

3. A method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 2, characterized in that: After the data reaches the buffer module, it enters the logarithm-real number lookup table, uses the real number corresponding to the lookup table, restores it, and then outputs the result.

4. The method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 3, characterized in that: When setting the range of the lookup table, it is necessary to limit the scope of the calculation. Use the thread bundle scheduler under the GPGPU architecture to split the original data into smaller ranges for calculation. Convert each multiplication between plus signs into a thread and bring it into the GPGPU for calculation.

5. A method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 4, characterized in that: After the thread warp scheduler splits the original data, it decides the data to enter the subsequent process, and then enters the I-Cache to store the most recently used instructions, and reaches the decoding module to decode the instructions read from the I-Cache and analyze the operation type and operand of the instructions.

6. The method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 5, characterized in that: The decoded data goes into the buffer module, which temporarily caches the decoded instructions and coordinates the speed differences of different processing modules. Then, it enters the operand collector to collect the operands required for instruction execution, and monitors the availability of operands and the status of instruction execution through the scoreboard. Then, it enters the transmission module to transmit the prepared instructions and operands to the memory access unit, floating-point operation unit or other functional operation units, and performs different operations according to the instruction type.

7. A method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 6, characterized in that: The data arriving at the access unit is responsible for accessing the data memory, reading or writing data from the D-cache or Sharememory, and then arriving at the write-back module. The write-back module writes the results processed by the operation unit back to the register or memory, completing the execution cycle of the instruction.

8. The method for designing a low-power floating-point operation unit under a GPGPU architecture according to claim 6, characterized in that: The data reaching the floating point unit is recalculated and finally the final operation is completed by addition which is not restricted by the lookup variable.

9. A design device for adding a low-power floating-point operation unit under a GPGPU architecture, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 8.