CUDA-based fast exponential function approximate calculation method

By using floating-point multiplication-addition fusion instructions and parallel computing rules in a GPU environment, efficient approximate calculation of the exponential function is achieved, which solves the problems of computational efficiency and accuracy in existing technologies and improves the efficiency and reliability of approximate calculation of the exponential function.

CN120670719APending Publication Date: 2025-09-19GUANGZHOU WERIDE TECH LTD CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510575431.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

In the existing technology, in the GPU parallel computing environment, the exponential function calculation method is difficult to fully utilize the parallel computing capabilities of CUDA, and there are problems such as difficulty in ensuring accuracy and high instruction overhead.

Method used

A single floating-point multiplication-addition fusion instruction simultaneously completes multiplication, type conversion, and addition operations, generates intermediate calculation results, and performs floating-point to integer numerical conversion and element-by-element difference calculation to determine whether the relative error is within the preset tolerance range.

Benefits of technology

It significantly reduces instruction overhead and calculation steps, improves the efficiency and reliability of approximate calculations of exponential functions, optimizes the numerical conversion process, and ensures calculation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670719A_ABST
    Figure CN120670719A_ABST
Patent Text Reader

Abstract

The invention relates to the field of computer data processing, and discloses a CUDA (compute unified device architecture)-based fast exponential function approximate calculation method, which is used for improving the calculation efficiency while ensuring the precision. The CUDA-based fast exponential function approximate calculation method comprises the following steps: acquiring a to-be-calculated input floating-point number set; for each input floating-point number x, executing floating-point multiplication and addition fusion operation, simultaneously completing multiplication operation of the input floating-point number x and a preset coefficient a, type conversion from a floating point to an integer and addition operation of adding a preset offset b through a single floating-point multiplication and addition instruction, and generating a corresponding intermediate calculation result; performing numerical value conversion operation from a floating point to an integer on each intermediate calculation result to obtain a corresponding approximate exponential function value; and performing element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp (x), and judging whether a relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range or not.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer data processing, and in particular to a CUDA-based fast exponential function approximate calculation method. Background Art

[0002] Exponential function calculation has wide applications in scientific computing, machine learning, signal processing and other fields.

[0003] In the prior art, to improve the computational efficiency of exponential functions, several approximate computation methods have been proposed, such as those based on Taylor series expansion, table lookup, or piecewise linear approximation. However, these methods still have shortcomings in the GPU parallel computing environment: Taylor series expansion requires multiple multiplication and addition operations, making it difficult to fully utilize the parallel computing capabilities of CUDA; table lookup is limited by memory access latency and storage overhead; and piecewise linear approximation, while ensuring accuracy, may introduce additional branch decisions, affecting the execution efficiency of thread bundles. Furthermore, existing methods typically require multiple instructions to complete multiplication, addition, and type conversion operations, failing to fully utilize the advantages of floating-point multiply-add (FMA) operations in modern GPU architectures. Summary of the Invention

[0004] The present invention provides a CUDA-based fast exponential function approximate calculation method to solve the technical problem in the prior art that it is difficult to perform exponential function approximate calculation efficiently while ensuring accuracy.

[0005] The first aspect of the present invention provides a CUDA-based fast exponential function approximate calculation method, comprising: obtaining an input floating-point number set to be calculated, the input floating-point number set containing at least one input floating-point number x; for each input floating-point number x, performing a floating-point multiplication-addition fusion operation, and simultaneously completing the multiplication operation of the input floating-point number x and a preset coefficient a, the floating-point to integer type conversion, and the addition operation of adding a preset offset b through a single floating-point multiplication-addition instruction to generate a corresponding intermediate calculation result; performing a floating-point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value; performing element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and judging whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

[0006] In a feasible embodiment, the step of simultaneously completing the multiplication operation of the input floating-point number x and the preset coefficient a, the floating-point to integer type conversion, and the addition operation of adding the preset offset b through a single floating-point multiplication-addition instruction to obtain the corresponding intermediate calculation result includes: using CUDA's floating-point multiplication-addition fusion instruction to multiply the input floating-point number x and the preset coefficient a; based on the multiplication result, simultaneously performing the floating-point to integer type conversion operation; adding the result after type conversion to the preset offset b to obtain the corresponding intermediate calculation result.

[0007] In a feasible implementation, performing a floating-point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value includes: checking whether the numerical value of each intermediate calculation result is within a valid range that can be converted to an integer; if the data is within the valid range that can be converted to an integer, using a floating-point to integer conversion function to convert the corresponding intermediate calculation result into an integer form; and performing a bit operation or scaling operation on the converted integer to obtain an approximate exponential function value that meets the requirements.

[0008] In a feasible implementation, when the input floating-point number set contains multiple input floating-point numbers x, the element-by-element difference calculation is performed on each approximate exponential function value and the corresponding standard exponential function exp(x), and whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range is determined, including: using CUDA's parallel computing rules to simultaneously calculate the difference between multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x); calculating the ratio of each difference to the absolute value of the corresponding standard exponential function exp(x) to obtain each relative error; comparing the calculated relative error with the preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirement.

[0009] In a feasible implementation, the method utilizes CUDA's parallel computing rules to simultaneously calculate the differences between multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x), including: processing the input floating-point number x to be calculated in batches, each batch containing multiple floating-point numbers; allocating CUDA thread blocks to each batch of floating-point numbers, each thread block being responsible for calculating the differences between the approximate exponential function values ​​of one or more floating-point numbers and the standard exponential function exp(x); and within each thread block, utilizing CUDA's thread-level parallelism to simultaneously calculate multiple differences.

[0010] In a feasible implementation, the ratio of each difference value to the absolute value of the standard exponential function exp(x) is used to obtain each relative error, including: calculating the difference between each approximate exponential function value and the standard exponential function exp(x), and synchronously obtaining the absolute value of the corresponding standard exponential function exp(x); using the parallel division operation instructions provided by CUDA to divide each difference value by the absolute value of the corresponding standard exponential function exp(x), and completing the division operations of all differences simultaneously in a parallel computing environment to obtain each relative error.

[0011] In a feasible implementation manner, the calculated relative error is compared with a preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirement, including: comparing each relative error with the upper and lower limits of the preset tolerance range; based on the comparison result, determining whether each approximate exponential function value meets the accuracy requirement, and outputting a corresponding judgment result.

[0012] The second aspect of the present invention provides a CUDA-based fast exponential function approximate calculation device, including: an acquisition module, used to obtain an input floating-point number set to be calculated, wherein the input floating-point number set includes at least one input floating-point number x; a generation module, used to perform a floating-point multiplication and addition fusion operation for each input floating-point number x, and simultaneously complete the multiplication operation of the input floating-point number x and a preset coefficient a, the type conversion of the floating-point to integer, and the addition operation of adding a preset offset b through a single floating-point multiplication and addition instruction to generate a corresponding intermediate calculation result; a conversion module, used to perform a floating-point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value; a judgment module, used to perform element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and judge whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

[0013] In a feasible implementation, the generation module is specifically used to: use CUDA's floating-point multiplication and addition fusion instruction to multiply the input floating-point number x with the preset coefficient a; based on the multiplication result, simultaneously perform a floating-point to integer type conversion operation; add the result after type conversion to the preset offset b to obtain the corresponding intermediate calculation result.

[0014] In a feasible embodiment, the conversion module is specifically used to: check whether the numerical value of each intermediate calculation result is within the valid range that can be converted to an integer; if the data is within the valid range that can be converted to an integer, use the floating-point to integer conversion function to convert the corresponding intermediate calculation result into an integer form; perform bit operations or scaling operations on the converted integer to obtain an approximate exponential function value that meets the requirements.

[0015] In a feasible embodiment, the judgment module includes: a first calculation unit, used to use CUDA's parallel computing rules to simultaneously calculate the difference between multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x); a second calculation unit, used to obtain each relative error by the ratio of each difference to the absolute value of the standard exponential function exp(x); and a judgment unit, used to compare the calculated relative error with a preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirements.

[0016] In a feasible embodiment, the first computing unit is used to: process the input floating-point number x to be calculated in batches, each batch containing multiple floating-point numbers; assign a CUDA thread block to each batch of floating-point numbers, each thread block is responsible for calculating the difference between the approximate exponential function value of one or more floating-point numbers and the standard exponential function exp(x); within each thread block, utilize CUDA's thread-level parallelism to simultaneously calculate multiple differences.

[0017] In a feasible embodiment, the second computing unit is specifically used to: calculate the difference between each approximate exponential function value and the standard exponential function exp(x), and simultaneously obtain the absolute value of the corresponding standard exponential function exp(x); use the parallel division operation instructions provided by CUDA to divide each difference by the absolute value of the corresponding standard exponential function exp(x), and complete the division operation of all differences simultaneously in a parallel computing environment to obtain each relative error.

[0018] In a feasible implementation, the judgment unit is specifically used to: compare each relative error with the upper and lower limits of a preset tolerance range; based on the comparison result, judge whether each approximate exponential function value meets the accuracy requirement, and output a corresponding judgment result.

[0019] A third aspect of the present invention provides an electronic device comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the electronic device executes the above-mentioned CUDA-based fast exponential function approximation calculation method.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned CUDA-based fast exponential function approximation calculation method.

[0021] In the technical solution provided by the present invention, an input floating-point number set to be calculated is obtained, and the input floating-point number set includes at least one input floating-point number x; for each input floating-point number x, a floating-point multiplication-addition fusion operation is performed, and the multiplication operation of the input floating-point number x and a preset coefficient a, the floating-point to integer type conversion, and the addition operation with a preset offset b are simultaneously completed through a single floating-point multiplication-addition instruction to generate a corresponding intermediate calculation result; a floating-point to integer numerical conversion operation is performed on each intermediate calculation result to obtain a corresponding approximate exponential function value; an element-by-element difference calculation is performed on each approximate exponential function value and the corresponding standard exponential function exp(x), and it is determined whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range. In an embodiment of the present invention, a single floating-point multiply-add fusion instruction is used to implement the combined operation of multiplication, type conversion and addition, which significantly reduces the instruction overhead and calculation steps, and completes the rapid generation of approximate exponential functions with a simpler instruction stream. While ensuring the calculation accuracy, the numerical conversion process is optimized, the loss of calculation efficiency caused by additional operations is reduced, and the reliability of the processing flow is further improved by directly performing numerical conversion and error judgment on the intermediate results. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of an embodiment of a method for fast exponential function approximation calculation based on CUDA in an embodiment of the present invention;

[0023] Figure 2 1 is a schematic diagram of another embodiment of a method for fast exponential function approximation calculation based on CUDA in an embodiment of the present invention;

[0024] Figure 3 Schematic diagram of an embodiment of a CUDA-based fast exponential function approximate calculation device in accordance with an embodiment of the present invention;

[0025] Figure 4 1 is a schematic diagram of another embodiment of a CUDA-based fast exponential function approximate calculation device according to an embodiment of the present invention;

[0026] Figure 5 FIG. 1 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] An embodiment of the present invention provides a CUDA-based fast exponential function approximate calculation method, which simultaneously completes multiplication, type conversion and addition operations through a single floating-point multiplication-addition fusion instruction, reduces instruction overhead and calculation steps, and improves the efficiency of exponential function approximate calculation while ensuring accuracy.

[0028] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0029] It is understood that the execution subject of the present invention can be a CUDA-based fast exponential function approximation calculation device, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking the server as the execution subject as an example.

[0030] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 In one embodiment of the present invention, a method for fast exponential function approximation calculation based on CUDA includes:

[0031] 101. Obtain an input floating-point number set to be calculated, where the input floating-point number set includes at least one input floating-point number x;

[0032] The set of input floating-point numbers to be calculated can be obtained through a user interaction interface. Specifically, a graphical user interface or command line interface can be developed to allow users to directly input floating-point numbers or import floating-point data into the program by pasting, dragging, etc. The data entered by the user can be verified in real time and stored in a data structure in memory to form an input floating-point number set.

[0033] 102. For each input floating-point number x, perform a fused floating-point multiplication-addition operation. This operation uses a single floating-point multiplication-addition instruction to simultaneously perform the multiplication of the input floating-point number x by a preset coefficient a, convert the floating-point number to an integer, and perform the addition operation by adding the preset offset b, thereby generating a corresponding intermediate calculation result.

[0034] In the CUDA programming environment, the __constant__ memory modifier is first used to store the preset coefficient a and the preset offset b in the constant memory of the CUDA device. Next, for each floating-point number x in the input floating-point number set, the CUDA intrinsic function __fmaf_rn(x,a,b) is called. This function performs the floating-point multiplication of x and a in a single instruction, and then performs a floating-point addition on the multiplication result and the offset b. Before the addition, if implicit floating-point to integer conversion is required, it is achieved by adjusting the values ​​of a and b or utilizing the hardware truncation / rounding features. Finally, the intermediate calculation result is generated and stored.

[0035] 103. Perform a floating point to integer conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value;

[0036] Directly call the intrinsic function provided by CUDA to perform the conversion. Specifically, for the intermediate calculation results of single-precision floating-point numbers, use CUDA's built-in function to convert the floating-point number to an integer. This function determines the corresponding integer value based on the binary representation of the floating-point number, especially its exponent and mantissa bits, according to the nearest rounding rule. During the conversion, the CUDA hardware automatically handles special floating-point value cases such as NaN and infinity to ensure that the conversion process does not cause exceptions. The converted integer value is directly used as the approximate exponential function value.

[0037] 104. Perform element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and determine whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

[0038] For each approximate exponential function value, the corresponding standard exponential function value exp(x) is calculated by calling the exp function in the CUDA math library; then, an element-by-element difference calculation is performed, that is, the difference between the approximate exponential function value and exp(x) is obtained by subtraction; then, in order to obtain the relative error, the calculated difference is divided by the absolute value of exp(x), and a floating-point division instruction is used here to ensure the accuracy of the calculation; after obtaining the relative error, it is compared with a preset tolerance range, and the calculated relative error is compared with the preset tolerance range to determine whether the approximate exponential function value meets the accuracy requirements. If the relative error is within the preset tolerance range, the approximate value is considered valid; otherwise, it is marked as invalid or further processed.

[0039] In an embodiment of the present invention, a set of input floating-point numbers to be calculated is obtained, where the input floating-point number set includes at least one input floating-point number x; for each input floating-point number x, a floating-point multiplication-addition fusion operation is performed, and the multiplication operation of the input floating-point number x and a preset coefficient a, the floating-point to integer type conversion, and the addition operation of the preset offset b are simultaneously completed through a single floating-point multiplication-addition instruction to generate a corresponding intermediate calculation result; a floating-point to integer numerical conversion operation is performed on each intermediate calculation result to obtain a corresponding approximate exponential function value; an element-by-element difference calculation is performed on each approximate exponential function value and the corresponding standard exponential function exp(x), and it is determined whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range. This improves the efficiency and reliability of the exponential function approximate calculation while ensuring accuracy.

[0040] See also Figure 2 Another embodiment of the CUDA-based fast exponential function approximate calculation method in the embodiment of the present invention includes:

[0041] 201. Obtain an input floating-point number set to be calculated, where the input floating-point number set includes at least one input floating-point number x;

[0042] Build a data input module that is compatible with various data sources, such as files, network streams, or direct user input. For file input, design a file parser that supports common formats such as CSV and TXT, reads data line by line, and parses it into the input floating-point number x. For network streams, implement a network data receiving interface to capture and convert transmitted data in real time.

[0043] 202. For each input floating-point number x, perform a fused floating-point multiplication-addition operation, using a single floating-point multiplication-addition instruction to simultaneously complete the multiplication of the input floating-point number x by a preset coefficient a, the floating-point to integer type conversion, and the addition of the preset offset b, to generate a corresponding intermediate calculation result.

[0044] Using CUDA's floating-point multiplication-addition fusion instruction, the input floating-point number x is multiplied by the preset coefficient a; based on the multiplication result, the floating-point to integer type conversion operation is performed at the same time; the result after type conversion is added to the preset offset b to obtain the corresponding intermediate calculation result.

[0045] For each input floating-point number x, call the floating-point multiplication and addition fusion instruction provided by CUDA, for example

[0046] __fmaf_rn, this instruction accepts three parameters: an input floating-point number x, a preset coefficient a, and a preset offset b. When the instruction is executed, a floating-point multiplication operation is first performed on x and a to obtain the product result. Subsequently, no explicit floating-point to integer type conversion is performed directly. Instead, the type conversion is implicitly processed by leveraging the characteristics of the CUDA hardware and subsequent instructions or operations. For example, by designing specific scaling factors and offset strategies, the product result is made suitable for the subsequent integer representation in terms of numerical range, or in subsequent steps, the type conversion effect is achieved through bit operations, rounding functions, etc. Based on the multiplication operation and implicit type conversion, the processed result is further subjected to floating-point addition operation with the preset offset b, and finally the corresponding intermediate calculation result is obtained.

[0047] 203. Perform a floating point to integer conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value;

[0048] Check whether the value of each intermediate calculation result is within the valid range that can be converted to an integer. If the data is within the valid range that can be converted to an integer, use the floating-point to integer conversion function to convert the corresponding intermediate calculation result into integer form. Perform bit operations or scaling operations on the converted integers to obtain an approximate exponential function value that meets the requirements.

[0049] Each intermediate calculation result is traversed, and a series of comparison operations are performed to check whether its value is within the valid range that can be converted to an integer. This range is usually determined by the representation capability of the integer type. For example, for a 32-bit integer, the valid range may be from -2 to the power of 31 to 2 to the power of 31 minus 1, while excluding special values ​​such as NaN and infinity. If the intermediate calculation result is within the valid range, the floating-point to integer conversion function provided by CUDA is called, such as __float2int_rn. This function uses the nearest rounding mode to convert the floating-point number to an integer. During the conversion process, the CUDA hardware automatically processes the exponent and mantissa bits of the floating-point number to ensure the accuracy of the conversion. After obtaining the intermediate result in integer form, the integer is subjected to bit operations or scaling operations according to the preset scaling factor and offset. The bit operations may include left or right shifts to adjust the precision or range of the value; the scaling operation is performed by multiplying or dividing by a specific factor to make the integer result meet the requirements of the approximate exponential function value. These operations need to be designed based on specific application scenarios and accuracy requirements. For example, in the approximate calculation of exponential functions, it may be necessary to select an appropriate scaling factor based on the characteristics of the exponential function to ensure the accuracy of the approximation.

[0050] 204. When the input floating-point number set includes multiple input floating-point numbers x, using the parallel computing rules of CUDA, the differences between the multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x) are simultaneously calculated;

[0051] The input floating-point number x to be calculated is divided into batches, each batch contains multiple floating-point numbers; a CUDA thread block is assigned to each batch of floating-point numbers, and each thread block is responsible for calculating the difference between the approximate exponential function value of one or more floating-point numbers and the standard exponential function exp(x); within each thread block, multiple differences are calculated simultaneously using CUDA's thread-level parallelism.

[0052] Based on the size of the input floating-point number set and the GPU's computing power, the input floating-point number x is divided into several batches, each containing multiple floating-point numbers, to ensure full utilization of the GPU's parallel computing resources. For each batch of floating-point numbers, a corresponding number of thread blocks are allocated in CUDA, with each thread block responsible for processing one or more floating-point numbers. The allocation of thread blocks must take into account the GPU's hardware architecture and thread block size limitations to optimize computational efficiency. Within each thread block, CUDA's thread-level parallelism is leveraged to allocate one or more threads to each floating-point number, ensuring that multiple difference calculations can be performed simultaneously. Each thread first reads its corresponding input floating-point number x and calls the exp function in the CUDA math library to calculate the standard exponential function value exp(x). Simultaneously, the thread reads a pre-computed approximate exponential function value from global memory or shared memory. Next, the thread performs a subtraction operation to calculate the difference between the approximate exponential function value and exp(x). To further improve computational efficiency, shared memory can be used within the thread block to cache frequently accessed data and reduce global memory access latency. In addition, attention must be paid to inter-thread synchronization and avoiding memory access conflicts to ensure that each thread can correctly read and process its corresponding data. Through batch processing and thread-level parallelism, the scheme can efficiently process a large number of input floating-point numbers and realize parallel calculation of multiple differences.

[0053] 205. Calculate the ratio of each difference to the absolute value of its corresponding standard exponential function exp(x) to obtain the relative error;

[0054] Calculate the difference between each approximate exponential function value and the standard exponential function exp(x), and simultaneously obtain the absolute value of the corresponding standard exponential function exp(x); use the parallel division operation instructions provided by CUDA to divide each difference by the absolute value of the corresponding standard exponential function exp(x), and complete the division operation of all differences simultaneously in a parallel computing environment to obtain the relative errors.

[0055] The absolute value of each standard exponential function, exp(x), is calculated and stored. This can be done synchronously with the calculation of exp(x) or by using an additional kernel function for batch processing. To ensure efficient parallel computing, the data layout should be organized so that each difference and its corresponding absolute value of exp(x) are stored contiguously in memory, facilitating fast thread access. Subsequently, a CUDA kernel function is launched, assigning a thread to each difference. CUDA's thread-level parallelism is leveraged for parallel processing. In each thread, the thread index is used to locate the corresponding difference and absolute value of exp(x). The parallel division instruction provided by CUDA is then called to perform the division operation and obtain the relative error. CUDA's division instruction is optimized for efficient execution in a parallel computing environment, ensuring that the division operation for all differences is completed simultaneously.

[0056] 206. Compare the calculated relative error with a preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirement.

[0057] Compare each relative error with the upper and lower limits of the preset tolerance range; based on the comparison results, determine whether each approximate exponential function value meets the accuracy requirements and output the corresponding judgment result.

[0058] The upper and lower limits of the preset tolerance range should be set as constants or configuration parameters and loaded into constant memory or texture memory during program initialization for fast access. A CUDA kernel function is called to allocate a thread for each relative error, leveraging CUDA's thread-level parallelism for parallel comparison. In each thread, the corresponding relative error is located using the thread index and then compared sequentially with the upper and lower limits of the preset tolerance range. The comparison operation can be implemented using the conditional judgment instructions provided by CUDA to ensure efficient execution. If the relative error is less than or equal to the upper limit and greater than or equal to the lower limit, the approximate exponential function value is judged to meet the accuracy requirements; otherwise, it is judged to not meet the requirements. The judgment result can be stored in global memory via the thread's output variable or atomic operation for subsequent processing or output. By calling parallel comparison and conditional judgment instructions, this solution can efficiently and accurately determine whether each approximate exponential function value meets the accuracy requirements.

[0059] In the embodiment of the present invention, efficient approximate calculation is achieved by obtaining a set of input floating-point numbers and performing a floating-point multiplication-addition fusion operation on each floating-point number to generate an intermediate calculation result, and then obtaining an approximate exponential function value through numerical conversion. When processing multiple input floating-point numbers, the parallel computing rules of CUDA are used to simultaneously calculate the differences and relative errors between multiple approximate values ​​and the standard exponential function, which greatly improves the calculation efficiency. The relative error is compared with the preset tolerance range and the judgment result is output. It can accurately evaluate whether the accuracy of the approximate value meets the requirements. While ensuring the accuracy, the efficiency and reliability of the approximate calculation of the exponential function are improved.

[0060] The above describes the fast exponential function approximate calculation method based on CUDA in the embodiment of the present invention. The following describes the fast exponential function approximate calculation device based on CUDA in the embodiment of the present invention. Figure 3 In one embodiment of the present invention, a fast exponential function approximate calculation device based on CUDA includes:

[0061] An acquisition module 301 is configured to acquire an input floating-point number set to be calculated, where the input floating-point number set includes at least one input floating-point number x;

[0062] A generation module 302 is configured to perform a fused floating-point multiplication-addition operation for each input floating-point number x, performing a multiplication operation of the input floating-point number x by a preset coefficient a, performing a floating-point to integer type conversion, and performing an addition operation by adding a preset offset b, using a single floating-point multiplication-addition instruction, thereby generating a corresponding intermediate calculation result.

[0063] The conversion module 303 is used to perform a floating point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value;

[0064] The judgment module 304 is used to perform element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and judge whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

[0065] In an embodiment of the present invention, a set of input floating-point numbers to be calculated is obtained, where the input floating-point number set includes at least one input floating-point number x; for each input floating-point number x, a floating-point multiplication-addition fusion operation is performed, and the multiplication operation of the input floating-point number x and a preset coefficient a, the floating-point to integer type conversion, and the addition operation of the preset offset b are simultaneously completed through a single floating-point multiplication-addition instruction to generate a corresponding intermediate calculation result; a floating-point to integer numerical conversion operation is performed on each intermediate calculation result to obtain a corresponding approximate exponential function value; an element-by-element difference calculation is performed on each approximate exponential function value and the corresponding standard exponential function exp(x), and it is determined whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range. This improves the efficiency and reliability of the exponential function approximate calculation while ensuring accuracy.

[0066] See also Figure 4 Another embodiment of the CUDA-based fast exponential function approximate calculation device in the embodiment of the present invention includes:

[0067] An acquisition module 301 is configured to acquire an input floating-point number set to be calculated, where the input floating-point number set includes at least one input floating-point number x;

[0068] A generation module 302 is configured to perform a fused floating-point multiplication-addition operation for each input floating-point number x, performing a multiplication operation of the input floating-point number x by a preset coefficient a, performing a floating-point to integer type conversion, and performing an addition operation by adding a preset offset b, using a single floating-point multiplication-addition instruction, thereby generating a corresponding intermediate calculation result.

[0069] The conversion module 303 is used to perform a floating point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value;

[0070] The judgment module 304 is used to perform element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and judge whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

[0071] Optionally, the generating module 302 may be specifically configured to:

[0072] Using CUDA's floating-point multiplication-addition fusion instruction, the input floating-point number x is multiplied by the preset coefficient a; based on the multiplication result, the floating-point to integer type conversion operation is performed at the same time; the result after type conversion is added to the preset offset b to obtain the corresponding intermediate calculation result.

[0073] Optionally, the conversion module 303 may be specifically configured to:

[0074] Check whether the value of each intermediate calculation result is within the valid range that can be converted to an integer. If the data is within the valid range that can be converted to an integer, use the floating-point to integer conversion function to convert the corresponding intermediate calculation result into integer form. Perform bit operations or scaling operations on the converted integers to obtain an approximate exponential function value that meets the requirements.

[0075] Optionally, the judgment module 304 includes:

[0076] A first computing unit 3041 is configured to simultaneously compute differences between a plurality of approximate exponential function values ​​and a corresponding standard exponential function exp(x) using CUDA's parallel computing rules;

[0077] The second calculation unit 3042 is used to calculate the ratio of each difference value to the absolute value of the standard exponential function exp(x) to obtain each relative error;

[0078] The judging unit 3043 is configured to compare the calculated relative error with a preset tolerance range to judge whether each approximate exponential function value meets the accuracy requirement.

[0079] Optionally, the first computing unit 3041 can be specifically used to: process the input floating-point number x to be calculated in batches, each batch containing multiple floating-point numbers; assign a CUDA thread block to each batch of floating-point numbers, each thread block is responsible for calculating the difference between the approximate exponential function value of one or more floating-point numbers and the standard exponential function exp(x); within each thread block, utilize the thread-level parallelism of CUDA to simultaneously calculate multiple differences.

[0080] Optionally, the second computing unit 3042 can be specifically used to: calculate the difference between each approximate exponential function value and the standard exponential function exp(x), and simultaneously obtain the absolute value of the corresponding standard exponential function exp(x); use the parallel division operation instructions provided by CUDA to divide each difference by the absolute value of the corresponding standard exponential function exp(x), and complete the division operation of all differences simultaneously in a parallel computing environment to obtain each relative error.

[0081] Optionally, the judgment unit 3043 can be specifically used to: compare each relative error with the upper limit and lower limit of a preset tolerance range; based on the comparison result, judge whether each approximate exponential function value meets the accuracy requirement, and output a corresponding judgment result.

[0082] In the embodiment of the present invention, efficient approximate calculation is achieved by obtaining a set of input floating-point numbers and performing a floating-point multiplication-addition fusion operation on each floating-point number to generate an intermediate calculation result, and then obtaining an approximate exponential function value through numerical conversion. When processing multiple input floating-point numbers, the parallel computing rules of CUDA are used to simultaneously calculate the differences and relative errors between multiple approximate values ​​and the standard exponential function, which greatly improves the calculation efficiency. The relative error is compared with the preset tolerance range and the judgment result is output. It can accurately evaluate whether the accuracy of the approximate value meets the requirements. While ensuring the accuracy, the efficiency and reliability of the approximate calculation of the exponential function are improved.

[0083] above Figure 3 and Figure 4 The CUDA-based fast exponential function approximation calculation device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The electronic device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0084] See also Figure 5 As shown, the electronic device includes a processor 500 and a memory 501 , wherein the memory 501 stores machine executable instructions that can be executed by the processor 500 , and the processor 500 executes the machine executable instructions to implement the above-mentioned CUDA-based fast exponential function approximation calculation method.

[0085] Further, Figure 5 The electronic device shown further includes a bus 502 and a communication interface 503 , and the processor 500 , the communication interface 503 and the memory 501 are connected via the bus 502 .

[0086] Among them, the memory 501 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), for example, at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 503 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 502 can be an ISA bus, a PCI bus, or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0087] The processor 500 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in the processor 500. The processor 500 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present disclosure may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 501 , and the processor 500 reads the information in the memory 501 and completes the method steps of the aforementioned embodiment in combination with its hardware.

[0088] The present invention also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the CUDA-based fast exponential function approximation calculation method in the above-mentioned embodiments.

[0089] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the steps of the CUDA-based fast exponential function approximation calculation method.

[0090] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0091] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.

[0092] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A CUDA-based fast exponential function approximate calculation method, characterized in that: The CUDA-based fast exponential function approximate calculation method includes: Obtain an input floating-point number set to be calculated, wherein the input floating-point number set includes at least one input floating-point number x; For each input floating-point number x, perform a fused floating-point multiplication-addition operation. This multiplication of the input floating-point number x by the preset coefficient a, the floating-point to integer conversion, and the addition of the preset offset b are all completed using a single floating-point multiplication-addition instruction, generating the corresponding intermediate calculation result. Perform a floating point to integer conversion operation on each intermediate calculation result to obtain the corresponding approximate exponential function value; An element-by-element difference calculation is performed on each approximate exponential function value and the corresponding standard exponential function exp(x), and it is determined whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

2. The CUDA-based fast exponential function approximate calculation method according to claim 1, characterized in that: The step of simultaneously completing the multiplication operation of the input floating-point number x and the preset coefficient a, the floating-point to integer type conversion, and the addition operation of the preset offset b by a single floating-point multiplication-addition instruction to obtain the corresponding intermediate calculation result includes: Use CUDA's floating-point multiplication and addition fusion instruction to multiply the input floating-point number x by the preset coefficient a; Based on the multiplication result, the floating point to integer type conversion operation is performed at the same time; Add the result after type conversion to the preset offset b to obtain the corresponding intermediate calculation result.

3. The CUDA-based fast exponential function approximate calculation method according to claim 1, characterized in that: The performing of a floating point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value includes: Check whether the value of each intermediate calculation result is within the valid range that can be converted to an integer; If the data is in the valid range that can be converted to an integer, the corresponding intermediate calculation result is converted to integer form using the floating-point to integer conversion function; Perform bit manipulation or scaling operations on the converted integer to obtain an approximate exponential function value that meets the requirements.

4. The CUDA-based fast exponential function approximate calculation method according to claim 1, characterized in that: When the input floating-point number set includes multiple input floating-point numbers x, performing element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and determining whether a relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range, includes: Using CUDA's parallel computing rules, the differences between multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x) are calculated simultaneously; Calculate the ratio of each difference to the absolute value of its corresponding standard exponential function exp(x) to obtain the relative error; The calculated relative error is compared with the preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirement.

5. The CUDA-based fast exponential function approximate calculation method according to claim 4, characterized in that: The method utilizes CUDA's parallel computing rules to simultaneously calculate the differences between multiple approximate exponential function values ​​and the corresponding standard exponential function exp(x), including: The input floating-point number x to be calculated is divided into batches, each batch contains multiple floating-point numbers; Assign a CUDA thread block to each batch of floating-point numbers. Each thread block is responsible for calculating the difference between the approximate exponential function value of one or more floating-point numbers and the standard exponential function exp(x). Within each thread block, multiple differences are calculated simultaneously using CUDA's thread-level parallelism.

6. The CUDA-based fast exponential function approximate calculation method according to claim 4, characterized in that: The calculation of the ratio of each difference to the absolute value of the corresponding standard exponential function exp(x) to obtain each relative error includes: Calculate the difference between each approximate exponential function value and the standard exponential function exp(x), and simultaneously obtain the absolute value of the corresponding standard exponential function exp(x); Using the parallel division operation instructions provided by CUDA, each difference is divided by the absolute value of the corresponding standard exponential function exp(x). The division operation of all differences is completed simultaneously in a parallel computing environment to obtain the relative errors.

7. The CUDA-based fast exponential function approximate calculation method according to claim 4, characterized in that: The step of comparing the calculated relative error with a preset tolerance range to determine whether each approximate exponential function value meets the accuracy requirement includes: Comparing each relative error with the upper and lower limits of a preset tolerance range; According to the comparison results, it is judged whether each approximate exponential function value meets the accuracy requirements and the corresponding judgment result is output.

8. A CUDA-based fast exponential function approximate calculation device, characterized in that: The CUDA-based fast exponential function approximate calculation device includes: An acquisition module, configured to acquire an input floating-point number set to be calculated, wherein the input floating-point number set includes at least one input floating-point number x; A generation module is used to perform a floating-point multiplication-addition fusion operation for each input floating-point number x, and to simultaneously complete the multiplication operation of the input floating-point number x and a preset coefficient a, the floating-point to integer type conversion, and the addition operation of the preset offset b through a single floating-point multiplication-addition instruction, and generate a corresponding intermediate calculation result; A conversion module, configured to perform a floating point to integer numerical conversion operation on each intermediate calculation result to obtain a corresponding approximate exponential function value; The judgment module is used to perform element-by-element difference calculation on each approximate exponential function value and the corresponding standard exponential function exp(x), and judge whether the relative error between each approximate exponential function value and the corresponding standard exponential function value is within a preset tolerance range.

9. An electronic device, characterized in that: The electronic device comprises: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instruction in the memory to enable the electronic device to execute the CUDA-based fast exponential function approximation calculation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by a processor, a CUDA-based fast exponential function approximation calculation method according to any one of claims 1 to 7 is implemented.