GPU-based parallel sliding window DFT implementation method and GPU
By dividing data blocks on the GPU and iteratively computing the sliding window DFT, the sliding window DFT algorithm in the prior art has solved the problem of large memory space occupied on the GPU and high data alignment requirements, and more efficient GPU resource utilization is achieved.
Patent Information
- Application Number
- CN202510219367.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
When the existing sliding window DFT algorithm is implemented on GPU, it requires a large video memory space and has high requirements for data alignment and window size, resulting in low utilization of GPU computing resources.
By dividing the data to be computed into multiple data blocks, and at each calculation, according to the Fourier transform results of the first window and the newly added data of each window, the Fourier transform results of the remaining windows on each data block is iteratively calculated, so as to achieve the multiplexing of the existing FFT results and reduce the data loading and calculation amount.
It reduces the usage of GPU video memory space, reduces the requirements for data alignment and window size, and improves the utilization rate of GPU computing resources.
Smart Images

Figure CN120067504A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method for implementing parallel sliding window DFT based on GPU and a GPU. Background Art
[0002] Sliding window discrete Fourier transform (DFT) is a technology widely used in the fields of signal processing, speech analysis, radar signal processing, etc. Its core idea is to divide the signal into multiple overlapping or non-overlapping windows, and perform DFT calculations within each window.
[0003] Traditional sliding window DFT algorithms are often implemented through a digital signal processor (DSP) or a central processing unit (CPU). With the increase in data scale and the improvement of real-time requirements, DSP and CPU cannot meet the application needs. Therefore, it is necessary to consider the implementation of the sliding window DFT algorithm on a graphics processing unit (GPU).
[0004] The GPU uses a general-purpose parallel computing architecture and has a large number of arithmetic and logic units (ALUs) and high-throughput memory units. Since the GPU memory throughput is much smaller than the GPU floating-point operation data throughput, when implementing the sliding window DFT on the GPU using the traditional method, the FFT in the cuFFT library is usually called to calculate different windows; and the sliding window DFT algorithm belongs to an I / O-intensive task. In the GPU computing architecture, overlapping windows means that some data needs to be read repeatedly. Since the computing power of the GPU device is often much greater than its memory bandwidth, there is a problem of inefficient use of computing resources when executing the sliding window DFT algorithm, that is, the GPU is limited by the slow speed of input / output operation data, and most of the computing resources of the GPU are in a waiting state and the utilization rate is low.
[0005] Therefore, another sliding window DFT algorithm is implemented using the parallel FFT method, that is, the data of multiple windows are batch-loaded into the GPU, and the FFT library is called once for calculation. This method reduces the number of I / O operations, but requires a large amount of video memory space and has high requirements for data alignment and window size. Summary of the Invention
[0006] The embodiments of the present invention provide a method for implementing parallel sliding window DFT based on GPU and a GPU, which can solve the problems that the current sliding window DFT algorithm requires a large amount of video memory space and has high requirements for data alignment and window size.
[0007] In a first aspect, an implementation method of parallel sliding window DFT based on GPU provided by an embodiment of the present invention includes:
[0008] Divide the data to be calculated into multiple data blocks;
[0009] Calculate the Fourier transform result of the first window on each data block;
[0010] Based on the DFT iterative model, according to the Fourier transform result of the first window and the new data of each window, iteratively calculate the Fourier transform results of the remaining windows on each data block.
[0011] In a second aspect, an embodiment of the present invention provides a GPU, including:
[0012] A data division unit, which is used to divide the data to be calculated into multiple data blocks;
[0013] A first calculation unit, which is used to calculate the Fourier transform result of the first window on each data block;
[0014] A second calculation unit, which is used to iteratively calculate the Fourier transform results of the remaining windows on each data block based on the DFT iterative model, according to the Fourier transform result of the first window and the new data of each window.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including a GPU and a memory. The memory is used to store a computer program; the GPU can be used to execute the calculator program (instructions) stored in the memory to implement the method in the first aspect above.
[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed, the method in the first aspect above can be implemented.
[0017] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows: According to the method provided by the present invention, by calculating the FFT result of the current window based on the FFT result of the first window and the new data of each window during each calculation, instead of using all the data of the current window for FFT calculation each time, the reuse of existing FFT results can be achieved, the amount of data read each time, the amount of calculation for each calculation, and the number of data loading times can be reduced, thereby reducing the video memory space occupied by the GPU; since only the new data is read each time, the requirements for data alignment and window size are relatively low. Description of the Drawings
[0018] Figure 1The implementation flowchart of a method for implementing parallel sliding window DFT based on GPU provided by an embodiment of the present invention;
[0019] Figure 2 The schematic diagram of a storage structure of a GPU provided by an embodiment of the present invention;
[0020] Figure 3 The schematic diagram of a scenario for calculating the Fourier transform results of each data block based on the DFT iterative model provided by an embodiment of the present invention;
[0021] Figure 4 The schematic diagram of a structure of a GPU provided by an embodiment of the present invention;
[0022] Figure 5 The comparison schematic diagram of the implementation process of a sliding window FFT method provided by an embodiment of the present invention;
[0023] Figure 6 The comparison schematic diagram of an acceleration effect provided by an embodiment of the present invention;
[0024] Figure 7 The schematic diagram of a structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0025] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0026] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0027] It should also be understood that the term " / and" as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0028] As used in the specification of the present invention and the appended claims, the term "if" may be construed, depending on the context, as "when", "once", "in response to determining", or "in response to detecting". Similarly, the phrases "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, to mean "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]".
[0029] In addition, in the description of the specification of the present invention and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0030] Reference to "one embodiment" or "some embodiments" or the like described in the specification of the present invention means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0031] The present invention will be further described in detail below in conjunction with specific embodiments, but the implementation manners of the present invention are not limited thereto.
[0032] Figure 1 The figure shows an implementation flowchart of a method for implementing parallel sliding window DFT based on GPU provided by an embodiment of the present invention. By way of example and not limitation, this method can be applied to an electronic device including a GPU, such as a mobile phone, a laptop computer, a supercomputer, etc. This method may include steps S101 - S103, and each step will be described below:
[0033] S101, divide the data to be calculated into multiple data blocks.
[0034] In one possible implementation manner, referring to Figure 2 , after receiving the data to be calculated, a memory block may be first allocated in the global memory to store the data to be calculated and the Fourier transform result calculated later. Since the I / O speed of the global memory is limited, after dividing the data to be calculated into multiple data blocks, the data blocks may be stored in the shared memory of different cores.
[0035] In one example, the cudaMemcpy() function can be used to transfer the data to be processed from the host memory to the global memory of the GPU.
[0036] In one possible implementation, the data to be processed can be divided into multiple data blocks according to the number of currently available cores, the data length of the data to be processed, the window size, and the sliding step.
[0037] In one example, the data length of each data block can satisfy the following formula:
[0038]
[0039] where Dk is the data length of the k-th data block, floor(·) is the floor function, L is the sliding step, K is the total number of data blocks (i.e., the total number of available cores), M is the data length of the data to be processed; x satisfies M / K ≤ N + xL ≤ floor(M / K), and N is the window size.
[0040] Exemplarily, to avoid missing some calculation results due to data division, there will be partial overlap in the data content of adjacent two data blocks.
[0041] For example, if the window size is 16 and the sliding step is 4, and the first data block includes the data from the 1st to the 100th bits of the data to be processed, then the second data block can include the data from the 88th to the 188th bits of the data to be processed.
[0042] S102, Calculate the Fourier transform result of the first window on each data block.
[0043] In one possible implementation, the cuFFTX function library can be called to generate FFT tasks to calculate the Fourier transform result of the first window on each data block.
[0044] Specifically, the cuFFTX function library can be called to generate FFT tasks, generate the FFT parameters required for each window calculation, such as the phase rotation factor, and then use the cuFFTX function to calculate the Fourier transform result of the first N points of data in each data block and save it to the GPU video memory.
[0045] In one example, the FFT parameters can also be stored in the shared memory for convenient call during subsequent calculations.
[0046] S103, Based on the DFT iterative model, iterate and calculate the Fourier transform results of the remaining windows on each data block according to the Fourier transform result of the first window and the new data of each window.
[0047] In some embodiments, each data block can be further subdivided into multiple windows according to the sliding window size and the sliding step, and there is partial overlapping data and partial newly added data between two adjacent windows. Read the (i + 1)-th newly added data on the k-th data block; then, based on the DFT iterative model, calculate the Fourier transform result of the (i + 1)-th window on the k-th data block according to the Fourier transform result of the i-th window and the (i + 1)-th newly added data.
[0048] Exemplarily, the data of each window can be expressed as xw[m] = x[n + m], where m = 0, 1, ..., N - 1, and n is the starting position of the window.
[0049] Exemplarily, the (i + 1)-th newly added data is the newly added data of the (i + 1)-th sliding window on the k-th data block compared to the i-th sliding window.
[0050] Exemplarily, i is a positive integer less than or equal to Nkmax, and Nkmax is the total number of windows on the k-th data block.
[0051] In a possible implementation, the DFT iterative model can be obtained from the DFT transformation formula.
[0052] Exemplarily, the DFT transformation formula can satisfy the following formula:
[0053]
[0054] where Xi[m] is the Fourier transform result of the m-th point data of the i-th window.
[0055] Exemplarily, the DFT iterative model can satisfy the following formula:
[0056]
[0057] where Xi+1[m] is the Fourier transform result of the m-th point data in the (i + 1)-th window, Xi[m] is the Fourier transform result of the m-th point data in the i-th window, x[iL] is the sampling point entering the (i + 1)-th window (which is a repeated data between the i-th window and the (i + 1)-th window), x[iL + N] is the sampling point leaving the (i + 1)-th window (which is one of the (i + 1)-th newly added data), is the phase rotation factor, which is used to compensate for the frequency-domain phase change caused by window sliding, and m is a non-negative integer less than or equal to N.
[0058] In an example, refer to Figure 3, multiple threads can be created in the kernel, and registers can be instantiated simultaneously. The Fourier transform results of the i-th window are read from the shared memory, and the sampling points leaving and entering the (i + 1)-th window are transferred to the registers of the k-th kernel. The registers use the Fourier transform results of the i-th window and the frequency domain components of the sampling points entering the (i + 1)-th window, subtract the frequency domain components of the sampling points leaving the (i + 1)-th window, and obtain the Fourier transform results of the (i + 1)-th window. The above steps are repeated until all data processing is completed.
[0059] Specifically, N threads can be created in each kernel, and 3L + 1 registers can be instantiated for each thread. L of these registers are used to store phase rotation factors. Then, the FFT results of N points in the i-th window are read from the shared memory into the registers of N threads, and the sampling points leaving and entering the window are read into the registers of N threads in each kernel. The registers calculate the FFT results of N points in the (i + 1)-th window based on the above DFT iteration model, and then return the Fourier transform results of the (i + 1)-th window to the GPU video memory to complete a sliding window DFT.
[0060] According to the method provided by the present invention, by calculating the FFT results of the current window based on the FFT results of the first window and the new data of each window during each calculation, instead of performing FFT calculations on all the data of the current window each time, the reuse of existing FFT results can be achieved, the amount of data read each time and the amount of calculation and the number of data loading times for each calculation can be reduced, thereby reducing the occupied GPU video memory space; since only new data is read each time, the requirements for data alignment and window size can be lowered.
[0061] Furthermore, storing the FFT parameters required for FFT calculations, the intermediate FFT results obtained, and the divided data blocks in the shared memory instead of the global memory can avoid the problems of low calculation efficiency and high access latency caused by the limited I / O speed of the global memory. Through the DFT iteration model, the FFT results of the current window can be obtained by simply subtracting the frequency domain components of the sampling points when entering and leaving the previous window from the FFT results of the previous window; compared with the current method of performing FFT calculations on the data of each window in batches, the calculation amount of the present invention is smaller, and the occupied video memory space can be further reduced.
[0062] Figure 4 The figure shows a schematic structural diagram of a GPU provided by an embodiment of the present invention. By way of example and not limitation, the GPU may include a data division unit 410, a first calculation unit 420, and a second calculation unit 430.
[0063] Exemplarily, the data partitioning unit 410 is configured to partition the data to be computed into multiple data blocks; the first computing unit 420 is configured to compute the Fourier transform results of the first window on each data block; and the second computing unit 430 is configured to iteratively compute the Fourier transform results of the remaining windows on each data block based on the DFT iterative model, according to the Fourier transform results of the first window and the newly added data of each window.
[0064] To better illustrate the beneficial effects of the present invention, a comparative experiment is conducted between the present invention and a traditional method.
[0065] Exemplarily, referring to Figure 5 , after partitioning the data to be computed into multiple data blocks, the traditional method reads all the data of one data block's windows at a time for batch FFT computation. In contrast, after storing the data to be computed in the shared memory, the present invention performs sliding window FFT calculation based on the above method.
[0066] Figure 6 Shown is a comparative schematic diagram of the acceleration effect provided by an embodiment of the present invention.
[0067] Referring to Figure 6 , it can be seen that the method proposed by the present invention has a higher acceleration ratio than the traditional method under different input data sizes (120000*1, 240000*1, 120000*10, 240000*10), with a sliding window size of 128 and a sliding step of 4.
[0068] Therefore, by loading data into the shared memory of the GPU, the present invention can reduce the number of global memory accesses, reduce the I / O overhead, and improve the utilization rate of computing resources; by utilizing the data overlap feature between windows and reusing data in the shared memory, the present invention can reduce the number of data loading times; by performing FFT calculation in registers, the present invention can reduce the dependence on global memory and reduce the memory access latency; by implementing sliding window DFT instead of batch FFT through the kernel-level FFT algorithm and the cuFFTX library, the present invention can reduce the video memory occupancy and improve the processing ability of large-scale data.
[0069] Figure 7 Shown is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. As Figure 7 shown, the electronic device 700 may include: at least one GPU 710 ( Figure 7 only one GPU is shown in
[0070] The electronic device 700 may be a processing device such as a robot that can implement the above method. The embodiments of the present invention do not impose any restrictions on the specific type of the electronic device.
[0071] Those skilled in the art can understand that Figure 7 This is only an example of the electronic device 700 and does not constitute a limitation on the electronic device. It may include more or fewer components than those shown in the figure, or combine some components, or different components. For example, the electronic device 700 may further include an input / output interface.
[0072] In some embodiments, the memory 720 may be an internal storage unit, such as a hard disk or a memory. In other embodiments, the memory 720 may also be an external storage device, such as a plug-in hard disk, a Smart Memory Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 720 may also include both an internal storage unit and an external storage device. The memory 720 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as the program code of the computer program. The memory 720 may also be used to temporarily store data that has been output or will be output.
[0073] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0074] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0075] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program, which when executed by a GPU, can implement the steps in each of the above method embodiments.
[0076] An embodiment of the present invention provides a computer program product, which when running on an electronic device, enables the electronic device to execute the steps in each of the above method embodiments.
[0077] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above method embodiments of the present invention, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a GPU, the steps in each of the above method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable medium may not be an electrical carrier signal and a telecommunication signal.
[0078] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0079] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
Claims
1. A parallel sliding window DFT implementation method based on GPU, characterized in that: include: Divide the data to be calculated into multiple data blocks; Calculate the Fourier transform result of the first window on each data block; Based on the DFT iteration model, the Fourier transform results of the remaining windows on each data block are iteratively calculated according to the Fourier transform result of the first window and the newly added data of each window.
2. The method according to claim 1, characterized in that The step of dividing the data to be operated into a plurality of data blocks comprises: The data to be calculated is divided into a plurality of data blocks according to the data length, sliding window size and sliding step size of the data to be calculated.
3. The method according to claim 2, characterized in that The data length of the data block satisfies the following formula: Wherein, Dk is the data length of the kth data block, floor(·) is the floor rounding function, L is the sliding step size, K is the total number of the data blocks, M is the data length of the data to be calculated; x satisfies M / K≤N+xL≤floor(M / K), and N is the sliding window size.
4. The method according to claim 1, characterized in that: The calculating of the Fourier transform result of the first window on each data block includes: The cuFFTX function library is called to generate an FFT task to calculate the Fourier transform result of the first window on each data block.
5. The method according to claim 1, characterized in that The DFT iterative model is based on which, according to the Fourier transform result of the first window and the newly added data of each window, the Fourier transform results of the remaining windows on each data block are iteratively calculated, including: Read the i+1th newly added data on the kth data block, where i is a positive integer less than or equal to Nkmax, Nkmax is the total number of windows on the kth data block, and the i+1th newly added data is the newly added data of the i+1th window on the kth data block compared with the ith window; Based on the DFT iterative model, the Fourier transform result of the i+1th window on the kth data block is calculated according to the Fourier transform result of the i-th window and the i+1th newly added data.
6. The method according to claim 5, characterized in that The DFT iteration model satisfies the following formula: Wherein, Xi+1[m] is the Fourier transform result of the m-th point data in the i+1-th window, Xi[m] is the Fourier transform result of the m-th point data in the i-th window, x[iL] is the sampling point entering the i+1-th window, and x[iL+N] is the sampling point leaving the i+1-th window. is a phase rotation factor used to compensate for the frequency domain phase change caused by window sliding, and m is a non-negative integer less than or equal to N.
7. The method according to claim 6, characterized in that The data block, the Fourier transform result of each sliding window and the phase rotation factor used in each calculation are all stored in the shared memory of the GPU.
8. A GPU, characterized in that: include: A data partitioning unit, the data partitioning unit is used to partition the data to be operated into a plurality of data blocks; A first calculation unit, the first calculation unit is used to calculate the Fourier transform result of the first window on each data block; The second calculation unit is used to iteratively calculate the Fourier transform results of the remaining windows on each data block based on the DFT iterative model according to the Fourier transform result of the first window and the newly added data of each window.
9. An electronic device, characterized in that: The method comprises a memory, a GPU and a computer program stored in the memory, wherein the GPU implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the electronic device, the method according to any one of claims 1 to 7 is implemented.