Method and apparatus for executing FFT computation, and computing device
By dividing the FFT calculation into multiple calculation stages and using the merge matrix for calculation, the problem of low efficiency of rotation factor multiplication in FFT calculation is solved, and more efficient FFT calculation is achieved.
Patent Information
- Application Number
- PCT/CN2024/100733
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-06-21
- Publication Date
- 2025-06-05
AI Technical Summary
The current FFT calculation involves multiplication calculation of a large number of rotation factors, resulting in low computational efficiency. The existing technology optimizes through vector multiplication or matrix multiplication, but there are still efficiency bottlenecks.
By dividing the FFT calculation into multiple calculation stages, and determining the merge matrix in each calculation stage, the matrix is the product of the DFT coefficient matrix and the rotation factor, and then the calculation of the target calculation stage is achieved through matrix multiplication.
This method reduces the number of rotation factor calculations, reduces the amount of calculations, and improves the execution efficiency of FFT calculations.
Smart Images

Figure CN2024100733_05062025_PF_FP_ABST
Abstract
Description
Method, device and computing equipment for performing FFT calculation
[0001] This application claims priority to Chinese patent application No. 202311638516.4 filed on November 30, 2023, entitled “Method, Apparatus and Computing Device for Performing FFT Calculations,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a method, apparatus, and computing device for performing FFT calculations. Background Art
[0003] Fast Fourier Transform (FFT) is an efficient algorithm for implementing Discrete Fourier Transform (DFT).
[0004] The current FFT calculation process involves the multiplication of a large number of rotation factors. To improve the efficiency of FFT calculations, related technologies can implement the multiplication of rotation factors through vector multiplication or matrix multiplication. However, since the multiplication of rotation factors accounts for a large proportion of the entire FFT calculation, even if it is implemented through vector multiplication or matrix multiplication, it still requires a large amount of calculation, resulting in the current FFT calculation efficiency is still low.
[0005] Summary of the Invention
[0006] The present application provides a method, apparatus, and computing device for performing FFT calculations, which can improve the efficiency of performing FFT calculations. The corresponding technical solutions are as follows:
[0007] In a first aspect, a method for performing FFT calculation is provided. The method may be executed by a processor, and the method includes:
[0008] The processor obtains a data sequence to be subjected to FFT calculation in response to a Fast Fourier Transform (FFT) calculation request. The FFT calculation is divided into multiple calculation stages based on the length of the data sequence. For at least one target calculation stage among the multiple calculation stages, a merging matrix corresponding to the target calculation stage is determined based on the bases corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, wherein the merging matrix is the product of the Discrete Fourier Transform (DFT) coefficient matrix and the rotation factor in the target calculation stage. The multiple calculation stages are executed sequentially, wherein, when executing the target calculation stage, the product of the input sequence of the target calculation stage and the merging matrix is determined as the output sequence of the target calculation stage. The calculation result of performing the FFT calculation on the data sequence is determined based on the output sequence of the last calculation stage among the multiple calculation stages.
[0009] In the solution described in this application, when executing the target calculation phase of the FFT calculation, a combined matrix can be obtained after multiplying the DFT coefficient matrix and the rotation factor, and then the calculation of the target calculation phase can be performed using the combined matrix. The target calculation phase can be any calculation phase in the FFT calculation process. By performing a single matrix operation on the input data matrix composed of the combined matrix and the input sequence, the DFT calculation and the rotation factor calculation can be replaced, thereby improving the execution efficiency of the FFT.
[0010] In one achievable manner, when the decimation method corresponding to the FFT calculation is time domain decimation, the input sequence of the first calculation stage among the multiple calculation stages is changed to a data sequence. When the decimation method corresponding to the FFT calculation is frequency domain decimation, the input sequence of the first calculation stage is changed to a sequence obtained by rearranging the data sequence, wherein the rearrangement processing refers to determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to the multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, wherein the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
[0011] In the scheme shown in the present application, by changing the input sequence of the first calculation stage, the elements corresponding to the same rotation factor in the input sequence of each calculation stage in the FFT calculation can be arranged continuously, thereby avoiding the discontinuous reading of the elements in the input sequence during the FFT calculation process, thereby improving the execution efficiency of the FFT.
[0012] In one achievable embodiment, determining a merge matrix corresponding to a target calculation stage based on bases corresponding to the multiple calculation stages and the execution order of the target calculation stage in the multiple calculation stages includes: determining a reordering number corresponding to each block included in the target calculation stage, wherein the reordering number indicates the arrangement order of each block in the target calculation stage when the input sequence of the first calculation stage is unchanged. Determining a merge matrix corresponding to each block in the target calculation stage based on the reordering number of each block.
[0013] In one achievable embodiment, when the decimation mode of the FFT calculation is frequency domain decimation, determining the reordering number corresponding to each block included in the target calculation stage includes: obtaining a third multidimensional array, wherein the order of dimensions of each level of the third multidimensional array is opposite to the order of dimensions of each level of the fourth multidimensional array, wherein the elements in the fourth multidimensional array are sequentially arranged starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multidimensional array are, in sequence, the bases corresponding to the calculation stages subsequent to the target calculation stage. Determining the elements in the third multidimensional array as the reordering number corresponding to each block included in the target calculation stage in sequence.
[0014] In one achievable embodiment, when the decimation method of the FFT calculation is time-domain decimation, determining the reordering number corresponding to each block included in the target calculation stage includes: obtaining a fifth multidimensional array, wherein the order of dimensions of each level of the fifth multidimensional array is opposite to the order of dimensions of each level of the sixth multidimensional array, wherein elements in the sixth multidimensional array are sequentially arranged starting from 0 and have a length equal to the number of blocks included in the target calculation stage, and wherein the dimensions of the sixth multidimensional array are, in sequence, the bases corresponding to the calculation stages preceding the target calculation stage. Determining the elements in the fifth multidimensional array as the reordering number corresponding to each block included in the target calculation stage in sequence.
[0015] In one achievable manner, when the decimation mode of the FFT calculation is time domain decimation, determining the reordering number corresponding to each block included in the target calculation stage includes: obtaining a first sequence matrix, the first sequence matrix being the transposed matrix of a second sequence matrix, the second sequence matrix being arranged sequentially starting from 0 in a row-major manner, the number of columns of the second sequence matrix being equal to the basis of the calculation stage before the target calculation stage, and the number of rows of the second sequence matrix being equal to the product of the basis of the calculation stage before the previous calculation stage. Elements in the first sequence matrix are sequentially determined as the reordering number corresponding to each block included in the target calculation stage.
[0016] In one achievable embodiment, the method further includes: for each calculation stage between the second calculation stage and the penultimate calculation stage in the plurality of calculation stages, obtaining a seventh multidimensional array, wherein the order of the dimensions after the first-level dimension of the seventh multidimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multidimensional array, the first-level dimension of the seventh multidimensional array is equal to the last-level dimension of the eighth multidimensional array, the elements in the eighth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the calculation stage, the last-level dimension is the basis corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last level are the basis corresponding to the calculation stages before the previous calculation stage, respectively. Determining the seventh multidimensional array as the storage order of the blocks in each calculation stage; and storing the calculation results of the blocks in each calculation stage based on the storage order.
[0017] In one implementable manner, determining the merge matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage includes: determining the imaginary data and the real data of each element in the merge matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merge matrix, wherein the rearrangement number, the position information, and the imaginary data and the real data of each element satisfy the following formula: k=(Alp+Bpq)modN;
[0018] Where l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are the row and column numbers of the elements in the merged matrix respectively. The real data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage.
[0019] In the solution presented in this application, the corresponding combined matrix can be obtained by substituting the rearranged number of each block in each calculation stage into the corresponding formula. Compared to the prior art method of first determining the rotation factor matrix and the DFT matrix and then performing matrix multiplication of the rotation factor matrix and the DFT matrix to obtain the combined matrix, this method can reduce the amount of computation required to obtain the combined matrix, thereby improving the efficiency of performing FFT calculations.
[0020] In one achievable embodiment, determining the output sequence of the target calculation stage as the product of the input sequence of the target calculation stage and the merging matrix includes: arranging each odd-numbered row in the merging matrix to each even-numbered row, or arranging each even-numbered row in the merging matrix to each odd-numbered row, to obtain a rearranged merging matrix. Determining the output sequence of the target calculation stage as the product of the input sequence of the target calculation stage and the rearranged merging matrix is.
[0021] In the scheme shown in the present application, the odd and even rows of the merged matrix are first rearranged, and then calculations are performed based on the rearranged merged matrix and the input data matrix. The real data of the complex numbers and the imaginary data of the complex numbers in the result matrix are arranged continuously. This avoids the non-continuous reading of the real data and the imaginary data in the result matrix when converting the result matrix into an output sequence, which can improve the efficiency of obtaining the output sequence and further improve the efficiency of executing the FFT calculation.
[0022] In one achievable method, the FFT calculation is divided into multiple calculation stages based on the length of the data sequence, including: dividing the FFT calculation into multiple calculation stages based on the length of the data sequence and the priorities of multiple candidate bases, and determining a basis corresponding to each calculation stage, wherein the candidate basis is less than or equal to one-half the calculation size of the matrix multiplication unit performing the FFT calculation, and the priority of the candidate basis is proportional to the numerical size of the candidate basis. Setting the size of the basis corresponding to each calculation stage based on the calculation size of the matrix multiplication unit can improve the utilization of the matrix multiplication unit, thereby improving the efficiency of performing the FFT calculation.
[0023] In an implementable manner, the at least one target computing stage is another computing stage except the first computing stage and the last computing stage in the multiple computing stages.
[0024] In a second aspect, a device for performing FFT calculation is provided, the device comprising:
[0025] An acquisition module, configured to obtain a data sequence to be executed for FFT calculation in response to a Fast Fourier Transform (FFT) calculation request;
[0026] A partitioning module is used to divide the FFT calculation into multiple calculation stages based on the length of the data sequence;
[0027] a determination module, configured to determine, for at least one target calculation stage among the multiple calculation stages, a merging matrix corresponding to the target calculation stage based on bases corresponding to the multiple calculation stages and an execution order of the target calculation stage in the multiple calculation stages, wherein the merging matrix is a product of a discrete Fourier transform (DFT) coefficient matrix in the target calculation stage and a rotation factor;
[0028] an execution module, configured to sequentially execute a plurality of calculation phases, wherein, when executing a target calculation phase, a product of an input sequence of the target calculation phase and the merge matrix is determined as an output sequence of the target calculation phase;
[0029] The determination module is used to determine a calculation result of performing FFT calculation on the data sequence based on an output sequence of a last calculation stage in the multiple calculation stages.
[0030] In one feasible manner, when the extraction method corresponding to the FFT calculation is time domain extraction, the input sequence of the first calculation stage in multiple calculation stages is changed to a data sequence; when the extraction method corresponding to the FFT calculation is frequency domain extraction, the input sequence of the first calculation stage is changed to a sequence after the data sequence is rearranged, wherein the rearrangement processing refers to determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, and the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
[0031] In one implementable manner, the determination module is used to: determine a reordering number corresponding to each block included in the target calculation stage, wherein the reordering number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage; and determine a merging matrix corresponding to each block based on the reordering number of each block in the target calculation stage.
[0032] In one feasible manner, when the extraction method of the FFT calculation is frequency domain extraction, the determination module is used to: obtain a third multidimensional array, the order of the dimensions of each level of the third multidimensional array is opposite to the order of the dimensions of each level of the fourth multidimensional array, the elements in the fourth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multidimensional array are respectively the bases corresponding to the calculation stages after the target calculation stage; and determine the elements in the third multidimensional array as the rearranged numbers corresponding to each block included in the target calculation stage.
[0033] In one feasible manner, when the extraction method of the FFT calculation is time domain extraction, the determination module is used to: obtain a fifth multidimensional array, the order of the dimensions of each level of the fifth multidimensional array is opposite to the order of the dimensions of each level of the sixth multidimensional array, the elements in the sixth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multidimensional array are respectively the bases corresponding to the calculation stages before the target calculation stage; and determine the elements in the fifth multidimensional array as the rearranged numbers corresponding to each block included in the target calculation stage.
[0034] In one implementable manner, when the extraction method of the FFT calculation is time domain extraction, the determination module is used to: obtain a first sequence matrix, the first sequence matrix is the transposed matrix of the second sequence matrix, the second sequence matrix is arranged in sequence starting from 0 in a row-first manner, the number of columns of the second sequence matrix is equal to the basis of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the corresponding bases of the calculation stages before the previous calculation stage; and determine the elements in the first sequence matrix in sequence as the rearranged numbers corresponding to each block included in the target calculation stage.
[0035] In one achievable manner, the above-mentioned device also includes a storage module, which is used to: for each calculation stage between the second calculation stage and the penultimate calculation stage in the multiple calculation stages, obtain a seventh multidimensional array, the order of the dimensions after the first-level dimension of the seventh multidimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multidimensional array, the dimension of the first level of the seventh multidimensional array is equal to the dimension of the last level of the eighth multidimensional array, the elements in the eighth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the calculation stage, the dimension of the last level is the basis corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last level are respectively the basis corresponding to the calculation stages before the previous calculation stage; determine the seventh multidimensional array as the storage order of each block in each calculation stage; and store the calculation results of each block in each calculation stage based on the storage order.
[0036] In one implementable manner, the determination module is configured to: determine the imaginary data and real data of each element in the merged matrix corresponding to each block based on the rearrangement number of each block in the target calculation phase and the position information of each element in the merged matrix, wherein the rearrangement number, the position information, and the imaginary data and real data of each element satisfy the following formula: k=(Alp+Bpq)modN;
[0037] Where l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are the row and column numbers of the elements in the merged matrix respectively. The real data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage.
[0038] In one achievable manner, the determining module is configured to:
[0039] After arranging the odd rows in the merge matrix to the even rows, or arranging the even rows in the merge matrix to the odd rows, a rearranged merge matrix is obtained; the product of the input sequence of the target calculation stage and the rearranged merge matrix is determined as the output sequence of the target calculation stage.
[0040] In one achievable manner, the division module is used to: divide the FFT calculation into multiple calculation stages based on the length of the data sequence and the priorities corresponding to multiple candidate bases, and determine the basis corresponding to each calculation stage, wherein the candidate basis is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate basis is proportional to the numerical size of the candidate basis.
[0041] In an implementable manner, the at least one target computing stage is another computing stage except the first computing stage and the last computing stage in the multiple computing stages.
[0042] In a third aspect, a computing device is provided, comprising a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the computing device executes the method described in the first aspect above.
[0043] In a fourth aspect, a computer program product comprising instructions is provided. When the instructions are executed by a computing device, the computing device is caused to execute the method as described in the first aspect above.
[0044] In a fifth aspect, a computer-readable storage medium is provided, which includes computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method described in the first aspect above. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] FIG1 is a flow chart of an 18-point FFT calculation implemented by Cooley-Tukey;
[0046] FIG2 is a schematic diagram of a method for performing FFT calculations provided in an embodiment of the present application;
[0047] FIG3 is a butterfly network diagram for implementing 8-point FFT calculation based on frequency domain extraction and frequency domain extraction in the related art;
[0048] FIG4 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0049] FIG5 is a flow chart of a method for performing FFT calculations provided in an embodiment of the present application;
[0050] FIG6 is a flow chart of a method for processing a data sequence provided by an embodiment of the present application;
[0051] FIG7 is a schematic diagram of the structure of a device for performing FFT calculations provided in an embodiment of the present application. DETAILED DESCRIPTION
[0052] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0053] The Discrete Fourier Transform (DFT) is a Fourier transform that presents discrete forms in both time and frequency domains. It is used to transform discrete time-domain sampled data into discrete frequency-domain sampled data. In terms of data form, the DFT's input data (discrete time-domain sampled data) and output data (discrete frequency-domain data) are complex sequences of finite length, and the two complex sequences are of equal length.
[0054] Fast Fourier Transform (FFT): An algorithm that quickly computes the discrete Fourier transform (DFT) or its inverse transform (IDFT). By recursively decomposing the DFT of a long complex sequence into the DFTs of shorter complex sequences, the original O(N) 2 ) is reduced to O(NlogN), where N is the length of the long complex sequence.
[0055] The Cooley-Tukey algorithm is a commonly used FFT algorithm. Based on a divide-and-conquer strategy, it decomposes the DFT of a complex sequence of length N into N1 DFTs of length N2, followed by complex multiplications with rotation factors, where N = N1*N2.
[0056] Figure 1 is a flowchart for implementing an 18-point FFT calculation using Cooley-Tukey, which can be referred to as a butterfly network. As shown in Figure 1, the FFT calculation process can be decomposed into multiple stages. The number of stages is equal to the number of radixes divided by the length of the data sequence on which the FFT calculation is performed. Where the data sequence is a complex number sequence, an FFT calculation for a data sequence of length N can be referred to as an N-point FFT calculation. For example, assuming that N can be decomposed into N1, N2, ..., Ni (N = N1 * N2 * ... * Ni), the N-point FFT calculation process can include i stages. N1, N2, ..., Ni can be referred to as radixes. N1 is the radix corresponding to the first stage, N2 is the radix corresponding to the second stage, and Ni is the radix corresponding to the i-th stage.
[0057] In FFT calculation, the input sequence of the first calculation stage can be obtained by performing the data sequence corresponding to the FFT calculation. For other calculation stages after the first calculation stage, the input sequence of each calculation stage is obtained from the output sequence of the previous calculation stage, and the length of the input sequence of each calculation stage is the same. A calculation stage can be divided into at least one block (section), and a block can be divided into at least one butterfly unit (butterfly). By performing butterfly calculation on each butterfly unit in the calculation stage, the output sequence of the calculation stage can be obtained. Among them, a calculation stage includes N / Ni butterfly units, where N is the length of the input data of the calculation stage, and Ni is the basis corresponding to the calculation stage. The butterfly calculation of the butterfly unit can be divided into rotation factor calculation and DFT calculation. In one example, the rotation factor calculation can be a complex vector multiplication calculation of the input data of the butterfly unit and the rotation factor, and the DFT calculation can be a complex matrix multiplication calculation of the rotated input data (input data after the complex multiplication calculation of the rotation factor) and the DFT coefficient matrix corresponding to the butterfly unit. The DFT coefficient matrix corresponding to the butterfly unit is related to the structure of the butterfly unit, and the DFT coefficient matrix corresponding to each butterfly unit in each calculation stage is the same. Alternatively, in another example, the DFT calculation can be a complex matrix multiplication calculation of the input data of the butterfly unit and the DFT coefficient matrix, and the rotation factor calculation can be a complex vector multiplication calculation of the calculation result corresponding to the rotation factor and the complex matrix multiplication calculation. The DFT coefficient matrix corresponding to the butterfly unit is related to the structure of the butterfly unit, and the DFT coefficient matrix corresponding to each butterfly unit in each calculation stage is the same. There may be multiple butterfly units corresponding to the same rotation factor in each calculation stage.
[0058] Each calculation stage in the FFT calculation includes a large number of rotation factor calculations and DFT calculations. Although the rotation factor calculations can be implemented through vector operations or matrix operations to improve the execution efficiency of the FFT calculation, the current optimization effect on the execution efficiency of the FFT calculation is still not ideal.
[0059] An embodiment of the present application provides a calculation method for performing FFT, which can determine the product of the rotation factor and the corresponding DFT coefficient matrix in each calculation stage of the FFT calculation before performing the FFT calculation. This product can be called a combination matrix in the embodiment of the present application. In this way, in the process of performing the FFT calculation, it is only necessary to multiply the input sequence with the corresponding combination matrix in each calculation stage to obtain the output sequence of each calculation stage. In this way, the multiplication calculation of the rotation factor in the process of performing the FFT calculation is avoided, thereby improving the execution efficiency of the FFT calculation.
[0060] FIG2 is a schematic diagram of a method for performing FFT calculations provided by an embodiment of the present application. As shown in FIG2 , a calculation stage may originally include a matrix multiplication operation of an input data matrix composed of a DFT matrix and the input data of each butterfly unit in the calculation stage, and a vector multiplication operation of the result corresponding to the matrix multiplication and the rotation factor of each butterfly unit. Among them, if the rotation factors corresponding to multiple butterfly units are the same, the rotation factors corresponding to the multiple butterfly units can be converted into a diagonal matrix (which can be called a rotation factor matrix), and a matrix operation is performed on the result corresponding to the matrix multiplication. Then, using the binding rate of matrix multiplication, the calculation in the calculation stage may include a matrix multiplication operation of the rotation factor matrix and the DFT matrix, and then a matrix multiplication operation with the input data matrix. Since the DFT matrix and the rotation factor matrix are fixed in each calculation stage, the result of the matrix multiplication operation performed by the DFT matrix and the rotation factor matrix is fixed, that is, the binding matrix corresponding to each calculation stage is fixed. Therefore, in the process of executing FFT, the calculation corresponding to each calculation stage can be performed by obtaining the combined matrix of each calculation stage, which can avoid the calculation of the rotation factor in the calculation stage, thereby reducing the amount of calculation corresponding to each calculation stage and improving the efficiency of executing FFT calculation.
[0061] To facilitate understanding of the embodiments of the present application, some terms involved in the embodiments of the present application are introduced below:
[0062] Multidimensional array: An array consisting of at least two dimensions. In one example, a two-dimensional array is a matrix. The number of rows in the matrix is the first dimension of the two-dimensional data, also known as the first-level dimension, and the number of columns in the matrix is the second dimension of the two-dimensional data, also known as the second-level dimension. Accordingly, a three-dimensional array consists of three levels of dimensions, and an n-dimensional array consists of n levels of dimensions.
[0063] Frequency domain extraction and time domain extraction: two methods for implementing FFT calculations. As shown in Figure 3, Figure 3 is a butterfly network diagram for implementing 8-point FFT calculations based on frequency domain extraction and time domain extraction in related technologies.
[0064] FIG4 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in FIG4 , the computing device 400 may include: a bus 402, a processor 404, and a memory 406. Optionally, the computing device 400 may also include a communication interface 408. The processor 404, the memory 406, and the communication interface 408 communicate with each other via the bus 402. The computing device 400 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 400. The computing device 400 may be a device for running a model, a terminal, or a server. When the computing device 400 is a terminal, the computing device 400 includes but is not limited to a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 400 is a server, the computing device 400 may be a separate server, or a server cluster consisting of multiple servers, or a physical machine, or a virtual machine or container virtualized by virtualization technology.
[0065] Bus 402 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG4 illustrates a single bus line, but this does not imply a single bus or type of bus. Bus 402 may include a path for transmitting information between various components of computing device 400 (e.g., memory 406, processor 404, and communication interface 408).
[0066] The processor 404 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). The processor 404 may further include a matrix multiplication unit that can be used to perform matrix multiplication operations involved in the FFT calculation process.
[0067] The memory 406 may include a volatile memory (volatile memory), such as a random access memory (RAM). The memory 406 may also include a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid state drive (SSD). The memory 406 may store a program code, and the processor 404 may execute the program code to implement the method for performing FFT calculations provided in an embodiment of the present application. For example, in response to a fast Fourier transform (FFT) calculation request, a data sequence to be performed on the FFT calculation is obtained. Based on the length of the data sequence, the FFT calculation is divided into a plurality of calculation stages. Multiple calculation stages are executed sequentially, wherein, when executing the target calculation stage, the product of the input sequence of the target calculation stage and the merge matrix is determined as the output sequence of the target calculation stage. Based on the output sequence of the last calculation stage in the plurality of calculation stages, the calculation result of performing the FFT calculation on the data sequence is determined.
[0068] The communication interface 408 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 400 and other devices or a communication network.
[0069] FIG5 is a flow chart of a method for performing FFT calculations provided in an embodiment of the present application. The method may be performed by the computing device 400 shown in FIG4 , and further may be performed by the processor 404 in the computing device 400. As shown in FIG5 , the method includes:
[0070] Step 501: The processor obtains a data sequence to be subjected to FFT calculation in response to a Fast Fourier Transform (FFT) calculation request.
[0071] In implementation, a computing device may run an application that performs FFT calculations, such as a high-performance computing (HPC) application or an artificial intelligence (AI) application. During operation, when an FFT calculation is required, the application may send an FFT execution request to the processor. For example, the application may be a VASP (Vienna Ab-initio Simulation Package). When a wave function needs to be solved in the VASP, a wave function solution request may be sent to the processor. This wave function solution request is an FFT execution request. The wave function is a common function in quantum mechanics, and the wave function can be solved by performing an FFT calculation on the input data of the wave function to obtain the solution result of the wave function.
[0072] After receiving the execution request of the FFT calculation, the processor may obtain the data sequence to be subjected to the FFT calculation and execute subsequent steps 502 to 505. The data sequence to be subjected to the FFT calculation may be carried in the execution request of the FFT calculation.
[0073] In one example, for the FFT execution method provided herein, a technician can write the corresponding processing program as an FFT processing function and add it to a mathematical library. The mathematical library can be stored on a computing device executing an application and includes a large number of processing functions that can be used to implement various mathematical calculations involved in the application. When an FFT calculation is required in the application, an FFT execution request can be sent to the processor, which then calls and executes the FFT processing function to implement the following steps.
[0074] Step 502: Based on the length of the data sequence, divide the FFT calculation into multiple calculation stages.
[0075] In one example, after receiving an execution request for FFT calculation, the processor can obtain the length of the input data to be FFT calculated and decompose the FFT calculation into multiple calculation stages. For example, the processor can decompose the length of the input data according to the Cooley-Tukey algorithm to obtain the various calculation stages for performing FFT calculation on the input data and the basis corresponding to each calculation stage. For example, if the number of data elements in the input data is 8, 8 can be decomposed into 2*2*2, and the process of performing FFT on the input data includes three calculation stages, and the basis corresponding to each calculation stage is 2. For example, when the number of data elements in the input data is 27, 27 can be decomposed into 3*3*3, and the process of performing FFT on the input data includes three calculation stages, and the basis corresponding to each calculation stage is 3.
[0076] Step 503: For at least one target calculation stage among the multiple calculation stages, determine a merging matrix corresponding to the target calculation stage based on the bases corresponding to the multiple calculation stages and the execution order of the target calculation stage in the multiple calculation stages, wherein the merging matrix is the product of the discrete Fourier transform DFT coefficient matrix in the target calculation stage and the rotation factor.
[0077] In FFT calculations, once the number of calculation stages and the basis for each stage are determined, the corresponding DFT sparse matrix for each stage and the corresponding rotation factors for each element in the input sequence for each stage are also determined. In other words, in FFT calculations, once the number of calculation stages and the basis for each stage are determined, the corresponding binding matrix for each stage can also be determined.
[0078] The target calculation phase refers to the calculation phase in the FFT calculation process in which the DFT matrix calculation and the rotation factor calculation are implemented by combining matrices. The target calculation phase can be any calculation phase among the multiple calculation phases included in the FFT calculation, for example, it can be each calculation phase among the multiple calculation phases. In one example, the number of multiple calculation phases, the execution order of the calculation phases among the multiple calculation phases, the corresponding basis of the calculation phases, and the corresponding relationship with the combining matrix can be pre-stored in the computing device. Before performing the FFT calculation, the combining matrix corresponding to the target calculation phase can be determined based on this corresponding relationship.
[0079] Step 504: Execute multiple calculation stages sequentially, wherein, when executing the target calculation stage, the product of the input sequence of the target calculation stage and the merge matrix is determined as the output sequence of the target calculation stage.
[0080] After determining the binding matrix corresponding to the target calculation stage, multiple calculation stages of the FFT calculation can be executed sequentially. Each time the target calculation stage is executed, the input sequence of the target calculation stage can be converted into an input data matrix, and then the multiplication operation of the input data matrix and the binding matrix is performed by the matrix multiplication unit included in the processor.
[0081] Step 505: Determine a calculation result of performing the FFT calculation on the data sequence based on an output sequence of the last calculation stage among the multiple calculation stages.
[0082] After each calculation stage in the FFT calculation is executed in sequence, the output sequence of the last calculation stage can be determined as the calculation result of executing the FFT calculation, and the result can be returned to the application that sent the FFT calculation request.
[0083] In an embodiment of the present application, only one matrix multiplication operation needs to be performed to complete the rotation factor calculation and DFT calculation included in the target calculation stage, which can improve the execution efficiency of the target calculation stage and further improve the efficiency of executing FFT calculation.
[0084] FFT calculation can be divided into calculation based on time domain extraction method and calculation based on frequency domain extraction method. The following is a further introduction to the FFT calculation method provided in the embodiment of the present application in combination with the extraction method for implementing FFT calculation:
[0085] Solution 1: Perform FFT calculation based on frequency domain extraction
[0086] Step R1: Obtain a data sequence to be processed by FFT calculation.
[0087] Step R2: Based on the length of the data sequence, the FFT calculation is divided into multiple calculation stages.
[0088] In implementation, the FFT calculation may be divided into multiple calculation stages based on the length of the data sequence and the calculation size of a matrix multiplication unit in the processor that performs the matrix multiplication operation.
[0089] In each calculation phase, the size of the combined matrix is equal to the size of the DFT coefficient matrix, and the size of the DFT coefficient matrix is related to the basis of the calculation phase. For example, if the basis of the calculation phase is 8, the size of the DFT coefficient matrix and the combined matrix is 8*8. In actual calculations, the real and imaginary data in the combined matrix are split into 2*2 matrices, so in actual calculations, the size of the combined matrix is 16*16. In other words, when the basis of a calculation phase is n, the size of the combined matrix corresponding to that calculation phase is 2n*2n.
[0090] The calculation size of a matrix multiplication unit can be used to indicate the maximum matrix size supported by the matrix multiplication unit when performing a matrix multiplication operation. For example, if the maximum matrix size supported by the matrix multiplication unit when performing a matrix multiplication operation is M*M, the calculation size of the matrix multiplication unit can be M. When the matrix multiplication unit performs a matrix multiplication operation, the closer the size of the matrix performing the matrix multiplication operation is to the maximum matrix size corresponding to the matrix multiplication unit, the higher the utilization rate of the matrix multiplication unit. Therefore, when determining the basis corresponding to each calculation stage, the basis corresponding to each calculation stage can be made as equal to half of the calculation size of the matrix multiplication unit as possible, thereby improving the utilization rate of the matrix multiplication unit. For example, when the calculation size of the matrix multiplication unit is 16 and the basis corresponding to the calculation stage is 8, the size of the combined matrix corresponding to the calculation stage is 16*16. Therefore, when the matrix multiplication unit performs the operation of the calculation stage, the combined matrix can fully occupy the matrix multiplication unit, thereby fully utilizing the performance of the matrix multiplication unit.
[0091] In one example, multiple candidate bases can be set in advance and a corresponding priority can be set for each candidate base. The multiple candidate bases are less than or equal to half of the calculation size of the matrix multiplication unit that performs FFT calculations, and the priority of the candidate base is proportional to the numerical size of the candidate base.
[0092] When dividing the FFT calculation into multiple calculation stages, the FFT calculation can be divided into multiple calculation stages based on multiple candidate bases set, and a basis for each calculation stage can be determined. For example, if the calculation size of the matrix multiplication unit is 16, the multiple candidate bases can be 8, 7, 6, 5, 4, 3, and 2, respectively, and the priorities of the multiple candidate bases can decrease in sequence.
[0093] FIG6 is a flow chart of a method for processing a data sequence according to an embodiment of the present application. As shown in FIG6 , the method includes:
[0094] Step R21, according to the priority of multiple candidate bases, set the order list corresponding to the candidate set, n = [8, 7, 6, 5, 4, 3, 2], initialize k = 0, and initialize the decomposition list of the basis of each calculation stage to empty, that is, base = [].
[0095] Step R22: Divide the length N of the data sequence on which the FFT calculation is performed by n[k].
[0096] Step R23: Determine whether N can divide n[k].
[0097] If it is divisible, execute step R24; if it is not divisible, execute step R25.
[0098] Step R24: add n[k] to the decomposition list base, update N=N / n[k], and go to step R22.
[0099] Step R25: Determine whether n[k] is the last element in the sequence list n.
[0100] If not, execute step R26; if yes, execute step R28.
[0101] Step R26: Set k=k+1 and transpose step R22.
[0102] Step R27: Add N to the decomposition list base.
[0103] After completing the execution of step R27, the number of elements included in the decomposition list base is the number of calculation stages in the FFT calculation, and the elements included in the decomposition list base are the bases corresponding to each calculation stage in turn.
[0104] In this way, through the method of decomposing the data sequence shown in Figure 6, the basis of each calculation stage of the FFT calculation can be made as equal as possible to half the calculation size of the matrix multiplication unit, thereby improving the utilization rate of the matrix multiplication unit and improving the efficiency of executing the FFT calculation.
[0105] Step R3: Rearrange the data sequence for FFT calculation, and determine the rearranged data sequence as the input sequence for the first calculation stage in the FFT calculation.
[0106] In the traditional frequency-domain extraction-based FFT calculation process, the input sequence of the first calculation stage in the FFT calculation is the data sequence for performing the FFT calculation. After the calculation of the last calculation stage is completed, the output sequence of the last calculation stage needs to be rearranged to obtain the result sequence of the FFT calculation. However, in the traditional frequency-domain extraction-based FFT calculation process, the corresponding elements of the same rotation factor are not arranged continuously in the input sequence of the calculation stage. Therefore, when the input sequence is loaded into the vector register or matrix register in the form of an input data matrix, the input sequence needs to be read non-continuously.
[0107] In order to avoid the occurrence of discontinuous reading problems, the embodiment of the present application, compared with the traditional FFT calculation process based on frequency domain extraction, modifies the arrangement order of elements in the input sequence of the first calculation stage, so that when the output sequence corresponding to each calculation stage is used as the input sequence corresponding to the next calculation stage, the elements corresponding to the same rotation factor of the next calculation stage can be arranged continuously in the input sequence.
[0108] Rearranging the elements in the input sequence of the first calculation stage may include: determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, and the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
[0109] For example, the FFT calculation includes m calculation stages, and the bases corresponding to the m calculation stages are N1, N2, ...Nm respectively. In implementation, the data sequence for the FFT calculation can be converted into a first multidimensional array with dimensions of N1, N2, ...Nm at each level. Then, the first multidimensional array is transposed into a second multidimensional array with dimensions of Nm, Nm-1, ...N1 at each level. The order of the elements in the second multidimensional array is the order of the elements in the input sequence of the first calculation stage after the rearrangement. Among them, the input sequence of the first calculation stage after the rearrangement is actually the input sequence of the first calculation stage in the process of implementing FFT calculation in the traditional time domain extraction method.
[0110] Taking an 8-point FFT calculation consisting of three calculation stages as an example, for the input sequence N = [0, 1, 2, 3, 4, 5, 6, 7] in the first calculation stage before rearrangement, the input sequence N = [0, 4, 2, 6, 1, 5, 3, 7] in the first calculation stage after rearrangement. As shown in Figure 3, the elements corresponding to the rotation factor W2 / 8 in the second calculation stage before rearrangement are at the 4th and 8th positions of the input sequence, respectively. The elements corresponding to the rotation factor W2 / 8 in the second calculation stage after rearrangement are at the 7th and 8th positions of the input sequence, respectively.
[0111] As can be seen, in the embodiment of the present application, after the input sequence of the first calculation stage is rearranged, the elements corresponding to the same rotation factor in the input sequence of each calculation stage are arranged continuously. Furthermore, in the embodiment of the present application, the elements corresponding to the same rotation factor in the input sequence can be read and stored as each row of the input data matrix in a vector register or matrix register, thereby avoiding the discontinuous reading of the input sequence and improving the efficiency of executing the FFT calculation.
[0112] Step R4: Determine the combination matrix corresponding to the target calculation stage in the multiple calculation stages.
[0113] The elements and rotation factors in the DFT coefficient matrix in the calculation stage can be obtained by calculating the trigonometric function with a period of N, where N is the length of the data sequence for performing the FFT calculation. Therefore, for each element in the combined matrix, it can also be calculated by a trigonometric function with a period of N. Therefore, the present application exemplarily provides a method for determining the combined matrix based on a trigonometric function with a period of N, as follows:
[0114] Step R41: Determine the reordering number corresponding to each block included in the target calculation stage, wherein the reordering number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0115] Because the input sequence of the first calculation stage is rearranged in step R2 above, the order of the blocks in each calculation stage differs from the order of the blocks in each calculation stage in traditional FFT calculations. Therefore, in the embodiment of the present application, the order of each block in the target calculation stage can be determined first without changing the input sequence, and then the combination matrix in the target calculation stage can be determined based on the rearranged order.
[0116] Wherein, when the decimation mode of the FFT calculation is frequency domain decimation, determining the reordering number corresponding to each block included in the target calculation stage includes: obtaining a third multidimensional array, wherein the order of dimensions of each level of the third multidimensional array is opposite to the order of dimensions of each level of the fourth multidimensional array, elements in the fourth multidimensional array are sequentially arranged starting from 0 and having a length equal to the number of blocks included in the target calculation stage, and the dimensions of the fourth multidimensional array are, in sequence, bases corresponding to calculation stages subsequent to the target calculation stage. Determining the elements in the third multidimensional array as the reordering number corresponding to each block included in the target calculation stage in sequence.
[0117] In implementation, the third multidimensional array Q3 can be used to indicate the rearrangement number corresponding to each block included in the target calculation stage. The third multidimensional array Q3 can be obtained by transposing the fourth multidimensional array Q4. In one example, the FFT calculation includes m calculation stages, the target calculation stage is the nth calculation stage, and includes r blocks, then the fourth multidimensional array Q4 = [0, 1, 2, ... r], and the corresponding dimensions of each level are N in sequence. n+1 、N n+2 ,…N m Correspondingly, the third multidimensional array Q3 is the fourth multidimensional array Q4 transposed, and the corresponding dimensions at each level are N m ,…N n+1 、N n+1 The value of the first element in the third multidimensional array Q3 is the reordering number corresponding to the first block in the current target calculation stage.
[0118] Step R42: Based on the rearranged number of each block in the target calculation stage, determine the merge matrix corresponding to each block.
[0119] After determining the reordering number corresponding to each block, the reordering number of each block can be substituted into the following formula to obtain the imaginary data and real data of each element in the merged matrix. k = (Alp + Bpq) mod N;
[0120] Among them, l is the rearrangement number, N is the length of the data sequence, p and q are the row and column numbers of the elements in the merged matrix respectively. The real data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in row p, column q of the merged matrix corresponding to block l in the nth computation phase. A and B are preset coefficients, where A is equal to the product of the bases corresponding to the computation phase preceding the target computation phase, and B is equal to the ratio of the data sequence length N to the base corresponding to the target computation phase.
[0121] After obtaining the real and imaginary data of each element in the merging matrix corresponding to the target calculation stage, the real and imaginary data of each element can be used to calculate the merging matrix of the target calculation stage. It should be noted that the number of merging matrices in the target calculation stage is the same as the number of blocks. A merging matrix can be mapped to multiple butterfly units corresponding to the same rotation factor, and the input data matrix multiplied by the merging matrix is composed of the input data of the multiple butterfly units.
[0122] Step R5: Execute multiple calculation stages in sequence to obtain the calculation result of FFT calculation.
[0123] After obtaining the combination matrix corresponding to each block in the target calculation stage, each calculation stage included in the FFT calculation can be executed sequentially. When executing to the target calculation stage, the input data of the target calculation stage can be respectively composed into an input data matrix multiplied by the combination matrix corresponding to each block, and the multiplication operation of the combination matrix and the input data matrix is implemented based on the matrix multiplication unit, thereby realizing the calculation of the target calculation stage. For example, in the target calculation stage, the combination matrix corresponding to the lth block is When performing calculations, the The input data matrix composed of the input data of each butterfly unit in the lth block is multiplied to obtain the calculation result corresponding to the lth block. In this way, when executing the target calculation phase in the embodiment of the present application, only the multiplication operation of the combination matrix and the input data matrix is required to realize the calculation of the target calculation phase, thereby improving the execution efficiency of the target calculation phase and further improving the execution efficiency of the FFT calculation.
[0124] In one achievable manner, since the first calculation stage does not have the same rotation factor and the rotation factor of the last calculation stage is 1, the target calculation stage can be a calculation stage between the first calculation stage and the last calculation stage. The calculation of the first calculation stage and the last calculation stage can be performed using a traditional calculation method, which will not be described in detail in this application. In addition, an embodiment of the present application also provides a method for determining the rotation factor matrix corresponding to the first calculation stage based on a trigonometric function with a period of N, as well as a method for determining the DFT system matrix in the first and last calculation stages, as follows:
[0125] For the rotation factor matrix of the first calculation stage, the rearrangement number corresponding to each block in the first calculation stage can be obtained. The method for obtaining the rearrangement number corresponding to each block corresponding to the first calculation stage can refer to the above step R41 and will not be repeated here. After obtaining the rearrangement number corresponding to each block corresponding to the first calculation stage, the real data and imaginary data corresponding to each element in the rotation factor matrix can be determined by the following calculation formula. k = lj mod N;
[0126] Where l is the rearrangement number, N is the length of the data sequence, i and j are the row and column numbers of the elements in the rotation factor matrix R1, R1_r[i,j] is the real data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R1, and R1_i[i,j] is the imaginary data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R1. i∈[0,N2N3…N m ), j∈[0,N1), N1 is the basis corresponding to the first calculation stage, N2N3…N m It is the product of the bases corresponding to each calculation stage after the first calculation stage.
[0127] For the DFT coefficient matrix of the first calculation stage, i, j∈[0,N1), calculate k=Nij / N1mod N, and get the real and imaginary parts. Among them, W1_r[i,j] is the real data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W1 in the first calculation stage, and W1_i[i,j] is the imaginary data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W1 in the first calculation stage.
[0128] For the DFT coefficient matrix of the last calculation stage, i,j∈[0,N m ), calculate k = Nij / N m mod N, get the real and imaginary parts W m _r[i,j]=C[k],W m_i[i,j]=S[k]. Where, W m _r[i,j] is the DFT coefficient matrix W of the last calculation stage m The real data corresponding to the element in row i and column j in W m _r[i,j] is the DFT coefficient matrix W of the last calculation stage m The imaginary data corresponding to the element in the i-th row and j-th column.
[0129] Solution 2: A method based on time domain extraction to perform FFT calculations
[0130] Step S1: Obtain a data sequence to be subjected to FFT calculation.
[0131] Step S2: Based on the length of the data sequence, the FFT calculation is divided into multiple calculation stages.
[0132] The processing of step S1 and step S2 is the same as that of the above-mentioned step R1 and step R2, and will not be repeated here.
[0133] Step S3: Determine the data sequence calculated by FFT as the input sequence of the first calculation stage in the FFT calculation.
[0134] In a conventional FFT calculation process based on time-domain extraction, the input sequence for the first calculation stage of the FFT calculation is obtained by rearranging the data sequence for performing the FFT calculation, wherein the result after rearrangement is the same as the result after rearrangement in step R3 described above. However, in a conventional FFT calculation process based on time-domain extraction, the input sequence corresponding to the same rotation factor elements in the calculation stage are not arranged continuously. Therefore, when the input sequence is loaded into a vector register or a matrix register in the form of an input data matrix, the input sequence needs to be read non-continuously.
[0135] To avoid the issue of discontinuous reading, the present embodiment, unlike the traditional time-domain decimation-based FFT calculation process, uses the data sequence calculated by the FFT as the input sequence for the first calculation stage without reordering the data sequence. This ensures that when the output sequence corresponding to each calculation stage serves as the input sequence for the next calculation stage, the elements corresponding to the same rotation factors for the next calculation stage can be arranged continuously in the input sequence.
[0136] Taking an 8-point FFT calculation including three calculation stages as an example, for the input sequence N = [0, 1, 2, 3, 4, 5, 6, 7] in the first calculation stage before reordering, and the input sequence N = [0, 4, 2, 6, 1, 5, 3, 7] in the first calculation stage after reordering, as shown in Figure 3, the elements corresponding to the rotation factor W2 / 8 in the second calculation stage before reordering are at the 4th and 8th positions of the input sequence, respectively, and the elements corresponding to the rotation factor W2 / 8 in the second calculation stage after reordering are at the 7th and 8th positions of the input sequence, respectively.
[0137] As can be seen, in the embodiment of the present application, after the input sequence of the first calculation stage is changed, the elements corresponding to the same rotation factor in the input sequence of each calculation stage are arranged continuously. Furthermore, in the embodiment of the present application, the elements corresponding to the same rotation factor in the input sequence can be read and stored as each row of the input data matrix in a vector register or matrix register, thereby avoiding the discontinuous reading of the input sequence and improving the efficiency of executing the FFT calculation.
[0138] Step S4: Determine a combination matrix corresponding to a target calculation stage in the plurality of calculation stages.
[0139] Similar to step R4 above, in solution 2, the embodiment of the present application also provides a method for determining a combination matrix based on a trigonometric function with a period of N, including:
[0140] Step S41: Determine a reordering number corresponding to each block included in the target calculation stage, wherein the reordering number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0141] Because the input sequence of the first calculation stage is rearranged in step S2 above, the order of the blocks in each calculation stage differs from the order of the blocks in each calculation stage in traditional FFT calculations. Therefore, in this embodiment of the present application, the order of each block in the target calculation stage can be determined first without changing the input sequence, and then the combined matrix in the target calculation stage can be determined based on the rearranged order.
[0142] Wherein, when the decimation mode of the FFT calculation is time domain decimation, determining the reordering number corresponding to each block included in the target calculation stage includes: obtaining a fifth multidimensional array, wherein the order of the dimensions of each level of the fifth multidimensional array is opposite to the order of the dimensions of each level of the sixth multidimensional array, wherein the elements in the sixth multidimensional array are sequentially arranged starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multidimensional array are, in sequence, the bases corresponding to the calculation stages before the target calculation stage. Determining the elements in the fifth multidimensional array as the reordering number corresponding to each block included in the target calculation stage in sequence.
[0143] In implementation, the fifth multidimensional array Q5 can be used to indicate the rearrangement number corresponding to each block included in the target calculation stage. The fifth multidimensional array Q5 can be obtained by transposing the sixth multidimensional array Q6. In one example, the FFT calculation includes m calculation stages, the target calculation stage is the nth calculation stage, and includes r blocks, then the sixth multidimensional array Q6 = [0, 1, 2, ... r], and the corresponding dimensions are N1, N2, ... N n-1 Correspondingly, the fifth multidimensional array Q5 is the transposed sixth multidimensional array Q6, and the corresponding dimensions at each level are N n-1 ,…N2, N1. Among them, N1, N2,…N n-1 are the bases corresponding to the first n-1 computation stages, respectively. The sequentially arranged elements of the fifth multidimensional array Q5 are the reordering numbers corresponding to each block in the target computation stage. For example, the value of the first element in the fifth multidimensional array Q5 is the reordering number corresponding to the first block in the current target computation stage. For example, the value of the second element in the fifth multidimensional array Q5 is the reordering number corresponding to the second block in the current target computation stage.
[0144] Step S42: Determine the merge matrix corresponding to each block based on the rearranged number of each block in the target calculation stage.
[0145] The processing of step S42 is the same as that of the above-mentioned step R42, and will not be repeated here.
[0146] Step S5: Execute multiple calculation stages in sequence to obtain the calculation result of FFT calculation.
[0147] After obtaining the combination matrix corresponding to each block in the target calculation stage, each calculation stage included in the FFT calculation can be executed sequentially. When executing to the target calculation stage, the input data of the target calculation stage can be respectively composed into an input data matrix multiplied by the combination matrix corresponding to each block, and the multiplication operation of the combination matrix and the input data matrix is implemented based on the matrix multiplication unit, thereby realizing the calculation of the target calculation stage. For example, in the target calculation stage, the combination matrix corresponding to the lth block is When performing calculations, the The input data matrix composed of the input data of each butterfly unit in the lth block is multiplied to obtain the calculation result corresponding to the lth block. In this way, when executing the target calculation phase in the embodiment of the present application, only the multiplication operation of the combination matrix and the input data matrix is required to realize the calculation of the target calculation phase, thereby improving the execution efficiency of the target calculation phase and further improving the execution efficiency of the FFT calculation.
[0148] In one achievable manner, similar to the above-mentioned solution 1, since the rotation factors of the first calculation stage are all 1 and there are no identical rotation factors in the last calculation stage, the above-mentioned target calculation stage can be a calculation stage between the first calculation stage and the last calculation stage. In an embodiment of the present application, a method for determining the DFT system matrix in the last calculation stage based on a trigonometric function with a period of N is also provided, as follows:
[0149] For the rotation factor matrix of the last calculation stage, the rearrangement number corresponding to each block in the last calculation stage can be obtained. The method for obtaining the rearrangement number corresponding to each block corresponding to the last calculation stage can refer to the above step S41 and will not be repeated here. After obtaining the rearrangement number corresponding to each block corresponding to the last calculation stage, the real data and imaginary data corresponding to each element in the rotation factor matrix can be determined by the following calculation formula. k = lj mod N;
[0150] Among them, l is the rearrangement number, N is the length of the data sequence, i and j are the rotation factor matrices R m The row and column numbers of the elements in R m _r[i,j] is the rotation factor matrix R m The real data corresponding to the element in row i and column j in R m _i[i,j] is the rotation factor matrix R m The element in row i and column j in the numerator corresponds to the imaginary data. i∈[0,N1N2…N m-1 ), j∈[0,N m ), N m is the basis corresponding to the last calculation stage, N1N2…N m-1 It is the product of the bases corresponding to each calculation stage after the last calculation stage.
[0151] For the DFT coefficient matrix of the first calculation stage and the DFT coefficient matrix of the last calculation stage, please refer to the calculation method of the above-mentioned solution 1, which will not be repeated here.
[0152] Solution 3: Another way to perform FFT calculation based on time domain extraction
[0153] Step T1: Obtain a data sequence to be processed by FFT calculation.
[0154] Step T2: Based on the length of the data sequence, the FFT calculation is divided into multiple calculation stages.
[0155] Step T3: Determine the data sequence calculated by FFT as the input sequence of the first calculation stage in the FFT calculation.
[0156] The processing of steps T1 to T3 is the same as that of steps R1 to R3 above and will not be repeated here. The difference between the second solution and the third solution is that the second solution rearranges the output sequence of the last calculation stage in the FFT calculation to obtain the FFT calculation result, while the third solution rearranges the output sequence of each calculation stage after the second calculation stage in the FFT calculation to obtain the FFT calculation result.
[0157] Step T4: Determine the combination matrix corresponding to the target calculation stage in the multiple calculation stages.
[0158] Similar to steps R4 and S4 above, in solution three, the embodiment of the present application also provides a method for determining a combination matrix based on a trigonometric function with a period of N, including:
[0159] Step T41: Determine the reordering number corresponding to each block included in the target calculation stage, where the reordering number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0160] Since the output sequence of each calculation stage after each second calculation stage is rearranged in scheme three, the order of blocks in each calculation stage will be different from the order of blocks in each calculation stage in the traditional FFT calculation. Therefore, in the embodiment of the present application, the arrangement order of each block in the target calculation stage can be determined first without retaking the output sequence, and then the combination matrix in the target calculation stage can be determined based on the rearranged number. Among them, the process of rearranging the output sequence of each calculation stage after the second calculation stage can be referred to step T5, which will not be introduced here.
[0161] In the third solution, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a first sequence matrix, the first sequence matrix being the transposed matrix of the second sequence matrix, the second sequence matrix being arranged in row-first order starting from 0, the number of columns of the second sequence matrix being equal to the basis of the calculation stage immediately preceding the target calculation stage, and the number of rows of the second sequence matrix being equal to the product of the respective bases of the calculation stage immediately preceding the previous calculation stage. Elements in the first sequence matrix are sequentially determined as the rearrangement number corresponding to each block included in the target calculation stage.
[0162] In implementation, the first sequence matrix P1 can be used to indicate the rearrangement number corresponding to each block included in the target calculation stage. The first sequence matrix P1 can be obtained by transposing the second sequence matrix P2. In one example, the FFT calculation includes m calculation stages, and the corresponding dimensions are N1, N2, ...N. m , the target calculation stage is the nth calculation stage, and includes r blocks. The elements in the second sequence matrix P2 can be arranged in a row-first manner as [0, 1, 2, ...r]. The number of columns in the second sequence matrix is equal to N n-1 , the number of rows is equal to N1*N2*…*N n- 2. Accordingly, the first sequence matrix P1 is the transposed second sequence matrix P2, and the number of rows of the first sequence matrix P1 is equal to N n-1 , the number of columns is equal to N1*N2*…*N n-2 . Each element in the first sequence matrix P1 is arranged in rows, which is the rearrangement number corresponding to each block included in the target calculation stage. Among them, the first sequence matrix and the second sequence matrix can also be multidimensional arrays respectively, and the dimensions of each level before the last level of the multidimensional array corresponding to the second sequence matrix are respectively the basis corresponding to the calculation stage before the previous calculation stage, and the last level dimension of the multidimensional array corresponding to the second sequence matrix is the basis corresponding to the previous calculation stage, that is, the dimensions of each level of the multidimensional array corresponding to the second sequence matrix are N1, N2, ...N n-2 、N n-1 Correspondingly, the dimensions of the multidimensional array corresponding to the first sequence matrix are N n-1 , N1, N2, ...N n-2 .
[0163] Step T42: Determine the merge matrix corresponding to each block based on the rearranged number of each block in the target calculation stage.
[0164] The processing of step T42 is the same as that of the above-mentioned step R42, and will not be repeated here.
[0165] Step T5: Execute multiple calculation stages in sequence to obtain the calculation result of FFT calculation.
[0166] After obtaining the combination matrix corresponding to each block corresponding to the target calculation stage, each calculation stage included in the FFT calculation can be executed sequentially. When executing to the target calculation stage, the input data of the target calculation stage can be respectively composed into an input data matrix that is multiplied by the combination matrix corresponding to each block, and the multiplication operation of the combination matrix and the input data matrix is implemented based on the matrix multiplication unit, thereby realizing the calculation of the target calculation stage.
[0167] In this third solution, for each calculation stage after the second of the multiple calculation stages, the calculation results corresponding to each calculation stage can be stored according to a predetermined storage order. The calculation results corresponding to each block are a short output sequence, and the calculation results corresponding to each block included in a calculation stage can form the output sequence corresponding to that calculation stage. The following describes the storage order for each calculation stage after the second calculation stage.
[0168] For each computational stage between the second computational stage and the penultimate computational stage: a seventh multidimensional array may be obtained, wherein the order of the dimensions after the first dimension of the seventh multidimensional array is the same as the order of the dimensions before the last dimension of the eighth multidimensional array, the first dimension of the seventh multidimensional array is equal to the last dimension of the eighth multidimensional array, the elements in the eighth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the computational stage, the last dimension is the basis corresponding to the previous computational stage of the computational stage, and the dimensions before the last dimension are the basis corresponding to the computational stages before the previous computational stage, respectively. The seventh multidimensional array is determined to be the storage order of the blocks in each computational stage.
[0169] In implementation, the storage order corresponding to each block in each calculation stage can be indicated by the seventh multidimensional array Q7. The seventh multidimensional array Q7 is obtained by transposing the eighth multidimensional array Q8. The elements in the eighth multidimensional array Q8 are arranged in sequence starting from 0 and the corresponding length is the number of blocks included in the calculation stage. For example, the calculation stage is the nth calculation stage in the FFT calculation, and the number of blocks included is m, wherein the nth calculation stage is between the second calculation stage and the second to last calculation stage. Correspondingly, the elements in the eighth multidimensional array Q8 can be arranged in sequence from 0 to m, and the corresponding dimensions at each level are N. n-2 、N n-3 ,…N1,N n-1 , where N1, ...N n-3 、N n-2 、N n-1 , which are the basis of n-1 calculation stages before the nth calculation stage.
[0170] After determining the eighth multidimensional array Q8, the eighth multidimensional array Q8 can be processed to obtain the seventh multidimensional array Q7. For example, the eighth multidimensional array Q8 is transposed into the seventh multidimensional array Q7. The dimensions of each level of the seventh multidimensional array Q7 are N n-1 、N n-2 , ...N1. After obtaining the seventh multidimensional array Q7, the storage order corresponding to the calculation results of each block in the calculation stage can be indicated according to the value of each element in the seventh multidimensional array Q7. For example, if the elements in the seventh multidimensional array Q7 are 0, 3, 1, 4, 2, and 5, it can be indicated that in the corresponding calculation stage, the calculation result of the first block is stored in the first position, the calculation result of the third block is arranged after the calculation result of the first block, the calculation result of the fifth block is arranged after the calculation result of the third block, the calculation result of the second block is arranged after the calculation result of the fifth block, the calculation result of the fourth block is arranged after the calculation result of the second block, and the calculation result of the sixth block is arranged after the calculation result of the fourth block.
[0171] For the penultimate calculation stage, the storage order of the calculation results corresponding to each block in the penultimate calculation stage can be the same as the storage order corresponding to the calculation stages between the second and penultimate calculation stages. However, in the penultimate calculation stage, the calculation results corresponding to the same block are stored with intervals of X elements, where X is equal to the number of blocks included in the penultimate calculation stage. Furthermore, in the penultimate calculation stage, elements in the same position in the calculation results corresponding to multiple blocks are stored according to the corresponding calculation order.
[0172] For the last calculation phase: After the last calculation phase completes, Y elements can be read sequentially, where the value of Y is equal to the product of the bases corresponding to the calculation phases before the penultimate calculation phase. The elements read each time are then stored in the storage order indicated by the third sequence matrix. The third sequence matrix is the transposed matrix of the fourth sequence matrix, which is arranged in row-major order starting from 0. The number of columns in the fourth sequence matrix is equal to the basis corresponding to the last calculation phase, and the number of rows in the fourth sequence matrix is equal to the basis corresponding to the penultimate calculation phase.
[0173] After storing the calculation results of each block included in each calculation stage according to the above storage method, the calculation results of performing the FFT calculation on the data sequence can be obtained. In this way, when executing the target calculation stage in the embodiment of the present application, only the multiplication operation of the combination matrix and the input data matrix is required to realize the calculation of the target calculation stage, thereby improving the execution efficiency of the target calculation stage and further improving the execution efficiency of the FFT calculation.
[0174] In one achievable manner, similar to the above-mentioned solution 2, since the rotation factors of the first calculation stage are all 1 and there are no identical rotation factors in the last calculation stage, the target calculation stage can be a calculation stage between the first and last calculation stages. In an embodiment of the present application, a method for determining the DFT system matrix in the last calculation stage based on a trigonometric function with a period of N is also provided, as follows:
[0175] For the rotation factor matrix of the last calculation stage, we can use i∈[0,N m ),j∈[0,N1N2…N m-1 ), calculate k = ij mod N, and get the real and imaginary parts Among them, N is the length of the data sequence, i and j are the rotation factor matrices R m The row and column numbers of the elements in R m _r[i,j] is the rotation factor matrix R m The real data corresponding to the element in row i and column j in R m _i[i,j] is the rotation factor matrix R m The element in row i and column j in N corresponds to the imaginary data. m is the basis corresponding to the last calculation stage, N1N2…N m-1 is the product of the bases corresponding to the various calculation stages after the last calculation stage. In addition, the rotation factor matrix R of the last calculation stage is obtained based on the calculation method provided by this application. m After that, you need to check R m Perform the conversion process, including R m Reshape (convert) to N dimensions m 、N m-1 、N m-2 ...N1 multidimensional array, and then swap the coordinate axes to convert the multidimensional array into a multidimensional array with N dimensions. m-1 、N m 、N m-2 ...N1 multidimensional array, and then reshape the resulting multidimensional array into a multidimensional array with N dimensions. m-1 N m 、N m-2 ...a two-dimensional array of N1. This two-dimensional data is the rotation factor matrix of the final calculation stage.
[0176] For the DFT coefficient matrix of the first calculation stage and the DFT coefficient matrix of the last calculation stage, please refer to the calculation method of the above-mentioned solution 1, which will not be repeated here.
[0177] The present application also provides a matrix operation method that can be applied to the matrix multiplication operation of the input data matrix and the combined matrix in the above embodiment. The method is as follows: after arranging each odd row in the combined matrix to each even row, or after arranging each even row in the combined matrix to each odd row, a rearranged combined matrix is obtained. The product of the input sequence of the target calculation stage and the rearranged combined matrix is determined as the output sequence of the target calculation stage.
[0178] During matrix multiplication, the input data matrix is composed of complex real and imaginary data arranged alternately in rows. Each row and column of the combined matrix includes both complex real and imaginary data, and the real and imaginary data are arranged alternately in each row and column. In the result matrix of performing matrix multiplication on the input data matrix and the combined matrix, the complex real and imaginary data are also arranged alternately in rows. Thus, when reshaping the result matrix into the input sequence for the next calculation stage, it is necessary to read each row of real data and each row of imaginary data in the result matrix at intervals, resulting in discontinuous reading and reducing the computational efficiency of the FFT.
[0179] The matrix operation method provided in the embodiment of the present application can arrange each odd row in the merged matrix to each even row, or arrange each odd row in the merged matrix to each even row, to obtain a rearranged merged matrix. In this way, in the result matrix obtained by performing matrix multiplication operation based on the rearranged merged matrix and the input data matrix, the real part data of the complex numbers are arranged continuously in rows, and the imaginary part data of the complex numbers are arranged continuously in rows. In this way, when the result matrix is reshaped into the input sequence of the next calculation stage, there is no need to read each row of real part data and each row of imaginary part data in the result matrix at intervals, thereby avoiding non-continuous reading and improving the calculation efficiency of FFT.
[0180] FIG7 is a device for performing FFT calculation provided by an embodiment of the present application. The device may be the processor in the above embodiment, and the device includes:
[0181] The acquisition module 710 is used to respond to a Fast Fourier Transform (FFT) calculation request and acquire a data sequence to be subjected to FFT calculation. Specifically, it can be used to implement the acquisition functions of the above step 501 and the implicit steps.
[0182] The division module 720 is used to divide the FFT calculation into multiple calculation stages based on the length of the data sequence, and can be used to implement the division function of the above step 502 and the implicit steps.
[0183] The determination module 730 is used to determine, for at least one target calculation stage among the multiple calculation stages, a merging matrix corresponding to the target calculation stage based on the bases corresponding to the multiple calculation stages and the execution order of the target calculation stage in the multiple calculation stages, wherein the merging matrix is the product of the discrete Fourier transform DFT coefficient matrix in the target calculation stage and the rotation factor, which can be specifically used to implement the determination function of the above-mentioned step 503 and the implicit step.
[0184] The execution module 740 is used to sequentially execute the multiple calculation stages, wherein, when executing the target calculation stage, the product of the input sequence of the target calculation stage and the merge matrix is determined as the output sequence of the target calculation stage, which can be specifically used to implement the execution function of the above-mentioned step 504 and the implicit step.
[0185] The determination module 720 is used to determine the calculation result of performing the FFT calculation on the data sequence based on the output sequence of the last calculation stage in the multiple calculation stages, and can specifically be used to implement the determination function of the above-mentioned step 505 and implicit steps.
[0186] In one implementable manner, when the decimation mode corresponding to the FFT calculation is time domain decimation, the input sequence of the first calculation stage in the multiple calculation stages is changed to the data sequence;
[0187] When the extraction method corresponding to the FFT calculation is frequency domain extraction, the input sequence of the first calculation stage is changed to a sequence after the data sequence is rearranged, wherein the rearrangement processing refers to determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to the multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, and the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
[0188] In one achievable manner, the determining module is configured to:
[0189] Determine a reordering number corresponding to each block included in the target computing stage, wherein the reordering number is used to indicate an arrangement order of each block in the target computing stage when an input sequence of the first computing stage is not changed;
[0190] Based on the rearranged number of each block in the target calculation stage, a merging matrix corresponding to each block is determined.
[0191] In one implementable manner, when the decimation mode of the FFT calculation is frequency domain decimation, the determining module is configured to:
[0192] Obtaining a third multidimensional array, wherein the order of dimensions of the third multidimensional array is opposite to the order of dimensions of the fourth multidimensional array, wherein the elements of the fourth multidimensional array are arranged sequentially starting from 0 and the length is the number of blocks included in the target computation stage, and the dimensions of the fourth multidimensional array are, in order, bases corresponding to the computation stages subsequent to the target computation stage;
[0193] The elements in the third multi-dimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
[0194] In one implementable manner, when the decimation mode of the FFT calculation is time domain decimation, the determining module is configured to:
[0195] Obtaining a fifth multidimensional array, wherein the order of dimensions of the fifth multidimensional array is opposite to the order of dimensions of the sixth multidimensional array, wherein elements in the sixth multidimensional array are arranged sequentially starting from 0 and have a length equal to the number of blocks included in the target computation stage, and wherein the dimensions of the sixth multidimensional array are, in order, bases corresponding to the computation stages preceding the target computation stage;
[0196] The elements in the fifth multi-dimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
[0197] In one implementable manner, when the decimation mode of the FFT calculation is time domain decimation, the determining module is configured to:
[0198] Obtain a first sequence matrix, where the first sequence matrix is a transposed matrix of the second sequence matrix, the second sequence matrix is arranged in row-first order starting from 0, the number of columns of the second sequence matrix is equal to a basis of a calculation stage before the target calculation stage, and the number of rows of the second sequence matrix is equal to a product of respective bases of the calculation stage before the previous calculation stage;
[0199] The elements in the first sequence matrix are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
[0200] In one achievable manner, the device further includes a storage module, configured to:
[0201] For each calculation stage between a second calculation stage and a penultimate calculation stage in the plurality of calculation stages, obtaining a seventh multidimensional array, wherein the order of dimensions after a first-level dimension of the seventh multidimensional array is the same as the order of dimensions before a last-level dimension of the eighth multidimensional array, the dimensions of the first level of the seventh multidimensional array are equal to the dimensions of the last level of the eighth multidimensional array, elements in the eighth multidimensional array are arranged sequentially starting from 0 and a length is the number of blocks included in the calculation stage, the dimensions of the last level are a basis corresponding to a calculation stage immediately preceding the calculation stage, and the dimensions before the last level are, in order, bases corresponding to the calculation stages immediately preceding the previous calculation stage;
[0202] Determining the seventh multidimensional array as the storage order of each block in each calculation stage;
[0203] Based on the storage order, the calculation results of each block in each calculation stage are stored.
[0204] In one achievable manner, the determining module is configured to:
[0205] Based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merged matrix, the imaginary data and the real data of each element in the merged matrix corresponding to each block are determined, wherein the rearrangement number, the position information, and the imaginary data and the real data of each element satisfy the following formula: k=(Alp+Bpq)modN;
[0206] Wherein, l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are the row number and column number of the element in the merge matrix respectively, The real data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in the p-th row and q-th column of the merged matrix corresponding to the l-th block in the n-th calculation stage.
[0207] In one achievable manner, the determining module is configured to:
[0208] Arranging each odd-numbered row in the merged matrix to each even-numbered row, or arranging each even-numbered row in the merged matrix to each odd-numbered row, to obtain a rearranged merged matrix;
[0209] The product of the input sequence of the target calculation stage and the rearranged merge matrix is determined as the output sequence of the target calculation stage.
[0210] In one achievable manner, the partitioning module is configured to:
[0211] Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, the FFT calculation is divided into multiple calculation stages, and a basis corresponding to each calculation stage is determined, wherein the candidate basis is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate basis is proportional to the numerical size of the candidate basis.
[0212] In an implementable manner, the at least one target computing stage is another computing stage among the multiple computing stages except the first computing stage and the last computing stage.
[0213] The division of modules in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, the functional modules in the various embodiments of the present application can be integrated into one processor, or they can exist physically separately, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the device for performing FFT calculations provided in the above embodiments and the method embodiment for performing FFT calculations belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.
[0214] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions for enabling a terminal device (which can be a personal computer, mobile phone, or network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.
[0215] The present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a power management device or stored on any available medium. When the computer program product is run on the power management device, it causes at least one computing device to execute the method for performing FFT calculations provided in the present application.
[0216] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the method for performing FFT calculations provided in the present application.
[0217] In this application, the terms "first", "second", etc. are used to distinguish between identical or similar items having substantially the same effects and functions. It should be understood that there is no logical or temporal dependency between "first" and "second", nor is the quantity and execution order limited. It should also be understood that although the following description uses the terms first, second, etc. to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first multidimensional array can be referred to as the second multidimensional array, and similarly, the second multidimensional array can be referred to as the first multidimensional array. Both the first multidimensional array and the multidimensional array can be collectively referred to as multidimensional arrays, and in some cases, can be separate and different multidimensional arrays.
[0218] The term "at least one" in this application means one or more, and the term "plurality" in this application means two or more.
[0219] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for performing FFT calculation, characterized in that, The method comprises: In response to a Fast Fourier Transform (FFT) calculation request, obtaining a data sequence to be performed on the FFT calculation; Based on the length of the data sequence, dividing the FFT calculation into a plurality of calculation stages; For at least one target calculation stage among the multiple calculation stages, determining a merge matrix corresponding to the target calculation stage based on bases corresponding to the multiple calculation stages respectively and an execution order of the target calculation stage in the multiple calculation stages, wherein the merge matrix is a product of a discrete Fourier transform DFT coefficient matrix in the target calculation stage and a rotation factor; Executing the plurality of calculation stages sequentially, wherein, when executing the target calculation stage, determining the product of the input sequence of the target calculation stage and the merge matrix as the output sequence of the target calculation stage; A calculation result of performing the FFT calculation on the data sequence is determined based on an output sequence of a last calculation stage in the multiple calculation stages.
2. The method according to claim 1, characterized in that When the extraction method corresponding to the FFT calculation is time domain extraction, changing the input sequence of the first calculation stage in the multiple calculation stages to the data sequence; In the case where the extraction method corresponding to the FFT calculation is frequency domain extraction, the input sequence of the first calculation stage is changed to a sequence after the data sequence is rearranged, wherein the rearrangement refers to determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to the multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, wherein the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
3. The method according to claim 2, characterized in that The determining, based on bases respectively corresponding to the plurality of calculation stages and the execution order of the target calculation stage in the plurality of calculation stages, a merge matrix corresponding to the target calculation stage comprises: Determine a reordering number corresponding to each block included in the target computing stage, wherein the reordering number is used to indicate the arrangement order of each block in the target computing stage without changing the input sequence of the first computing stage; Based on the rearrangement number of each block in the target calculation stage, a merge matrix corresponding to each block is determined.
4. The method according to claim 3, characterized in that In the case where the extraction method of the FFT calculation is frequency domain extraction, the determining of the rearrangement number corresponding to each block included in the target calculation stage includes: Obtain a third multidimensional array, wherein the order of dimensions of each level of the third multidimensional array is opposite to the order of dimensions of each level of the fourth multidimensional array, elements in the fourth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multidimensional array are, in sequence, bases corresponding to the calculation stages after the target calculation stage; The elements in the third multidimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
5. The method according to claim 3, characterized in that: In the case where the extraction method of the FFT calculation is time domain extraction, the determining of the reordering number corresponding to each block included in the target calculation stage includes: Obtain a fifth multidimensional array, wherein the order of dimensions of each level of the fifth multidimensional array is opposite to the order of dimensions of each level of the sixth multidimensional array, elements in the sixth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multidimensional array are, in sequence, bases corresponding to the calculation stages before the target calculation stage; The elements in the fifth multidimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
6. The method according to claim 3, characterized in that In the case where the extraction method of the FFT calculation is time domain extraction, the determining of the reordering number corresponding to each block included in the target calculation stage includes: Acquire a first sequence matrix, where the first sequence matrix is a transposed matrix of the second sequence matrix, the second sequence matrix is arranged in order from 0 in a row-first manner, the number of columns of the second sequence matrix is equal to a basis of a calculation stage before the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the bases respectively corresponding to the calculation stages before the previous calculation stage; The elements in the first sequence matrix are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
7. The method according to claim 6, characterized in that The method further comprises: For each calculation stage between the second calculation stage and the penultimate calculation stage in the multiple calculation stages, a seventh multidimensional array is obtained, the order of the dimensions after the first-level dimension of the seventh multidimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multidimensional array, the first-level dimension of the seventh multidimensional array is equal to the last-level dimension of the eighth multidimensional array, the elements in the eighth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage, the last-level dimension is a basis corresponding to a previous calculation stage of the calculation stage, and the dimensions before the last level are, in sequence, bases corresponding to the calculation stages before the previous calculation stage; Determine the seventh multidimensional array as the storage order of each block in each calculation stage; Based on the storage order, the calculation results of each block in each calculation stage are stored.
8. The method according to any one of claims 3 to 7, characterized in that: The step of determining a merge matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage includes: Based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merged matrix, the imaginary data and the real data of each element in the merged matrix corresponding to each block are determined, wherein the rearrangement number, the position information, and the imaginary data and the real data of each element satisfy the following formula: k=(Alp+Bpq)modN; Wherein, l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are the row number and column number of the elements in the merge matrix, respectively. The real data corresponding to the element in the p-th row and q-th column of the merge matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in the p-th row and q-th column of the merge matrix corresponding to the l-th block in the n-th calculation stage.
9. The method according to any one of claims 1 to 8, characterized in that: The step of determining the product of the input sequence of the target calculation stage and the merge matrix as the output sequence of the target calculation stage comprises: After arranging each odd-numbered row in the merged matrix to each even-numbered row, or arranging each even-numbered row in the merged matrix to each odd-numbered row, a rearranged merged matrix is obtained; The product of the input sequence of the target calculation stage and the rearranged merge matrix is determined as the output sequence of the target calculation stage.
10. The method according to any one of claims 1 to 9, characterized in that: The FFT calculation is divided into a plurality of calculation stages based on the length of the data sequence, including: Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, the FFT calculation is divided into multiple calculation stages, and a basis corresponding to each calculation stage is determined, wherein the candidate basis is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate basis is proportional to the numerical size of the candidate basis.
11. The method according to any one of claims 1 to 10, characterized in that: The at least one target computing stage is other computing stages among the multiple computing stages except the first computing stage and the last computing stage.
12. A device for performing FFT calculation, characterized in that: The device comprises: An acquisition module, used for acquiring a data sequence to be executed for FFT calculation in response to a Fast Fourier Transform (FFT) calculation request; A division module, used for dividing the FFT calculation into multiple calculation stages based on the length of the data sequence; A determination module, configured to determine, for at least one target calculation stage among the multiple calculation stages, a merging matrix corresponding to the target calculation stage based on bases corresponding to the multiple calculation stages and an execution order of the target calculation stage in the multiple calculation stages, wherein the merging matrix is a product of a discrete Fourier transform DFT coefficient matrix in the target calculation stage and a rotation factor; an execution module, configured to sequentially execute the plurality of calculation stages, wherein, when executing the target calculation stage, a product of an input sequence of the target calculation stage and the merging matrix is determined as an output sequence of the target calculation stage; The determination module is used to determine the calculation result of performing the FFT calculation on the data sequence based on the output sequence of the last calculation stage in the multiple calculation stages.
13. The device according to claim 12, characterized in that When the extraction method corresponding to the FFT calculation is time domain extraction, changing the input sequence of the first calculation stage in the multiple calculation stages to the data sequence; In the case where the extraction method corresponding to the FFT calculation is frequency domain extraction, the input sequence of the first calculation stage is changed to a sequence after the data sequence is rearranged, wherein the rearrangement refers to determining the data sequence as a first multidimensional array whose dimensions at each level are bases corresponding to the multiple calculation stages in sequence, transposing the first multidimensional array to obtain a second multidimensional array, wherein the order of the dimensions at each level corresponding to the second multidimensional array is opposite to the order of the dimensions at each level corresponding to the first multidimensional array.
14. The device according to claim 13, characterized in that The determining module is used to: Determine a reordering number corresponding to each block included in the target computing stage, wherein the reordering number is used to indicate the arrangement order of each block in the target computing stage without changing the input sequence of the first computing stage; Based on the rearrangement number of each block in the target calculation stage, a merge matrix corresponding to each block is determined.
15. The device according to claim 14, characterized in that When the extraction method of the FFT calculation is frequency domain extraction, the determining module is used to: Obtain a third multidimensional array, wherein the order of dimensions of each level of the third multidimensional array is opposite to the order of dimensions of each level of the fourth multidimensional array, elements in the fourth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multidimensional array are, in sequence, bases corresponding to the calculation stages after the target calculation stage; The elements in the third multidimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
16. The device according to claim 14, characterized in that When the extraction method of the FFT calculation is time domain extraction, the determining module is used to: Obtain a fifth multidimensional array, wherein the order of dimensions of each level of the fifth multidimensional array is opposite to the order of dimensions of each level of the sixth multidimensional array, elements in the sixth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multidimensional array are, in sequence, bases corresponding to the calculation stages before the target calculation stage; The elements in the fifth multidimensional array are sequentially determined as the reordering numbers corresponding to each block included in the target calculation stage.
17. The device according to claim 14, characterized in that When the extraction method of the FFT calculation is time domain extraction, the determining module is used to: Acquire a first sequence matrix, where the first sequence matrix is a transposed matrix of the second sequence matrix, the second sequence matrix is arranged in order from 0 in a row-first manner, the number of columns of the second sequence matrix is equal to a basis of a calculation stage before the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the bases respectively corresponding to the calculation stages before the previous calculation stage; The elements in the first sequence matrix are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
18. The device according to claim 17, characterized in that The device also includes a storage module, which is used to: For each calculation stage between the second calculation stage and the penultimate calculation stage in the multiple calculation stages, a seventh multidimensional array is obtained, the order of the dimensions after the first-level dimension of the seventh multidimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multidimensional array, the first-level dimension of the seventh multidimensional array is equal to the last-level dimension of the eighth multidimensional array, the elements in the eighth multidimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage, the last-level dimension is a basis corresponding to a previous calculation stage of the calculation stage, and the dimensions before the last level are, in sequence, bases corresponding to the calculation stages before the previous calculation stage; Determine the seventh multidimensional array as the storage order of each block in each calculation stage; Based on the storage order, the calculation results of each block in each calculation stage are stored.
19. The device according to any one of claims 14 to 18, characterized in that The determining module is used to: Based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merged matrix, the imaginary data and the real data of each element in the merged matrix corresponding to each block are determined, wherein the rearrangement number, the position information, and the imaginary data and the real data of each element satisfy the following formula: k=(Alp+Bpq)modN; Wherein, l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are the row number and column number of the elements in the merge matrix, respectively. The real data corresponding to the element in the p-th row and q-th column of the merge matrix corresponding to the l-th block in the n-th calculation stage, The imaginary data corresponding to the element in the p-th row and q-th column of the merge matrix corresponding to the l-th block in the n-th calculation stage.
20. The device according to any one of claims 12 to 19, characterized in that The determining module is used to: After arranging each odd-numbered row in the merged matrix to each even-numbered row, or arranging each even-numbered row in the merged matrix to each odd-numbered row, a rearranged merged matrix is obtained; The product of the input sequence of the target calculation stage and the rearranged merge matrix is determined as the output sequence of the target calculation stage.
21. The device according to any one of claims 12 to 20, characterized in that The partitioning module is used for: Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, the FFT calculation is divided into multiple calculation stages, and a basis corresponding to each calculation stage is determined, wherein the candidate basis is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate basis is proportional to the numerical size of the candidate basis.
22. The device according to any one of claims 12 to 21, characterized in that The at least one target computing stage is other computing stages among the multiple computing stages except the first computing stage and the last computing stage.
23. A computing device, characterized in that The computing device includes a processor and a memory; The processor is configured to execute instructions stored in the memory so that the computing device performs the method according to claims 1 to 11.
24. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device, the computing device is caused to perform the method according to claims 1 to 11.
25. A computer-readable storage medium, characterized in that: The method comprises computer program instructions which, when executed by a computing device, cause the computing device to perform the method as claimed in claims 1 to 11.
Citation Information
Patent Citations
Method and device for executing FFT (Fast Fourier Transform) calculation and calculation equipment
CN120067502A
Multi-element LDPC code high-speed parallel decoder based on GPU, and decoding method thereof
CN108462495A
Resonant frequency determination device for servo system and servo system
CN115313956A
Information processing method and device, electronic equipment, storage medium and program product
CN115774992A
Method, device and equipment for executing FFT (Fast Fourier Transform)
CN115859003A
Cited By
Signal processing method, device, equipment and medium
CN121030166A