Method and device for executing FFT (Fast Fourier Transform) calculation and calculation equipment
By dividing the FFT calculation into multiple calculation stages and determining the merge matrix at each stage, the problem of low efficiency in rotation factor multiplication in FFT calculation is solved, and a more efficient calculation process is achieved.
Patent Information
- Application Number
- CN202311638516.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The current FFT calculation involves multiplication calculation of a large number of rotation factors, resulting in low computational efficiency. Although optimization is performed by vector multiplication or matrix multiplication, it still requires a large amount of calculation.
By dividing the FFT calculation into multiple calculation stages, and determining the merge matrix in each target calculation stage, the matrix is the product of the DFT coefficient matrix and the rotation factor, and then the calculation is realized through matrix operations.
The calculation amount of each calculation stage is reduced, the overall efficiency of FFT calculation is improved, the rotation factor calculation is avoided, and the efficiency of calculation is enhanced.
Smart Images

Figure CN120067502A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, an apparatus, and a computing device for performing FFT calculations. Background Art
[0002] Fast Fourier Transform (FFT) is an efficient algorithm for implementing Discrete Fourier Transform (DFT).
[0003] During the current FFT calculation process, a large number of multiplications of twiddle factors are involved. To improve the efficiency of FFT calculation, in related technologies, the multiplication of twiddle factors can be implemented through vector multiplication or matrix multiplication. However, since the multiplication of twiddle factors accounts for a relatively large proportion in the entire FFT calculation, even if the multiplication of twiddle factors is implemented through vector multiplication or matrix multiplication, a large amount of computation is still required, resulting in low efficiency of the current FFT calculation. Summary of the Invention
[0004] Embodiments of the present application provide a method, an apparatus, and a computing device for performing FFT calculations, which can improve the efficiency of performing FFT calculations. The corresponding technical solutions are as follows:
[0005] In a first aspect, a method for performing FFT calculations is provided. The method can be executed by a processor and includes:
[0006] The processor obtains a data sequence to be subjected to FFT calculation in response to a Fast Fourier Transform (FFT) calculation request. Based on the length of the data sequence, the FFT calculation is divided into multiple calculation stages. For at least one target calculation stage among the multiple calculation stages, based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, a combined matrix corresponding to the target calculation stage is determined, where the combined matrix is the product of the Discrete Fourier Transform (DFT) coefficient matrix and the twiddle factor in the target calculation stage. The multiple calculation stages are sequentially executed, and when the target calculation stage is executed, the product of the input sequence of the target calculation stage and the combined matrix is determined as the output sequence of the target calculation stage. Based on the output sequence of the last calculation stage among the multiple calculation stages, the calculation result of performing FFT calculation on the data sequence is determined.
[0007] In the solution shown in this application, during the target calculation stage of performing FFT calculation, the combined matrix after multiplying the DFT coefficient matrix by the twiddle factor can be obtained, and then the calculation of the target calculation stage can be achieved through the combined matrix. Among them, the target calculation stage can be any calculation stage in the FFT calculation process. Through a matrix operation of the combined matrix and the input data matrix composed of the input sequence, the DFT calculation and the twiddle factor calculation can be replaced, thereby improving the execution efficiency of the FFT.
[0008] In an implementable manner, when the decimation method corresponding to the FFT calculation is time-domain decimation, the input sequence of the first calculation stage among multiple calculation stages is changed to the data sequence. When the decimation method corresponding to the FFT calculation is frequency-domain decimation, the input sequence of the first calculation stage is changed to the sequence obtained by rearranging the data sequence. Among them, the rearrangement process refers to determining the data sequence as the first multi-dimensional array with the levels of each dimension being the bases corresponding to multiple calculation stages in turn, transposing the first multi-dimensional array to obtain the second multi-dimensional array, and the order of the levels of each dimension corresponding to the second multi-dimensional array is opposite to the order of the levels of each dimension corresponding to the first multi-dimensional array.
[0009] In the solution shown in this application, by changing the input sequence of the first calculation stage, the elements corresponding to the same twiddle factor in the input sequence of each calculation stage in the FFT calculation can be continuously arranged, thereby avoiding non-continuous reading of the elements in the input sequence during the FFT calculation process, and thus improving the execution efficiency of the FFT.
[0010] In an implementable manner, based on the bases corresponding to multiple calculation stages respectively and the execution order of the target calculation stage among multiple calculation stages, determining the combined matrix corresponding to the target calculation stage includes: determining the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage when the input sequence of the first calculation stage is not changed. Based on the rearrangement numbers of each block in the target calculation stage, determining the combined matrix corresponding to each block.
[0011] In an implementable manner, when the decimation method of the FFT calculation is frequency-domain decimation, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a third multi-dimensional array, the order of the levels of the third multi-dimensional array is opposite to the order of the levels of the fourth multi-dimensional array, the elements in the fourth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multi-dimensional array are the bases corresponding to the calculation stages after the target calculation stage respectively. Determining the elements in the third multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in turn.
[0012] In an implementable manner, when the decimation method of FFT calculation is decimation in time domain, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a fifth multi-dimensional array, where the order of each level of dimensions of the fifth multi-dimensional array is opposite to the order of each level of dimensions of a sixth multi-dimensional array. The elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage. The dimensions of the sixth multi-dimensional array are respectively the bases corresponding to the calculation stages before the target calculation stage in sequence. Determine the elements in the fifth multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0013] In an implementable manner, when the decimation method of FFT calculation is decimation in time domain, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a first sequence matrix, where the first sequence matrix is the transposed matrix of a second sequence matrix. The second sequence matrix is arranged in sequence starting from 0 in a row-major manner. The number of columns of the second sequence matrix is equal to the base of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the bases respectively corresponding to the calculation stages before the previous calculation stage. Determine the elements in the first sequence matrix as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0014] In an implementable manner, the method further includes: for each calculation stage between the second calculation stage and the penultimate calculation stage among multiple calculation stages, obtaining a seventh multi-dimensional array, where the order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of an eighth multi-dimensional array. The dimension of the first level of the seventh multi-dimensional array is equal to the dimension of the last level of the eighth multi-dimensional array. The elements in the eighth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage. The dimension of the last level is the base corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last level are respectively the bases corresponding to the previous calculation stages in sequence. Determine the seventh multi-dimensional array as the storage order of each block in each calculation stage; based on the storage order, store the calculation results of each block in each calculation stage.
[0015] In an implementable manner, based on the rearrangement number of each block in the target calculation stage, determining the merge matrix corresponding to each block includes: based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merge matrix, determining the imaginary part data and real part data of each element in the merge matrix corresponding to each block, where the rearrangement number, position information, and the imaginary part data and real part data of each element satisfy the following formula:
[0016] k = (Alp + Bpq) mod N;
[0017]
[0018]
[0019] Wherein, l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are respectively the row number and column number of the elements in the merging matrix, The real part data corresponding to the element in the p-th row and q-th column of the merging matrix corresponding to the l-th block in the n-th calculation stage, The imaginary part data corresponding to the element in the p-th row and q-th column of the merging matrix corresponding to the l-th block in the n-th calculation stage.
[0020] In the solution shown in the present application, by substituting the rearrangement number of each block in each calculation stage into the corresponding formula, the corresponding combination matrix can be obtained. Compared with the prior art in which the rotation factor matrix and the DFT matrix are first determined and then the matrix multiplication operation of the rotation factor matrix and the DFT matrix is performed to obtain the combination matrix, the calculation amount of obtaining the combination matrix can be reduced, thereby improving the efficiency of performing FFT calculation.
[0021] In a realizable manner, determining the product of the input sequence of the target calculation stage and the merging matrix as the output sequence of the target calculation stage includes: arranging each odd row in the merging matrix after each even row, or arranging each even row in the merging matrix after each odd row to obtain the rearranged merging matrix. Determining the product of the input sequence of the target calculation stage and the rearranged merging matrix as the output sequence of the target calculation stage.
[0022] In the solution shown in the present application, the odd rows and even rows of the merging matrix are rearranged first, and then the rearranged merging matrix and the input data matrix are calculated. In the obtained result matrix, the real part data of the complex numbers are continuously arranged and the imaginary part data of the complex numbers are continuously arranged. In this way, when converting the result matrix into the output sequence, the discontinuous reading of the real part data and the imaginary part data in the result matrix can be avoided, and the efficiency of obtaining the output sequence can be improved, and further the efficiency of performing FFT calculation can be improved.
[0023] In a realizable manner, based on the length of the data sequence, dividing the FFT calculation into multiple calculation stages includes: dividing the FFT calculation into multiple calculation stages based on the length of the data sequence and the priorities corresponding to multiple candidate bases, and determining the base corresponding to each calculation stage, wherein the candidate base is less than or equal to half of the calculation size of the matrix multiplication unit for performing FFT calculation, and the priority of the candidate base is proportional to the numerical size of the candidate base. In this way, by setting the size of the base corresponding to each calculation stage according to the calculation size of the matrix multiplication unit, the utilization rate of the matrix multiplication unit can be improved, thereby improving the efficiency of performing FFT calculation.
[0024] In an implementable manner, at least one target calculation stage is other calculation stages among the multiple calculation stages except the first calculation stage and the last calculation stage.
[0025] In a second aspect, a device for performing FFT calculation is provided. The device includes:
[0026] An acquisition module, configured to acquire a data sequence to be subjected to FFT calculation in response to a fast Fourier transform (FFT) calculation request;
[0027] A division module, configured to divide the FFT calculation into multiple calculation stages based on the length of the data sequence;
[0028] A determination module, configured to, for at least one target calculation stage among the multiple calculation stages, determine a merging matrix corresponding to the target calculation stage based on bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, where the merging matrix is a product of a discrete Fourier transform (DFT) coefficient matrix and a rotation factor in the target calculation stage;
[0029] An execution module, configured to sequentially execute the multiple calculation stages, where when executing the target calculation stage, a product of an input sequence of the target calculation stage and the merging matrix is determined as an output sequence of the target calculation stage;
[0030] A determination module, configured to determine a calculation result of performing FFT calculation on the data sequence based on an output sequence of the last calculation stage among the multiple calculation stages.
[0031] In an implementable manner, when the decimation method corresponding to the FFT calculation is time-domain decimation, the input sequence of the first calculation stage among the multiple calculation stages is changed to the data sequence; when the decimation method corresponding to the FFT calculation is frequency-domain decimation, the input sequence of the first calculation stage is changed to a sequence obtained by rearranging the data sequence, where the rearrangement process refers to determining the data sequence as a first multi-dimensional array with levels of dimensions being bases respectively corresponding to the multiple calculation stages, performing a transpose process on the first multi-dimensional array to obtain a second multi-dimensional array, and the order of levels of dimensions corresponding to the second multi-dimensional array is opposite to the order of levels of dimensions corresponding to the first multi-dimensional array.
[0032] In an implementable manner, the determination module is configured to: determine a rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage when the input sequence of the first calculation stage is not changed; determine a merging matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage.
[0033] In an implementable manner, when the decimation method of FFT calculation is decimation in frequency domain, the determining module is configured to: obtain a third multi-dimensional array, where the order of each level of dimensions of the third multi-dimensional array is opposite to the order of each level of dimensions of a fourth multi-dimensional array, the elements in the fourth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multi-dimensional array are respectively the radixes corresponding to the calculation stages after the target calculation stage in sequence; determine the elements in the third multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0034] In an implementable manner, when the decimation method of FFT calculation is decimation in time domain, the determining module is configured to: obtain a fifth multi-dimensional array, where the order of each level of dimensions of the fifth multi-dimensional array is opposite to the order of each level of dimensions of a sixth multi-dimensional array, the elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multi-dimensional array are respectively the radixes corresponding to the calculation stages before the target calculation stage in sequence; determine the elements in the fifth multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0035] In an implementable manner, when the decimation method of FFT calculation is decimation in time domain, the determining module is configured to: obtain a first sequence matrix, where the first sequence matrix is the transpose matrix of a second sequence matrix, the second sequence matrix is arranged in sequence starting from 0 in a row-major manner, the number of columns of the second sequence matrix is equal to the radix of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the radixes respectively corresponding to the calculation stages before the previous calculation stage; determine the elements in the first sequence matrix as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0036] In an implementable manner, the above device further includes a storage module, configured to: for each calculation stage between the second calculation stage and the penultimate calculation stage among multiple calculation stages, obtain a seventh multi-dimensional array, where the order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of an eighth multi-dimensional array, the first-level dimension of the seventh multi-dimensional array is equal to the last-level dimension of the eighth multi-dimensional array, the elements in the eighth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage, the last-level dimension is the radix corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last-level dimension are respectively the radixes corresponding to the respective calculation stages before the previous calculation stage in sequence; determine the seventh multi-dimensional array as the storage order of each block in each calculation stage; and store the calculation results of each block in each calculation stage based on the storage order.
[0037] In one possible implementation, the determining module is configured to: determine the imaginary part data and real part data of each element in the merging matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merging matrix, where the rearrangement number, the position information, and the imaginary part data and real part data of each element satisfy the following formula:
[0038] k = (Alp + Bpq) mod N;
[0039]
[0040]
[0041] where l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are respectively the row number and column number of the element in the merging matrix, the real part data of the element corresponding to the p-th row and q-th column in the merging matrix corresponding to the l-th block in the n-th calculation stage, the imaginary part data of the element corresponding to the p-th row and q-th column in the merging matrix corresponding to the l-th block in the n-th calculation stage.
[0042] In one possible implementation, the determining module is configured to:
[0043] After arranging each odd row in the merging matrix to each even row, or arranging each even row in the merging matrix to each odd row, obtain the rearranged merging matrix; determine the product of the input sequence of the target calculation stage and the rearranged merging matrix as the output sequence of the target calculation stage.
[0044] In one possible implementation, the partitioning module is configured to: divide the FFT calculation into multiple calculation stages based on the length of the data sequence and the priorities corresponding to multiple candidate bases, and determine the base corresponding to each calculation stage, where the candidate base is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate base is proportional to the numerical size of the candidate base.
[0045] In one possible implementation, at least one target calculation stage is other calculation stages in the multiple calculation stages except the first calculation stage and the last calculation stage.
[0046] In a third aspect, a computing device is provided. The computing device includes a processor and a memory; the processor is configured to execute instructions stored in the memory so that the computing device executes the method described in the first aspect above.
[0047] In a fourth aspect, a computer program product containing instructions is provided. When the instructions are run on a computing device, the computing device is caused to execute the method described in the first aspect above.
[0048] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes computer program instructions. When the computer program instructions are executed by a computing device, the computing device executes the method described in the first aspect above. Description of the Drawings
[0049] Figure 1 is a flowchart for implementing an 18-point FFT calculation through Cooley-Tukey;
[0050] Figure 2 is a schematic diagram of a method for performing an FFT calculation provided by an embodiment of the present application;
[0051] Figure 3 is a butterfly network diagram for implementing an 8-point FFT calculation based on frequency domain decimation and time domain decimation respectively in the related art;
[0052] Figure 4 is a schematic structural diagram of a computing device provided by an embodiment of the present application;
[0053] Figure 5 is a flowchart of a method for performing an FFT calculation provided by an embodiment of the present application;
[0054] Figure 6 is a flowchart of a method for a data sequence provided by an embodiment of the present application;
[0055] Figure 7 is a schematic structural diagram of a device for performing an FFT calculation provided by an embodiment of the present application. Detailed Embodiments
[0056] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0057] The discrete Fourier transform (DFT) is a form in which the Fourier transform is discrete both in the time domain data and the frequency domain data, and is used to transform discrete time domain sampling data into discrete frequency domain sampling data. In terms of data form, the input data (discrete time domain sampling data) and output data (discrete frequency domain data) of the DFT are finite-length complex number sequences, and the lengths of the two complex number sequences are equal.
[0058] Fast Fourier Transform (FFT): An algorithm for quickly calculating the discrete Fourier transform (DFT) or its inverse transform (IDFT). By recursively decomposing the DFT of a long complex number sequence into the DFTs of shorter complex number sequences, it can reduce the original O(N 2) The computational complexity is reduced to O(NlogN), where N is the length of the long complex sequence.
[0059] The Cooley-Tukey algorithm is a commonly used FFT algorithm. Based on the divide-and-conquer strategy, this algorithm can decompose the DFT of a complex sequence with length N into N1 DFTs of length N2 and complex multiplications with twiddle factors, where N = N1 * N2.
[0060] Figure 1 is a flowchart for implementing the 18-point FFT calculation through Cooley-Tukey, and this flowchart can be called a butterfly network. As Figure 1 shown, the calculation process of FFT calculation can be decomposed into multiple calculation stages. The number of calculation stages is equal to the number of times the length of the data sequence for FFT calculation is divided by the radix. Among them, the data sequence is a complex sequence. For the FFT calculation with the length of the data sequence being N, it can be called the N-point FFT calculation. For example, for the N-point FFT calculation, assuming that N can be decomposed into N1, N2,..., Ni (N = N1 * N2 *... * Ni), then the calculation process of the N-point FFT can include i calculation stages. N1, N2,..., Ni can be called the radix. N1 is the radix corresponding to the first calculation stage, N2 is the radix corresponding to the second calculation stage, and Ni is the radix corresponding to the i-th calculation stage.
[0061] In the FFT calculation, the input sequence of the first calculation stage can be obtained from the data sequence corresponding to the execution of the FFT calculation. For other calculation stages after the first calculation stage, the input sequence of each calculation stage is obtained from the output sequence of the previous calculation stage, and the lengths of the input sequences of each calculation stage are the same. A calculation stage can be divided into at least one section, and a section can be divided into at least one butterfly. By performing butterfly calculations on each butterfly in the calculation stage, the output sequence of the calculation stage can be obtained. Among them, there are N / Ni butterflies in a calculation stage, where N is the length of the input data of the calculation stage and Ni is the base corresponding to the calculation stage. The butterfly calculation of the butterfly can be divided into twiddle factor calculation and DFT calculation. In one example, the twiddle factor calculation can be the complex vector multiplication calculation of the input data of the butterfly and the twiddle factor, and the DFT calculation can be the complex matrix multiplication calculation of the rotated input data (the input data after the complex multiplication calculation of the twiddle factor) and the DFT coefficient matrix corresponding to the butterfly. Among them, the DFT coefficient matrix corresponding to the butterfly is related to the structure of the butterfly, and the DFT coefficient matrices corresponding to each butterfly in each calculation stage are the same. Or, in another example, the DFT calculation can be the complex matrix multiplication calculation of the input data of the butterfly and the DFT coefficient matrix, and the twiddle factor calculation can be the complex vector multiplication calculation of the twiddle factor and the calculation result corresponding to the complex matrix multiplication calculation. Among them, the DFT coefficient matrix corresponding to the butterfly is related to the structure of the butterfly, and the DFT coefficient matrices corresponding to each butterfly in each calculation stage are the same. There may be multiple butterflies corresponding to the same twiddle factor in each calculation stage.
[0062] Each calculation stage in the FFT calculation includes a large number of twiddle factor calculations and DFT calculations. Although the twiddle factor calculation can currently be implemented through vector operations or matrix operations to improve the execution efficiency of the FFT calculation, the current optimization effect on the execution efficiency of the FFT calculation is still not ideal.
[0063] The embodiment of the present application provides a method for performing FFT calculation. This method can determine the product of the twiddle factor and the corresponding DFT coefficient matrix in each calculation stage of the FFT calculation before performing the FFT calculation. This product can be called the combined matrix in the embodiment of the present application. In this way, during the process of performing the FFT calculation, in each calculation stage, only by multiplying the input sequence with the corresponding combined matrix, the output sequence of each calculation stage can be obtained. In this way, the multiplication calculation of the twiddle factor during the FFT calculation is avoided, and thus the execution efficiency of the FFT calculation can be improved.
[0064] Figure 2It is a schematic diagram of a method for performing FFT calculation provided by an embodiment of the present application. As Figure 2 shown, originally, a calculation stage may include a matrix multiplication operation of a DFT matrix and an input data matrix composed of the input data of each butterfly unit in the calculation stage, and a vector multiplication operation of the result corresponding to the matrix multiplication and the rotation factor of each butterfly unit. Among them, if the rotation factors corresponding to multiple butterfly units are the same, the rotation factors corresponding to the multiple butterfly units can be converted into a diagonal matrix (which can be called a rotation factor matrix) and perform a matrix operation with the result corresponding to the matrix multiplication. Then, by using the associativity of matrix multiplication, the calculation in this calculation stage may include a matrix multiplication operation of the rotation factor matrix and the DFT matrix, and then a matrix multiplication operation with the input data matrix. Since the DFT matrix and the rotation factor matrix are fixed in each calculation stage, the result of the matrix multiplication operation performed by the DFT matrix and the rotation factor matrix is fixed, that is, the combined matrix corresponding to each calculation stage is fixed. Therefore, during the process of performing FFT, by obtaining the combined matrix of each calculation stage, the calculation corresponding to each calculation stage can be performed, the rotation factor calculation in the calculation stage can be avoided, and thus the calculation amount corresponding to each calculation stage can be reduced, improving the efficiency of performing FFT calculation.
[0065] For the convenience of understanding the embodiments of the present application, some terms involved in the embodiments of the present application are introduced below:
[0066] Multi-dimensional array: An array including at least two dimensions. In one example, a two-dimensional data is a matrix, the number of rows in the matrix is the first dimension in the two-dimensional data, also called the first-level dimension, and the number of columns in the matrix is the second dimension in the two-dimensional data, also called the second-level dimension. Correspondingly, a three-dimensional array includes three levels of dimensions, and an n-dimensional array includes n levels of dimensions.
[0067] Frequency-domain decimation and time-domain decimation: Two methods for implementing FFT calculation. As Figure 3 shown, Figure 3 are the butterfly network diagrams for implementing 8-point FFT calculation based on frequency-domain decimation and frequency-domain decimation respectively in the related art.
[0068] Figure 4 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application. As Figure 4As shown, the computing device 400 may include: a bus 402, a processor 404, a memory 406. Optionally, the computing device 400 may further include a communication interface 408. The processor 404, the memory 406, and the communication interface 408 communicate with each other via the bus 402. The computing device 400 may be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 400. The computing device 400 may be a device for running a model, which may be a terminal or a server. When the computing device 400 is a terminal, the computing device 400 includes, but is not limited to, a desktop computer, a mobile phone, a notebook, a tablet computer, etc. When the computing device 400 is a server, the computing device 400 may be a single server or a server cluster composed of multiple servers, which may be a physical machine or a virtual machine or a container virtualized by virtual technology.
[0069] The bus 402 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only one line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The bus 402 may include a path for transmitting information between various components of the computing device 400 (for example, the memory 406, the processor 404, the communication interface 408).
[0070] The processor 404 may include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP). Among them, the processor 404 may further include a matrix multiplication unit, which can be used to perform matrix multiplication operations involved in the FFT calculation process.
[0071] The memory 406 may include volatile memory, such as random access memory (RAM). The memory 406 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD). The memory 406 may store program code, and the processor 404 may implement the method for performing FFT calculations provided by the embodiments of the present application by executing the program code. For example, in response to a fast Fourier transform (FFT) calculation request, a data sequence to be subjected to the FFT calculation is obtained. Based on the length of the data sequence, the FFT calculation is divided into multiple calculation stages. The multiple calculation stages are sequentially executed. Among them, when executing the target calculation stage, the product of the input sequence of the target calculation stage and the merge matrix is determined as the output sequence of the target calculation stage. Based on the output sequence of the last calculation stage among the multiple calculation stages, the calculation result of performing the FFT calculation on the data sequence is determined, etc.
[0072] The communication interface 408 uses a transceiver module, such as but not limited to a network interface card or a transceiver, to implement communication between the computing device 400 and other devices or communication networks.
[0073] Figure 5 is a flowchart of a method for performing FFT calculations provided by the embodiments of the present application. This method can be executed by the computing device 400 as described above Figure 4 shown, and can further be executed by the processor 404 in the computing device 400. As Figure 5 shown, the method includes:
[0074] Step 501: The processor, in response to a fast Fourier transform (FFT) calculation request, obtains a data sequence to be subjected to the FFT calculation.
[0075] In implementation, an application program capable of performing FFT calculations can run on a computing device. For example, the application program can be a High Performance Computing (HPC) application, an Artificial Intelligence (AI) application, etc. During the running of the application program, when an FFT calculation is required, an execution request for the FFT can be sent to the processor. For example, the application program can be VASP (Vienna Ab-initio Simulation Package). When wave function solving is required in VASP, a wave function solving request can be sent to the processor, and this wave function solving request is the execution request for the FFT. Among them, the wave function is a common function in the field of quantum mechanics, and the solution of the wave function can be obtained by performing an FFT calculation on the input data of the wave function.
[0076] After receiving the execution request for the FFT calculation, the processor can obtain the data sequence to be subjected to the FFT calculation and execute the subsequent steps 502 to 505. Among them, the data sequence to be subjected to the FFT calculation can be carried in the execution request for the FFT calculation.
[0077] In one example, for the method for executing FFT provided in this application, a technician can write the corresponding processing program as an FFT processing function and add it to the math library. Among them, a large number of processing functions can be stored in the math library in the computing device for executing the application program, and these processing functions can be used to implement various mathematical calculations involved in the application program. When an FFT calculation needs to be executed in the application program, an execution request for the FFT can be sent to the processor, and then the processor can call and execute this FFT processing function to implement the processing of the following steps.
[0078] Step 502: Divide the FFT calculation into multiple calculation stages based on the length of the data sequence.
[0079] In one example, after receiving the execution request for the FFT calculation, the processor can obtain the length of the input data to be subjected to the FFT calculation and decompose the FFT calculation into multiple calculation stages. For example, the processor can decompose the length of the input data according to the Cooley-Tukey algorithm to obtain each calculation stage for performing the FFT calculation on the input data and the corresponding radix for each calculation stage. For example, if the number of data elements in the input data is 8, then 8 can be decomposed into 2*2*2, and there are three calculation stages in the process of performing the FFT on the input data, and the corresponding radix for each calculation stage is 2. For example, when the number of data elements in the input data is 27, 27 can be decomposed into 3*3*3, and there are three calculation stages in the process of performing the FFT on the input data, and the corresponding radix for each calculation stage is 3.
[0080] Step 503: For at least one target calculation stage among multiple calculation stages, based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, determine the combined matrix corresponding to the target calculation stage, where the combined matrix is the product of the discrete Fourier transform (DFT) coefficient matrix and the rotation factor in the target calculation stage.
[0081] In FFT calculation, after the number of calculation stages and the base of each calculation stage are determined, the DFT sparse matrix corresponding to each calculation stage and the rotation factor corresponding to each element at each position in the input sequence of each calculation stage are also determined. That is to say, in FFT calculation, after the number of calculation stages and the base of each calculation stage are determined, the combined matrix corresponding to each calculation stage can also be determined.
[0082] The target calculation stage refers to the calculation stage in the FFT calculation process that realizes the DFT matrix calculation and the rotation factor calculation through the combined matrix. The target calculation stage can be any calculation stage among the multiple calculation stages included in the FFT calculation, for example, it can be each calculation stage among the multiple calculation stages. In one example, the number of multiple calculation stages, the execution order of the calculation stage among the multiple calculation stages, the base corresponding to the calculation stage, and the corresponding relationship with the combined matrix can be pre-stored in the computing device. Before performing the FFT calculation, the combined matrix corresponding to the target calculation stage can be determined according to this corresponding relationship.
[0083] Step 504: Sequentially execute the multiple calculation stages, where when executing the target calculation stage, determine the output sequence of the target calculation stage as the product of the input sequence of the target calculation stage and the combined matrix.
[0084] After determining the combined matrix corresponding to the target calculation stage, the multiple calculation stages of the FFT calculation can be sequentially executed. When each time the target calculation stage is executed, the input sequence of the target calculation stage can be converted into an input data matrix, and then the multiplication operation of the input data matrix and the combined matrix can be performed through the matrix multiplication unit included in the processor.
[0085] Step 505: Based on the output sequence of the last calculation stage among the multiple calculation stages, determine the calculation result of performing the FFT calculation on the data sequence.
[0086] After sequentially executing each calculation stage in the FFT calculation, the output sequence of the last calculation stage can be determined as the calculation result of performing the FFT calculation, and this result can be returned to the application program that sends the FFT calculation request.
[0087] In the embodiments of the present application, only one matrix multiplication operation needs to be performed to complete the rotation factor calculation and DFT calculation included in the target calculation stage. In this way, the execution efficiency of the target calculation stage can be improved, and further the efficiency of performing FFT calculation can be improved.
[0088] FFT calculation can be divided into calculation based on time-domain decimation and calculation based on frequency-domain decimation. Next, the embodiments of the present application will be further introduced in combination with the decimation method for implementing FFT calculation to the method for performing FFT calculation provided by the embodiments of the present application:
[0089] Solution 1: Perform FFT calculation based on the frequency-domain decimation method
[0090] Step R1: Obtain the data sequence to be subjected to FFT calculation.
[0091] Step R2: Divide the FFT calculation into multiple calculation stages based on the length of the data sequence.
[0092] In implementation, the FFT calculation can be divided into multiple calculation stages in combination with the length of the data sequence and the calculation size of the matrix multiplication unit for performing matrix multiplication operations in the processor.
[0093] Since in each calculation stage, the size of the combined matrix is equal to the size of the DFT coefficient matrix, and the size of the DFT coefficient matrix is related to the base corresponding to the calculation stage. For example, if the base corresponding to the calculation stage is 8, then the sizes of the DFT coefficient matrix and the combined matrix for this calculation are 8*8. Also, since in actual calculation, the real part data and imaginary part data in the combined matrix will be split into 2*2 matrices, the size of the combined matrix in actual calculation is 16*16. That is to say, when the base of a calculation stage is n, in actual calculation, the combined size corresponding to this calculation stage is 2n*2n.
[0094] The computational size of the matrix multiplication unit can be used to indicate the maximum matrix size supported by the matrix multiplication unit for performing a single matrix multiplication operation. For example, if the maximum matrix size supported by the matrix multiplication unit for performing a single matrix multiplication operation is M*M, then the computational size of this matrix multiplication unit can be M. During the execution of a single matrix multiplication operation by the matrix multiplication unit, the closer the size of the matrix for performing the matrix multiplication operation is to the maximum matrix size corresponding to the matrix multiplication unit, the higher the utilization rate of the matrix multiplication unit. Therefore, when determining the basis corresponding to each computational stage, the basis corresponding to each computational stage can be made as close as possible to half of the computational size of the matrix multiplication unit, thereby improving the utilization rate of the matrix multiplication unit. For example, when the computational size of the matrix multiplication unit is 16 and the basis corresponding to the computational stage is 8, the size of the combined matrix corresponding to this computational stage is 16*16. Therefore, when the matrix multiplication unit executes the operation of this computational stage, this combined matrix can fill the matrix multiplication unit, and thus the performance of the matrix multiplication unit can be fully utilized.
[0095] In one example, multiple candidate bases can be preset, and corresponding priorities can be set for each candidate base. The multiple candidate bases are less than or equal to half of the computational size of the matrix multiplication unit for performing FFT calculations, and the priority of the candidate base is proportional to the numerical value of the candidate base.
[0096] When dividing the FFT calculation into multiple computational stages, according to the set multiple candidate bases, the FFT calculation can be divided into multiple computational stages, and the basis for each computational stage can be determined. For example, the computational size of the matrix multiplication unit is 16, and the multiple candidate bases can be 8, 7, 6, 5, 4, 3, 2 respectively, and the priorities of these multiple candidate bases decrease in turn.
[0097] Figure 6 It is a flowchart of a method for a data sequence provided by an embodiment of the present application. As Figure 6 shown, the method includes:
[0098] Step R21: Set the order list corresponding to the candidate set according to the priorities of the multiple candidate bases, n = [8, 7, 6, 5, 4, 3, 2], initialize k = 0, and initialize the decomposition list of the basis for each computational stage to be empty, that is, base = [].
[0099] Step R22: Divide the length N of the data sequence for performing the FFT calculation by n[k].
[0100] Step R23: Determine whether N can be divided evenly by n[k].
[0101] If it can be divided evenly, execute step R24; if it cannot be divided evenly, execute step R25.
[0102] Step R24: Add n[k] to the decomposition list base, update N = N / n[k], and go to execute Step R22.
[0103] Step R25: Determine whether n[k] is the last element in the sequential list n.
[0104] If not, execute Step R26; if so, execute Step R28.
[0105] Step R26: Let k = k + 1, and transpose to execute Step R22.
[0106] Step R27: Add N to the decomposition list base.
[0107] After the execution of Step R27 is completed, the number of each element included in the decomposition list base is the number of calculation stages in the FFT calculation, and each element included in the decomposition list base is the base corresponding to each calculation stage in turn.
[0108] In this way, through Figure 6 the method of decomposing the data sequence shown, it can be made that the base of each calculation stage in the FFT calculation is as close as possible to half of the calculation size of the matrix multiplication unit, thereby improving the utilization rate of the matrix multiplication unit and the efficiency of executing the FFT calculation.
[0109] Step R3: Rearrange the data sequence for the FFT calculation, and determine the rearranged data sequence as the input sequence of the first calculation stage in the FFT calculation.
[0110] In the process of implementing the FFT calculation in the traditional frequency-domain decimation-based method, the input sequence of the first calculation stage in the FFT calculation is the data sequence for executing the FFT calculation. After the calculation of the last calculation stage is completed, it is necessary to rearrange the output sequence of the last calculation stage to obtain the result sequence of the FFT calculation. However, in the process of implementing the FFT calculation in the traditional frequency-domain decimation-based method, the elements corresponding to the same rotation factor are not continuously arranged in the input sequence of the calculation stage. Thus, in the process of loading the input sequence into the vector register or matrix register in the form of an input data matrix, it is necessary to perform non-continuous reading of the input sequence.
[0111] To avoid the occurrence of the non-continuous reading problem, in the embodiments of the present application, compared with the process of implementing the FFT calculation in the traditional frequency-domain decimation-based method, the arrangement order of the elements in the input sequence of the first calculation stage is modified, so that when the output sequence corresponding to each calculation stage is used as the input sequence corresponding to the next calculation stage, the elements corresponding to the same rotation factor in the next calculation stage can be continuously arranged in this input sequence.
[0112] Rearranging the elements in the input sequence for the first calculation stage may include: determining the data sequence as a first multi-dimensional array with dimensions at each level being the bases corresponding to multiple calculation stages in sequence, performing a transpose operation on the first multi-dimensional array to obtain a second multi-dimensional array, where the order of the dimensions at each level of the second multi-dimensional array is opposite to the order of the dimensions at each level of the first multi-dimensional array.
[0113] For example, in an FFT calculation, there are m calculation stages, and the bases corresponding to the m calculation stages are N1, N2, … Nm respectively. In implementation, the data sequence for the FFT calculation can be converted into a first multi-dimensional array with dimensions at each level being N1, N2, … Nm in sequence. Then, the first multi-dimensional array is transposed into a second multi-dimensional array with dimensions at each level being Nm, Nm - 1, … N1 in sequence. The order of the elements in the second multi-dimensional array is the order of the elements in the input sequence for the first calculation stage after rearrangement. Among them, the input sequence for the first calculation stage after rearrangement is actually also the input sequence for the first calculation stage in the process of implementing FFT calculation in the traditional time-domain decimation method.
[0114] Taking the 8-point FFT calculation as an example, where the 8-point FFT calculation includes 3 calculation stages, for the input sequence N = [0, 1, 2, 3, 4, 5, 6, 7] for the first calculation stage before rearrangement, the input sequence N = [0, 4, 2, 6, 1, 5, 3, 7] for the first calculation stage after rearrangement. As Figure 3 shown, the elements corresponding to the rotation factor W2 / 8 in the second calculation stage before rearrangement are respectively at the 4th and 8th positions in the input sequence. The elements corresponding to the rotation factor W2 / 8 in the second calculation stage after rearrangement are respectively at the 7th and 8th positions in the input sequence.
[0115] It can be seen that in the embodiments of the present application, after rearranging the input sequence for the first calculation stage, the elements corresponding to the same rotation factor in the input sequence for each calculation stage are continuously arranged. Furthermore, in the embodiments of the present application, the elements corresponding to the same rotation factor in the input sequence can be read and stored as each row of the input data matrix into the vector register or matrix register, which can avoid non-continuous reading of the input sequence, and thus can improve the efficiency of performing FFT calculation.
[0116] Step R4: Determine the combination matrix corresponding to the target calculation stage among multiple calculation stages.
[0117] The elements and rotation factors in the DFT coefficient matrix during the calculation stage can all be obtained by calculating trigonometric functions with a period of N, where N is the length of the data sequence for performing FFT calculation. Therefore, for each element in the combination matrix, it can also be obtained by calculating trigonometric functions with a period of N. Therefore, the present application exemplarily provides a method for determining the combination matrix based on trigonometric functions with a period of N, as follows:
[0118] Step R41: Determine the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0119] Since in the above step R2, the input sequence of the first calculation stage is rearranged, resulting in a difference in the order of the blocks in each calculation stage compared to the order of the blocks in each calculation stage in traditional FFT calculation. Therefore, in the embodiments of the present application, the arrangement order of each block in the target calculation stage without changing the input sequence can be determined first, and then the combination matrix in the target calculation stage can be determined according to the rearrangement number.
[0120] Wherein, when the decimation-in-frequency method is used in FFT calculation, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a third multi-dimensional array, the order of the levels of the third multi-dimensional array is opposite to the order of the levels of the fourth multi-dimensional array, the elements in the fourth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multi-dimensional array are the bases corresponding to the calculation stages after the target calculation stage respectively. The elements in the third multi-dimensional array are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
[0121] In implementation, the rearrangement number corresponding to each block included in the target calculation stage can be indicated by a third multi-dimensional array Q3. Wherein, the third multi-dimensional array Q3 can be obtained by transposing a fourth multi-dimensional array Q4. In an example, there are m calculation stages in FFT calculation, the target calculation stage is the nth calculation stage, and it includes r blocks, then the fourth multi-dimensional array Q4 = [0, 1, 2,... r], and the corresponding levels of dimensions are N n+1 、N n+2 、…N m 。 Correspondingly, after the fourth multi-dimensional array Q4 is transposed to obtain the third multi-dimensional array Q3, the corresponding levels of dimensions are N m 、…N n+1 、N n+1 。 Wherein, the value of the first element in the third multi-dimensional array Q3 is the rearrangement number corresponding to the first block in the current target calculation stage.
[0122] Step R42: Determine the merging matrix corresponding to each block in the target calculation stage based on the rearrangement numbers of each block.
[0123] After determining the rearrangement number corresponding to each block, the rearrangement number of each block can be substituted into the following formula to obtain the imaginary part data and real part data of each element in the merging matrix.
[0124] k = (Alp + Bpq) mod N;
[0125]
[0126]
[0127] Where l is the rearrangement number, N is the length of the data sequence, p and q are respectively the row number and column number of the element in the merging matrix, The real part data of the element in the p-th row and q-th column of the merging matrix corresponding to the l-th block in the n-th calculation stage, The imaginary part data of the element in the p-th row and q-th column of the merging matrix corresponding to the l-th block in the n-th calculation stage. A and B are preset coefficients. A is equal to the product of each basis corresponding to the calculation stage before the target calculation stage, and B is equal to the ratio of the data sequence length N to the basis corresponding to the target calculation stage.
[0128] After obtaining the real part data and imaginary part data of each element in the merging matrix corresponding to the target calculation stage, the real part data and imaginary part data of each element can be used for the merging matrix for calculation in the target calculation stage. It should be noted that the number of merging matrices in the target calculation stage is the same as the number of blocks. A merging matrix can be multiplied by multiple butterfly units corresponding to the same rotation factor, and the input data matrix multiplied by the merging matrix is composed of the input data of these multiple butterfly units.
[0129] Step R5: Execute multiple calculation stages in sequence to obtain the calculation result of the FFT calculation.
[0130] After obtaining the combination matrix corresponding to each block in the target calculation stage, each calculation stage included in the FFT calculation can be executed sequentially. When executing to the target calculation stage, the input data of the target calculation stage can be respectively formed into an input data matrix for multiplying with the combination matrix corresponding to each block, and the multiplication operation of the combination matrix and the input data matrix can be implemented based on the matrix multiplication unit, thereby implementing the calculation of the target calculation stage. For example, in the target calculation stage, the combination matrix corresponding to the l-th block is When performing the calculation, this Perform a multiplication operation with the input data matrix composed of the input data of each butterfly unit in the l-th block, and then obtain the calculation result corresponding to the l-th block. In this way, when the target calculation stage is executed in the embodiment of the present application, only the multiplication operation of the combined matrix and the input data matrix needs to be executed to achieve the calculation of the target calculation stage, thereby improving the execution efficiency of the target calculation stage and further improving the execution efficiency of the FFT calculation.
[0131] In an implementable manner, since there are no identical rotation factors in the first calculation stage and the rotation factor in the last calculation stage is 1, the above target calculation stage can be the calculation stages between the first calculation stage and the last calculation stage among multiple calculation stages. The calculations for the first calculation stage and the last calculation stage can adopt traditional calculation methods, which will not be introduced in detail in this application. In addition, in the embodiment of the present application, a method for determining the rotation factor matrix corresponding to the first calculation stage based on trigonometric functions with a period of N, and a method for the DFT system matrix in the first and last calculation stages are also provided as follows:
[0132] For the rotation factor matrix of the first calculation stage, the rearrangement numbers corresponding to each block in the first calculation stage can be obtained. The method for obtaining the rearrangement numbers corresponding to each block in the first calculation stage can refer to the above step R41 and will not be elaborated here. After obtaining the rearrangement numbers corresponding to each block in the first calculation stage, the real part data and imaginary part data corresponding to each element in the rotation factor matrix can be determined through the following calculation formula.
[0133] k = lj mod N;
[0134]
[0135] Where l is the rearrangement number, N is the length of the data sequence, i and j are respectively the row number and column number of the element in the rotation factor matrix R 1 , R 1 _r[i, j] is the real part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R 1 , R 1 _i[i, j] is the imaginary part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R 1 . i ∈ [0, N 2 N 3 …N m ), j ∈ [0, N 1 ), N 1 is the base corresponding to the first calculation stage, and N 2 N 3 …N m is the product of the bases corresponding to the subsequent calculation stages after the first calculation stage.
[0136] For the DFT coefficient matrix in the first calculation stage, for i, j ∈ [0, N1), calculate k = Nij / N 1 mod N, to obtain the real and imaginary parts where, W 1 _r[i, j] is the real part data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W 1 in the first calculation stage, and W 1 _i[i, j] is the imaginary part data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W 1 in the first calculation stage.
[0137] For the DFT coefficient matrix in the last calculation stage, for i, j ∈ [0, N m ), calculate k = Nij / N m mod N, to obtain the real and imaginary parts W m _r[i, j] = C[k], W m _i[i, j] = S[k]. Where, W m _r[i, j] is the real part data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W m in the last calculation stage, and W m _r[i, j] is the imaginary part data corresponding to the element in the i-th row and j-th column of the DFT coefficient matrix W m in the last calculation stage.
[0138] Solution 2: A method for performing FFT calculation based on time-domain decimation
[0139] Step S1: Obtain the data sequence to be subjected to FFT calculation.
[0140] Step S2: Divide the FFT calculation into multiple calculation stages based on the length of the data sequence.
[0141] Among them, the processing of Step S1 and Step S2 is the same as the processing of the above-mentioned Step R1 and Step R2, and will not be elaborated here.
[0142] Step S3: Determine the data sequence of the FFT calculation as the input sequence in the first calculation stage of the FFT calculation.
[0143] In the process of implementing FFT calculation in the traditional time-domain decimation-based method, the input sequence in the first calculation stage of the FFT calculation needs to be obtained by rearranging the data sequence for performing the FFT calculation. Among them, the result after rearrangement is the same as the result after rearrangement in step R3 above. However, in the process of implementing FFT calculation in the traditional time-domain decimation-based method, the elements corresponding to the same rotation factor in the input sequence of the calculation stage are not continuously arranged. Thus, in the process of loading the input sequence into the vector register or matrix register in the form of an input data matrix, the input sequence needs to be read discontinuously.
[0144] In order to avoid the occurrence of discontinuous reading problems, in the embodiment of the present application, compared with the process of implementing FFT calculation in the traditional time-domain decimation-based method, the data sequence of the FFT calculation is used as the input sequence of the first calculation stage, and the data sequence of the FFT calculation is not rearranged. Thus, when the output sequence corresponding to each calculation stage is used as the input sequence corresponding to the next calculation stage, the elements corresponding to the same rotation factor in the next calculation stage can be continuously arranged in this input sequence.
[0145] Taking the 8-point FFT calculation as an example, where the 8-point FFT calculation includes 3 calculation stages, for the input sequence N = [0, 1, 2, 3, 4, 5, 6, 7] of the first calculation stage before rearrangement, the input sequence N = [0, 4, 2, 6, 1, 5, 3, 7] of the first calculation stage after rearrangement. As Figure 3 shown, the elements corresponding to the rotation factor W2 / 8 in the second calculation stage before rearrangement are respectively at the 4th position and the 8th position in the input sequence, and the elements corresponding to the rotation factor W2 / 8 in the second calculation stage after rearrangement are respectively at the 7th position and the 8th position in the input sequence.
[0146] It can be seen that in the embodiment of the present application, after changing the input sequence of the first calculation stage, the elements corresponding to the same rotation factor in the input sequence of each calculation stage are continuously arranged. Furthermore, in the embodiment of the present application, the elements corresponding to the same rotation factor in the input sequence can be read and stored as each row of the input data matrix into the vector register or matrix register, which can avoid the discontinuous reading of the input sequence, and thus can improve the efficiency of performing FFT calculation.
[0147] Step S4, determine the combination matrix corresponding to the target calculation stage among multiple calculation stages.
[0148] Similar to step R4 above, in Solution 2, the embodiment of the present application also provides a method for determining the combination matrix based on trigonometric functions with a period of N, including:
[0149] Step S41: Determine the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0150] Since in the above Step S2, the input sequence of the first calculation stage is rearranged, resulting in a difference in the order of blocks in each calculation stage compared to the order of blocks in each calculation stage in traditional FFT calculations. Therefore, in the embodiments of the present application, the arrangement order of each block in the target calculation stage without changing the input sequence can be determined first, and then the combination matrix in the target calculation stage can be determined according to this rearrangement number.
[0151] Among them, when the decimation method in FFT calculation is decimation in time domain, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a fifth multi-dimensional array, where the order of each level of dimensions of the fifth multi-dimensional array is opposite to the order of each level of dimensions of the sixth multi-dimensional array. The elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage. The dimensions of the sixth multi-dimensional array are the bases corresponding to the calculation stages before the target calculation stage respectively. The elements in the fifth multi-dimensional array are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
[0152] In implementation, the rearrangement number corresponding to each block included in the target calculation stage can be indicated by a fifth multi-dimensional array Q5. Among them, the fifth multi-dimensional array Q5 can be obtained by transposing the sixth multi-dimensional array Q6. In an example, there are m calculation stages in FFT calculation, the target calculation stage is the nth calculation stage, and it includes r blocks. Then the sixth multi-dimensional array Q6 = [0, 1, 2,... r], and the corresponding levels of dimensions are N 1 、N 2 、…N n-1 respectively. Correspondingly, after the sixth multi-dimensional array Q6 is transposed, the fifth multi-dimensional array Q5 has corresponding levels of dimensions N n-1 、…N 2 、N 1 respectively. Among them, N 1 、N 2 、…N n-1 are the bases corresponding to the first n - 1 calculation stages respectively. The elements arranged in sequence in the fifth multi-dimensional array Q5 are the rearrangement numbers corresponding to each block in the target calculation stage. For example, the value of the first element in the fifth multi-dimensional array Q5 is the rearrangement number corresponding to the first block in the current target calculation stage. For example, the value of the second element in the fifth multi-dimensional array Q5 is the rearrangement number corresponding to the second block in the current target calculation stage.
[0153] Step S42: Based on the rearrangement numbers of each block in the target calculation stage, determine the combination matrix corresponding to each block.
[0154] Among them, the processing of step S42 is the same as that of the above step R42, which will not be elaborated here.
[0155] Step S5: Sequentially execute multiple calculation stages to obtain the calculation result of the FFT calculation.
[0156] After obtaining the combination matrix corresponding to each block corresponding to the target calculation stage, each calculation stage included in the FFT calculation can be sequentially executed. When executing to the target calculation stage, the input data of the target calculation stage can be respectively composed into an input data matrix for multiplying with the combination matrix corresponding to each block, and the multiplication operation of the combination matrix and the input data matrix can be implemented based on the matrix multiplication unit, thereby implementing the calculation of the target calculation stage. For example, in the target calculation stage, the combination matrix corresponding to the l-th block is When performing the calculation, this can be multiplied with the input data matrix composed of the input data of each butterfly unit in the l-th block to obtain the calculation result corresponding to the l-th block. In this way, when the embodiment of the present application executes the target calculation stage, only the multiplication operation of the combination matrix and the input data matrix needs to be executed to implement the calculation of the target calculation stage, thereby improving the execution efficiency of the target calculation stage and further improving the execution efficiency of the FFT calculation.
[0157] In an implementable manner, similar to the above solution 1, since the rotation factors of the first calculation stage are all 1 and there are no identical rotation factors in the last calculation stage, the above target calculation stage can be the calculation stages between the first calculation stage and the last calculation stage. In the embodiment of the present application, a method for determining the DFT system matrix in the last calculation stage based on trigonometric functions with a period of N is also provided as follows:
[0158] For the rotation factor matrix of the last calculation stage, the rearrangement numbers corresponding to each block in the last calculation stage can be obtained. The obtaining method of the rearrangement numbers corresponding to each block corresponding to the last calculation stage can refer to the above step S41, which will not be elaborated here. After obtaining the rearrangement numbers corresponding to each block corresponding to the last calculation stage, the real part data and imaginary part data corresponding to each element in the rotation factor matrix can be determined through the following calculation formula.
[0159] k = lj mod N;
[0160]
[0161] Wherein, l is the rearrangement number, N is the length of the data sequence, and i and j are respectively the row number and column number of the elements in the rotation factor matrix R m ; R m _r[i, j] is the real part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R m ; R m _i[i, j] is the imaginary part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R m . i ∈ [0, N 1 N 2 …N m-1 ), j ∈ [0, N m ), N m is the base corresponding to the last calculation stage, and N 1 N 2 ...N m-1 is the product of the bases corresponding to each calculation stage after the last calculation stage.
[0162] For the DFT coefficient matrix of the first calculation stage and the DFT coefficient matrix of the last calculation stage, refer to the calculation method of the above Scheme 1, which will not be elaborated here.
[0163] Scheme 3: Another way to perform FFT calculation based on time-domain decimation
[0164] Step T1: Obtain the data sequence to be subjected to FFT calculation.
[0165] Step T2: Divide the FFT calculation into multiple calculation stages based on the length of the data sequence.
[0166] Step T3: Determine the data sequence of the FFT calculation as the input sequence of the first calculation stage in the FFT calculation.
[0167] Among them, the processing of steps T1 to T3 is the same as the processing of steps R1 to R3 above, which will not be elaborated here. The difference between this Scheme 2 and Scheme 3 is that Scheme 2 rearranges the output sequence of the last calculation stage in the FFT calculation to obtain the calculation result of the FFT, and Scheme 3 rearranges the output sequence of each calculation stage after the second calculation stage in the FFT calculation, and then obtains the calculation result of the FFT.
[0168] Step T4: Determine the combination matrix corresponding to the target calculation stage in multiple calculation stages.
[0169] Same as steps R4 and S4 above, in Scheme 3, the embodiments of the present application also provide a method for determining the combination matrix based on trigonometric functions with a period of N, including:
[0170] Step T41: Determine the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage.
[0171] Since in Solution 3, the output sequence of each calculation stage after each second calculation stage is rearranged, the order of the blocks in each calculation stage is different from the order of the blocks in each calculation stage in the traditional FFT calculation. Therefore, in the embodiments of the present application, the arrangement order of each block in the target calculation stage without rearranging the output sequence can be determined first, and then the combination matrix in the target calculation stage can be determined according to the rearrangement number. The process of rearranging the output sequence of each calculation stage after each second calculation stage can be referred to Step T5 and will not be introduced here.
[0172] Among them, in Solution 3, determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a first sequence matrix, where the first sequence matrix is the transpose matrix of a second sequence matrix. The second sequence matrix is sequentially arranged from 0 in a row-first manner. The number of columns of the second sequence matrix is equal to the radix of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the respective radixes of the calculation stages before the previous calculation stage. The elements in the first sequence matrix are sequentially determined as the rearrangement numbers corresponding to each block included in the target calculation stage.
[0173] In implementation, the rearrangement numbers corresponding to each block included in the target calculation stage can be indicated by a first sequence matrix P1. Among them, the first sequence matrix P1 can be obtained by transposing a second sequence matrix P2. In an example, there are m calculation stages in the FFT calculation, and the corresponding dimensions at each level are N 1 、N 2 、…N m , the target calculation stage is the nth calculation stage and includes r blocks. The elements in the second sequence matrix P2 can be arranged in a row-first manner as [0, 1, 2,... r]. The number of columns of the second sequence matrix is equal to N n-1 , and the number of rows is equal to N 1 *N 2 *…*N n-2 . Correspondingly, after the second sequence matrix P2 is transposed to obtain the first sequence matrix P1, the number of rows of the first sequence matrix P1 is equal to N n-1 , and the number of columns is equal to N 1 *N 2 *…*N n-2。Each element in the first sequence matrix P1 is arranged row by row, which is the rearrangement number corresponding to each block included in the target calculation stage. Among them, the first sequence matrix and the second sequence matrix can also be multi-dimensional arrays respectively. The dimensions of each level before the last level of the multi-dimensional array corresponding to the second sequence matrix are respectively the bases corresponding to the calculation stages before the previous calculation stage, and the dimension of the last level of the multi-dimensional array corresponding to the second sequence matrix is the base corresponding to the previous calculation stage, that is, the dimensions of each level of the multi-dimensional array corresponding to the second sequence matrix are N 1 、N 2 、…N n-2 、N n-1 。Correspondingly, the dimensions of each level of the multi-dimensional array corresponding to the first sequence matrix are N n-1 、N 1 、N 2 、…N n-2 。
[0174] Step T42: Based on the rearrangement numbers of each block in the target calculation stage, determine the combined matrix corresponding to each block.
[0175] Among them, the processing of step T42 is the same as that of the above step R42, and will not be elaborated here.
[0176] Step T5: Sequentially execute multiple calculation stages to obtain the calculation result of the FFT calculation.
[0177] After obtaining the combined matrix corresponding to each block corresponding to the target calculation stage, each calculation stage included in the FFT calculation can be sequentially executed. When executing to the target calculation stage, the input data of the target calculation stage can be respectively composed into an input data matrix for multiplying with the combined matrix corresponding to each block, and the multiplication operation of the combined matrix and the input data matrix can be implemented based on the matrix multiplication unit, thereby implementing the calculation of the target calculation stage.
[0178] In Solution 3 of this scheme, for each calculation stage after the second calculation stage among multiple calculation stages, the calculation results corresponding to each calculation stage can be stored according to the set storage order. Among them, the calculation result corresponding to each block is a short output sequence, and the calculation results corresponding to each block included in a calculation stage can form the output sequence corresponding to this calculation stage. The following introduces each calculation stage after the second calculation stage according to the corresponding storage order.
[0179] For each computational stage between the second computational stage and the penultimate computational stage: The seventh multi-dimensional array can be obtained. The order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multi-dimensional array. The dimension of the first level of the seventh multi-dimensional array is equal to the dimension of the last level of the eighth multi-dimensional array. The elements in the eighth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the computational stage. The dimension of the last level is the basis corresponding to the previous computational stage of the computational stage, and the dimensions before the last level are the bases respectively corresponding to the previous computational stages before the previous computational stage in sequence. The seventh multi-dimensional array is determined as the storage order of each block in each computational stage.
[0180] In implementation, the storage order corresponding to each block in each computational stage can be indicated by the seventh multi-dimensional array Q7. Among them, the seventh multi-dimensional array Q7 is obtained by transposing the eighth multi-dimensional array Q8. The elements in the eighth multi-dimensional array Q8 are arranged in sequence starting from 0 and the corresponding length is the number of blocks included in the computational stage. For example, the computational stage is the nth computational stage in the FFT calculation, and the number of included blocks is m, where the nth computational stage is between the second computational stage and the penultimate second computational stage. Correspondingly, the elements in the eighth multi-dimensional array Q8 can be arranged in the order from 0 to m, and the corresponding levels of dimensions are N n-2 、N n-3 、…N 1 、N n-1 ,where N 1 、…N n-3 、N n-2 、N n-1 are the bases of the previous n - 1 computational stages before the nth computational stage in sequence.
[0181] After determining the eighth multi-dimensional array Q8, device processing can be performed on the eighth multi-dimensional array Q8 to obtain the seventh multi-dimensional array Q7. For example, transpose the above eighth multi-dimensional array Q8 to obtain the seventh multi-dimensional array Q7. The levels of dimensions of this seventh multi-dimensional array Q7 are N n-1 、N n-2 、…N 1 in sequence. After obtaining the seventh multi-dimensional array Q7, the storage order corresponding to the calculation results of each block in the computational stage can be indicated according to the values of the elements in the seventh multi-dimensional array Q7. For example, if the elements in the seventh multi-dimensional array Q7 are 0, 3, 1, 4, 2, 5 in sequence, it can indicate that in the corresponding computational stage, the calculation result of the first block is stored in the first position, the calculation result of the third block is arranged after the calculation result of the first block, the calculation result of the fifth block is arranged after the calculation result of the third block, the calculation result of the second block is arranged after the calculation result of the fifth block, the calculation result of the fourth block is arranged after the calculation result of the second block, and the calculation result of the sixth block is arranged after the calculation result of the fourth block.
[0182] For the penultimate calculation stage: The storage order of the calculation results corresponding to each block in the penultimate calculation stage can be the same as the storage order corresponding to the calculation stage between the second calculation stage and the penultimate calculation stage described above. The difference is that in the penultimate calculation stage, the calculation results corresponding to the same block are stored with an interval of X elements, and the value of X is equal to the number of blocks included in the penultimate calculation stage. And in the penultimate calculation stage, among the calculation results corresponding to multiple blocks, the elements at the same position are stored according to the corresponding calculation order.
[0183] For the last calculation stage: After the calculation in the last calculation stage is completed, Y elements can be read in sequence, and the value of Y is equal to the product of the bases corresponding to the respective calculation stages before the penultimate calculation stage. Then, the elements read each time are stored according to the storage order indicated by the third sequence matrix. Among them, the third sequence matrix is the transpose matrix of the fourth sequence matrix. The fourth sequence matrix is arranged sequentially starting from 0 in row-major order. The number of columns of the fourth sequence matrix is equal to the base corresponding to the last calculation stage, and the number of rows of the fourth sequence matrix is equal to the base corresponding to the penultimate calculation stage.
[0184] After storing the calculation results of each block included in each calculation stage according to the above storage method, the calculation results of performing FFT calculation on the data sequence can be obtained. In this way, when the target calculation stage is executed in the embodiment of the present application, only the multiplication operation of combining the matrix and the input data matrix needs to be executed to achieve the calculation of the target calculation stage, and thus the execution efficiency of the target calculation stage can be improved, and further the execution efficiency of the FFT calculation can be improved.
[0185] In a feasible manner, the same as the above solution two, since the rotation factors in the first calculation stage are all 1 and there are no identical rotation factors in the last calculation stage, the above target calculation stage can be the calculation stages between the first calculation stage and the last calculation stage. In the embodiment of the present application, a method for determining the DFT system matrix in the last calculation stage based on trigonometric functions with a period of N is also provided as follows:
[0186] For the rotation factor matrix of the last calculation stage, let i ∈ [0, N m ) and j ∈ [0, N 1 N 2 …N m-1 ), calculate k = ij mod N to obtain the real and imaginary parts where N is the length of the data sequence, and i and j are the row number and column number of the elements in the rotation factor matrix R m respectively, and R m_r[i, j] is the real part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R m and R m _i[i, j] is the imaginary part data corresponding to the element in the i-th row and j-th column of the rotation factor matrix R m and N m is the basis corresponding to the last calculation stage, and N 1 N 2 …N m-1 is the product of the bases corresponding to each calculation stage after the last calculation stage. Additionally, after obtaining the rotation factor matrix R of the last calculation stage based on the calculation method provided in this application m , it is also necessary to perform a conversion process on R m , including reshaping (converting) R m into a multi-dimensional array with dimensions N m , N m-1 , N m-2 …N 1 in sequence, and then swapping the coordinate axes to convert the multi-dimensional array into a multi-dimensional array with dimensions N m-1 , N m , N m-2 …N 1 in sequence, and then reshaping the obtained multi-dimensional array into a two-dimensional array with dimensions N m-1 N m , N m-2 …N 1 in sequence. This two-dimensional data is the rotation factor matrix of the last calculation stage
[0187] For the DFT coefficient matrix of the first calculation stage and the DFT coefficient matrix of the last calculation stage, refer to the calculation method of the above-mentioned Solution 1, which will not be elaborated here
[0188] The embodiments of this application also provide a method for matrix operation. This method can be applied to the matrix multiplication operation of the input data matrix and the combined matrix in the above embodiments. The method is as follows: After arranging each odd row in the combined matrix behind each even row, or arranging each even row in the combined matrix behind each odd row, a rearranged combined matrix is obtained. The product of the input sequence of the target calculation stage and the rearranged combined matrix is determined as the output sequence of the target calculation stage
[0189] During the matrix multiplication operation, the input data matrices are composed of the real part data and the imaginary part data of complex numbers arranged alternately by rows. In the combined matrix, each row and each column include both the real part data and the imaginary part data of complex numbers, and the real part data and the imaginary part data are arranged alternately in each row and each column. In the result matrix obtained by performing matrix multiplication on the input data matrix and the combined matrix, the real part data and the imaginary part data of complex numbers are also arranged alternately by rows. Thus, when reshaping the result matrix into the input sequence for the next calculation stage, it is necessary to read the real part data of each row in the result matrix at intervals and read the imaginary part data of each row in the result matrix at intervals, which brings discontinuous reading and reduces the calculation efficiency of the FFT.
[0190] The matrix operation method provided by the embodiments of the present application can arrange the odd rows in the combined matrix after the even rows, or arrange the odd rows in the combined matrix after the even rows to obtain the rearranged combined matrix. In this way, in the result matrix obtained by performing matrix multiplication on the rearranged combined matrix and the input data matrix, the real part data of complex numbers are arranged continuously by rows, and the imaginary part data of complex numbers are arranged continuously by rows. Thus, when reshaping the result matrix into the input sequence for the next calculation stage, there is no need to read the real part data of each row and the imaginary part data of each row in the result matrix at intervals, thereby avoiding discontinuous reading and improving the calculation efficiency of the FFT.
[0191] Figure 7 This is a device for performing FFT calculation provided by the embodiments of the present application. The device may be the processor in the above embodiments. The device includes:
[0192] An acquisition module 710, configured to acquire a data sequence to be subjected to FFT calculation in response to a fast Fourier transform (FFT) calculation request, and specifically may be used to implement the acquisition function of the above step 501 and the implicit steps.
[0193] A division module 720, configured to divide the FFT calculation into multiple calculation stages based on the length of the data sequence, and specifically may be used to implement the division function of the above step 502 and the implicit steps.
[0194] A determination module 730, configured to determine, for at least one target calculation stage among the multiple calculation stages, a combined matrix corresponding to the target calculation stage based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, where the combined matrix is the product of the discrete Fourier transform (DFT) coefficient matrix and the rotation factor in the target calculation stage, and specifically may be used to implement the determination function of the above step 503 and the implicit steps.
[0195] The execution module 740 is used to sequentially execute the multiple calculation stages. When executing the target calculation stage, the product of the input sequence of the target calculation stage and the merging matrix is determined as the output sequence of the target calculation stage, which can be specifically used to implement the execution functions of the above step 504 and the implicit steps.
[0196] The determination module 720 is used to determine the calculation result of performing the FFT calculation on the data sequence based on the output sequence of the last calculation stage in the multiple calculation stages, which can be specifically used to implement the determination functions of the above step 505 and the implicit steps.
[0197] In an implementable manner, when the decimation method corresponding to the FFT calculation is time-domain decimation, the input sequence of the first calculation stage in the multiple calculation stages is changed to the data sequence;
[0198] In the case where the decimation method corresponding to the FFT calculation is frequency-domain decimation, the input sequence of the first calculation stage is changed to the sequence obtained by rearranging the data sequence. The rearrangement process refers to determining the data sequence as a first multi-dimensional array with the levels of each dimension being the bases corresponding to the multiple calculation stages in sequence, performing a transpose process on the first multi-dimensional array to obtain a second multi-dimensional array, and the order of the levels of each dimension corresponding to the second multi-dimensional array is opposite to the order of the levels of each dimension corresponding to the first multi-dimensional array.
[0199] In an implementable manner, the determination module is used to:
[0200] Determine the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage when the input sequence of the first calculation stage is not changed;
[0201] Based on the rearrangement numbers of each block in the target calculation stage, determine the merging matrix corresponding to each block.
[0202] In an implementable manner, when the decimation method of the FFT calculation is frequency-domain decimation, the determination module is used to:
[0203] Obtain a third multi-dimensional array, the order of the levels of the third multi-dimensional array is opposite to the order of the levels of the fourth multi-dimensional array. The elements in the fourth multi-dimensional array are sequentially arranged starting from 0 and the length is the number of blocks included in the target calculation stage. The dimensions of the fourth multi-dimensional array are the bases corresponding to the calculation stages after the target calculation stage respectively;
[0204] Determine the elements in the third multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0205] In an implementable manner, when the decimation method of the FFT calculation is decimation in time, the determining module is configured to:
[0206] Obtain a fifth multi-dimensional array, where the order of each level of dimensions of the fifth multi-dimensional array is opposite to the order of each level of dimensions of the sixth multi-dimensional array. The elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage. The dimensions of the sixth multi-dimensional array are the bases corresponding to the calculation stages before the target calculation stage respectively;
[0207] Determine the elements in the fifth multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0208] In an implementable manner, when the decimation method of the FFT calculation is decimation in time, the determining module is configured to:
[0209] Obtain a first sequence matrix, where the first sequence matrix is the transpose matrix of the second sequence matrix. The second sequence matrix is arranged in sequence starting from 0 in row-major order. The number of columns of the second sequence matrix is equal to the base of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the bases corresponding to the calculation stages before the previous calculation stage respectively;
[0210] Determine the elements in the first sequence matrix as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
[0211] In an implementable manner, the apparatus further includes a storage module, configured to:
[0212] For each calculation stage between the second calculation stage and the penultimate calculation stage among the multiple calculation stages, obtain a seventh multi-dimensional array, where the order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multi-dimensional array. The first-level dimension of the seventh multi-dimensional array is equal to the last-level dimension of the eighth multi-dimensional array. The elements in the eighth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage. The last-level dimension is the base corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last-level are the bases corresponding to the respective calculation stages before the previous calculation stage in sequence;
[0213] Determine the seventh multi-dimensional array as the storage order of each block in each calculation stage;
[0214] Based on the storage order, store the calculation results of each block in each calculation stage.
[0215] In an implementable manner, the determining module is configured to:
[0216] Based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merging matrix, determine the imaginary part data and real part data of each element in the merging matrix corresponding to each block, where the rearrangement number, the position information, and the imaginary part data and real part data of each element satisfy the following formula:
[0217] k = (Alp + Bpq) mod N;
[0218]
[0219]
[0220] where l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, p and q are respectively the row number and column number of the element in the merging matrix, The real part data of the element corresponding to the p-th row and q-th column in the merging matrix corresponding to the l-th block in the n-th calculation stage. The imaginary part data of the element corresponding to the p-th row and q-th column in the merging matrix corresponding to the l-th block in the n-th calculation stage.
[0221] In an implementable manner, the determining module is configured to:
[0222] After arranging each odd row in the merging matrix to each even row, or after arranging each even row in the merging matrix to each odd row, obtain the rearranged merging matrix;
[0223] Determine the product of the input sequence of the target calculation stage and the rearranged merging matrix as the output sequence of the target calculation stage.
[0224] In an implementable manner, the partitioning module is configured to:
[0225] Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, divide the FFT calculation into multiple calculation stages, and determine the base corresponding to each calculation stage, where the candidate base is less than or equal to half of the calculation size of the matrix multiplication unit performing the FFT calculation, and the priority of the candidate base is proportional to the numerical size of the candidate base.
[0226] In an implementable manner, the at least one target calculation stage is other calculation stages except the first calculation stage and the last calculation stage among the multiple calculation stages.
[0227] In the embodiments of the present application, the division of modules is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional module can be integrated in a processor, or can exist physically alone, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. In addition, the device for performing FFT calculation provided in the above embodiments and the method embodiments for performing FFT calculation belong to the same concept. The specific implementation process can be found in the method embodiments and will not be elaborated here.
[0228] If the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to enable a terminal device (which can be a personal computer, a mobile phone, or a network device, etc.) or a processor to execute all or part of the steps of the method in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0229] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be a software or program product containing instructions that can run on a power management device or be stored in any available medium. When the computer program product runs on the power management device, it enables at least one computing device to execute the method for performing FFT calculation provided in the embodiments of the present application.
[0230] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method for performing FFT calculation provided in the embodiments of the present application.
[0231] In this application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and effects. It should be understood that there is no logical or temporal dependency between "first" and "second", nor are the quantity and execution order limited. It should also be understood that although the following description uses terms such as first and second to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of various examples, the first multi-dimensional array can be referred to as the second multi-dimensional array, and similarly, the second multi-dimensional array can be referred to as the first multi-dimensional array. The first multi-dimensional array and the multi-dimensional array can both be collectively referred to as the multi-dimensional array, and in some cases, they can be separate and different multi-dimensional arrays.
[0232] In this application, the meaning of the term "at least one" refers to one or more, and the meaning of the term "a plurality of" refers to two or more.
[0233] The above description is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art in the technical field disclosed in this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.
Claims
1. A method for performing FFT calculation, characterized in that, the method includes: responding to a fast Fourier transform (FFT) calculation request, obtaining a data sequence to be subjected to FFT calculation; dividing the FFT calculation into multiple calculation stages based on the length of the data sequence; for at least one target calculation stage among the multiple calculation stages, determining a merging matrix corresponding to the target calculation stage based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages, wherein the merging matrix is the product of a discrete Fourier transform (DFT) coefficient matrix and a rotation factor in the target calculation stage; sequentially executing the multiple calculation stages, wherein when executing the target calculation stage, determining the product of the input sequence of the target calculation stage and the merging matrix as the output sequence of the target calculation stage; determining a calculation result of performing the FFT calculation on the data sequence based on the output sequence of the last calculation stage among the multiple calculation stages.
2. The method according to claim 1, characterized in that, when the decimation method corresponding to the FFT calculation is time-domain decimation, changing the input sequence of the first calculation stage among the multiple calculation stages to the data sequence; when the decimation method corresponding to the FFT calculation is frequency-domain decimation, changing the input sequence of the first calculation stage to a sequence obtained by rearranging the data sequence, wherein the rearrangement process refers to determining the data sequence as a first multi-dimensional array with each level dimension being the bases respectively corresponding to the multiple calculation stages, performing a transpose process on the first multi-dimensional array to obtain a second multi-dimensional array, and the order of each level dimension corresponding to the second multi-dimensional array is opposite to the order of each level dimension corresponding to the first multi-dimensional array.
3. The method according to claim 2, characterized in that, the determining the merging matrix corresponding to the target calculation stage based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage among the multiple calculation stages includes: determining a rearrangement number corresponding to each block included in the target calculation stage, wherein the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage without changing the input sequence of the first calculation stage; determining a merging matrix corresponding to each block based on the rearrangement number of each block in the target calculation stage.
4. The method according to claim 3, characterized in that, when the decimation method of the FFT calculation is frequency-domain decimation, the determining the rearrangement number corresponding to each block included in the target calculation stage includes: obtaining a third multi-dimensional array, the order of each level dimension of the third multi-dimensional array being opposite to the order of each level dimension of a fourth multi-dimensional array, the elements in the fourth multi-dimensional array being sequentially arranged starting from 0 and the length being the number of blocks included in the target calculation stage, and the dimensions of the fourth multi-dimensional array being the bases respectively corresponding to the calculation stages after the target calculation stage; Determine the rearrangement numbers corresponding to each block included in the target calculation stage in sequence from the elements in the third multi-dimensional array.
5. The method according to claim 3, wherein, when the decimation method of the FFT calculation is decimation in time, the determination of the rearrangement numbers corresponding to each block included in the target calculation stage includes: Obtain a fifth multi-dimensional array, the order of the dimensions at all levels of the fifth multi-dimensional array is opposite to the order of the dimensions at all levels of the sixth multi-dimensional array, the elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multi-dimensional array are respectively the bases corresponding to the calculation stages before the target calculation stage; Determine the rearrangement numbers corresponding to each block included in the target calculation stage in sequence from the elements in the fifth multi-dimensional array.
6. The method according to claim 3, wherein, when the decimation method of the FFT calculation is decimation in time, the determination of the rearrangement numbers corresponding to each block included in the target calculation stage includes: Obtain a first sequence matrix, the first sequence matrix is the transpose matrix of the second sequence matrix, the second sequence matrix is arranged in sequence starting from 0 in row-major order, the number of columns of the second sequence matrix is equal to the base of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the bases respectively corresponding to the calculation stages before the previous calculation stage; Determine the rearrangement numbers corresponding to each block included in the target calculation stage in sequence from the elements in the first sequence matrix.
7. The method according to claim 6, wherein, the method further includes: For each calculation stage between the second calculation stage and the penultimate calculation stage among the multiple calculation stages, obtain a seventh multi-dimensional array, the order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multi-dimensional array, the dimension of the first level of the seventh multi-dimensional array is equal to the dimension of the last level of the eighth multi-dimensional array, the elements in the eighth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the calculation stage, the dimension of the last level is the base corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last level are respectively the bases corresponding to the previous calculation stages before the previous calculation stage; Determine the storage order of each block in each calculation stage as the seventh multi-dimensional array; Based on the storage order, store the calculation results of each block in each calculation stage.
8. The method according to any one of claims 3 to 7, wherein, the determination of the merge matrix corresponding to each block based on the rearrangement numbers of each block in the target calculation stage includes: Based on the rearrangement numbers of each block in the target calculation stage and the position information of each element in the merging matrix, determine the imaginary part data and real part data of each element in the merging matrix corresponding to each block, where the rearrangement numbers, the position information, and the imaginary part data and real part data of each element satisfy the following formula: k = (Alp + Bpq) mod N; Wherein, l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, and p and q are respectively the row number and column number of the elements in the merging matrix. The real part data corresponding to the element in the q-th column and p-th row of the merging matrix corresponding to the l-th block in the n-th calculation stage. The imaginary part data corresponding to the element in the q-th column and p-th row of the merging matrix corresponding to the l-th block in the n-th calculation stage.
9. The method according to any one of claims 1 to 8, characterized in that, the multiplication of the input sequence of the target calculation stage and the merging matrix is determined as the output sequence of the target calculation stage, including: After arranging each odd row in the merging matrix to each even row, or arranging each even row in the merging matrix to each odd row, a rearranged merging matrix is obtained; The multiplication of the input sequence of the target calculation stage and the rearranged merging matrix is determined as the output sequence of the target calculation stage.
10. The method according to any one of claims 1 to 9, characterized in that, the dividing the FFT calculation into multiple calculation stages based on the length of the data sequence includes: Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, divide the FFT calculation into multiple calculation stages, and determine the base corresponding to each calculation stage, where the candidate base is less than or equal to half of the calculation size of the matrix multiplication unit that performs the FFT calculation, and the priority of the candidate base is proportional to the numerical size of the candidate base.
11. The method according to any one of claims 1 to 10, characterized in that, the at least one target calculation stage is other calculation stages except the first calculation stage and the last calculation stage in the multiple calculation stages.
12. An apparatus for performing FFT calculation, characterized in that, the apparatus includes: An acquisition module, configured to acquire a data sequence to be subjected to FFT calculation in response to a fast Fourier transform (FFT) calculation request; A division module, configured to divide the FFT calculation into multiple calculation stages based on the length of the data sequence; A determination module, configured to, for at least one target calculation stage in the multiple calculation stages, determine the merging matrix corresponding to the target calculation stage based on the bases respectively corresponding to the multiple calculation stages and the execution order of the target calculation stage in the multiple calculation stages, where the merging matrix is the product of the discrete Fourier transform (DFT) coefficient matrix and the rotation factor in the target calculation stage; An execution module, configured to sequentially execute the multiple calculation stages, where when executing the target calculation stage, the multiplication of the input sequence of the target calculation stage and the merging matrix is determined as the output sequence of the target calculation stage; The determination module is configured to determine the calculation result of performing the FFT calculation on the data sequence based on the output sequence of the last calculation stage in the multiple calculation stages.
13. The apparatus according to claim 12, characterized in that, When the decimation method corresponding to the FFT calculation is decimation in time, change the input sequence of the first calculation stage in the multiple calculation stages to the data sequence; When the decimation method corresponding to the FFT calculation is decimation in frequency, change the input sequence of the first calculation stage to the sequence obtained by rearranging the data sequence, where the rearrangement process refers to determining the data sequence as a first multi-dimensional array with each level of dimension being the radix corresponding to the multiple calculation stages in turn, performing a transpose process on the first multi-dimensional array to obtain a second multi-dimensional array, and the order of each level of dimension corresponding to the second multi-dimensional array is opposite to the order of each level of dimension corresponding to the first multi-dimensional array.
14. The apparatus according to claim 13, wherein, the determining module is configured to: determine the rearrangement number corresponding to each block included in the target calculation stage, where the rearrangement number is used to indicate the arrangement order of each block in the target calculation stage when the input sequence of the first calculation stage is not changed; based on the rearrangement numbers of each block in the target calculation stage, determine the merge matrix corresponding to each block.
15. The apparatus according to claim 14, wherein, when the decimation method of the FFT calculation is decimation in frequency, the determining module is configured to: obtain a third multi-dimensional array, the order of each level of dimension of the third multi-dimensional array is opposite to the order of each level of dimension of the fourth multi-dimensional array, the elements in the fourth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the fourth multi-dimensional array are the radices corresponding to the calculation stages after the target calculation stage respectively; successively determine the elements in the third multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage.
16. The apparatus according to claim 14, wherein, when the decimation method of the FFT calculation is decimation in time, the determining module is configured to: obtain a fifth multi-dimensional array, the order of each level of dimension of the fifth multi-dimensional array is opposite to the order of each level of dimension of the sixth multi-dimensional array, the elements in the sixth multi-dimensional array are arranged in sequence starting from 0 and the length is the number of blocks included in the target calculation stage, and the dimensions of the sixth multi-dimensional array are the radices corresponding to the calculation stages before the target calculation stage respectively; successively determine the elements in the fifth multi-dimensional array as the rearrangement numbers corresponding to each block included in the target calculation stage.
17. The apparatus according to claim 14, wherein, when the decimation method of the FFT calculation is decimation in time, the determining module is configured to: Obtain a first sequence matrix, where the first sequence matrix is the transpose matrix of the second sequence matrix. The second sequence matrix is sequentially arranged starting from 0 in row-major order. The number of columns of the second sequence matrix is equal to the base of the previous calculation stage of the target calculation stage, and the number of rows of the second sequence matrix is equal to the product of the respective corresponding bases of the calculation stages before the previous calculation stage. Determine the elements in the first sequence matrix as the rearrangement numbers corresponding to each block included in the target calculation stage in sequence.
18. The apparatus according to claim 17, wherein, the apparatus further includes a storage module, configured to: For each calculation stage between the second calculation stage and the penultimate calculation stage among the multiple calculation stages, obtain a seventh multi-dimensional array. The order of the dimensions after the first-level dimension of the seventh multi-dimensional array is the same as the order of the dimensions before the last-level dimension of the eighth multi-dimensional array. The dimension of the first level of the seventh multi-dimensional array is equal to the dimension of the last level of the eighth multi-dimensional array. The elements in the eighth multi-dimensional array are sequentially arranged starting from 0 and the length is the number of blocks included in the calculation stage. The last-level dimension is the base corresponding to the previous calculation stage of the calculation stage, and the dimensions before the last level are the respective corresponding bases of the previous calculation stages before the previous calculation stage in sequence; Determine the seventh multi-dimensional array as the storage order of each block in each calculation stage; Based on the storage order, store the calculation results of each block in each calculation stage.
19. The apparatus according to any one of claims 14 to 18, wherein, the determination module is configured to: Based on the rearrangement number of each block in the target calculation stage and the position information of each element in the merge matrix, determine the imaginary part data and real part data of each element in the merge matrix corresponding to each block, where the rearrangement number, the position information, and the imaginary part data and real part data of each element satisfy the following formula: k = (Alp + Bpq) mod N; where l is the rearrangement number, A and B are preset coefficients, N is the length of the data sequence, and p and q are the row number and column number of the elements in the merging matrix, respectively. The real part data corresponding to the element in the q-th column and p-th row of the merging matrix corresponding to the l-th block in the n-th calculation stage. The imaginary part data corresponding to the element in the q-th column and p-th row of the merging matrix corresponding to the l-th block in the n-th calculation stage.
20. The apparatus according to any one of claims 12 to 19, wherein, the determination module is configured to: After arranging the odd rows in the merge matrix to the even rows, or after arranging the even rows in the merge matrix to the odd rows, obtain a rearranged merge matrix; Determine the product of the input sequence of the target calculation stage and the rearranged merge matrix as the output sequence of the target calculation stage.
21. The apparatus according to any one of claims 12 to 20, wherein, the division module is configured to: Based on the length of the data sequence and the priorities corresponding to multiple candidate bases, divide the FFT calculation into multiple calculation stages and determine the base corresponding to each calculation stage, where the candidate base is less than or equal to half of the calculation size of the matrix multiplication unit performing the FFT calculation, and the priority of the candidate base is proportional to the numerical size of the candidate base.
22. The apparatus according to any one of claims 12 to 21, wherein, The at least one target computing stage is other computing stages in the plurality of computing stages except the first computing stage and the last computing stage.
23. A computing device, wherein, the computing device includes a processor and a memory; the processor is configured to execute instructions stored in the memory, so that the computing device executes the method according to claims 1 to 11.
24. A computer program product containing instructions, wherein, when the instructions are run by a computing device, the computing device is caused to execute the method according to claims 1 to 11.
25. A computer-readable storage medium, wherein, it includes computer program instructions, and when the computer program instructions are executed by a computing device, the computing device executes the method according to claims 1 to 11.
Citation Information
Cited By
Method and apparatus for executing FFT computation, and computing device
EP4804053A1
Method and apparatus for executing FFT computation, and computing device
WO2025112485A1