A calculation method and system for reducing the storage overhead of rotation transformation in FFT

By optimizing the FFT calculation process, n2 discrete Fourier transforms with length n1 are first calculated, and the required data is stored in the on-chip cache of the processor, and then the rotation factor is calculated and the rotation transformation is performed, which solves the problem of large overhead of the rotation factor calculation cache in FFT, achieving more efficient computing performance.

CN115033839BActive Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210506402.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-06-10
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

In fast Fourier transform (FFT), the cache overhead required for rotation factor calculation is large, especially in the case of large calculation scale, resulting in waste of resources and degradation of performance.

Method used

Through the optimization of the calculation process, n2 discrete Fourier transforms with length n1 are first calculated, and the required data is stored in the on-chip cache, then the rotation factor is calculated and the rotation transformation is performed, and finally n1 discrete Fourier transform with length n2 is calculated, effectively reducing cache overhead.

Benefits of technology

While maintaining the calculation amount low, the cache overhead required for rotation factor calculation is significantly reduced, thereby improving the computational efficiency and performance of FFT.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115033839B_ABST
    Figure CN115033839B_ABST
Patent Text Reader

Abstract

The present invention discloses a calculation method and system for reducing the storage overhead of the rotation transformation in FFT. The method includes: Step 1, calculating n2 discrete Fourier transforms of length n1, where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms of length n1, and m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of length n1; Step 2, calculating the rotation factor, performing the rotation transformation, and calculating n1 discrete Fourier transforms of length n2, where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms of length n2. While maintaining a relatively low computational amount, the present invention effectively reduces the cache overhead required for calculating the rotation factor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a calculation method and system for reducing the storage overhead of rotation transformation in FFT (Fast Fourier Transform). Background Art

[0002] The Fast Fourier Transform (FFT) algorithm is listed as one of the "Top Ten Algorithms of the 20th Century" and plays an important role in many large-scale scientific computing fields such as solving differential equations and molecular dynamics. As a component of the HPC Challenge benchmark test program, it can also be used to evaluate the architecture and overall performance of supercomputers. The efficient and scalable FFT algorithm library (FFT-DCP) in large-scale multi-core heterogeneous systems has been listed as a research content in the Exascale Computing Project (ECP) led by the U.S. National Nuclear Security Administration.

[0003] FFT is an efficient implementation of the DFT (Discrete Fourier Transform). Directly calculating an N-point DFT requires approximately O(N2) complex arithmetic operations, where one arithmetic operation includes one multiplication and one addition. The introduction of efficient algorithms can greatly reduce the computational complexity. The key to reducing the computational complexity lies in that among the N2 elements of the DFT matrix (size NxN), only N are different, and a DFT with a composite length can be calculated through multiple shorter DFTs. These algorithms for reducing complexity are collectively referred to as FFT.

[0004] Currently, most FFT algorithms adopt the divide-and-conquer method. For example, converting a DFT of length n=n 1 *n 2 is converted into the following three steps:

[0005] Step 1: Calculate n 2 DFTs of length n 1 , where n is the transformation length, and n 1 and n 2 are factors of n;

[0006] Step 2: Multiply the calculation result in Step 1 by a rotation factor;

[0007] Step 3: Using the output of Step 2 as the input, calculate n 1 DFTs of length n 2 .

[0008] The above process can be performed recursively (continuing to decompose the DFTs in Step 1 and Step 3 using a similar method), and the fast algorithm FFT of the DFT is obtained.

[0009] In the above step 2, the n-point FFT requires a total of (n 1 -1) (n 2 -1) different twiddle factors , where is the unit imaginary number, , are integers, and the calculation of each factor corresponds to a call to the sin and cos functions. When the calculation scale is large, it will consume a large amount of computing resources and cache space.

[0010] Therefore, when performing FFT, how to effectively reduce the cache overhead required for twiddle factor calculation while maintaining a low computational load is an urgent problem to be solved. Summary of the Invention

[0011] In view of this, the present invention provides a calculation method and system for reducing the storage overhead of the rotation transformation in FFT, which can effectively reduce the cache overhead required for twiddle factor calculation while maintaining a low computational load.

[0012] The present invention provides a calculation method for reducing the storage overhead of the rotation transformation in FFT, including:

[0013] Step 1, calculate n 2 discrete Fourier transforms of length n 1 , where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms of length n 1 , where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of length n 1 ;

[0014] Step 2, calculate the twiddle factors, perform the rotation transformation, and calculate n 1 discrete Fourier transforms of length n 2 , where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms of length n 2 .

[0015] Preferably, the calculation of n 2 discrete Fourier transforms of length n 1 includes:

[0016] Step 1.1, load into the on-chip cache of the processor, where is a one-dimensional vector, is the subscript of, is the value range of the data loaded at one time;

[0017] Step 1.2: Divide the data in the processor's on-chip cache into groups according to different subscript values, and perform discrete Fourier transforms on each group respectively, where the length of each group is n 1 ;

[0018] Step 1.3: Output the transformation result of Step 1.2 from the processor's on-chip cache to external memory;

[0019] Step 1.4: If < , update , and transfer to Step 1.1; otherwise, enter Step 2.

[0020] Preferably, calculating the rotation factor, performing rotation transformation, and calculating n 1 discrete Fourier transforms of length n 2 include:

[0021] Step 2.1: Load into the processor's on-chip cache, where is the range of subscript values of the data loaded at one time;

[0022] Step 2.2: Calculate the rotation factor using trigonometric functions;

[0023] Step 2.3: Calculate the rotation factor using the recurrence formula;

[0024] Step 2.4: Calculate the remaining rotation factors and perform rotation transformation;

[0025] Step 2.5: Divide the data in the processor's on-chip cache into groups according to different subscript values, and perform discrete Fourier transforms on each group respectively, where the length of each group is n 2 ;

[0026] Step 2.6: Output the transformation result of Step 2.5 from the processor's on-chip cache to external memory;

[0027] Step 2.7: If < , let , and transfer to Step 2.1; otherwise, the calculation is completed and exit.

[0028] Preferably, the calculating the rotation factor using trigonometric functions includes:

[0029] Step 2.2.1: Calculate the subscript set S of the rotation factors that need to be calculated by calling trigonometric functions.

[0030] Preferably, calculating the subscript set S of the rotation factors that need to be calculated by calling trigonometric functions includes:

[0031] Step 2.2.1.1: Initialize s = +1, S = {}. Where , represents the ceiling function of the real number *, where s is a temporary variable and its value is the subscript of the rotation factor;

[0032] Step 2.2.1.2: If > 1, add to the set S, let s = , and loop to execute Step 2.1.1.2; if = 1, add to the set S, and execute Step 2.2.1.3;

[0033] Step 2.2.1.3: Call the processor's trigonometric function calculation instruction to calculate , and store the calculated rotation factor in the processor's on-chip cache.

[0034] Preferably, calculating the rotation factor using the recurrence formula includes:

[0035] Step 2.3.1: Denote the elements in the set S arranged in ascending order as: , initialize u = 2, v = , where U is the number of elements in the set S, v is a temporary variable, calculate = f( ), , , and then store the value of in the processor's on-chip cache, where f is the trigonometric product-to-sum formula, S u is the u-th element in the set S;

[0036] Step 2.3.2: Update ; calculate = g( ), , and then store the value of in the processor's on-chip cache, where g is the trigonometric sum-to-product formula;

[0037] Step 2.3.3: If u < U - 2, update u += 1, go to Step 2.3.1, otherwise go to Step 2.4.

[0038] Preferably, calculating the remaining rotation factors and performing rotation transformation includes:

[0039] Step 2.4.1. For each group perform rotation transformation operations.

[0040] Preferably, for each group performing rotation transformation operations includes:

[0041] Step 2.4.1.1. If or , that is, the corresponding rotation factors have been calculated, update ; Otherwise, go to Step 2.4.1.2, where the initialized ;

[0042] Step 2.4.1.2. If , calculate the rotation factor , and update , otherwise go to Step 2.4.1.3;

[0043] Step 2.4.1.3. If , calculate the rotation factor , and update , otherwise go to Step 2.4.1.4;

[0044] Step 2.4.1.4. Update , if , go to Step 2.4.1.1.

[0045] A computing system for reducing the storage overhead of rotation transformation in FFT includes:

[0046] A first computing module for calculating n 2 discrete Fourier transforms of length n 1 , where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms of length n 1 , where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of length n 1 ;

[0047] A second computing module for calculating rotation factors, performing rotation transformation, and calculating n 1 discrete Fourier transforms of length n 2 , where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms of length n 2 .

[0048] In summary, the present invention discloses a computing method for reducing the storage overhead of rotation transformation in FFT. First, calculate n2 a discrete Fourier transform of length n 1 where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms of length n 1 where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of length n 1 ; then calculate the twiddle factors, perform the rotation transformation, and calculate n 1 discrete Fourier transforms of length n 2 where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms of length n 2 The present invention can effectively reduce the cache overhead required for calculating the twiddle factors while maintaining a low computational amount. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 It is a schematic flowchart of a calculation method for reducing the storage overhead of the rotation transformation in FFT disclosed in an embodiment of the present invention;

[0051] Figure 2 It is a schematic structural diagram of a calculation system for reducing the storage overhead of the rotation transformation in FFT disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0053] As Figure 1 shown, it is a flowchart of a calculation method for reducing the storage overhead of the rotation transformation in FFT disclosed in an embodiment of the present invention. The hardware suitable for this method has the characteristic that the on-chip cache can be completely controlled by the program, such as DSP, FPGA, etc., and can include the following steps:

[0054] Step 1, calculate n 2 discrete Fourier transforms of length n 1The discrete Fourier transform, where the on-chip cache of the processor can store at most the data required for calculating the discrete Fourier transform of m sequences of length n 1 where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of sequences of length n 1 cache required for calculating the discrete Fourier transform of sequences of length n

[0055] Specifically, the calculation of n in step 1 2 sequences of length n 1 for the discrete Fourier transform may include the following steps:

[0056] Step 1.1: Load into the on-chip cache of the processor, where is a one-dimensional vector, is the subscript, is the range of values of the data loaded at one time;

[0057] Step 1.2: Divide the data in the on-chip cache of the processor into groups according to the different values of the subscript and perform the discrete Fourier transform on each group respectively, where the length of each group is n 1 ;

[0058] Step 1.3: Output the transformation result of step 1.2 from the on-chip cache of the processor to the external memory;

[0059] Step 1.4: If < , update and transfer to step 1.1; otherwise, enter step 2.

[0060] Step 2: Calculate the rotation factor, perform the rotation transformation, and calculate the discrete Fourier transform of n 1 sequences of length n 2 where the on-chip cache of the processor can store at most the data required for calculating the discrete Fourier transform of t sequences of length n 2 for the discrete Fourier transform.

[0061] Specifically, the calculation of the rotation factor, the performance of the rotation transformation, and the calculation of the discrete Fourier transform of n 1 sequences of length n 2 for the discrete Fourier transform may include the following steps:

[0062] Step 2.1: Load into the on-chip cache of the processor, where is the subscript of the data loaded at one time range of values;

[0063] Step 2.2: Calculate the rotation factor using trigonometric functions;

[0064] Specifically, calculating the rotation factor using trigonometric functions may include the following steps:

[0065] Step 2.2.1: Calculate the subscript set S of the rotation factor that needs to be calculated by calling trigonometric functions.

[0066] Specifically, calculating the subscript set S of the rotation factor that needs to be calculated by calling trigonometric functions may include the following steps:

[0067] Step 2.2.1.1: Initialize s = +1, S = {}. Where , represents the ceiling of the real number *, where s is a temporary variable and its value is the subscript of the rotation factor;

[0068] Step 2.2.1.2: If > 1, add to the set S, let s = , and loop to execute Step 2.1.1.2; if = 1, add to the set S and execute Step 2.2.1.3;

[0069] Step 2.2.1.3: Call the processor's trigonometric function calculation instruction to calculate , and store the calculated rotation factor in the processor's on-chip cache.

[0070] Step 2.3: Calculate the rotation factor using the recurrence formula;

[0071] Specifically, calculating the rotation factor using the recurrence formula may include the following steps:

[0072] Step 2.3.1: Denote the elements in the set S arranged in ascending order as: , initialize u = 2, v = , where U is the number of elements in the set S and v is a temporary variable, calculate = f( ), , , and then store the value of in the processor's on-chip cache, where f is the trigonometric product-to-sum formula, S u is the u-th element in the set S;

[0073] Step 2.3.2: Update ; calculate = g( ), , and then The value is stored in the on-chip cache of the processor, where g is the product-to-sum formula of trigonometric functions;

[0074] Step 2.3.3: If u < U - 2, update u += 1 and go to Step 2.3.1; otherwise, go to Step 2.4.

[0075] It should be noted that the input of the function f in the above Steps 2.3.1 and 2.3.2 has been calculated previously and stored in the on-chip cache of the processor. Each transformation with a length of n in Steps 2.2 and 2.3.2 2 The maximum number of cached is, that is, about 1 / 3 of the rotation factors of the transformation length.

[0076] Step 2.4: Calculate the remaining rotation factors and perform rotation transformation;

[0077] Specifically, calculating the remaining rotation factors and performing rotation transformation may include the following steps:

[0078] Step 2.4.1: For each group of , perform rotation transformation operations.

[0079] Specifically, for each group of , performing rotation transformation operations may include the following steps:

[0080] Step 2.4.1.1: If or , that is, the corresponding rotation factor has been calculated, update ; otherwise, go to Step 2.4.1.2, where the initialized ;

[0081] Step 2.4.1.2: If , calculate the rotation factor , and update , otherwise, go to Step 2.4.1.3;

[0082] Step 2.4.1.3: If , calculate the rotation factor , and update , otherwise, go to Step 2.4.1.4;

[0083] Step 2.4.1.4: Update , if , go to Step 2.4.1.1.

[0084] It should be noted that the rotation factors calculated in the above Steps 2.4.1.2 and 2.4.1.3 (about Those that do not need to be put into the on-chip cache can be immediately used for updating after being calculated the value, and then can be discarded.

[0085] Step 2.5: The data in the on-chip cache of the processor , according to the subscript is divided into groups according to different values, and the discrete Fourier transform is performed on each group respectively, where the length of each group is n 2 ;

[0086] Step 2.6: Output the transformation result of Step 2.5 from the on-chip cache of the processor to the external memory;

[0087] Step 2.7: If < , let , and transfer to Step 2.1; otherwise, the calculation is completed and exit.

[0088] In summary, the present invention adopts a method from large to small, and adaptively backtracks the rotation factors that need to be pre-called for trigonometric function calculation according to the transformation length, rather than for all lengths of transformation. Using this method, the on-chip cache overhead required for the rotation transformation calculation can be reduced from 1 / 2 of the size of the input data to 1 / 3 of the size of the input data. While maintaining a relatively low computational complexity, the cache overhead required for the rotation factor calculation is effectively reduced.

[0089] As Figure 2 shown, it is a schematic structural diagram of a computing system for reducing the storage overhead of the rotation transformation in the FFT disclosed in the embodiment of the present invention, which may include:

[0090] The first computing module 201 is used to calculate n 2 discrete Fourier transforms with a length of n 1 , where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms with a length of n 1 , where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform with a length of n 1 ;

[0091] The second computing module 202 is used to calculate the rotation factors, perform the rotation transformation, and calculate n 1 discrete Fourier transforms with a length of n 2 , where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms with a length of n 2 .

[0092] The working principle of the computing system for reducing the storage overhead of the rotation transformation in the FFT disclosed in this embodiment is the same as that of the computing method for reducing the storage overhead of the rotation transformation in the FFT, and will not be elaborated again here.

[0093] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0094] Those skilled in the art can further realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of the examples have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0095] The steps of the methods or algorithms described in combination with the embodiments disclosed in this document can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0096] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A calculation method for reducing the storage overhead of the rotation transformation in FFT, characterized in that, comprising: Step 1, calculate n 2 discrete Fourier transforms of length n 1 where the on-chip cache of the processor can store at most the data required to calculate m discrete Fourier transforms of length n 1 where m = on-chip cache of the processor / cache required to calculate the discrete Fourier transform of length n 1 of the discrete Fourier transform; Step 2: Calculate the rotation factor, perform rotation transformation, and calculate n 1 discrete Fourier transforms of length n 2 where the on-chip cache of the processor can store at most the data required for calculating t discrete Fourier transforms of length n 2 ; Calculating the rotation factor, performing rotation transformation, and calculating the discrete Fourier transform of n 1 with a length of n 2 includes: Step 2.

1. Load into the on-chip cache of the processor, where is the subscript of the data loaded at one time is the value range. Step 2.2: Calculate the rotation factor using trigonometric functions; The calculation of the rotation factor using trigonometric functions includes: Step 2.2.1: Calculate the subscript set S of the rotation factors that need to be calculated by calling trigonometric functions; Step 2.3: Calculate the rotation factor using a recurrence formula; The calculation of the rotation factor using a recurrence formula includes: Step 2.3.1: Denote the elements in set S arranged in ascending order as: , initialize u = 2, v = , where U is the number of elements in set S, v is a temporary variable, calculate = f( ), , , and then store the value of in the on-chip cache of the processor, where f is the product-to-sum formula for trigonometric functions, S u is the u-th element in set S; Step 2.3.2, Update ; Calculate = g( ), , and then store the value of in the on-chip cache of the processor, where g is the product-to-sum formula of trigonometric functions; Step 2.3.3: If u < U - 2, update u += 1, and go to Step 2.3.1, otherwise go to Step 2.4; Step 2.4: Calculate the remaining rotation factors and perform the rotation transformation; Step 2.5, the data in the on-chip cache of the processor , according to the subscript , is divided into groups, and the discrete Fourier transform is performed on each group respectively, where the length of each group is n 2 ; Step 2.

6. Output the transformation result of Step 2.5 from the on-chip cache of the processor to the external memory; Step 2.7, if < , let , go to Step 2.1; otherwise, the calculation is completed and exit.

2. The method according to claim 1, characterized in that, The calculation of n 2 discrete Fourier transforms of length n 1 includes: Step 1.

1. Load into the on-chip cache of the processor, where is a one-dimensional vector, is the subscript of the number of data loaded at one time the value range of Step 1.2: Divide the data in the on-chip cache of the processor into groups according to different values, and perform discrete Fourier transform on each group respectively, where the length of each group is n 1 ; Step 1.

3. Output the transformation result of Step 1.2 from the on-chip cache of the processor to the external memory; Step 1.

4. If < , update , go to step 1.1; otherwise, go to step 2.

3. The method according to claim 2, characterized in that, The calculation of the subscript set S of the rotation factors that need to be calculated by calling trigonometric functions includes: Step 2.2.1.1, initialize s = +1, S = {}, where , represents the ceiling function of the real number *, where s is a temporary variable and its value is the subscript of the rotation factor; Step 2.2.1.2, if > 1, add to set S, let s = , and loop to execute Step 2.1.1.2; if = 1, add to set S and execute Step 2.2.1.3; Step 2.2.1.3: Call the processor's trigonometric function calculation instruction to calculate , store the calculated rotation factor in the on-chip cache of the processor.

4. The method according to claim 3, characterized in that, The calculation of the remaining rotation factors and performing the rotation transformation includes: Step 2.4.1: For each group , perform a rotation transformation operation.

5. The method according to claim 4, characterized in that, Each group Perform a rotation transformation operation, including: Step 2.4.1.

1. If or , that is, the corresponding rotation factor has been calculated, update ; Otherwise, go to Step 2.4.1.2, where the initialized ; Step 2.4.1.2, if , calculate the rotation factor , and update , otherwise go to Step 2.4.1.3; Step 2.4.1.3, if , calculate the rotation factor , and update , otherwise go to Step 2.4.1.4; Step 2.4.1.4, Update If , go to Step 2.4.1.

1.

6. A calculation system for reducing the storage overhead of the rotation transformation in FFT, which is used to implement the method described in any one of claims 1 to 5, characterized in that, comprising: A first calculation module for calculating n 2 discrete Fourier transforms of length n 1 , where the on-chip cache of the processor can store at most the data required for calculating m discrete Fourier transforms of length n 1 , where m = on-chip cache of the processor / cache required for calculating the discrete Fourier transform of length n 1 ; A second computing module, configured to calculate rotation factors, perform rotation transformation, and calculate the discrete Fourier transform of n 1 with a length of n 2 wherein the on-chip cache of the processor can store at most the data required for calculating the discrete Fourier transform of t with a length of n 2 .

Citation Information

Patent Citations

  • Signal processing method, data processing method and device

    CN101930425A

  • Apparatus for realizing fast Fourier transform

    CN107291660A