FPGA-based 256K point fast Fourier transform method and device, and medium
By decomposing the 256K point FFT into a 22-level basic processing unit and using the 18-level delay and butterfly calculation method, the problems of low FFT calculation efficiency and high storage resources in the prior art are solved, and efficient FFT processing and resource utilization are achieved.
Patent Information
- Application Number
- CN202510594628.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-09
AI Technical Summary
The existing FPGA-based FFT method has low computational efficiency and high storage resources when processing large points, making it difficult to meet the real-time computing needs with configurable accuracy.
By decomposing the FFT of 256K points into 210×28 and further decomposing it into a 22-level basic processing unit, the 18-level delay and butterfly calculation method is adopted, and the rotation factor is used to calculate and store it in advance to achieve efficient data processing.
It reduces the storage usage of the rotation factor, reduces the use of multiplier and storage resources, improves the computing efficiency and utilization of storage resources, and adapts to the two precisions of 16-bit and 32-bit.
Smart Images

Figure CN120104933A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital signal processing, and in particular to a 256K-point fast Fourier transform method, device and medium based on FPGA. Background Art
[0002] DFT (Discrete Fourier Transform) is one of the most effective digital signal processing methods, and is widely used in various fields such as astronomical signal processing, radar signal processing, and mobile communications. In 1965, Cooley and Tukey proposed the fast Fourier transform algorithm using the symmetry of the rotation factor, which greatly improved the calculation speed of DFT and reduced the calculation complexity of DFT. Since then, FFT (Fast Fourier Transform) has become one of the commonly used digital signal processing methods. With the development of front-end collectors, the sampling rate is getting higher and higher, and there is a further demand for the speed and depth of FFT processing. FFT with a large number of points has received more and more attention. The increase in FFT depth means that more storage space, more multipliers, and more calculation delays are required.
[0003] Currently, the main high-performance computing processors for FFT are Digital Signal Processor (DSP), Application Specific Integrated Circuit (ASIC) and Field Programmable Gate Array (FPGA). Among them, FPGA is more flexible than DSP and ASIC, and integrates a large number of parallel computing resources and storage resources internally, which is highly compatible with the pipeline FFT processing structure. Therefore, FPGA-based FFT methods have been widely studied and applied.
[0004] At present, many scholars have conducted a lot of research on the FFT method based on FPGA. However, in order to meet the real-time computing requirements of FFT with configurable precision for a large number of points, the current FFT method has the problems of low computing efficiency and high storage resources. Summary of the invention
[0005] The purpose of the present invention is to provide a 256K-point fast Fourier transform method, device and medium based on FPGA to address the deficiencies of the prior art. The specific technical solution is as follows:
[0006] A 256K-point fast Fourier transform method based on FPGA comprises the following steps:
[0007] Step 1: Determine that the calculation accuracy is 16 bits or 32 bits, calculate the rotation factors of 16, 64, 256, 1024, and 256K points in advance and store them in storage units; instantiate rotation factor storage units of different bit widths according to the calculation accuracy;
[0008] Step 2: Use base 2 2 The butterfly basic processing unit first decomposes the 256K point data stream into 2 10 ×2 8 , and then 2 10 Decompose into 2 2 ×2 8 , 2 8 Decompose into 2 2 ×2 6 , and then 2 6 Decompose into 2 2 ×2 4 , and then 2 4 Decompose into 2 2 ×2 2 , the last 2 2 Decompose into 2×2, perform 18 levels of delay and butterfly calculation, that is, the 18-level rotation factor basis is 2 2 , 2 10 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 18 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 0 .
[0009] Furthermore, in step 2, 18 levels of delay and butterfly calculation are performed, wherein:
[0010] The specific operation of odd-numbered level decomposition in 18 levels is as follows:
[0011] The S level is divided into 2 S-1 group operation, S represents the series, S is an odd number between 1 and 18; in order 2 19-S Points are grouped together, and the first 2 18-S points to cache; in the latter 2 18-S When you click Enter, the first 2 18-S The points are read out in sequence, with an interval of 2 18-SThe two points of the mth group are operated in pairs by butterfly operation; the two data of the nth pair of butterflies of the mth group are added to obtain the n+2th pair of butterflies. 19-S ×(m-1) point output, the two data of the nth pair of butterflies in the mth group are subtracted to get the n+2th pair of butterflies 19-S ×(m-1)+ 2 18-S Point output; the result of adding two data is directly output; the result of subtracting two data is written into the storage unit in place, and waits for the second 18-S After the butterfly calculation is completed, it is read out from the first storage address in sequence; then the subsequent groups of data are calculated in the same way; the last 2 of each group of data output 17-S The point data is multiplied by -j, where j*j=-1, that is, the sign bit of the real part is inverted to become the imaginary part, and the imaginary part becomes the real part;
[0012] The specific operation of even-numbered level decomposition in 18 levels is as follows:
[0013] The S level is divided into 2 S-1 group operation, S represents the series, S is an even number between 1 and 18; in order 2 19-S Points are grouped together, and the first 2 18-S points to cache; in the latter 2 18-S When you click Enter, the first 2 18-S The points are read out in sequence, with an interval of 2 18-S The two points of the mth group are operated in pairs by butterfly operation; the nth pair of butterflies of the mth group, the two data are added to get the n+2th pair of butterflies. 19-S ×(m-1) point output, subtract the two data to get the n+2 19-S ×(m-1)+ 2 18-S Point output; the result of adding two data is directly output; the result of subtracting two data is written into the storage unit in place, and waits for the second 18-S After the butterfly calculation is completed, it is read out from the first storage address in sequence; then the subsequent groups of data are calculated in the same way; the 256K points output by the first sixteen levels are multiplied by the calculated rotation factors respectively, and the calculated data is intercepted with high-order data according to the precision configuration, and the bit width is kept unchanged, and the eighteenth level data is directly output to obtain the final FFT result.
[0014] Furthermore, in the step 1, the rotation factors of 16, 64, and 256 points are calculated in advance and stored in registers respectively; the rotation factors of 1024 and 256K points are calculated in advance and stored in ROM respectively.
[0015] Furthermore, when the depth of the delay unit performing the delay calculation is not less than 256, the delay is implemented using a memory; otherwise, the delay is implemented using an on-chip logic unit.
[0016] Furthermore, when performing 18 levels of delay and butterfly calculation, the multipliers only appear at the 2nd, 4th, 6th, 8th, 10th, 12th, 14th, and 16th levels.
[0017] Furthermore, when the multiplier performs complex multiplication, it is implemented by three real number multiplications and five real number additions.
[0018] Furthermore, the storage of the rotation factors only stores 1 / 8 of the length.
[0019] Furthermore, the rotation factors of even levels are obtained as follows:
[0020] The 256K points calculated by each level of butterfly are multiplied by the rotation factors respectively, and the output index is represented by 18-bit binary index[17:0];
[0021] The second level rotation factors are ;
[0022] The fourth level rotation factor is ;
[0023] The sixth level rotation factor is ;
[0024] The eighth level rotation factor is ;
[0025] The tenth level rotation factor is ;
[0026] The twiddle factor of the twiddle level is ;
[0027] The fourteenth level rotation factors are ;
[0028] The sixteenth level rotation factor is .
[0029] A 256K-point fast Fourier transform device based on FPGA, including 18 butterfly computing units, 8 rotation factor generation units, and a depth of 2 i The delay unit, i is an integer from 0 to 17;
[0030] The butterfly computing unit is a radix-2 butterfly computing unit, which is used to implement butterfly operations on data and complete data interleaving;
[0031] The 8 rotation factor generation units are located in the 2nd, 4th, 6th, 8th, 10th, 12th, 14th, and 16th level calculations, respectively.
[0032] The second-level rotation factor generation unit is used to generate 1024-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0033] The fourth-level rotation factor generation unit and the twelfth-level rotation factor generation unit are both used to generate 256-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0034] The sixth-level rotation factor generation unit and the fourteenth-level rotation factor generation unit are both used to generate 64-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0035] The eighth-level rotation factor generation unit and the sixteenth-level rotation factor generation unit are both used to generate 16-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0036] The tenth-level rotation factor generation unit is used to generate 256k-point rotation factors with 16-bit or 32-bit precision respectively;
[0037] The delay unit includes a memory and an address control unit to achieve data delays of different lengths; the memory is used to store the data input of this stage and the output of the butterfly operation subtraction; the address control unit is used to control the read and write address of the memory and the control signal of the memory to achieve the delay function.
[0038] A computer-readable storage medium stores a program, which, when executed by a processor, is used to implement the 256K-point fast Fourier transform method based on FPGA.
[0039] The beneficial effects of the present invention are as follows:
[0040] (1) The present invention converts the FFT as a whole into 2 10 ×2 8 Decomposition can reduce the storage occupancy of the rotation factors.
[0041] (2) The present invention can use logic to implement delay when the delay is short, thereby reducing the use of on-chip storage resources and improving the utilization rate of storage resources; at the same time, only 8 levels of the 18-level processing unit require multipliers, thereby reducing the use of multipliers.
[0042] (3) The present invention can adapt to two different precisions: 16 bits and 32 bits. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is an overall architecture diagram of a 256K-point FFT based on FPGA according to an embodiment of the present invention;
[0044] Figure 2 is a schematic diagram of the computing structure of each level of an embodiment of the present invention;
[0045] Figure 3 is a schematic diagram of a butterfly computing structure according to an embodiment of the present invention;
[0046] Figure 4 is a binary tree decomposition graph, where Figure 4 (a) in the figure is a binary tree decomposition diagram of base 2. Figure 4 (b) in the formula is base 2 2 The binary tree decomposition diagram of Figure 4 (c) is a binary tree decomposition diagram used in an embodiment of the present invention.
[0047] Figure 5 Schematic diagram of a 256K-point fast Fourier transform device based on FPGA according to an embodiment of the present invention. DETAILED DESCRIPTION
[0048] The present invention will be described in detail below based on the accompanying drawings and preferred embodiments, and the purpose and effects of the present invention will become more clear. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0049] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0050] The terms used in the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "the" used in the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates other meanings. It should also be understood that the terms "and, or" used herein refer to and include any or all possible combinations of one or more associated listed items.
[0051] It should be understood that although the terms first, second, third, etc. may be used in the present invention to describe various information, these information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present invention, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0052] The present invention is described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the features of the following embodiments and implementations can be combined with each other.
[0053] like Figure 1As shown, the 256K-point fast Fourier transform method based on FPGA of the present invention is applied in astronomical signal processing, and includes the following steps:
[0054] Step 1: Determine whether the calculation precision is 16 bits or 32 bits, calculate the rotation factors of 16, 64, 256, 1024, and 256K points in advance and store them in storage units, calculate 16-bit and 32-bit precisions, and select rotation factor storage units with different bit widths according to the precision, that is, if the precision is 16 bits, instantiate a rotation factor storage unit with a precision of 16, and if the precision is 32 bits, instantiate a rotation factor storage unit with a precision of 32.
[0055] Step 1 includes the following sub-steps:
[0056] (1.1) Calculation , and enlarge the calculation results by 2 15 and 2 31 , get the results of 16-bit precision and 32-bit precision and store them in registers;
[0057] (1.2) Calculation , and enlarge the calculation results by 2 15 and 2 31 , get the results of 16-bit precision and 32-bit precision and store them in registers;
[0058] (1.3) Calculation , and enlarge the calculation results by 2 15 and 2 31 , get the results of 16-bit precision and 32-bit precision and store them in registers;
[0059] (1.4) Calculation , and enlarge the calculation results by 2 15 and 2 31 , get the results of 16-bit precision and 32-bit precision, and store them in read-only memory (ROM);
[0060] (1.5) Calculation , and enlarge the calculation results by 2 15 and 2 31 , get the results of 16-bit precision and 32-bit precision and store them in ROM;
[0061] (1.6) Determine whether the calculation precision is 16 bits or 32 bits, and instantiate rotation factor storage units of different bit widths according to the precision.
[0062] Step 2: Use base 2 2 The butterfly basic processing unit first decomposes the 256K point data stream into 2 10×2 8 , and then 2 10 Decompose into 2 2 ×2 8 , 2 8 Decompose into 2 2 ×2 6 , and then 2 6 Decompose into 2 2 ×2 4 , and then 2 4 Decompose into 2 2 ×2 2 , the last 2 2 Decompose into 2×2, perform 18 levels of delay and butterfly calculation, that is, the 18-level rotation factor basis is 2 2 , 2 10 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 18 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 0 .
[0063] The overall structure is as follows Figure 1 As shown in the figure, it is divided into 18 levels of computing units, and the size of the delay unit decreases gradually. The computing structure of each level is as follows Figure 2 As shown, the butterfly computing architecture is Figure 3 shown. Figure 4 (a) is the binary tree decomposition diagram of base 2, and (b) is the base 2 2 (c) is the binary tree decomposition diagram used in the present invention, and the numbers in the figure represent powers of 2. The decomposition method in the present invention can reduce the number of multipliers compared to the traditional radix 2 decomposition. 2 Decomposition reduces the number of large subscript values of large rotation factors, thereby reducing the storage requirements of the rotation factors, as shown in Tables 1 and 2 below.
[0064] Table 1 Figure 4 The number of adders required for the three different binary tree decomposition methods
[0065]
[0066] Table 2 Figure 4 Storage requirements for rotation factors of three different binary tree decomposition methods
[0067]
[0068] Step 2 specifically includes the following sub-steps:
[0069] (2.1) The first stage performs the following operations: the first 128K points of the input data are cached using the on-chip RAM. When the last 128K points are input, the first 128K points are read out in sequence. The butterfly operation is performed in pairs of two points separated by 128K, for a total of 128K pairs, 1 group. For the nth pair of butterflies, the two data are added to obtain the nth point output, and the two data are subtracted to obtain the (n+128K)th point output. The result of the addition of the two data is directly output to the second stage, and the result of the subtraction of the two data is written into the storage unit in situ. After the 128Kth butterfly calculation is completed, it is read out from the first storage address to the second stage in sequence, and the last 64K points of the output data are multiplied by -j (j*j=-1), that is, the real part sign bit is inverted to the imaginary part, and the imaginary part is converted to the real part, without using multiplier resources.
[0070] (2.2) The second stage performs the following operations: The second stage is divided into 2 groups of operations, with 128K points as a group in sequence. The first 64K points of the input data are cached and cached using the on-chip RAM. When the last 64K points are input, the first 64K points are read out in sequence, and butterfly operations are performed on two points separated by 64K in pairs, for a total of 64K pairs, 2 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+128K×(m-1)th point output, and the two data are subtracted to obtain the (n+128K×(m-1)+64K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 64Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; then the second group of data is calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is ,calculate , the address read from the rotation factor storage unit is W addr The result read is , the actual rotation factor is , the result is as shown in formula 1 according to the different intervals of addr. The calculated data is intercepted according to the precision configuration to keep the bit width unchanged.
[0071] (1)
[0072] The complex multiplication with the rotation factor is implemented using three multiplications and five additions to reduce the use of multipliers, as follows:
[0073] ;
[0074] (2.3) The third level performs the following operations: The third level is divided into 4 groups of operations, with 64K points as one group in order. The first 32K points of the input data are cached and cached using the on-chip RAM. When the last 32K points are input, the first 32K points are read out in order, and butterfly operations are performed in pairs of two points separated by 32K, for a total of 32K pairs, 4 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+64K×(m-1)th point output, and the two data are subtracted to obtain the (n+64K×(m-1)+32K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 32Kth butterfly calculation is completed, it is read out in sequence from the first storage address; and then the subsequent groups of data are calculated in the same way. The last 16K points of each group of data are multiplied by -j (j*j=-1).
[0075] (2.4) The fourth stage performs the following operations: The fourth stage is divided into 8 groups of operations, with 32K points as a group in sequence. The first 16K points of the input data are cached and cached using the on-chip RAM. When the last 16K points are input, the first 16K points are read out in sequence, and butterfly operations are performed on two points separated by 16K, for a total of 16K pairs and 8 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+32K×(m-1)th point output, and the two data are subtracted to obtain the (n+32K×(m-1)+16K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 16Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is ,calculate , the address read from the rotation factor storage unit is W addr The result read is , the actual rotation factor is , according to the different intervals of addr, the result is as shown in formula 2. The calculated data is intercepted according to the precision configuration to keep the bit width unchanged.
[0076] (2)
[0077] (2.5) The fifth stage performs the following operations: The fifth stage is divided into 16 groups of operations, with 16K points as one group in sequence. The first 8K points of the input data are cached and cached using the on-chip RAM. When the last 8K points are input, the first 8K points are read out in sequence, and butterfly operations are performed in pairs of two points separated by 8K, for a total of 8K pairs and 16 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+16K×(m-1)th point output, and the two data are subtracted to obtain the (n+16K×(m-1)+8K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 8Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The last 4K points of each group of data are multiplied by -j (j*j=-1);
[0078] (2.6) The sixth stage performs the following operations: The sixth stage is divided into 32 groups of operations, with 8K points as a group in sequence. The first 4K points of the input data are cached and cached using the on-chip RAM. When the last 4K points are input, the first 4K points are read out in sequence, and butterfly operations are performed on two points separated by 4K in pairs, for a total of 4K pairs and 32 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+8K×(m-1)th point output, and the two data are subtracted to obtain the (n+8K×(m-1)+4K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 4Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0079] (2.7) The seventh level performs the following operations: The seventh level is divided into 64 groups of operations, with 4K points as one group in sequence. The first 2K points of the input data are cached and cached using the on-chip RAM. When the last 2K points are input, the first 2K points are read out in sequence, and butterfly operations are performed in pairs of two points with a 2K interval, for a total of 2K pairs, 64 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+4K×(m-1)th point output, and the two data are subtracted to obtain the (n+4K×(m-1)+2K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 2Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The last 1K points of each group of data are multiplied by -j (j*j=-1);
[0080] (2.8) The eighth level performs the following operations: The eighth level is divided into 128 groups of operations, with 2K points as a group in sequence. The first 1K points of the input data are cached and cached using the on-chip RAM. When the last 1K points are input, the first 1K points are read out in sequence, and butterfly operations are performed on two points separated by 1K, for a total of 1K pairs and 128 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+2K×(m-1)th point output, and the two data are subtracted to obtain the (n+2K×(m-1)+1K)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 1Kth butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0081] (2.9) The ninth level performs the following operations: The ninth level is divided into 256 groups of operations, with 1K points as a group in sequence. The first 512 points of the input data are cached and cached using the on-chip RAM. When the last 512 points are input, the first 512 points are read out in sequence, and butterfly operations are performed in pairs of two points separated by 512, for a total of 512 pairs and 256 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+1K×(m-1)th point output, and the two data are subtracted to obtain the (n+1K×(m-1)+512)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 512th butterfly calculation is completed, it is read out in sequence starting from the first storage address; and then the subsequent groups of data are calculated in the same way. The last 256 points of each group of data are multiplied by -j (j*j=-1);
[0082] (2.10) The tenth stage performs the following operations: The tenth stage is divided into 512 groups of operations. In order, 512 points form a group. The first 256 points of the input data are input into the delay unit for caching. The on-chip RAM is used for caching. When the last 256 points are input, the first 256 points are read out in order. The two points separated by 256 are grouped into two groups for butterfly operations, for a total of 256 pairs and 512 groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+512×(m-1)th point output, and the two data are subtracted to obtain the (n+512×(m-1)+256)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is written into the storage unit in situ. After the 256th butterfly calculation is completed, it is read out from the first storage address in sequence; and then the subsequent groups of data are calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0083] (2.11) The eleventh level performs the following operations: The eleventh level is divided into 1K group operations, with 256 points as a group in order. The first 128 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 128 beats. When the last 128 points are input, they are grouped with the first 128 points for butterfly operations, for a total of 256 pairs, 1K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+256×(m-1)th point output, and the two data are subtracted to obtain the (n+256×(m-1)+128)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit. After the 128th butterfly calculation is completed, the delay unit data output is valid; and then the subsequent groups of data are calculated in the same way. The last 64 points of each group of data are multiplied by -j (j*j=-1);
[0084] (2.12) The twelfth stage performs the following operations: The twelfth stage is divided into 2K groups of operations, with 128 points as a group in sequence. The first 64 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 64 beats. When the last 64 points are input, they are grouped with the first 64 points for butterfly operations, for a total of 64 pairs, 2K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+128×(m-1)th point output, and the two data are subtracted to obtain the (n+128×(m-1)+64)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit. After the 64th butterfly calculation is completed, the data output of the delay unit is valid; the subsequent groups of data are then calculated in the same way. The 256K output points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0085] (2.13) The thirteenth level performs the following operations: The thirteenth level is divided into 4K groups of operations, with 64 points as a group in order. The first 32 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 32 beats. When the last 32 points are input, they are grouped with the first 32 points for butterfly operations, totaling 32 pairs, 4K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+64×(m-1)th point output, and the two data are subtracted to obtain the (n+64×(m-1)+32)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit. After the 32nd butterfly calculation is completed, the delay unit data output is valid; then the subsequent groups of data are calculated in the same way. The last 16 points of each group of data are multiplied by -j (j*j=-1);
[0086] (2.14) The fourteenth stage performs the following operations: The fourteenth stage is divided into 8K group operations, with 32 points as a group in sequence. The first 16 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 16 beats. When the last 16 points are input, they are grouped with the first 16 points for butterfly operations, for a total of 16 pairs, 8K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+32×(m-1)th point output, and the two data are subtracted to obtain the (n+32×(m-1)+16)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit in situ. After the 16th butterfly calculation is completed, the delay unit data output is valid; then the subsequent groups of data are calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output serial number index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0087] (2.15) The fifteenth stage performs the following operations: The fifteenth stage is divided into 16K group operations, with 16 points as a group in sequence. The first 8 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 8 beats. When the last 8 points are input, they are grouped with the first 8 points for butterfly operations, for a total of 8 pairs, 16K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+16×(m-1)th point output, and the two data are subtracted to obtain the (n+16×(m-1)+ 8)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit in situ. After the eighth butterfly calculation is completed, the delay unit data output is valid; then the subsequent groups of data are calculated in the same way. The last 4 points of each group of data are multiplied by -j (j*j=-1);
[0088] (2.16) The sixteenth stage performs the following operations: The sixteenth stage is divided into 32K group operations, with 8 points as a group in sequence. The first 4 points of the input data are input into the delay unit. The delay unit uses internal logic to delay 4 beats. When the last 4 points are input, they are grouped with the first 4 points for butterfly operations, for a total of 4 pairs, 32K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+8×(m-1)th point output, and the two data are subtracted to obtain the (n+8×(m-1)+4)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit in situ. After the fourth butterfly calculation is completed, the delay unit data output is valid; then the second group of data is calculated in the same way. The output 256K points are multiplied by the rotation factors respectively, and the output sequence index uses 18-bit binary representation index[17:0], and the rotation factor is , the calculation method is as above, and the calculated data is cut off according to the precision configuration to keep the bit width unchanged;
[0089] (2.17) The seventeenth stage performs the following operations: The seventeenth stage is divided into 64K group operations, with 4 points as a group in order. The first two points of the input data are input into the delay unit. The delay unit uses internal logic to delay 2 beats. When the last two points are input, they are grouped with the first two points for butterfly operations, with a total of 2 pairs, 64K groups. For the nth pair of butterflies in the mth group, the two data are added to obtain the n+4×(m-1)th point output, and the two data are subtracted to obtain the (n+4×(m-1)+2)th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is input into the delay unit in situ. After the 32Kth butterfly calculation is completed, the delay unit data output is valid; and then the subsequent groups of data are calculated in the same way. The last point of each group of data is multiplied by -j (j*j=-1);
[0090] (2.18) The eighteenth stage performs the following operations: The eighteenth stage is divided into 128K group operations, and 2 points are grouped in order. The first point of the input data is input into the delay unit, and the delay unit uses internal logic to cache 1 beat. When the latter point is input, it just forms a group of butterfly operations with the previous point. There are 1 pair, 128K groups in total. The two data of the mth group are added to obtain the 2m-1th point output, and the two data are subtracted to obtain the 2m-1th point output. The result of the addition of the two data is directly output, and the result of the subtraction of the two data is cached for 1 beat output. The final FFT result is obtained.
[0091] In addition, the embodiment of the present invention also provides a 256K-point fast Fourier transform device based on FPGA, including 18 butterfly computing units, 8 rotation factor generation units, and a depth of 2 i The delay unit, i is an integer from 0 to 17;
[0092] The butterfly computing unit is a radix-2 butterfly computing unit, which is used to implement butterfly operations on data and complete data interleaving;
[0093] The 8 rotation factor generation units are located in the 2nd, 4th, 6th, 8th, 10th, 12th, 14th, and 16th level calculations, respectively.
[0094] The second-level rotation factor generation unit is used to generate 1024-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0095] The fourth-level rotation factor generation unit and the twelfth-level rotation factor generation unit are both used to generate 256-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0096] The sixth-level rotation factor generation unit and the fourteenth-level rotation factor generation unit are both used to generate 64-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0097] The eighth-level rotation factor generation unit and the sixteenth-level rotation factor generation unit are both used to generate 16-point rotation factors with precisions of 16 bits or 32 bits respectively;
[0098] The tenth-level rotation factor generation unit is used to generate 256k-point rotation factors with 16-bit or 32-bit precision respectively;
[0099] The delay unit includes a memory and an address control unit to realize data delays of different lengths; the memory is used to store the data input of this stage and the output of the butterfly operation subtraction; the address control unit is used to control the read and write address of the memory and the control signal of the memory to realize the delay function.
[0100] like Figure 1 As shown, 128k, 64k, 32k, 16k, 8k, 4k, 2k, 1k, 512, 256, 128, 64, 32, 16, 8, 4, 2, 1 represent delay units of different depths, BF2 represents a butterfly computing unit, and W represents a rotation factor basis.
[0101] The 256K-point fast Fourier transform device based on FPGA of the embodiment of the present invention also includes one or more processors for implementing the 256K-point fast Fourier transform method based on FPGA. The 256K-point fast Fourier transform device based on FPGA of the embodiment of the present invention can be applied to any device with data processing capability, and the any device with data processing capability can be a device or apparatus such as a computer. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capability in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 5 As shown in the figure, it is a hardware structure diagram of any device with data processing capability where the 256K-point fast Fourier transform device based on FPGA of the present invention is located. Figure 5 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0102] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.
[0103] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0104] An embodiment of the present invention further provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the FPGA-based 256K-point fast Fourier transform method in the above embodiment is implemented.
[0105] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart memory card (SmartMedia card, SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capability and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.
[0106] Those skilled in the art can understand that the above are only preferred examples of the invention and are not intended to limit the invention. Although the invention is described in detail with reference to the above examples, those skilled in the art can still modify the technical solutions recorded in the above examples or replace some of the technical features therein with equivalents. Any modification, equivalent replacement, etc. made within the spirit and principle of the invention shall be included in the protection scope of the invention.
Claims
1. A 256K-point fast Fourier transform method based on FPGA, characterized in that: The following steps are involved: Step 1: Determine that the calculation accuracy is 16 bits or 32 bits, calculate the rotation factors of 16, 64, 256, 1024, and 256K points in advance and store them in storage units; instantiate rotation factor storage units of different bit widths according to the calculation accuracy; Step 2: Use base 2 2 The butterfly basic processing unit first decomposes the 256K point data stream into 2 10 ×2 8 , and then 2 10 Decompose into 2 2 ×2 8 , 2 8 Decompose into 2 2 ×2 6 , and then 2 6 Decompose into 2 2 ×2 4 , and then 2 4 Decompose into 2 2 ×2 2 , the last 2 2 Decompose into 2×2, perform 18 levels of delay and butterfly calculation, that is, the 18-level rotation factor basis is 2 2 , 2 10 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 18 , 2 2 , 2 8 , 2 2 , 2 6 , 2 2 , 2 4 , 2 2 , 2 0 .
2. The 256K-point fast Fourier transform method based on FPGA according to claim 1, characterized in that: In step 2, 18 levels of delay and butterfly calculation are performed, wherein: The specific operation of odd-numbered level decomposition in 18 levels is as follows: The S level is divided into 2 S-1 group operation, S represents the series, S is an odd number between 1 and 18; in order 2 19-S Points are grouped together, and the first 2 18-S points to cache; after 2 18-S When you click Enter, the first 2 18-S The points are read out in sequence, with an interval of 2 18-S The two points of the mth group are operated in pairs by butterfly operation; the two data of the nth pair of butterflies of the mth group are added to obtain the n+2th pair of butterflies. 19-S ×(m-1) point output, the two data of the nth pair of butterflies in the mth group are subtracted to get the n+2th pair of butterflies 19-S ×(m-1)+ 2 18-S Point output; the result of adding two data is directly output; the result of subtracting two data is written into the storage unit in place, and waits for the second 18-S After the butterfly calculation is completed, it is read out from the first storage address in sequence; then the subsequent groups of data are calculated in the same way; the last 2 of each group of data output 17-S The point data is multiplied by -j, where j*j=-1, that is, the sign bit of the real part is inverted to become the imaginary part, and the imaginary part becomes the real part; The specific operation of even-numbered level decomposition in 18 levels is as follows: The S level is divided into 2 S-1 group operation, S represents the series, S is an even number between 1 and 18; in order 2 19-S Points are grouped together, and the first 2 18-S points to cache; after 2 18-S When you click Enter, the first 2 18-S The points are read out in sequence, with an interval of 2 18-S The two points of the mth group are operated in pairs by butterfly operation; the nth pair of butterflies of the mth group, the two data are added to get the n+2th pair of butterflies. 19-S ×(m-1) point output, subtract the two data to get the n+2 19-S ×(m-1)+ 2 18-S Point output; the result of adding two data is directly output; the result of subtracting two data is written into the storage unit in place, and waits for the second 18-S After the butterfly calculation is completed, it is read out from the first storage address in sequence; then the subsequent groups of data are calculated in the same way; the 256K points output by the first sixteen levels are multiplied by the calculated rotation factors respectively, and the calculated data is intercepted with high-order data according to the precision configuration, and the bit width is kept unchanged, and the eighteenth level data is directly output to obtain the final FFT result.
3. The 256K-point fast Fourier transform method based on FPGA according to claim 1, characterized in that: In the step 1, the rotation factors of 16, 64, and 256 points are calculated in advance and stored in registers respectively; the rotation factors of 1024 and 256K points are calculated in advance and stored in ROM respectively.
4. The 256K-point fast Fourier transform method based on FPGA according to claim 1, characterized in that: When the depth of the delay unit for performing delay calculation is not less than 256, the delay is implemented using a memory; otherwise, the delay is implemented using an on-chip logic unit.
5. The 256K-point fast Fourier transform method based on FPGA according to claim 1, characterized in that: When performing 18 levels of delay and butterfly calculation, the multipliers only appear at the 2nd, 4th, 6th, 8th, 10th, 12th, 14th, and 16th levels.
6. The 256K-point fast Fourier transform method based on FPGA according to claim 5, characterized in that: When the multiplier performs complex number multiplication, it is implemented by three real number multiplications and five real number additions.
7. The 256K-point fast Fourier transform method based on FPGA according to claim 3, characterized in that: The storage of the twiddle factors only stores 1 / 8 of the length.
8. The 256K-point fast Fourier transform method based on FPGA according to claim 2, characterized in that: The twiddle factors for even levels are obtained as follows: The 256K points calculated by each level of butterfly are multiplied by the rotation factors respectively, and the output index is represented by 18-bit binary index[17:0]; The second level rotation factors are ; The fourth level rotation factor is ; The sixth level rotation factor is ; The eighth level rotation factor is ; The tenth level rotation factor is ; The twiddle factor of the twiddle level is ; The fourteenth level rotation factors are ; The sixteenth level rotation factor is .
9. A 256K-point fast Fourier transform device based on FPGA, characterized in that: It includes 18 butterfly computing units, 8 rotation factor generation units, and a depth of 2 i The delay unit, i is an integer from 0 to 17; The butterfly computing unit is a radix-2 butterfly computing unit, which is used to implement butterfly operations on data and complete data interleaving; The eight rotation factor generation units are respectively located in the 2nd, 4th, 6th, 8th, 10th, 12th, 14th and 16th level calculations, wherein the second level rotation factor generation unit is used to generate 1024-point rotation factors with precisions of 16 bits or 32 bits respectively; The fourth-level rotation factor generation unit and the twelfth-level rotation factor generation unit are both used to generate 256-point rotation factors with precisions of 16 bits or 32 bits respectively; The sixth-level rotation factor generation unit and the fourteenth-level rotation factor generation unit are both used to generate 64-point rotation factors with precisions of 16 bits or 32 bits respectively; The eighth-level rotation factor generation unit and the sixteenth-level rotation factor generation unit are both used to generate 16-point rotation factors with precisions of 16 bits or 32 bits respectively; The tenth-level rotation factor generation unit is used to generate 256k-point rotation factors with 16-bit or 32-bit precision respectively; The delay unit includes a memory and an address control unit to achieve data delays of different lengths; the memory is used to store the data input of this stage and the output of the butterfly operation subtraction; the address control unit is used to control the read and write address of the memory and the control signal of the memory to achieve the delay function.
10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, it is used to implement the 256K-point fast Fourier transform method based on FPGA as described in any one of claims 1 to 8.
Citation Information
Patent Citations
FFT processor implementing base 4FFT / IFFT operation
CN101354701A
Fast Fourier transform hardware design method based on base 2-2 algorithm
CN109522674A
Radix-2 fast Fourier transform hardware design method based on an FPGA
CN110765709A
Fast fourier transform device, digital filtering device, fast fourier transform method, and non-transitory computer-readable medium
US20230289397A1
Pipelined, high-precision fast fourier transform processor
US6081821A