Data shuffling method of mixed-base FFT (Fast Fourier Transform) based on vector processor
Through the hybrid-based FFT data shuffling method of vector processor, data access and parallel processing are optimized, and the problem of low computational efficiency of FFT algorithms in the prior art is solved, and higher computing performance and hardware utilization are achieved.
Patent Information
- Application Number
- CN202510268114.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art is difficult to effectively utilize the parallel processing capabilities of vector processors in FFT algorithms, resulting in data access delay and low computational efficiency.
Using a hybrid-based FFT-based data shuffling method based on a vector processor, the data shuffling method is efficiently shuffled and calculated by optimizing data access mode and parallel processing technology, and using the parallel processing capabilities of vector DSP.
It improves the computing efficiency and performance of the FFT algorithm, makes full use of the advantages of hardware performance, and improves instruction parallelism.
Smart Images

Figure CN120335870A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the optimization of the underlying code of a high-performance Digital Signal Processor (DSP) chip. More specifically, it relates to a data shuffling method for a mixed-radix FFT based on a vector processor. Background Art
[0002] A digital signal processor is a signal processing method that represents a signal as a digital or symbolic sequence and extracts useful information through hardware or software processing. Compared with traditional analog signals, digital signals have the advantages of strong reliability, high precision and flexibility, easy large-scale integration, time-division multiplexing, and high performance. Vector processors generally support single instruction multiple data stream (SIMD) operations, that is, under the control of the same vector instruction, all processing units simultaneously perform the same operation on the corresponding local registers to achieve data-level parallelism in developing application programs. In recent decades, with the rapid development of computer technology, the FFT algorithm has been widely applied in various fields of today's technology. For example, in image processing, using FFT for convolution calculation can restore an old image to its original state or remove stains on the picture. In radar signal processing, using FFT can implement key operations such as Doppler filtering. In speech signal processing, the FFT transform provides a more intuitive understanding of speech recognition and voice simulation by analyzing the speech spectrum. With the rapid development of various fields, the requirement for real-time performance is getting higher and higher, so there are also higher requirements for the performance of the FFT algorithm. The patent application with publication number CN 114116012A and publication date March 1, 2022 discloses a method and device for realizing the vectorization of the FFT bit-reversal algorithm based on a shuffle operation. This method describes the working process of the vector shuffle module, and can achieve the bit-reversal of FFT by shuffling vector data. The implementation method of this solution is different from this solution, and this solution does not specifically describe the shuffle mode of the butterfly operation. The patent application with publication number CN 103699516A and publication date April 2, 2014 discloses a method and device for parallel FFT / IFFT butterfly operations based on SIMD in a vector processor. This method introduces the operation method of parallel FFT / IFFT based on SIMD. The implementation method of this solution is different from this solution, and this solution does not specifically describe the shuffle mode of the butterfly operation. The present invention relates to a data shuffling method for mixed - radix FFT based on a vector processor. In terms of data access in the FFT algorithm, reducing memory access is an important factor in improving the algorithm performance. The vector shuffling technology reduces frequent data access and storage by optimizing the data access pattern. The parallel processing technology can effectively improve the algorithm performance by simultaneously processing multiple data and making full use of the parallel processing ability of the vector DSP. The condition for satisfying the mixed - radix FFT is that if the number of input data points N = N1 * N2 * … * N k *N k+1 , then N1, N2, …, N k is 4, and N k+1 is 2. For example: 128 = 4 * 4 * 4 * 2. Summary of the Invention
[0003] In order to solve the above - mentioned technical problems, the present invention proposes a data shuffling method for mixed - radix FFT based on a vector processor. The technical solution of the present invention is as follows: including the following steps, Step 1, the input points x(0) to x(n) correspond to the serial numbers X m 0~n . The subscript m represents the level where the input is located, starting from the 0th level. The superscripts 0 to n represent the serial numbers of the input points, and the total number of points satisfies the quantity of the mixed - radix FFT operation. The inputs of the butterfly are denoted as four components A, B, C, and D, and the outputs of the butterfly are denoted as four components E, F, G, and H. The output marker T is obtained through the theoretical FFT transformation m αβγδr , the subscript m represents the level where the output is located, and the superscripts αβγδ represent the serial numbers of the four input points participating in the operation. The superscript r represents the E, F, G, or H part of the output. At the same time, T m αβγδr corresponds to one of the inputs X m+1 0~n in the next level; Step 2, initially m = 0, that is, starting from the 0th level. According to the number of VPEs of the DSP vector processor, it is judged whether the input at this level of X m 0~n meets the vector calculation condition. If the data in the register meets the vector calculation condition, it directly jumps to Step 4. If it does not meet the vector calculation, the data in the register needs to be shuffled, and it jumps to Step 3; Step 3, perform a shuffling operation on the data at this level; Step 4, perform a DSP - FFT vector operation according to the FFT value in the current vector register. The first few levels all use the radix - 4 operation, and the last level uses the radix - 2 operation. After the vector calculation, X m 0~n →T mαβγδr →X m+1 0~n If the (m + 1)-th level is not the last level, then m = m + 1, and return to step 2. If this level is the last level, then go to step 5; Step 5, perform the final output shuffle and sorting. The specific manner of step 3 of the present invention includes the following steps Step 3.1, in the DSP-FFT vector operation, according to the output marker T m αβγδr Find the same output marker T in the FFT theoretical calculation diagram m αβγδr After finding it, copy the specific serial number in X corresponding to the output marker T in the FFT theoretical calculation m αβγδr to the corresponding position of the output marker T in the DSP-FFT. Finally, obtain an X m+1 0~n sequence. This sequence is the FFT result serial number corresponding to the m-th level DSP-FFT result and is also the input sequence of the (m + 1)-th level; m αβγδr m+1 0~n sequence, which is the FFT result serial number corresponding to the m-th level DSP-FFT result and is also the input sequence of the (m + 1)-th level; Step 3.2, according to the FFT theoretical calculation schematic diagram, divide the components of the butterfly of this level, re-divide the position of the input serial number of the m-th level, that is, obtain a new sequence List. The data of the new sequence satisfies the parallel operation of the vector DSP. Shuffle the sequence obtained in step 3.1 into the sequence List that satisfies the DSP-FFT vector operation, that is, complete the shuffle. The specific manner of step 5 of the present invention includes the following steps Step 5.1, in the FFT theoretical calculation diagram, the sequence output by the mixed-radix FFT is regular and in reverse order. The number of operation points N, N = 2 * 4 P , use the (P + 1)-bit quaternary-binary combined serial number to represent the positions of the input sequence and the output sequence values. According to the input-output comparison table, mark the final order y(0)~y(n) of the output data. This sequence is the order of the final result data obtained according to the FFT theoretical calculation; Step 5.2, in the DSP-FFT vector operation, according to the output marker T m αβγδr Find the same output marker T in the FFT theoretical calculation diagram m αβγδr After finding it, copy the corresponding y(n) of the output marker T in the FFT theoretical calculation to the output marker T in the DSP-FFT m αβγδr m αβγδr The corresponding positions, and finally, a y sequence is obtained, which is the order of the final result data calculated according to DSP-FFT; Step 5.3, if this sequence is exactly in order, no shuffling is required. If this sequence is out of order, then this sequence needs to be shuffled into an ordered y(0) to y(n), and thus all calculations are completed. Compared with the prior art, the beneficial effects of the present invention are: The present invention proposes a data shuffling method for mixed-radix FFT based on a vector processor for a vector processor DSP with a data shuffling instruction. The data shuffling method for the mixed-radix inverse FFT, that is, the mixed-radix IFFT, is the same as the data shuffling method for the mixed-radix FFT. For different numbers of points and different numbers of VPEs, data can be processed in parallel, making full use of the hardware performance advantages, improving the utilization efficiency of the computing unit, and achieving higher instruction parallelism. Description of the Drawings
[0004] Figure 1 It is a schematic diagram of the DSP vector addition operation of the present invention. Figure 2 It is a schematic diagram of data shuffling for obtaining data at even positions of the present invention. Figure 3 It is a radix-4 FFT butterfly structure diagram of the present invention. Figure 4 It is a schematic diagram of the theoretical calculation sequence numbers of the 4-point radix-4 FFT of the present invention. Figure 5 The radix-2 FFT butterfly structure diagram of the present invention. Figure 6 It is a schematic diagram of the theoretical calculation sequence numbers of the 2-point radix-2 FFT of the present invention. Figure 7 It is a schematic diagram of the theoretical calculation sequence numbers of the 8-point mixed-radix FFT of the present invention. Figure 8 The left figure is a schematic diagram of the theoretical calculation of the 8-point mixed-radix FFT of the present invention, and the right figure is a schematic diagram of the DSP processing sequence numbers of the 8-point mixed-radix FFT. Figure 9 It is a schematic diagram of the theoretical calculation of the 128-point mixed-radix FFT of the present invention. For the sake of clear image, the horizontal lines are eliminated. Figure 10 It is a schematic diagram of the theoretical calculation of the first 64-point mixed-radix FFT of the first, second, and third levels extracted from the schematic diagram of the theoretical calculation of the 128-point mixed-radix FFT of the present invention. Detailed Embodiments
[0005] The present invention will be further described below in conjunction with the drawings and embodiments. Refer to Figure 1, Schematic diagram of DSP vector addition operation of the present invention. A vector processor is a type of processor that can execute vector instructions and is not limited to scalars. For example, the instruction: ADDV V1, V2, V3 is a typical vector summation instruction. As shown in Figure 1 , V1, V2, and V3 are vectors composed of three scalars of the same type and length. As a vector register, the values in the V1 and V2 vector registers are summed and the result is placed in the V3 vector register. Refer to Figure 2 , Shuffle schematic diagram of obtaining data at even positions of the present invention. The DSP vector processor has data shuffle instructions, including word shuffle, which can configure different shuffle modes according to different data arrangement requirements and store the configured shuffle mode in the shuffle mode address register. For example, for the word shuffle shown in Figure 2 : Combine the data at even positions of SrcA and SrcB. Only need to configure its shuffle mode as: 12_02_10_00 16_06_14_04 1A_0A_18_08 1E_0E_1C_0C, where: 00 represents the first word, that is, the first word of vector SrcA; 10 represents the seventeenth word, that is, the first word of vector SrcB. Refer to Figure 3 , Radix-4 FFT butterfly structure diagram of the present invention, where A, B, C, and D are four input components participating in the butterfly calculation, E, F, G, and H are the corresponding output parts, and W is the rotation factor participating in the calculation. Refer to Figure 4 , Schematic diagram of the theoretical calculation sequence number of 4-point radix-4 FFT of the present invention. Assume that the four input data are x(0), x(1), x(2), x(3). For the convenience of identification, they are respectively represented by X0 0 , X0 1 , X0 2 , X0 3 , representing the 0th, 1st, 2nd, and 3rd data of the 0th level. T represents the butterfly calculation, and the output data of the 0th level is denoted as: T0 0 1 2 3E , T0 0 1 2 3F , T0 0 1 2 3G , T0 0 1 2 3H, the subscript represents the output of the 0th level, the superscript represents the serial numbers of the four input points participating in the operation, and E, F, G, and H behind represent the butterfly output part. At the same time, the output of the 0th level corresponds to the input X1 of the 1st level 0 , X1 1 , X1 2 , X1 3 . In the figure, y(0), y(1), y(2), and y(3) represent the final output results, corresponding to the radix-4 FFT butterfly structure diagram. The result data of y(0) is (x(0) + x(2)) + (x(1) + x(3)); the result data of y(1) is [(x(0) - x(2)) - j(x(1) - x(3))]W N n ; the result data of y(2) is [(x(0) + x(2)) - (x(1) + x(3))]W N 2n ; the result data of y(3) is [(x(0) - x(2)) + j(x(1) - x(3))]W N 3n . Reference Figure 5 , the radix-2 FFT butterfly structure diagram of the present invention, where A and B are two input components participating in the butterfly calculation, E and F are the corresponding output parts, and W is the rotation factor participating in the calculation. Reference Figure 6 , the schematic diagram of the serial number of the 2-point radix-2 FFT theoretical calculation of the present invention. Assume that the two input data are x(0) and x(1), For the convenience of identification, X0 0 , X0 1 are used to represent the 0th and 1st data of the 0th level respectively. T represents the butterfly calculation, and the output data of the 0th level is denoted as: T0 0 1E , T0 0 1F , the subscript represents the output of the 0th level, the superscript represents the serial numbers of the two input points participating in the operation, and E and F behind represent the butterfly output part. At the same time, the output of the 0th level corresponds to the input X1 of the 1st level 0 , X1 1 . In the figure, y(0) and y(1) represent the final output results, corresponding to the radix-2 FFT butterfly structure diagram. The result data of y(0) is A + B; the result data of y(1) is (A - B)*W N n . Reference Figure 7, Schematic diagram of the sequence number of the 8-point hybrid radix FFT theoretical calculation of the present invention. The sequence input at the beginning in the figure is a natural sequence, but after several levels of butterfly calculations, the output sequence is no longer a natural sequence. Due to the diversity of the factorization of the hybrid radix, the output sequences are also diverse. Here, the general rule of reverse order is given: Assume that N of the input sequence can be decomposed into the product of multiple factors r n …r2r12. Convert the input decimal sequence number into a quaternary-binary combined sequence number composed of n-bit quaternary in the front and 1-bit binary in the back. For example: 128 = 4 * 4 * 4 * 2. Reverse this sequence number, then the corresponding output sequence number becomes 2r1 r2…r n , at this time the sequence is 1-bit binary in the front and n-bit quaternary in the back. For example: 128 = 2 * 4 * 4 * 4. The table of the input sequence and the output sequence is shown in Table 1. Through Table 1, the sequence numbers of the final output results corresponding to the input points of this number of points can be obtained as y(0), y(4), y(1), y(5), y(2), y(6), y(3), y(7). Table 1: 8-point hybrid radix input-output comparison table Reference Figure 8 , The left figure is the schematic diagram of the 8-point hybrid radix FFT theoretical calculation of the present invention, and the right figure is the schematic diagram of the DSP processing sequence number of the 8-point hybrid radix FFT. For the data shuffling in the first few levels, taking 2 VPE vector processors that can process 2 data simultaneously as an example, at the 0th level of the 8-point calculation, exactly 2 consecutive data can be divided into one vector register to form one part of the butterfly calculation. srcA, srcB, srcC, srcD are respectively the four parts A, B, C, D participating in the butterfly operation. When the operation proceeds to the 1st level, a radix-2 operation is performed. The A parts of the four butterflies are X1 0 、X1 2 、X1 4 、X1 8 , and the B parts are X1 1 、X1 3 、X1 5 、X1 7 . At this time, if the data in the original register is used, vector calculation cannot be performed. In order to make full use of the characteristics of the vector register that can process in parallel and has high operation efficiency, the data in the vector register is shuffled into a parallel processing method. It is necessary to shuffle the same components of different butterflies into one vector register for vector calculation. After the shuffling is completed, the points in srcA are X1 0 、X1 2 , the points in srcB are X1 1 、X1 3 , and the points in srcC are X1 4 、X16 , the point in srcD is X1 5 、X1 7 If there are multiple levels of calculation, each level is shuffled according to this idea. For the order of the last level output, if the current level output is the last level, first mark the output data of the FFT theoretical calculation and the output data of the DSP-FFT parallel calculation respectively, and use the same output mark T m αβγδr Find the corresponding output data in DSP-FFT in the output data of FFT theoretical calculation, and record the serial number of the output data, for example, Figure 8 In DSP-FFT calculation, according to T1 0 1E The output markers in the FFT theory calculations find the same output markers as T1 0 1E The serial number is X2 0 , output data T1 2 3E Find the same output marker T1 in the FFT theory calculation 2 3E The serial number is X2 2 , output data T1 01F Find the same output marker T1 in FFT theory calculation 0 1F Corresponding to serial number X2 1 , output data T1 2 3F Find the same output marker T1 in FFT theory calculation 2 3F The corresponding number is X2 3 ...until the serial numbers corresponding to all output data are recorded in the DSP-FFT calculation results, and the serial numbers y(0)~y(n) of the last level output are obtained from the FFT theoretical calculation diagram. If the serial numbers are not in order, shuffle them to order, and the shuffling of the last level is completed. Figure 8 In the example, the 8-point mixed basis outputs y(0) to y(7) of the DSP operation are not in order and need to be shuffled further. refer to Figure 9, The 128-point mixed-radix FFT vector layout diagram of the present invention. For the sake of image clarity, the horizontal lines are removed. A 128-point FFT requires 4 levels of calculation in total. Among them, the first 3 levels are radix-four operations, and the last level is a radix-two operation. When the vector processor has 16 VPEs and can process 16 data simultaneously, the four components A, B, C, and D of the 0th-level butterfly group of the 128-point FFT each have 32, and together they form a butterfly group. At this time: registers VD11 and VD21 are part A of the butterfly group; registers VD31 and VD41 are part B of the butterfly group; registers VD12 and VD22 are part C of the butterfly group; registers VD32 and VD42 are part D of the butterfly group. When the 0th-level calculation is completed and the 1st-level calculation starts, the number of butterfly groups becomes 4 at this time. At this time, the data in one vector register cannot be used as part of a butterfly group, which is not conducive to vector calculation, and data shuffling is required through a shuffling operation. Through the following shuffling modes, shuffling mode 0: 03_02_01_00 07_06_05_04 13_12_11_10 17_16_15_14, shuffling mode 1: 0b_0a_09_08 0f_0e_0d_0c 1b_1a_19_18 1f_1e_1d_1c, as Figure 10 shown. After the first 64 points are shuffled, the data sequence numbers in VD11 are {0, 1, 2, 3, 4, 5, 6, 7, 32, 33, 34, 35, 36, 37, 38, 39}, the data sequence numbers in VD21 are {8, 9, 10, 11, 12, 13, 14, 15, 40, 41, 42, 43, 44, 45, 46, 47}, the data sequence numbers in VD31 are {16, 17, 18, 19, 20, 21, 22, 23, 48, 49, 50, 51, 52, 53, 54, 55}, the data sequence numbers in VD41 are {24, 25, 26, 27, 28, 29, 30, 31, 56, 57, 58, 59, 60, 61, 62, 63}. Refer to Figure 10, in the schematic diagram of the theoretical calculation of the 128-point mixed-radix FFT of the present invention, the schematic diagrams of the theoretical calculations of the first 64 points of the first, second, and third levels of the mixed-radix FFT are extracted. When the first-level calculation ends and the second-level calculation starts, at this time, the data in a vector register cannot be used as part of the butterfly group, which is not conducive to vector calculation, and data shuffling is required through a shuffle operation. Each shuffle at this level requires the participation of 4 registers to make the data in each register meet the butterfly operation of the next level. Taking the extraction of the first 64 points from the diagrams of the first, second, and third levels of the 128-point as an example. Through the following shuffle patterns, shuffle pattern 2: 11_10_01_00 13_12_03_02 19_18_09_08 1b_1a_0b_0a, shuffle pattern 3: 15_14_05_04 17_16_07_06 1d_1c_0d_0c 1f_1e_0f_0e, shuffle pattern 4: 03_02_01_00 13_12_11_10 0b_0a_09_08 1b_1a_19_18, shuffle pattern 5: 07_06_05_04 17_16_15_14 0f_0e_0d_0c 1f_1e_1d_1c, as Figure 10 shown, after the shuffle is completed, the data sequence numbers in VD11 are {0,1,8,9,16,17,24,25,32,33,40,41,48,49,56,57}, the data sequence numbers in VD21 are {2,3,10,11,18,19,26,27,34,35,42,43,50,51,58,59}, the data sequence numbers in VD31 are {4,5,12,13,20,21,28,29,36,37,44,45,52,53,60,61}, the data sequence numbers in VD41 are {6,7,14,15,22,23,30,31,38,39,46,47,54,55,62,63}. When the second-level calculation ends and the third-level calculation starts, the data in a register cannot be used as part of the butterfly group, which is not conducive to vector calculation, and data shuffling is required through a shuffle operation. Through the following shuffle patterns, shuffle pattern 6: 12_02_10_00 16_06_14_04 13_03_11_01 17_07_15_05, shuffle pattern 7: 1a_0a_18_08 1e_0e_1c_0c 1b_0b_19_09 1f_0f_1d_0d, shuffle mode 8: 11_10_01_00 13_12_03_02 15_14_05_04 17_16_07_06, shuffle mode 9: 19_18_09_08 1b_1a_0b_0a 1d_1c_0d_0c 1f_1e_0f_0e, as Figure 10 shown, after the shuffle is completed, the data sequence numbers in VD11 are {0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24, 26, 28, 30}, the data sequence numbers in VD21 are {32, 34, 36, 38, 40, 42, 44, 46, 48, 50, 52, 54, 56, 58, 60, 62}, the data sequence numbers in VD31 are {1, 3, 5, 7, 9, 11, 13, 15, 17, 19, 21, 23, 25, 27, 29, 31}, the data sequence numbers in VD41 are {33, 35, 37, 39, 41, 43, 45, 47, 49, 51, 53, 55, 57, 59, 61, 63}, and the rest of the points are shuffled according to this method. After the last - stage calculation is completed, the calculation result is not in order at this time, and a sorting operation is required. According to the method described above, find the corresponding output data in DSP - FFT from the output data of the FFT theoretical calculation and mark the position where the data should be. The shuffle modes include shuffle mode 10: 14_04_10_00 1c_0c_18_08 15_05_11_01 1d_0d_19_09, shuffle mode 11: 16_06_12_02 1e_0e_1a_0a 17_07_13_03 1f_0f_1b_0b, shuffle mode 12: 11_10_01_00 13_12_03_02 15_14_05_04 17_16_07_06, shuffle mode 13: 19_18_09_08 1b_1a_0b_0a 1d_1c_0d_0c 1f_1e_0f_0e. After shuffling through the above shuffle modes, the sorting operation of the output result can be realized. In summary, after those of ordinary skill in the art read the documents of the present invention, all other corresponding transformation schemes made without creative mental labor according to the technical solutions and technical concepts of the present invention fall within the scope protected by the present invention.
Claims
1. A data shuffling method for hybrid radix FFT based on a vector processor, characterized in that: The following steps are included: Step 1, input points x(0) to x(n) corresponding to serial numbers X m 0~n , where the subscript m represents the level where the input is located, starting from the 0th level, and the superscripts 0 to n represent the serial numbers of the input points, and the total number of points satisfies the quantity for the mixed-radix FFT operation. The inputs of the butterfly are denoted as four components A, B, C, D, and the outputs of the butterfly are denoted as four components E, F, G, H. The output marker T is obtained through the theoretical FFT transformation m αβγδr , where the subscript m represents the level where the output is located, the superscripts αβγδ represent the serial numbers of the four input points participating in the operation, the superscript r represents the E, F, G or H part of the output, and at the same time T m αβγδr corresponds to one of the inputs X m+1 0~n at the next level; Step 2, initially m = 0, that is, start from the 0th level. According to the number of VPEs of the DSP vector processor, judge whether the input at this level m 0~n meets the vector calculation conditions. If the data in the register meets the vector calculation conditions, directly jump to Step 4. If it does not meet the vector calculation, the data in the register needs to be shuffled, and jump to Step 3; Step 3, shuffle the data at this level; Step 4: Perform DSP-FFT vector operations based on the FFT values in the current vector register. The first few levels use radix-4 operations, and the last level uses radix-2 operations. After vector calculation, X m 0~n →T m αβγδr →X m+1 0~n , if the (m + 1)-th level is not the last level, then m = m + 1, and go back to Step 2. If this level is the last level, then go to Step 5; Step 5: Perform the final output shuffle and reorder.
2. The data shuffling method of a mixed-radix FFT based on a vector processor according to claim 1, wherein: The specific method of step 3 includes the following steps: Step 3.1, in DSP-FFT vector operation, according to the output mark T m αβγδr Find the same output label T in the FFT theory calculation diagram m αβγδr , after finding it, mark T as the output in the FFT theoretical calculation m αβγδr Corresponding X m+1 0~n The specific serial number in is copied to the DSP-FFT output label T m αβγδr The corresponding position, finally, gets an X m+1 0~n Sequence, this sequence is the FFT result sequence corresponding to the m-th level DSP-FFT result, and is also the input sequence of the m+1th level; Step 3.2, according to the FFT theoretical calculation diagram, divide the components of this level of butterfly, and re-divide the input sequence number position of the mth level, that is, obtain a new sequence List. The data of the new sequence meets the parallel operation of the vector DSP. Shuffle the sequence obtained in step 3.1 into the sequence List that meets the DSP-FFT vector operation, and the shuffling is completed.
3. A data shuffling method for hybrid radix FFT based on a vector processor according to claim 1, characterized in that: The specific method of step 5 includes the following steps: Step 5.1, in the FFT theoretical calculation diagram, the sequence output by the mixed-radix FFT is regular and in reverse order. The number of operation points N, N = 2 * 4 P , use the P + 1-bit quaternary-binary combined serial number to represent the positions of the input sequence and output sequence values. According to the input-output comparison table, mark the final order y(0) to y(n) of the output data. This sequence is the order of the final result data obtained according to the FFT theoretical calculation; Step 5.2, in DSP-FFT vector operation, according to the output mark T m αβγδr Find the same output label T in the FFT theory calculation diagram m αβγδr , after finding it, mark T as the output in the FFT theoretical calculation m αβγδr The corresponding y(n) is copied to the DSP-FFT output label T m αβγδr The corresponding positions, finally, a y sequence is obtained, which is the order of the final result data calculated by DSP-FFT; Step 5.3, if the sequence is in order, then no shuffling is required; if the sequence is out of order, then the sequence needs to be shuffled into the ordered y(0)~y(n). At this point, all calculations are completed.
Citation Information
Patent Citations
Single instruction multiple data (SIMD)-based parallel fast fourier transform / inverse fast fourier transform (FFT / IFFT) butterfly operation method and SIMD-based parallel FFT / IFFT butterfly operation device in vector processor
CN103699516A
FFT code bit inverse sequence algorithm vectorization implementation method and device based on shuffling operation
CN114116012A