An FFT processor and processing method supporting a non-base-2 point output sequence

By adopting non-integer point time domain interpolation technology and base 2FFT processing in the FFT processor, the problems of storage waste and hardware design complexity in the prior art are solved, and flexible configuration of arbitrary conversion lengths and efficient utilization of storage resources are achieved.

CN115422498BActive Publication Date: 2025-06-13BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210875299.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-22
Publication Date
2025-06-13
Estimated Expiration
2042-07-22

AI Technical Summary

Technical Problem

Existing FFT processing methods and processors have storage waste and hardware design complexity problems when dealing with non-integer power conversion lengths, and they cannot flexibly configure conversion lengths.

Method used

The non-integer point time domain interpolation technology is adopted, and the input sequence is converted into a time domain interpolation sequence of the base 2 points through the smooth interpolation module, and the base 2 FFT processing is performed in the FFT processing module. Finally, the output sequence of the non-base 2 points is obtained by the compensation interpolation module.

Benefits of technology

It realizes flexible configuration of arbitrary conversion lengths, reduces waste of storage resources, and simplifies the hardware design and development process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115422498B_ABST
    Figure CN115422498B_ABST
Patent Text Reader

Abstract

The present application provides an FFT processor that supports a non - base - 2 number of output sequences, including: a smoothing interpolation module for performing non - integer - point interpolation on an input sequence to obtain a time - domain interpolation sequence with a base - 2 number of points; the non - integer - point interpolation enables the output sequence obtained by passing the input sequence through the FFT processor to be a sequence obtained by zero - padding the output sequence of the DFT processing of the input sequence in the frequency domain before truncation processing; an FFT processing module for performing base - 2 FFT processing on the time - domain interpolation sequence with a base - 2 number of points to obtain a frequency - domain sequence with a base - 2 number of points; a compensation truncation module for performing phase compensation on the frequency - domain sequence with a base - 2 number of points and performing truncation processing to obtain a non - base - 2 number of output sequences as the output sequences of the FFT processor. The output sequences of the FFT processor in the present application will be transferred to the cache and memory of the system for subsequent signal processing. Since the output sequences can be of any length, a large amount of on - chip cache capacity and off - chip memory bandwidth can be saved for the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing, and particularly to an FFT processor and a processing method that support non-base-2 point output sequences. Background Art

[0002] Synthetic Aperture Radar (SAR) is a typical application scenario of the FFT algorithm. It is a two-dimensional imaging radar with the ability to work all day and all weather. Spaceborne synthetic aperture radar is an important part of the space earth observation field and is widely used in many important fields of national defense and people's livelihood such as earth remote sensing, military reconnaissance, and resource exploration. The massive raw data obtained by traditional imaging satellites needs to be transmitted to the ground for processing. There are many system links, large data transmission volume, and complex application guarantee, resulting in insufficient timeliness of remote sensing information acquisition and difficulty in meeting the requirements of application scenarios such as post-disaster search and rescue.

[0003] Converting radar signal imaging, information processing, etc. to on-orbit processing and extracting the information of interest and transmitting it to the ground can effectively solve the above problems. In the past on-orbit SAR real-time processing, due to factors such as the increase in the satellite observation angle, the improvement of the resolution, and the increase in the imaging mode, the overall scale of the hardware platform has increased. Coupled with the strict requirements of the space environment for volume, weight, and power consumption, reducing the scale of the hardware platform and ensuring real-time processing is the core problem in on-orbit SAR processing.

[0004] As the most important arithmetic module in SAR real-time processing, the FFT (Fast Fourier Transform) / IFFT (Inverse Fast Fourier Transform) arithmetic core is limited by the base-2 decomposition idea of the Cooley-Tukey algorithm. If the actual conversion length is less than an integer power of 2, zero-padding operations need to be performed in the time domain to complete the FFT points to a power of 2. Taking a 1600-point FFT as an example, a 1024-point FFT cannot meet the processing requirements and needs to be completed to 2048 points. In the field of SAR imaging processing, generally, the number of points in the conversion sequence for FFT operation processing is more than ten thousand. Therefore, the inflexibility of the FFT leads the system to consume a large amount of additional internal high-speed storage (such as SRAM, Static Random-Access Memory) and external large-capacity storage (such as SARAM, synchronous dynamic random-access memory) to match the output granularity of the FFT processor. More seriously, the two-dimensional nature of the SAR imaging algorithm (the azimuth direction is the flight direction, the range direction is the scanning direction, and they are perpendicular to each other) causes the storage waste to increase in a square trend.

[0005] In order to solve the problem of inflexible FFT conversion length, other fixed-radix algorithms, such as radix-3, radix-5 and radix-6 algorithms, are used in some special application scenarios, but not all FFT points are integer powers of fixed factors. In more flexible scenarios, mixed-radix FFT processors that decompose FFT points into at least two factors are used. Mixed-radix FFT processors generally have two structures: placing multiple pipeline fixed-radix butterfly units in sequence, and selecting a number of pipeline stages to form a data path according to the actual number of FFT points; using a variable-radix butterfly structure, placing multiple variable-radix butterfly units in sequence, configuring each level of radix according to the actual number of FFT points, and selecting a number of pipeline stages to form a data path.

[0006] According to the analysis of the existing FFT processing methods and processors for non-integer power conversion lengths, it can be seen that there are many problems, mainly the following two points:

[0007] The processing method and the processor have limited configurable conversion lengths, and storage waste is serious. Mixing multiple fixed bases essentially provides more optional decomposition factors for FFT point decomposition, but the number of mixed bases is limited by the circuit scale and cannot be infinite, and the number of points that can be configured is still limited. In addition, the disposable points are densely distributed in a small point range, and the number of configurable points in a large point range is relatively small. The storage waste problem will only be prominent in the case of large points, so the existing solutions do not effectively solve the storage waste problem. Taking the 13-level mixed base (base 2 / base 3) FFT as an example, the longest configurable length is 1.5M, and the maximum configurable length is 104. When the actual processing length is 1.1M, it still needs to be padded to 1.5M points, which wastes 26% of storage resources.

[0008] The processor design is complex and the configuration is inflexible. In order to enrich the number of configurable points or adapt to different conversion length scenarios, it is necessary to design a variety of fixed-base structures or even multi-base multiplexing structures, and the hardware structure needs to be changed according to the scenario. The design and debugging difficulties brought by this are very unfavorable for rapid development. In addition, for an implemented mixed-base FFT processor, in actual processing, when faced with a new number of conversion points, a dedicated factor decomposition solution needs to be generated, and then the hardware configuration of each level of FFT is performed. The preprocessing is complex and difficult to implement. Summary of the invention

[0009] In response to the shortcomings of the prior art that the output granularity obtained by the FFT processing method or processor is an integer power of a fixed factor, and the FFT processor has a complex design and inflexible configuration, the present application proposes an FFT processing method or processor that outputs a non-integer power conversion length sequence in the frequency domain based on non-integer point time domain interpolation technology, while overcoming the defect that significant hardware structure changes are required for different conversion length scenarios.

[0010] On the one hand, the present application provides an FFT processor supporting a non - radix - 2 number output sequence, including:

[0011] A smoothing interpolation module, configured to perform non - integer - point interpolation on an input sequence to obtain a time - domain interpolation sequence with radix - 2 numbers; the non - integer - point interpolation enables the output sequence obtained by the input sequence passing through the FFT processor to be the sequence obtained by zero - padding in the frequency domain of the output sequence of the DFT processing of the input sequence before truncation processing;

[0012] An FFT processing module, configured to perform radix - 2 FFT processing on the time - domain interpolation sequence with radix - 2 numbers to obtain a frequency - domain sequence with radix - 2 numbers;

[0013] A compensation truncation module, configured to perform phase compensation on the frequency - domain sequence with radix - 2 numbers and perform the truncation processing to obtain a non - radix - 2 number output sequence as the output sequence of the FFT processor; the output sequence of the FFT processor is the same as the output sequence of the DFT processing of the input sequence.

[0014] Preferably, the FFT processing module includes a variable - point FFT module and a pipelined re - ordering module; the variable - point FFT module is configured to perform multi - stage iterative butterfly operations on the time - domain interpolation sequence to obtain a bit - reversed FFT frequency - domain sequence; the pipelined re - ordering module is configured to re - order the bit - reversed FFT frequency - domain sequence to obtain the frequency - domain sequence with radix - 2 numbers.

[0015] Preferably, the smoothing interpolation module includes a data controller, a shift register, a floating - point multiplier - adder, and an interpolation kernel lookup table; the variable - point FFT module includes multi - stage single - path delay feedback structure units and a decoding controller; the pipelined re - ordering module includes forward and reverse address registers, an address selector, and a random access memory; preferably, the compensation truncation module includes a factor multiplier and a calculation truncator.

[0016] Preferably, the output sequence of the FFT processor is transferred to the cache and memory of the system for subsequent signal processing.

[0017] Preferably, the length of the input sequence is independently provided as the only parameter to the smoothing interpolation module, the FFT processing module, and the compensation truncation module.

[0018] On the other hand, the present application provides an FFT processing method supporting a non - radix - 2 number output sequence, including:

[0019] Obtain an input sequence, and perform smooth interpolation on the input sequence according to the non-integer point interpolation technique to obtain a time-domain interpolation sequence with a base-2 number of points; the non-integer point interpolation technique enables the output sequence obtained by subjecting the input sequence to the FFT processing method to be a sequence obtained by zero-padding in the frequency domain of the output sequence obtained by performing DFT processing on the input sequence before the truncation processing;

[0020] Process the time-domain interpolation sequence with a base-2 number of points according to the base-2 FFT technique to obtain a frequency-domain sequence with a base-2 number of points;

[0021] Perform phase compensation on the frequency-domain sequence with a base-2 number of points, and perform the truncation processing to obtain an output sequence with a non-base-2 number of points as the output sequence of the FFT processing method; the output sequence of the FFT processing method is the same as the output sequence obtained by performing DFT processing on the input sequence.

[0022] Preferably, the base-2 FFT technique includes performing multi-stage iterative butterfly operations on the time-domain interpolation sequence with a base-2 number of points to obtain a bit-reversed FFT frequency-domain sequence; the base-2 FFT technique further includes reordering the bit-reversed FFT frequency-domain sequence to obtain the frequency-domain sequence with a base-2 number of points.

[0023] Preferably, the non-integer point interpolation technique performs data control, interpolation kernel lookup, shift register, and floating-point multiplication and addition on the input sequence; the base-2 FFT technique performs multi-stage iterative butterfly operations on the time-domain interpolation sequence; uses forward and reverse address registers and an address selector to reorder the bit-reversed FFT frequency-domain sequence; uses cyclic factor compensation and calculation truncation techniques to process the frequency-domain sequence with a base-2 number of points to obtain the output sequence of the FFT processing method.

[0024] Preferably, the output sequence of the FFT processing method is transferred to the cache and memory of the system for subsequent signal processing.

[0025] Preferably, the length of the input sequence is the only parameter for performing the non-integer point interpolation technique, the base-2 FFT technique, the phase compensation, and the truncation processing.

[0026] Compared with the prior art, the present invention can implement an FFT processing method and a processor with arbitrarily configurable conversion lengths, which can fully save the storage resource overhead in the application scenario of large-point FFT; adopts a parameter configuration scheme, and can realize the online configuration of the conversion length without changing the hardware structure according to the application scenario, simplifying the development process. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more simply illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0028] Figure 1 It is a diagram of an FFT butterfly operation unit in the prior art;

[0029] Figure 2 It is a flow chart of a DIF-FFT butterfly transformation in the prior art;

[0030] Figure 3 It is a schematic structural diagram of an FFT processor supporting a non-radix-2 point output sequence provided by an embodiment of the present application;

[0031] Figure 4 It is a flow chart of an FFT processing method supporting a non-radix-2 point output sequence provided by an embodiment of the present application;

[0032] Figure 5 It is a schematic structural diagram of a smoothing interpolation module provided by an embodiment of the present application;

[0033] Figure 6 It is a schematic structural diagram of a variable-point FFT module provided by an embodiment of the present application;

[0034] Figure 7 It is a schematic structural diagram of a pipelined reordering module provided by an embodiment of the present application;

[0035] Figure 8 It is a schematic structural diagram of a compensation truncation module provided by an embodiment of the present application. Specific Embodiments

[0036] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts in the embodiments of the present invention fall within the scope of protection of the present invention.

[0037] For the underlying hardware implementation process of FFT, it mainly includes: a butterfly operation unit (such as Figure 1The structure as shown, the generation of rotation factors, the storage structure (storage of input data, storage of intermediate operands, storage of rotation factors, etc.), the relevant addressing rules, and the hardware architecture. In terms of algorithms, FFT mainly has the following two characteristics: The amount of input data in the time domain = the amount of output data in the frequency domain. For example, for an eight-point decimation-in-frequency (DIF) FFT, each time 8 data are input, and correspondingly 8 data are output; the input and output are in a reverse order relationship. Here, the reverse order refers to the reverse order in binary. As Figure 2 shown, the input order (taking the first four numbers as an example) is: 000, 001, 010, 011 (0, 1, 2, 3), then the corresponding output is 000, 100, 010, 110 (0, 4, 2, 6). That is, if the input time-domain sequence order is x(0), x(1), x(2), x(3), then the corresponding output frequency-domain sequence is in reverse order X(0), X(4), X(2), X(6).

[0038] In typical SAR imaging algorithms, such as the range Doppler (RD) algorithm, the chirp scaling (CS) algorithm in operation, etc., FFT is an inevitable core operation step. Due to the special requirements of traditional FFT for the operation length (it is necessary to perform zero-padding operations in the time domain to complete the FFT points to a power of 2) and the two-dimensional nature of SAR processing (the azimuth direction is the flight direction, and the range direction is the scanning direction, which are perpendicular to each other), the data scale is often expanded by 2-3 times. In the hardware system, the operation results of FFT operations are often transferred to the cache and memory of the system for subsequent signal processing. In the current SAR imaging system, the on-chip cache capacity and off-chip memory bandwidth are the most precious hardware resources and are often the bottleneck factors restricting the system performance.

[0039] The FFT processor proposed in this application adopts a pipeline structure. The FFT processor with a pipeline structure can simultaneously achieve continuous data input and output, large data throughput capacity, and high real-time performance, and is the most widely used. The circuit structure with the least resource occupancy and the best performance in the pipeline structure is the single-path delay feedback (SDF) structure.

[0040] In Figures 3 - 7 this invention, a 16- to 1024-point arbitrary-point FFT processor (single-precision floating-point) is used as an example for illustration.

[0041] Figure 3A structural schematic diagram of an FFT processor provided by an embodiment of the present application. The hardware implementation adopts a data flow-driven structure, which is composed of a smoothing interpolation module 31, a variable-point FFT module 32, a pipelined reordering module 33, and a compensation truncation module 34 connected end to end.

[0042] In Figure 3 the non-integer power conversion length FFT processor sequentially completes smoothing interpolation after cyclic shift, N-point FFT, reordering, and cyclic factor compensation and M-point truncation. The conversion length is independently provided to each module as the only parameter.

[0043] The smoothing interpolation module 31 is used to perform non-integer point interpolation on the input sequence to obtain a time-domain interpolation sequence with a base-2 number of points; the non-integer point interpolation makes the output sequence obtained by the input sequence passing through the FFT processor be the sequence obtained by zero-padding the output sequence of the DFT processing of the input sequence in the frequency domain before the truncation process; as Figure 5 shown, the smoothing interpolation module includes a data controller, a shift register, a floating-point multiplier-adder, and an interpolation kernel lookup table.

[0044] The principle of non-integer point time-domain interpolation is as follows:

[0045] Suppose a discrete time-domain signal is x(m), where m = 0, 1, 2,..., M - 1. Its discrete spectrum is X(w), where w = 0, 1, 2,..., M - 1. X(w) can be obtained by the following DFT formula:

[0046]

[0047] The most classic FFT algorithm, as a fast algorithm for DFT, requires that the conversion length must be an integer power of 2. When M does not meet this condition, M needs to be extended to N = 2 s . For example, a new spectrum X′(k) can be obtained using the FFT principle. This spectrum is obtained by supplementing M - N zeros after X(w) calculated by the DFT formula, that is, zero-padding the frequency domain of X(w). If the result of zero-padding the frequency domain is used as the output of the FFT processor, the required spectrum of any length can be obtained through simple truncation. Therefore, how to preprocess x(m) to obtain an input x′(n) with a length of N, and the FFT result of this x′(n) is the zero-padding of the original spectrum, is the key point of the technical solution of the present invention. The algorithm derivation is as follows:

[0048] Zero-padding X(w) to X′(k), where k = 0, 1, 2,..., N - 1, is summarized as follows:

[0049]

[0050]

[0051] According to the DFT formula, we can get:

[0052]

[0053] Let:

[0054]

[0055] Substituting into the original formula, we can get:

[0056]

[0057] The last term in the above formula is a weight term similar to the sinc function distribution: where the range of θ is -π to π. Therefore, sin(θ / 2) is a sine signal with a half-period duration, and the oscillation frequency of sin(Mθ / 2) is M times that of sin(θ / 2), lasting for M / 2 periods in total. So the main energy of this weight is concentrated near θ = 0. According to the calculation, 90% of the energy is distributed in the range of -0.009 rad to 0.009 rad, and 99% of the energy is distributed in the range of -0.1 rad to 0.1 rad. From the energy perspective, we can narrow the weighting range to near θ = 0, and at the same time, we can replace sin(θ / 2) with θ / 2. The formula is as follows:

[0058]

[0059] The formula is converted into the form of sinc interpolation, which is consistent with the principle that zero-padding in the frequency domain corresponds to interpolation in the time domain in the time-frequency relationship. Assume that the length of the interpolation kernel is cl, and it is customary to take cl as an even number. Let:

[0060]

[0061] The range of m can be obtained as follows, where m takes integers and when n takes a single value, it corresponds to cl values of m:

[0062]

[0063] The original formula is transformed into:

[0064]

[0065] Considering that the truncation operation on the original weighted term in finite-length interpolation will introduce ringing effects, the sinc function is weighted by a Kaiser window, and the weighted sinc function is named sinc′, representing the smoothed interpolation function. Due to the characteristics of the FFT algorithm itself, the high frequencies of the calculated spectrum are distributed in the middle, and the low frequencies are distributed on both sides, that is, the spectrum axis is [0~fs / 2, -fs / 2~0). If zeros are directly padded to the left of X(w) to obtain X′(k), the low-frequency region in X(w) will enter the high-frequency region of X′(k). Since the Kaiser window has a short stopband in the high-frequency region, the above method will significantly weaken the high-frequency components of x′(n) and cause spectrum distortion. The above problem can be solved by inserting the zero-padding sequence of X(w) into the high-frequency region through time-domain phase shift. The specific method is as follows: Circularly shift X(w) to the right by M / 2 to obtain X 1 (w); Pad M - N zeros to the right end of X 1 (w) to obtain X 2 (w); Circularly shift X 2 (w) to the left by M / 2 to obtain X′(k). According to signal processing theory:

[0066]

[0067] The above three steps can be summarized as:

[0068]

[0069] Combining the three equations gives:

[0070]

[0071] In the above formula, x′(n) corresponds to the result of zero-padding in the high-frequency region. At the same time, it can be found that the formula is naturally simplified.

[0072] Considering that is a very small phase perturbation offset value, the error in the FFT calculation is much smaller than the error introduced by the finite interpolation kernel, and the cost of real-time calculation is large. Therefore, the formula is further simplified to:

[0073]

[0074] Considering that the first cl values of x′(n) may require the last few values of x(m) for calculation, from the perspective of real-time processing, the FFT processor inputs x″(n), and x″(n) is x′(n) circularly shifted to the left by cl:

[0075] x″(n) = x′(mod(n + cl, N)),

[0076] In this way, the calculation of the first cl values of x′(n) can be placed at the end, but the first few values of x(m) need to be cached.

[0077] According to signal processing theory:

[0078]

[0079] The final output spectrum is adjusted to:

[0080]

[0081] In summary, the algorithm flow is summarized as follows:

[0082]

[0083] The target spectrum X(w) can be obtained by intercepting X′(k) in the following way:

[0084] M is even:

[0085] M is odd:

[0086] The variable-point FFT module 32 is used to perform multi-stage iterative butterfly operations on the time-domain interpolation sequence of the radix-2 number of points to obtain the bit-reversed FFT frequency-domain sequence; as Figure 6 shown, the variable-point FFT module includes multi-stage single-path delay feedback structure units and a decoding controller.

[0087] In one embodiment, the time-domain interpolation sequence is subjected to several stages of iterative butterfly operations using the radix-2 DIF-FFT technique to obtain the bit-reversed FFT frequency-domain sequence, and the number of stages of most iterations is determined according to the conversion length.

[0088] Let the length of the time-domain sequence x(n) be N, N is even, that is, N = 2M. Group x(n) before and after, and the DFT can be decomposed as:

[0089]

[0090] It can be seen from the above formula that different values of k will result in different addition and subtraction operations between two data. Based on this, X(k) is grouped according to odd and even, that is, X(2r) and X(2r + 1), where r = 0, 1,..., N / 2 - 1. Therefore, the above formula can be further decomposed as:

[0091]

[0092]

[0093] Let:

[0094]

[0095] The radix-2 butterfly unit can be deduced as Figure 1 shown.

[0096] Similarly, it can be known that when N / 2 is still an even number, continue to decompose according to this butterfly unit. When N is an integer power of 2, continue to decompose until the number of points is 2, and the 2-point DFT is also implemented using a butterfly unit. As Figure 2 shown in the schematic diagram of 8-point DIF-FFT

[0097] The pipelined reordering module 33 is used to reorder the bit-reversed FFT frequency domain sequence to obtain the frequency domain sequence with the number of points of radix 2; as Figure 7 shown, the pipelined reordering module includes forward and reverse address registers, an address selector, and a random access memory

[0098] The compensation truncation module 34 is used to perform phase compensation on the frequency domain sequence with the number of points of radix 2, and perform truncation processing to obtain the output sequence with non-radix 2 points as the output sequence of the FFT processor; the output sequence of the FFT processor is the same as the output sequence of the DFT processing of the input sequence; as Figure 8 shown, the compensation truncation module includes a factor multiplier and a calculation truncator

[0099] Figure 4 The flowchart of an FFT processing method supporting an output sequence with non-radix 2 points provided by an embodiment of the present application includes the following steps

[0100] Step S410: Obtain an input sequence, and perform smooth interpolation on the input sequence according to non-integer point interpolation technology to obtain a time domain interpolation sequence with the number of points of radix 2; non-integer point interpolation technology makes the output sequence obtained by the input sequence through the FFT processing method, before truncation processing, be the sequence obtained by zero-padding in the frequency domain of the output sequence of the DFT processing of the input sequence

[0101] Step S420: Process the time domain interpolation sequence with the number of points of radix 2 according to the radix 2 FFT technology to obtain a frequency domain sequence with the number of points of radix 2

[0102] The radix 2 FFT technology includes performing multi-stage iterative butterfly operations on the time domain interpolation sequence with the number of points of radix 2 to obtain a bit-reversed FFT frequency domain sequence; the radix 2 FFT technology also includes reordering the bit-reversed FFT frequency domain sequence to obtain a frequency domain sequence with the number of points of radix 2

[0103] Step S430: Perform phase compensation on the frequency domain sequence with the number of points of radix 2, and perform truncation processing to obtain the output sequence with non-radix 2 points as the output sequence of the FFT processing method; the output sequence of the FFT processing method is the same as the output sequence of the DFT processing of the input sequence

[0104] The non-integer point interpolation technology performs data control, interpolation kernel lookup, shift register, and floating-point multiplication and addition on the input sequence; the radix-2 FFT technology performs multi-stage iterative butterfly operations on the time-domain interpolation sequence; the bit-reversed FFT frequency-domain sequence is reordered using forward and reverse address registers and an address selector; the radix-2 point frequency-domain sequence is processed using cyclic factor compensation and calculation truncation techniques to obtain the output sequence of the FFT processing method.

[0105] The output sequence of the FFT processing method will be transferred to the cache and memory of the system for subsequent signal processing.

[0106] The length of the input sequence is the only parameter for performing non-integer point interpolation technology, radix-2 FFT technology, phase compensation, and truncation processing.

[0107] Figure 5 The structure diagram of the smoothing interpolation module provided by the embodiment of the present application. The smoothing interpolation module includes a data controller, a shift register, a floating-point multiplier, and an interpolation kernel lookup table to perform non-integer point time-domain interpolation on the obtained discrete time-domain sequence that needs to be FFT processed.

[0108] In one embodiment,

[0109] The original data x(m) generates x″(n) through the interpolation operation of this module. The data stream protocol is the most common AXI4-Stream. The interpolation kernel length is selected as 16, the quantization displacement is 1 / 64, and the sidelobe parameter of the Kaiser window is 10.

[0110] The introduction of each unit in the smoothing interpolation module is as follows:

[0111] The data controller 311, the execution process is as follows:

[0112] (1) Temporarily store the first 8 data in the register bank and take them out when calculating the last 16 data of x″(n);

[0113] (2) When calculating a certain x″(n) according to the conversion length control, determine whether to receive and transfer new data to the shift register (according to the range of m, it can be known that as n increases, m also slides and increases. However, since M < N, the sliding speed of m is less than that of n, that is, there are multiple n values corresponding to the same group of m values. In this case, when calculating x″(n), there is no need to receive new x(m), and a blocking signal is sent through the AXI4-Stream protocol);

[0114] (3) Control the enabling of the floating-point multiplier;

[0115] (4) Provide an address to the interpolation kernel lookup table according to the conversion length to fetch the required 16×32bit smoothing interpolation kernel.

[0116] The shift register 312 has the following execution process:

[0117] (1) Sixteen 64-bit registers are connected end to end to store the 16 x(m)'s required for calculating a certain x″(n).

[0118] (2) Transfer the 16 x(m)'s to the floating-point multiply-accumulator.

[0119] The floating-point multiply-accumulator 313 simultaneously performs multiplication of 16 data and 16 smoothing interpolation kernels, and then adds the multiplication results and outputs them.

[0120] The interpolation kernel lookup table 314 has the following execution process:

[0121] (1) It is used to store the 64×16-point smoothing interpolation kernel and provide a 16-point interpolation kernel according to the address.

[0122] (2) It is implemented by a single-port ROM, with a data bit width of 16×32 = 512 bit, and the total space occupies 64×16×32 bit = 4 KB.

[0123] (3) The sinc′ smoothing interpolation kernel data is shown in Table 1 below:

[0124] Table 1:

[0125] 0.000 0.000 0.000 0.000 0.000 0.000 0.000 1.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 0.000 -0.001 0.002 -0.003 0.006 -0.015 1.000 0.015 -0.007 0.003 -0.002 0.001 0.000 0.000 0.000 0.000 0.001 -0.002 0.003 -0.007 0.013 -0.029 0.998 0.031 -0.013 0.007 -0.004 0.002 -0.001 0.000 0.000 0.000 0.001 -0.003 0.005 -0.010 0.019 -0.042 0.996 0.047 -0.020 0.010 -0.005 0.003 -0.001 0.000 0.000 -0.001 0.001 -0.003 0.007 -0.013 0.025 -0.055 0.994 0.064 -0.027 0.014 -0.007 0.004 -0.002 0.001 0.000 -0.001 0.002 -0.004 0.008 -0.016 0.030 -0.068 0.990 0.081 -0.034 0.018 -0.009 0.005 -0.002 0.001 0.000 -0.001 0.002 -0.005 0.010 -0.019 0.036 -0.080 0.985 0.098 -0.041 0.021 -0.011 0.006 -0.003 0.001 0.000 -0.001 0.002 -0.006 0.011 -0.022 0.041 -0.091 0.980 0.116 -0.048 0.025 -0.013 0.007 -0.003 0.001 0.000 -0.001 0.003 -0.006 0.013 -0.024 0.046 -0.102 0.974 0.134 -0.055 0.028 -0.015 0.008 -0.003 0.001 0.000 -0.001 0.003 -0.007 0.014 -0.027 0.051 -0.112 0.967 0.153 -0.062 0.032 -0.017 0.009 -0.004 0.002 0.000 -0.001 0.003 -0.008 0.015 -0.029 0.056 -0.122 0.960 0.172 -0.069 0.036 -0.019 0.009 -0.004 0.002 -0.001 -0.001 0.004 -0.008 0.017 -0.032 0.060 -0.131 0.951 0.191 -0.076 0.039 -0.021 0.010 -0.005 0.002 -0.001 -0.001 0.004 -0.009 0.018 -0.034 0.064 -0.139 0.942 0.211 -0.083 0.043 -0.023 0.011 -0.005 0.002 -0.001 -0.001 0.004 -0.009 0.019 -0.036 0.068 -0.147 0.932 0.231 -0.090 0.046 -0.025 0.012 -0.006 0.002 -0.001 -0.002 0.004 -0.010 0.020 -0.038 0.072 -0.154 0.922 0.251 -0.097 0.050 -0.026 0.013 -0.006 0.002 -0.001 -0.002 0.004 -0.010 0.021 -0.040 0.075 -0.161 0.910 0.272 -0.104 0.053 -0.028 0.014 -0.007 0.003 -0.001 -0.002 0.005 -0.011 0.022 -0.041 0.078 -0.167 0.898 0.292 -0.111 0.057 -0.030 0.015 -0.007 0.003 -0.001 -0.002 0.005 -0.011 0.022 -0.043 0.081 -0.173 0.885 0.313 -0.118 0.060 -0.032 0.016 -0.007 0.003 -0.001 -0.002 0.005 -0.011 0.023 -0.044 0.084 -0.178 0.872 0.334 -0.124 0.063 -0.034 0.017 -0.008 0.003 -0.001 -0.002 0.005 -0.012 0.024 -0.046 0.086 -0.182 0.858 0.355 -0.131 0.067 -0.035 0.018 -0.008 0.003 -0.001 -0.002 0.005 -0.012 0.024 -0.047 0.089 -0.186 0.844 0.377 -0.137 0.070 -0.037 0.019 -0.009 0.004 -0.001 -0.002 0.005 -0.012 0.025 -0.048 0.090 -0.189 0.828 0.398 -0.143 0.072 -0.038 0.020 -0.009 0.004 -0.001 -0.002 0.005 -0.012 0.025 -0.049 0.092 -0.192 0.813 0.419 -0.149 0.075 -0.040 0.020 -0.010 0.004 -0.001 -0.002 0.005 -0.012 0.026 -0.049 0.093 -0.194 0.796 0.441 -0.154 0.078 -0.041 0.021 -0.010 0.004 -0.001 -0.002 0.005 -0.012 0.026 -0.050 0.094 -0.196 0.780 0.462 -0.160 0.080 -0.043 0.022 -0.010 0.004 -0.001 -0.002 0.005 -0.013 0.026 -0.050 0.095 -0.197 0.762 0.484 -0.165 0.083 -0.044 0.022 -0.011 0.004 -0.001 -0.002 0.005 -0.013 0.026 -0.051 0.096 -0.197 0.745 0.505 -0.169 0.085 -0.045 0.023 -0.011 0.004 -0.002 -0.002 0.005 -0.013 0.026 -0.051 0.096 -0.197 0.726 0.526 -0.174 0.087 -0.046 0.024 -0.011 0.005 -0.002 -0.002 0.005 -0.013 0.026 -0.051 0.096 -0.197 0.708 0.547 -0.178 0.089 -0.047 0.024 -0.011 0.005 -0.002 -0.002 0.005 -0.013 0.026 -0.051 0.096 -0.196 0.689 0.568 -0.182 0.091 -0.048 0.025 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.026 -0.051 0.096 -0.195 0.669 0.589 -0.185 0.092 -0.049 0.025 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.026 -0.050 0.095 -0.193 0.650 0.609 -0.188 0.093 -0.049 0.025 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.026 -0.050 0.094 -0.191 0.630 0.630 -0.191 0.094 -0.050 0.026 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.025 -0.049 0.093 -0.188 0.609 0.650 -0.193 0.095 -0.050 0.026 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.025 -0.049 0.092 -0.185 0.589 0.669 -0.195 0.096 -0.051 0.026 -0.012 0.005 -0.002 -0.002 0.005 -0.012 0.025 -0.048 0.091 -0.182 0.568 0.689 -0.196 0.096 -0.051 0.026 -0.013 0.005 -0.002 -0.002 0.005 -0.011 0.024 -0.047 0.089 -0.178 0.547 0.708 -0.197 0.096 -0.051 0.026 -0.013 0.005 -0.002 -0.002 0.005 -0.011 0.024 -0.046 0.087 -0.174 0.526 0.726 -0.197 0.096 -0.051 0.026 -0.013 0.005 -0.002 -0.002 0.004 -0.011 0.023 -0.045 0.085 -0.169 0.505 0.745 -0.197 0.096 -0.051 0.026 -0.013 0.005 -0.002 -0.001 0.004 -0.011 0.022 -0.044 0.083 -0.165 0.484 0.762 -0.197 0.095 -0.050 0.026 -0.013 0.005 -0.002 -0.001 0.004 -0.010 0.022 -0.043 0.080 -0.160 0.462 0.780 -0.196 0.094 -0.050 0.026 -0.012 0.005 -0.002 -0.001 0.004 -0.010 0.021 -0.041 0.078 -0.154 0.441 0.796 -0.194 0.093 -0.049 0.026 -0.012 0.005 -0.002 -0.001 0.004 -0.010 0.020 -0.040 0.075 -0.149 0.419 0.813 -0.192 0.092 -0.049 0.025 -0.012 0.005 -0.002 -0.001 0.004 -0.009 0.020 -0.038 0.072 -0.143 0.398 0.828 -0.189 0.090 -0.048 0.025 -0.012 0.005 -0.002 -0.001 0.004 -0.009 0.019 -0.037 0.070 -0.137 0.377 0.844 -0.186 0.089 -0.047 0.024 -0.012 0.005 -0.002 -0.001 0.003 -0.008 0.018 -0.035 0.067 -0.131 0.355 0.858 -0.182 0.086 -0.046 0.024 -0.012 0.005 -0.002 -0.001 0.003 -0.008 0.017 -0.034 0.063 -0.124 0.334 0.872 -0.178 0.084 -0.044 0.023 -0.011 0.005 -0.002 -0.001 0.003 -0.007 0.016 -0.032 0.060 -0.118 0.313 0.885 -0.173 0.081 -0.043 0.022 -0.011 0.005 -0.002 -0.001 0.003 -0.007 0.015 -0.030 0.057 -0.111 0.292 0.898 -0.167 0.078 -0.041 0.022 -0.011 0.005 -0.002 -0.001 0.003 -0.007 0.014 -0.028 0.053 -0.104 0.272 0.910 -0.161 0.075 -0.040 0.021 -0.010 0.004 -0.002 -0.001 0.002 -0.006 0.013 -0.026 0.050 -0.097 0.251 0.922 -0.154 0.072 -0.038 0.020 -0.010 0.004 -0.002 -0.001 0.002 -0.006 0.012 -0.025 0.046 -0.090 0.231 0.932 -0.147 0.068 -0.036 0.019 -0.009 0.004 -0.001 -0.001 0.002 -0.005 0.011 -0.023 0.043 -0.083 0.211 0.942 -0.139 0.064 -0.034 0.018 -0.009 0.004 -0.001 -0.001 0.002 -0.005 0.010 -0.021 0.039 -0.076 0.191 0.951 -0.131 0.060 -0.032 0.017 -0.008 0.004 -0.001 -0.001 0.002 -0.004 0.009 -0.019 0.036 -0.069 0.172 0.960 -0.122 0.056 -0.029 0.015 -0.008 0.003 -0.001 0.000 0.002 -0.004 0.009 -0.017 0.032 -0.062 0.153 0.967 -0.112 0.051 -0.027 0.014 -0.007 0.003 -0.001 0.000 0.001 -0.003 0.008 -0.015 0.028 -0.055 0.134 0.974 -0.102 0.046 -0.024 0.013 -0.006 0.003 -0.001 0.000 0.001 -0.003 0.007 -0.013 0.025 -0.048 0.116 0.980 -0.091 0.041 -0.022 0.011 -0.006 0.002 -0.001 0.000 0.001 -0.003 0.006 -0.011 0.021 -0.041 0.098 0.985 -0.080 0.036 -0.019 0.010 -0.005 0.002 -0.001 0.000 0.001 -0.002 0.005 -0.009 0.018 -0.034 0.081 0.990 -0.068 0.030 -0.016 0.008 -0.004 0.002 -0.001 0.000 0.001 -0.002 0.004 -0.007 0.014 -0.027 0.064 0.994 -0.055 0.025 -0.013 0.007 -0.003 0.001 -0.001 0.000 0.000 -0.001 0.003 -0.005 0.010 -0.020 0.047 0.996 -0.042 0.019 -0.010 0.005 -0.003 0.001 0.000 0.000 0.000 -0.001 0.002 -0.004 0.007 -0.013 0.031 0.998 -0.029 0.013 -0.007 0.003 -0.002 0.001 0.000 0.000 0.000 0.000 0.001 -0.002 0.003 -0.007 0.015 1.000 -0.015 0.006 -0.003 0.002 -0.001 0.000 0.000

[0126] Considering the length of the interpolation kernel, in this embodiment, the number of points from 0 to 16 is not preferably used. This is because Figure 5 in the interpolation kernel length is set to 16, that is, when the length of x(m) is at least 16, the calculated result can be guaranteed in terms of accuracy, that is, x(n) is at least 16 points. If the length of x(n) is less than 16, it will cause the number of x(m)'s participating in the weighting to be too small, and the operation accuracy cannot be guaranteed. In addition, the FFT within 100 points occupies very little storage resources, and there is no need to perform targeted design for saving storage. This application mainly aims at large-point application scenarios similar to SAR.

[0127] Figure 6 is a schematic structural diagram of the variable-point FFT module provided by the embodiment of the present application. The variable-point FFT module includes a multi-stage single-path delay feedback structure unit and a decoding controller, and is used to perform several levels of iterative butterfly operations on the time-domain interpolation sequence to obtain the bit-reversed FFT frequency-domain sequence.

[0128] In one embodiment,

[0129] This module undertakes the operation of FFT. Considering that when M < N / 2, only an FFT of N / 2 points or smaller points needs to be performed, this module supports FFT calculations within 1024 points and at integer powers of 2. According to the basic principle of the FFT algorithm, the data stream output by this module is in bit-reversed order. The following figure shows the schematic structure of the variable-point FFT module. The variable-point FFT module consists of 10 levels of SDF structures connected end to end, and the data paths of each level are controlled by a decoding controller.

[0130] The introduction of each unit in the variable-point FFT module is as follows:

[0131] The SDF unit 321 has the following execution process:

[0132] (1) The most widely used pipelined FFT structure. Operating principle: Taking the 0th-level SDF unit as an example, when the current 512-bit data flows into this structure, the data is cached in the 0th-level memory. During the inflow of the subsequent 512-bit data, the corresponding data in the buffer is read and sent to the butterfly unit for operation, and the addition result of the butterfly unit is sent to the W-factor multiplication unit. At the same time, the subtraction result of the butterfly unit is fed back to the 0th-level memory to overwrite the used data. When all the subsequent 512-bit data has flowed in, the 0th-level memory is all the subtraction results of the butterfly unit. At this time, the SDF structure receives the first 512 bits of the next frame of data, and the subtraction results in the 0th-level memory are sent to the W-factor multiplication unit. Repeat this process to complete the first-level calculation of multiple frames of FFT. The result of the nth-level SDF is exactly the same as that of the 0th-level SDF structure, only the size of the nth-level memory and the span of the processed data gradually decrease by 1 / 2.

[0133] (2) The data distributor and data selector work together to determine whether the data stream passes through a certain SDF structure for processing or is directly sent to the next level;

[0134] The decoding controller 322 determines the number of SDF structures to be used according to the conversion length and controls all data distributors and data selectors to achieve the control of the data link. Try to use the last-level SDF structure to avoid dynamic configuration of the SDF structure. For example, when M = 500, use levels 1 - 9; when M = 50, use levels 4 - 9.

[0135] The structure processor of the variable-point FFT module provided in this embodiment has a simple design and flexible configuration. It can be richly configured with different numbers of points or adapted to different conversion length scenarios. Only by controlling the number of SDF structures to be used according to the decoding controller and reusing the SDF of the same structure, any conversion length scenario can be achieved without the need for major changes to the hardware structure, which is very conducive to rapid development.

[0136] Figure 7This is a schematic diagram of the pipelined reordering module provided by an embodiment of the present application. The pipelined reordering module includes forward and reverse address registers, an address selector, and a random access memory, and is used to reorder the bit-reversed FFT frequency domain sequence to obtain the bit-normal FFT frequency domain sequence. This module is responsible for completing the reordering of the bit-reversed FFT result in a pipelined manner.

[0137] In one embodiment, the units in the pipelined reordering module are introduced as follows:

[0138] The forward address register 331 and the reverse address register 332 generate two address logics according to the bit-reversal principle for the address selector to select.

[0139] The address selector 333 performs the following operations:

[0140] (1) Alternately select the forward and reverse address pairs to read and write to the memory to achieve pipelined reordering. The process is as follows: The first frame of data is read in sequentially and read out reversely. Since the read and write share the address, at the same time, the second frame of data is read in reversely. Since the second frame of data is read in reversely, it only needs to be read out sequentially, and at the same time, the third frame of data is read in sequentially. Repeating this process can achieve pipelined reordering.

[0141] (2) Intercept the address according to the conversion length. Considering that when M < N / 2, only an FFT of N / 2 points or smaller points is required. For example, when M = 500, the forward address is intercepted as [8:0], and the reverse address is intercepted as [9:1]; when M = 50, the forward address is intercepted as [5:0], and the reverse address is intercepted as [9:4].

[0142] The random access memory (RAM) 334 is used to cache one frame of complete data to complete the reordering. The read and write addresses are the same, that is, the data is overwritten by new data when it is read out. The total space occupies 1024 × 64 bit = 8 KB.

[0143] Figure 8 This is a schematic diagram of the compensation and interception module provided by an embodiment of the present application. The compensation and interception module includes a factor multiplier and a calculation interceptor, and is used to perform phase compensation on the bit-normal FFT frequency domain sequence and intercept the non-zero part of the bit-normal FFT frequency domain sequence as the FFT frequency domain sequence of the conversion length. This module is responsible for performing phase compensation on the reordered N-point FFT result and intercepting it to M points.

[0144] In one embodiment, the units in the compensation and interception module are introduced as follows:

[0145] The factor multiplier 341 is used to complete the multiplication of the FFT result and the compensation phase of.

[0146] The counting interceptor 342 intercepts the compensated X′(k) according to the interception formula to obtain X(w).

[0147] The FFT processor adopting the technical solution of the present invention to support non-radix-2 point output sequences can avoid expanding the data scale, can multiply reduce the cache and memory overheads, compress the system hardware scale, and at the same time can relieve the operation pressure of subsequent signal processing.

[0148] It should be noted that the method provided herein is not inherently related to any specific computer, virtual processor or other device. Various general-purpose processors can also be used together with the teachings based herein. The structure required to construct such a processor is obvious according to the above description. In addition, the present invention is not directed to any specific programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the descriptions of the calls to specific languages and system function modules above are only for disclosing the best implementation mode of the invention.

[0149] In the specification provided herein, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be implemented without these specific details. In some examples, well-known methods, structures and technologies are not shown in detail so as not to obscure the understanding of this specification.

[0150] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. An FFT processor supporting a non - radix - 2 point output sequence, comprising: a smoothing interpolation module for performing non - integer point interpolation on an input sequence to obtain a time - domain interpolation sequence of radix - 2 points; the non - integer point interpolation enables the output sequence obtained by passing the input sequence through the FFT processor to be the sequence obtained by zero - padding the output sequence of the DFT processing of the input sequence in the frequency domain before the truncation process; an FFT processing module for performing radix - 2 FFT processing on the time - domain interpolation sequence of radix - 2 points to obtain a frequency - domain sequence of radix - 2 points; the FFT processing module includes a variable - point FFT module and a pipelined re - ordering module; the variable - point FFT module is used for performing multi - stage iterative butterfly operations on the time - domain interpolation sequence to obtain a bit - reversed FFT frequency - domain sequence and determining the number of stages of multi - stage iteration according to the transformation length; the pipelined re - ordering module is used for re - ordering the bit - reversed FFT frequency - domain sequence to obtain the frequency - domain sequence of radix - 2 points; a compensation truncation module for performing phase compensation on the frequency - domain sequence of radix - 2 points and performing the truncation process to obtain a non - radix - 2 point output sequence as the output sequence of the FFT processor; the output sequence of the FFT processor is the same as the output sequence of the DFT processing of the input sequence.

2. The processor according to claim 1, wherein, the smoothing interpolation module includes a data controller, a shift register, a floating - point multiplier - adder, and an interpolation kernel lookup table; the variable - point FFT module includes multi - stage single - path delay feedback structure units and a decoding controller; the pipelined re - ordering module includes forward and reverse address registers, an address selector, and a random access memory; the compensation truncation module includes a factor multiplier and a calculation truncator.

3. The processor according to claim 1, wherein, the output sequence of the FFT processor is transferred to the cache and memory of the system for subsequent signal processing.

4. The processor according to claim 1, wherein, the length of the input sequence is independently provided as the only parameter to the smoothing interpolation module, the FFT processing module, and the compensation truncation module.

5. An FFT processing method supporting a non - radix - 2 point output sequence, comprising: obtaining an input sequence and performing smoothing interpolation on the input sequence according to non - integer point interpolation technology to obtain a time - domain interpolation sequence of radix - 2 points; the non - integer point interpolation technology enables the output sequence obtained by passing the input sequence through the FFT processing method to be the sequence obtained by zero - padding the output sequence of the DFT processing of the input sequence in the frequency domain before the truncation process; Process the time-domain interpolation sequence with a base-2 number of points according to the base-2 FFT technique to obtain a frequency-domain sequence with a base-2 number of points; the process of processing the time-domain interpolation sequence with a base-2 number of points according to the base-2 FFT technique to obtain a frequency-domain sequence with a base-2 number of points includes: performing a multi-stage iterative butterfly operation on the time-domain interpolation sequence to obtain a bit-reversed FFT frequency-domain sequence, and determining the number of stages of the multi-stage iteration according to the transformation length; reordering the bit-reversed FFT frequency-domain sequence to obtain the frequency-domain sequence with a base-2 number of points; Perform phase compensation on the frequency-domain sequence with a base-2 number of points, and perform the truncation process to obtain an output sequence with a non-base-2 number of points as the output sequence of the FFT processing method; the output sequence of the FFT processing method is the same as the output sequence of the DFT processing of the input sequence.

6. The method according to claim 5, wherein, The non-integer point interpolation technique performs data control, interpolation kernel lookup, shift register, and floating-point multiplication and addition on the input sequence; the base-2 FFT technique performs a multi-stage iterative butterfly operation on the time-domain interpolation sequence; the bit-reversed FFT frequency-domain sequence is reordered by using forward and reverse address registers and an address selector; The frequency-domain sequence with a base-2 number of points is processed by using cyclic factor compensation and calculation truncation techniques to obtain the output sequence of the FFT processing method.

7. The method according to claim 5, wherein, The output sequence of the FFT processing method will be transferred to the cache and memory of the system for subsequent signal processing.

8. The method according to claim 5, wherein, The length of the input sequence is the only parameter for performing the non-integer point interpolation technique, the base-2 FFT technique, the phase compensation, and the truncation process.

Citation Information

Patent Citations

  • Narrowband communication signal processing method based on multimode reconfigurable FFT

    CN114201725A