Frequency domain implementation method of low-resource high-order FIR filter of FPGA
By adopting the segmented convolution method on the FPGA platform, combined with the overlapping addition method and the overlapping retention method, the frequency domain implementation of high-order FIR filters under low resource conditions is achieved, the problems of resource limitation and computational complexity in the prior art are solved, and efficient high-order filter calculation is realized.
Patent Information
- Application Number
- CN202510153020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-05-30
AI Technical Summary
The existing FPGA platform is difficult to effectively implement high-order FIR filters, mainly due to resource limitations and the complexity of FFT calculations, and cannot meet the resource requirements of high-order filters.
Through segmented convolution, a combination of overlapping addition method and overlapping retention method is used to realize a higher-order FIR filter in the frequency domain, and a small amount of hardware resources are used for calculation, reducing the use of resources such as multipliers.
The frequency domain implementation of high-order FIR filters is realized under low resource conditions, reducing the computational complexity and resource consumption, and supporting the implementation of ultra-high-order FIR filters.
Smart Images

Figure CN120074449A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of signal processing, and particularly to a method for implementing a high-order FIR filter with low resources in the frequency domain on an FPGA. Background Art
[0002] In digital communication systems, filters are one of the important components, and the selection of filter characteristics directly affects the recovery of digital signals. However, as the characteristics of the processed signals and system requirements become higher and higher, the order of the filters is also getting larger. Currently, there are generally two schemes for implementing FIR filters based on FPGAs. One is the time-domain scheme, and the other is the frequency-domain scheme. Analyzing from the time domain, if a high-order filter is implemented, the occupation of multipliers and storage resources is huge, and there is no suitable chip to meet the requirements. Therefore, this scheme can only be used for the implementation of low-order filters. However, implementing based on FFT in the frequency domain can greatly reduce the computational complexity and reduce the resource usage. However, currently mainstream FPGAs, such as the FFT IP cores provided by Xilinx official, can only calculate FFT of up to 65536 points at most, and they also consume a large amount of resources. For matched filtering that requires filters of more than a hundred thousand or two hundred thousand orders, they cannot be directly implemented. Even the FPGA of Altera provides an FFT IP of up to 262146 points, and only the top-level chips can support the required resources. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for implementing a high-order FIR filter with low resources in the frequency domain on an FPGA, which realizes an ultra-high-order FIR filter in the frequency domain by means of segmented convolution using a small amount of hardware resources.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions:
[0005] The present invention provides a method for implementing a high-order FIR filter with low resources in the frequency domain on an FPGA, including the following steps:
[0006] S1. Initialize the values of a partial buffer area to 0 for subsequent zero-padding.
[0007] S2. Receive the impulse response coefficients with a total length of M, read 2L-length coefficients in an overlapping L-length manner, and store them in the response buffer after performing the fast Fourier transform (FFT).
[0008] S3. Receive the data of L-length of the original sequence, perform FFT on it, and use the overlap-save method to calculate the convolution between the L-length sequence and the M-length coefficients of the pre-stored FFT coefficients.
[0009] S4. After reading the convolution truncation result between the sequence of the previous L length and the coefficient of the M length, padding L length of 0s, adding it to the current calculation result by the overlap - add method, and outputting the result of L length, and storing the remaining intermediate result of M length in the corresponding cache area for the next calculation application.
[0010] Further, the S1 includes: When calculating the FFT, the zero - padding operation is directly achieved by initializing the corresponding cache area to 0, and then directly reading and writing according to the corresponding address block to complete the zero - padding operation.
[0011] Further, the S2 includes: After padding L length of 0s before and after the coefficient h(n), reading the coefficient of 2L length with an overlap of L length, and after performing the FFT of 2L length, putting it into the cache area, without performing the FFT calculation multiple times.
[0012] Further, the S3 includes: In the internal sub - process, the convolution of the sequence of L length and the impulse response of M length is implemented by the overlap - save method.
[0013] Further, the S4 is specifically as follows:
[0014] S401. Using the dual - area ping - pong operation to cache the previous calculation result;
[0015] S402. By the method of padding L length of 0s after the M - length data, back - filling the calculation result into the reading area to implement the overlap - add method of the sliding window. Description of the Drawings
[0016] Figure 1 It is a schematic diagram of the calculation process of the overlap - add method;
[0017] Figure 2 It is a schematic diagram of the computer process of the overlap - save method;
[0018] Figure 3 It is a schematic diagram of the calculation processes of the overlap - save method and the overlap - add method. Detailed Embodiment
[0019] In order to make the purpose, technical solution and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the drawings. It should be understood that the specific embodiments described here are only used to explain the present invention and are not used to limit the present invention.
[0020] A frequency - domain implementation method of a low - resource high - order FIR filter for FPGA includes the following steps:
[0021] S1. Initialize the values of some cache areas to 0 for subsequent zero - padding;
[0022] S2. Receive the impulse response coefficients with a total length of M, read 2L-length coefficients in an overlapping L-length manner, and store them in the response buffer after performing the Fast Fourier Transform (FFT).
[0023] S3. Receive data of length L of the original sequence, perform FFT on it, and convolve it with the pre-stored coefficients after FFT using the overlap-save method to calculate the convolution between the sequence of length L and the coefficients of length M.
[0024] S4. Read the convolution truncation result between the sequence of length L and the coefficients of length M from the previous time, append L-length 0s, add it to the current calculation result using the overlap-add method, and output the result of length L. Then store the remaining intermediate result of length M in the corresponding buffer area for use in the next calculation.
[0025] The above S1 includes: When calculating the FFT, the zero-padding operation is directly achieved by initializing the corresponding buffer area to 0, and subsequent zero-padding operations are directly completed by reading and writing according to the corresponding address blocks.
[0026] The above S2 includes: Append L-length 0s before and after the coefficient h(n), read 2L-length coefficients in an overlapping L-length manner, perform 2L-length FFT, and then put them into the buffer area, without performing the FFT calculation multiple times.
[0027] The above S3 includes: In the internal sub-process, use the overlap-save method to implement the convolution between the sequence of length L and the impulse response of length M.
[0028] The above S4 is specifically as follows:
[0029] S401. Use a dual-region ping-pong operation to cache the calculation result from the previous time.
[0030] S402. Use the method of appending L-length 0s after the data of length M to backfill the calculation result into the reading area to implement the overlap-add method of the sliding window.
[0031] In actual engineering applications, there are two methods to implement segmented convolution, namely the overlap-add method and the overlap-save method.
[0032] Overlap-Add Method
[0033] The idea of the overlap-add method is to divide the long sequence x(n) into sub finite-length sequences of length L. Each sub-sequence x k (n) is non-overlapping and is convolved with the impulse response coefficient h(n) of length M respectively to obtain the convolution result of each segment sequence as y k (n). The length of the convolution calculation result is L + M - 1. Among them, the latter M - 1 is the overlapping part of y k (n) and y k+1 (n). Finally, the overlapping parts are added together to obtain the final result y(n), that is, x(n) * h(n). The schematic diagram of the calculation process is asFigure 1 as shown
[0034] Overlap-save method
[0035] Analyzing from the method, the overlap-add method has non-overlapping inputs for each sub-segment but overlapping outputs. Different from the overlap-add method, the overlap-save method has overlapping inputs for each sub-segment but non-overlapping outputs. The principle is to use circular convolution instead of linear convolution and achieve it through the way of segmented convolution kernel periodic extension;
[0036] The basic steps are as follows: divide the long sequence x(n) into sub-segments with a length of L = N - M + 1, and add a reserved M - 1 data at the front of each segment;
[0037] Specifically: M - 1 zeros are added in front of the initial sub-segment, and zeros with a length of L are added at the end of the last sub-segment; zeros with a length of L are added at the end of the last sub-segment;
[0038] Then perform circular convolution with the impulse response coefficient h(n), and finally discard the first M - 1 points of the calculation result of each sub-segment and merge the adjacent sub-segment results, which is the final result y(n). The schematic diagram of the calculation process is as Figure 2 shown
[0039] Since the convolution in the time domain is equal to the product in the frequency domain, then
[0040] y(n) = x(n) * h(n) = ifft{fft[x(n)] × fft[h(n)]}
[0041] Therefore, the fast Fourier transform FFT can be used to reduce the complexity when calculating the above convolution.
[0042] The present invention combines two methods of segmented convolution. In the external large process, the traditional overlap-add method is used to calculate the linear convolution of the impulse response coefficient h(n) (length M) and x(n). However, when calculating the convolution of each sub-segment x k (n) (length L < M) and h(n), it is no longer directly filled with zeros for calculation. Instead, using the commutative property of convolution, the process of filtering x k (n) sub-segments by h(n) is regarded as the process of filtering h(n) by x k (n) sub-segments in reverse, and this process is implemented by the overlap-save method. In this way, the sub-segment x k (n) can be split very small, and the calculation length of one FFT will not be too long, reducing the use of resources such as multipliers, and there is no need to approximate the lengths of L and M anymore.
[0043] Embodiment 1
[0044] As Figure 3 shown
[0045] Initialize the DDR area to 0;
[0046] Receive the original coefficients, with a length of M = 131072 floating-point data, and a sliding window L = 8192. That is, h(n) reads out a sequence with a length of 16384 and an overlap of 8192 according to the overlap-save method (the front end of the first sequence is padded with 8192 0s, and the end of the last sequence is padded with 8192 0s), and enters the FFT IP core to perform a 16384-point Fourier transform. A total of M / L + 1 = 17 calculations are required, and the results are saved in the ddr and wait for the original data to perform complex multiplication;
[0047] Receive the original data, with a length of L = 8192 floating-point data. That is, x(n) performs a 16384-point Fourier transform in the FFT IP core after padding 8192 0s according to the overlap-add method. The results are respectively multiplied by the FFT results of the original coefficients 17 times with floating-point complex multiplication, and then flow into the FFT IP core to perform the inverse Fourier transform. According to the principle of the overlap-save method, the first 8192 points of each time are discarded, and the subsequent 8192 points of data are retained.
[0048] The above 17 8192-point data overlap 131072 points according to the principle of the overlap-add method. Therefore, first read the previously stored 131072 overlapping points (the initial data are all 0), add the initialized 8192 0 data (to implement the sliding window function), a total of 139264 points, and add them to these 17 8192 data in sequence. And take the output of the first 8192 points at the front end as the calculation result of a large process. The remaining 131072 data, that is, the operation results from the 2nd to 17th times, are stored in the DDR again and wait for the next data to arrive for calculation again.
[0049] The above-described embodiments only express the implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be based on the appended claims.
Claims
1. A frequency domain implementation method of a low-resource high-order FIR filter in FPGA, characterized in that: The following steps are involved: S1, initialize the value of some cache areas to 0, so as to facilitate subsequent filling of 0; S2, receiving impulse response coefficients, the total length is M, and reading 2L length coefficients according to the overlapping length L and storing them in the response buffer after performing fast Fourier transform FFT; S3, receiving the data of the original sequence of length L and performing FFT and pre-stored FFT coefficients, and calculating the convolution between the sequence of length L and the coefficients of length M by using the overlap-and-reserve method; S4. After reading the convolution truncation result between the last L-length sequence and the M-length coefficient and filling the L-length 0, add it to the current calculation result by overlapping and adding method, and then output the L-length result, and store the remaining M-length intermediate result into the corresponding cache area for the convenience of the next calculation application.
2. The frequency domain implementation method of a low-resource high-order FIR filter of an FPGA according to claim 1, characterized in that: The S1 includes: when calculating FFT, the zero-filling operation is directly completed by initializing the corresponding cache area bit 0, and then directly completing the zero-filling operation according to the corresponding address block reading and writing.
3. The frequency domain implementation method of a low-resource high-order FIR filter of an FPGA according to claim 2, characterized in that: The S2 includes: after padding the coefficient h(n) with L length 0, reading the 2L length coefficient according to the overlapping L length, and after completing the 2L length FFT, putting it into the buffer area without performing FFT calculation multiple times.
4. The frequency domain implementation method of a low-resource high-order FIR filter of an FPGA according to claim 3, characterized in that: The S3 includes: using the overlap-preservation method in the internal small process to realize the convolution of the L-length sequence and the M-length impulse response.
5. The frequency domain implementation method of a low-resource high-order FIR filter of an FPGA according to claim 4, characterized in that: The S4 is specifically: S401, using a dual-region ping-pong operation to cache the last calculation result; S402, using the method of supplementing L length 0 after M length data, the calculation result is backfilled into the reading area to implement the sliding window overlap-add method.