FPGA (field programmable gate array) implementation method of distributed parallel FIR (finite impulse response) filter

Through the FPGA implementation method of distributed parallel FIR filter, the multi-channel parallel FIR filter design and split anti-symmetric lookup table structure is used to solve the contradiction between high throughput and low resource consumption of FPGA filters, real-time processing and resource optimization of high-speed signals are realized.

CN120342386APending Publication Date: 2025-07-18EAST CHINA UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408564.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing FPGA-based FIR filters are difficult to balance between high throughput and low hardware resource consumption. The traditional serial structure cannot meet the needs of high-speed signal processing, and multiplication operations lead to prominent resource consumption and delay problems.

Method used

The FPGA implementation method adopts a distributed parallel FIR filter, and the design of multiple parallel FIR filters, split anti-symmetric lookup table structure and time-sharing multiplexing technology reduce multiplication operations, reduce hardware resource consumption, and improve processing speed.

Benefits of technology

It realizes improving filter throughput while keeping the sampling rate unchanged, reducing the internal operating frequency of FPGA, improving processing speed and reducing hardware resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005341920870000031
    Figure BDA0005341920870000031
  • Figure BDA0005341920870000034
    Figure BDA0005341920870000034
  • Figure BDA0005341920870000035
    Figure BDA0005341920870000035
Patent Text Reader

Abstract

The invention discloses an FPGA (Field Programmable Gate Array) implementation method of a distributed parallel FIR (Finite Impulse Response) filter, an FPGA is a main hardware platform for implementing the FIR filter, a parallel technology needs to be introduced to increase the throughput rate of the filter along with the improvement of digital signal processing requirements, but the consumption of hardware resources is also greatly increased along with the increase of the parallelism degree, so that the complexity of the FIR filter is greatly reduced. According to the method, the multi-path parallel FIR filter is realized based on the FPGA, a distributed algorithm is improved, a split antisymmetric lookup table structure is designed according to the coefficient relationship between the sub-filters, the processing speed of the filter is improved, the hardware resource consumption in the FPGA is greatly reduced, and the method can be widely applied to high-speed digital signal processing scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital signal processing, and in particular, to an FPGA implementation method for a distributed parallel FIR filter. Background Art

[0002] Digital signal processing systems, with their advantages of flexibility, high precision, strong stability, and large-scale integration, have been widely used in fields such as speech processing, computer vision, industrial control, and aerospace. Among them, the finite impulse response (FIR) filter occupies an important position in key digital signal processing links such as anti-aliasing filtering and signal reconstruction due to its inherent stability and precisely designed linear phase characteristics. As a reconfigurable hardware platform, the field programmable gate array (FPGA), with its parallel architecture advantages, configurable logic units, and high-throughput data processing capabilities, has become the preferred hardware carrier for implementing FIR filters, especially outstanding in fields with strict real-time requirements such as 5G communication and radar signal processing.

[0003] However, with the development of signal processing, the sampling rate of signals is getting higher and higher, and the requirements for the digital domain FIR filtering rate are also getting stricter. The sampling rate of signals in the analog domain (usually reaching the GHz level) is much higher than the maximum operating clock frequency of the hardware platform (the typical value is several hundred MHz). Traditional serial-structured filters cannot achieve high-speed signal processing. Therefore, parallel technology needs to be introduced to increase the throughput of the filter, but at the same time, the consumption of hardware resources also increases significantly with the increase in parallelism, limiting the practical application of parallel FIR filters.

[0004] In an FPGA, addition is a weak operation, which is simple and easy to implement and consumes few resources. Multiplication is a strong operation. Each time multiplication is performed, the dedicated DSP Slice resources in the FPGA need to be called. A large number of multipliers will consume a large amount of resources in the FPGA. Moreover, in a sequential circuit, the multiplier will also introduce corresponding delays, and the cascading of multiple multipliers will significantly extend the critical path delay, thereby restricting the maximum operating frequency of the filter. Therefore, it is very important to implement the FIR filter through a better method, reduce the number of multiply-accumulate operations in the filter structure, reduce its hardware resource consumption, and improve the processing speed. Summary of the Invention

[0005] The present invention proposes an FPGA implementation method for a distributed parallel FIR filter. Without changing the sampling rate, the throughput rate of the filter is increased and the internal working frequency of the FPGA is reduced through parallelization technology. The FPGA performs real-time filtering on the input high-speed signal stream, and by improving the distributed algorithm and designing a split antisymmetric lookup table structure, the processing speed of the filter is increased and the hardware resource consumption is reduced, solving the problem that the existing FPGA-based FIR filter cannot simultaneously meet the requirements of high throughput rate and low hardware resource consumption.

[0006] To achieve the above object, the method proposed by the present invention includes the following steps:

[0007] 1. According to the system requirements, design parameters such as filter coefficients and orders, construct a multi-channel parallel FIR filter based on the input sequence and the polyphase decomposition structure of the filter, and generate corresponding sub-filters;

[0008] 2. Construct corresponding sub-filter lookup table modules and antisymmetric modules in the FPGA according to the sub-filter coefficients, and perform multi-level splitting on each lookup table;

[0009] 3. The FPGA receives the high-speed digital signal x[n] output by the ADC and the accompanying clock CLK;

[0010] 4. The FPGA performs serial-to-parallel conversion and clock down-conversion on the received data stream, converts the serial data stream {x[0], x[1], x[2], …, x[k]} into L-channel parallel data streams {x[0], x[1], …, x[L-1]}, {x[L], x[L+1], …, x[2L-1]}, …, {x[Ln], x[Ln+1], x[Ln+2], …, x[Ln+L-1]}, and obtains the down-converted slow clock CLK_DIV;

[0011] 5. Store the multi-channel parallel data in the corresponding registers according to the number of channels respectively, perform multi-level delay on each channel of input data in the slow clock domain, and latch each level of delayed data to obtain the input sequences {x[n], x[n-1], x[n-2], …} of multiple sub-filters;

[0012] 6. Send the sub-filter input sequence into the data bit recombination module. According to the distributed algorithm, after decomposing the input data by bit, form the input address vector of the lookup table by splicing the same-bit data of different sampling points;

[0013] 7. Input the address vector into the data splitting module and group the reconstructed address vectors. Part of this address vector is input to the splitting lookup table in ascending order, while the other part is input to the splitting lookup table after its order is adjusted by the reverse order module. The two are time-division multiplexed for the same lookup table through the multiplexing module and are alternately input to the lookup table for lookup under the control of the selection signal;

[0014] 8. After each path of data is respectively input to its corresponding split sub-table, its output is obtained through the table lookup operation. After the corresponding second-power weighted summation, the partial filtering result of the sub-table can be obtained. Then, the output results of all sub-tables are summed to obtain the filtering result of the input data through the sub-filter;

[0015] 9. Input the obtained filtering result into the output selection module to distinguish the results of each sub-filter. Finally, input the distinguished results into the filtering operation module to perform corresponding delay and addition operations based on the multi-path parallel results, and finally the multi-path parallel output {y[Ln], y[Ln + 1], y[Ln + 2], …, y[Ln + L - 1]} of the multi-path parallel filter can be obtained. Description of the Drawings

[0016] Figure 1 It is a flowchart of the method embodiment of the present invention.

[0017] Figure 2 It is a structure diagram of a two-path parallel filter.

[0018] Figure 3 It is a structure diagram of the improved filter of the present invention.

[0019] Figure 4 It is a structure diagram of the splitting lookup table. Specific Embodiment

[0020] The following combines the drawings and embodiments to describe in detail the FPGA implementation method of the distributed parallel FIR filter of the present invention.

[0021] For easy understanding, when introducing the specific embodiment, it is introduced through the specific embodiment of a two-path parallel sixth-order FIR filter. The flowchart of the method embodiment of the present invention is as Figure 1 shown. When designing the filter, first calculate the coefficients of the filter according to the system requirements. This calculation can be designed through FDATool in MATLAB. The finally obtained filter tap coefficients are respectively {h[0], h[1], h[2], h[3], h[4], h[5]}. Then, the input sequence and the filter can be decomposed into polyphase components, and a multi-path parallel FIR filter can be constructed based on the corresponding filtering calculations, as Figure 2As shown, and corresponding sub - filters H0 and H1 are generated, where the tap coefficients of H0 are {h[0], h[2], h[4]}, and the tap coefficients of H1 are {h[1], h[3], h[5]}.

[0022] When implementing multiple sub - filters through FPGA, the present invention adopts a distributed structure for design. This structure decomposes the input data bit - by - bit through a distributed algorithm, changes the summation order, and converts the multiplication - accumulation result of the input sequence x[n] and the tap coefficient h[n] into the following form:

[0023]

[0024] In the formula, x b [n] is the b - th bit of x[n], denoted as Then

[0025] This algorithm transforms the original multiplication - accumulation operation into a look - up table mapping and accumulation operation. By constructing a corresponding look - up table in the FPGA in advance, all possible values of f(h[n], x b [n]) are pre - stored in the look - up table. Since the filter coefficient h[n] is usually a fixed constant and has been determined during design. Therefore, after decomposing the input data bit - by - bit and recombining it into {x b [0], x b [1], x b [2], …, x b [N - 1]}, using it as the input address of the look - up table and inputting it into the look - up table. After obtaining the corresponding output, through the corresponding quadratic - power weighting and accumulation, the final output y is obtained, which is the final result of one - time filtering. This method replaces the multiplication - addition operation with a look - up table operation, effectively avoiding the use of multipliers. It is not only very simple to implement but also can effectively improve the system speed, and the corresponding result can be found within one clock cycle.

[0026] When constructing the corresponding sub - filter look - up table module in the FPGA according to the sub - filter coefficients, since the depth of the look - up table in the distributed structure grows exponentially with the filter order, for an FIR filter with a length of N, the depth of its look - up table is 2 N . When the filter order is relatively large, the resources occupied by the look - up table will be extremely large, which will make it difficult for the hardware resources to meet the requirements. Moreover, the multi - path parallel filter contains multiple sub - filters, and each sub - filter needs to construct a look - up table, which will further expand its scale by a corresponding multiple.

[0027] To solve this problem, when constructing the filter lookup table in the present invention, according to the coefficient law of the parallel filter, a split anti-symmetric lookup table is designed to reduce the number of lookup tables and the scale of a single lookup table, greatly reducing the resource consumption for implementing a multi-channel parallel filter through an FPGA. In this embodiment, the tap coefficients of the designed filter are respectively {h[0], h[1], h[2], h[3], h[4], h[5]}. In the parallel structure, this filter is composed of two H0 sub-filters and two H1 sub-filters in total. Among them, the tap coefficients in H0 are the odd terms of the original filter coefficients {h[0], h[2], h[4]}, and the coefficients in H1 are the even terms of the original filter coefficients {h[1], h[3], h[5]}.

[0028] Based on this, the lookup tables of two filters can be constructed. The following table is the 3rd-order lookup table of filter H0: Table 1

[0029] The following table is the 3rd-order lookup table of filter H1: Table 2

[0030] Since for an Nth-order linear-phase FIR filter, its filter coefficients have a certain symmetry, the following formula can be obtained:

[0031] h[0] = h[5], h[1] = h[4], h[2] = h[3] (2)

[0032] By examining the lookup tables of the two combined filters, it can be found that in H0, the data 0 at address 000 is equal to the data 0 at address 0 in H1, the data h[0] at address 001 is equal to the data h[5] at address 100 in H1, the data h[2] at address 010 is equal to the data h[3] at address 010 in H1, the data h[0]+h[2] at address 011 is equal to the data h[5]+h[3] at address 110 in H1, the data h[4] at address 100 is equal to the data h[1] at address 001 in H1... According to this rule, it can be concluded that the data at address a2a1a0 in H0 is the same as the data at address a0a1a2 in H1. That is to say, in H0 and H1, for address units that are in reverse order, the stored data is the same. For an input data, if one wants to obtain the output result of passing it through the H1 filter, it is actually not necessary to know the structure of the lookup table of the H1 filter. One only needs to reverse the input data and then input it into the H0 lookup table. The obtained output result is the same as the result of passing it through H1 in the normal order. Therefore, the reverse module can replace the H1 filter. Based on this, when constructing the lookup table, it is not necessary to construct the lookup table of the sub-filter H1. One only needs to reverse the input to obtain the reversed data, and then time-division multiplex the normal-order data and the reversed data for the lookup table of the sub-filter H0. After corresponding table lookup operations, the filtering results of the input data passing through H0 and H1 can be obtained respectively. The structure is as Figure 3 shown. By this method, the total size of the lookup table can be reduced to half of the original, and there is no need to construct the lookup table of H1 separately. This method can also be extended to the design of L-channel parallel filters. By this method, half of the sub-filters can be constructed less, and the total size of the lookup table can be reduced from L 2 ×2 N / L to 1 / 2×L 2 ×2 N / L .

[0033] On the basis of reducing the number of lookup tables, for high-order filters, the present invention also reduces the size of the lookup table of a single sub-filter in a parallel filter by splitting the lookup table. For an FIR filter with a length of N, its coefficients can be divided into K groups, with each group having a coefficient length of L, where N = KL. Then, K groups of coefficients are used to construct a lookup table respectively, and the depth of each lookup table is 2 L . According to the following formula, it can be known that:

[0034]

[0035] A multiplication-accumulation operation can be decomposed into the sum of multiple multiplication-accumulation operations. Therefore, the output of a high-order filter can be divided into the sum of multiple low-order filter outputs. Therefore, when performing a lookup table operation, K depths of 2 can be used. L The sum of the outputs of the lookup table replaces the original one with a depth of 2 N The output of the lookup table is as follows: Figure 4 Taking a 16-order filter as an example, its coefficients can be split into 4 groups and then implemented through the above structure. The total depth of the lookup table is reduced from 2 to 3 at the cost of adding only 3 adders. 16 =65536 is reduced to 4×2 4 =64, the size of the lookup table is greatly reduced.

[0036] After the construction of the lookup table module is completed by the above method, the FPGA can receive the high-speed digital signal x[n] and the accompanying clock CLK output by the ADC, and perform serial-to-parallel conversion and clock down-conversion on the received data stream, converting the serial data stream {x[0], x[1], x[2], …, x[k]} into L parallel data streams {x[0], x[1], …, x[L-1]}, {x[L], x[L+1], …, x[2L-1]}, …, {x[Ln], x[Ln+1], x[Ln+2], …, x[Ln+L-1]}, and obtain the slow clock CLK_DIV after down-conversion.

[0037] In this embodiment, a two-way parallel structure is used, so the input parallel data is first divided into two ways of data and stored in corresponding registers, xin_reg_0 and xin_reg_1. Since the filtering operation is the convolution of the input data sequence and the filter coefficient, each input data is then delayed at multiple levels, and each level of delayed data is latched to obtain the input sequence {x[n], x[n-1], x[n-2], ...}.

[0038] Because the design adopts a distributed structure, the data should be sent to the data bit reorganization module. After the input data is decomposed by bit, the input address vector of the lookup table is formed by splicing the same bit data of different sampling points. One part of the address vector is input to the lookup table in positive order, while the other part is input to the lookup table after adjusting the order through the reverse module. The two are time-division multiplexed on the same lookup table through the multiplexing module, and are input into the lookup table for search in turn according to the control of the selection signal.

[0039] The data of each group are respectively input into their corresponding split sub-tables, and their outputs are obtained through look-up table operations. After weighted summation of the corresponding second power, the partial filtering results of the sub-tables can be obtained. Then, the output results of all sub-tables are summed to obtain the filtering results of the input data through the sub-filter. Due to time-division multiplexing, an output selection module is also designed to distinguish the results passing through H0 from the results passing through H1 to obtain X0H0 and X0H1. According to the same idea, X1H0 and X1H1 can also be obtained. Finally, they are output to the filtering operation module, and through the corresponding delay and addition operations, the two-way parallel outputs Y1 and Y2 of the parallel filter can be obtained.

[0040] The above embodiments are only used to explain the specific implementation of the present invention, rather than limiting the present invention. Therefore, all changes made according to the shape and principle of the present invention should be covered within the scope of the present invention.

[0041] In summary, the present invention aims at high-speed data streams, designs a multi-way parallel filter based on the polyphase decomposition structure of the FIR filter, converts the filtering system into a multi-input multi-output system through parallel technology, improves the throughput rate of the filter, and realizes the real-time processing of high-speed signals. For the additional consumption of hardware resources, the present invention introduces a distributed algorithm to design multiple sub-filters, converts the multiply-accumulate structure into a look-up table structure, and avoids the use of multiplication resources. To address the problem of the large scale of the look-up table in the distributed algorithm, an optimized method of designing split anti-symmetry is adopted to reduce the number of look-up tables and the scale of a single look-up table without affecting the speed, significantly reducing its hardware resource consumption, and finally realizing low-resource and high-speed digital signal processing based on FPGA.

Claims

1. An FPGA implementation method for a distributed parallel FIR filter, characterized in that the FPGA implementation method comprises the following steps:

1. Design parameters such as filter coefficients and order according to system requirements. Construct a multi-channel parallel FIR filter based on the input sequence and the polyphase decomposition structure of the filter, and generate corresponding sub-filters.

2. Construct the corresponding sub-filter lookup table module and anti-symmetric module in the FPGA according to the sub-filter coefficients, and split each lookup table into multiple levels.

3. The FPGA receives the high-speed digital signal x[n] output by the ADC and the accompanying clock CLK.

4. The FPGA performs serial-to-parallel conversion and clock downscaling on the received data stream, converting the serial data stream {x[0], x[1], x[2], …, x[k]} into L-channel parallel data streams {x[0], x[1], …, x[L-1]}, {x[L], x[L+1], …, x[2L-1]}, …, {x[Ln], x[Ln+1], x[Ln+2], …, x[Ln+L-1]}, and obtaining the downscaled slow clock CLK_DIV.

5. Store the multi-channel parallel data in the corresponding registers according to the number of channels respectively, perform multi-level delay on each channel of input data in the slow clock domain, and latch each level of delayed data to obtain the input sequence {x[n], x[n-1], x[n-2], …} of the multi-channel sub-filters.

6. Send the sub-filter input sequence into the data bit recombination module. According to the distributed algorithm, after decomposing the input data by bit, form the input address vector of the lookup table by splicing the same-bit data of different sampling points.

7. Input the address vector into the data splitting module to group the recombined address vector. Part of this address vector is input to the split lookup table in positive order, while the other part is input to the split lookup table after adjusting the order through the reverse order module. The two are time-division multiplexed for the same lookup table through the multiplexing module, and are alternately input to the lookup table for lookup under the control of the selection signal.

8. After each channel of data is respectively input to its corresponding split sub-table, obtain its output through table lookup operation. After corresponding quadratic power weighted summation, the partial filtering result of the sub-table can be obtained. Then sum the output results of all sub-tables to obtain the filtering result of the input data through the sub-filter.

9. Input the obtained filtering result into the output selection module to distinguish the results of each sub-filter. Finally, input the distinguished results into the filtering operation module, perform corresponding delay and addition operations according to the multi-channel parallel results, and finally obtain the multi-channel parallel output {y[Ln], y[Ln+1], y[Ln+2], …, y[Ln+L-1]} of the multi-channel parallel filter.