A method for implementing a multiplication-optimized low-pass FIR filter
By optimizing the bit width and structural optimization of the FIR filter multiplication circuit, the problem of large area of the multiplication circuit is solved, and the area of the multiplication circuit is reduced, and the optimization effect is enhanced with the increase of the bit width of the filter coefficient.
Patent Information
- Application Number
- CN202210497553.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-05-09
AI Technical Summary
During the implementation of the FIR filter, as the filter order increases, the consumption of the multiplier also increases, resulting in a larger area of the multiplication circuit.
By optimizing the bit width and structural optimization of input parameters for the multiplication circuit of the traditional direct I type FIR filter, a small number of registers are added to reduce the area of the multiplication circuit.
The effect of reducing the circuit area of the multiplication method is achieved, and the larger the bit width of the filter coefficient, the more obvious the optimization effect is. Verified by the FPGA embodiment, although DFF resource consumption is increased by about 4.45%, the LUT resource consumption is reduced by 13.12%.
Smart Images

Figure CN114744982B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of digital filter design and relates to a multiplication optimized low-pass FIR filter implementation method. Background Art
[0002] In the field of digital signal processing, FIR filters are widely used in speech and image processing, spectrum analysis, radar signal processing, and pattern recognition. Compared with analog filters, digital filters have the advantages of high precision, good stability, small size and flexible configuration. In addition, they can effectively avoid problems such as temperature drift and noise. However, in the implementation process of FIR filters, as the order of the filter increases, the consumption of multipliers also increases. In reality, in order to obtain FIR filters with narrow transition band and superior performance, high-order filters are often used, which leads to large consumption of multiplication resources and large multiplication circuit area. In order to effectively solve the above problems, the present invention proposes a multiplication optimized low-pass FIR filter implementation method, which takes the traditional direct I-type FIR filter as the basis, optimizes the bit width and structure of the input parameters for the multiplication circuit, and only needs to add a small number of registers to greatly reduce the area of the multiplication circuit, and the larger the bit width of the filter coefficient, the more obvious the optimization effect. Summary of the invention
[0003] In view of the problem that the multiplication circuit area is large during the implementation of the traditional direct type I FIR filter, the present invention proposes a method for implementing a multiplication-optimized low-pass FIR filter. The method optimizes the bit width and structure of the input parameters for the multiplication circuit. Only a small number of registers need to be added to greatly reduce the area of the multiplication circuit, and the larger the bit width of the filter coefficient, the more obvious the optimization effect. The technical solution of the present invention is a method for implementing a multiplication-optimized low-pass FIR filter. The multiplication-optimized low-pass FIR filter implemented by the method includes N-1 general computing units with the same circuit structure cascaded in the form of a pipeline, an initial computing unit, and a summing circuit, wherein N-1 is the filter order;
[0004] The initial calculation unit includes a register group REG0, a register group REG1 and a multiplier, wherein the register group REG0 performs a delay register on the current input signal, and the multiplier performs a product operation on the current input signal and the first filter coefficient, and stores the operation result in the register group REG1;
[0005] The general calculation unit includes a register group REG00, a register group REG11, an adder group and a general adder, wherein the register group REG00 delays the input signal from the previous stage, and the adder group cooperates with the shift operation and the general adder to calculate formula (1):
[0006] res i=(a i -a i-1 )*x(ni)+res i-1 , i∈[1,N-1] and i∈Z (1)
[0007] res0 is the output data from the register group REG1 in the initial calculation unit, a i is the pre-generated i-th filter coefficient, x(n) is the input signal at time n, and it should be noted that res0 is calculated by formula (2):
[0008] res0=a0*x(n) (2)
[0009] Register group REG11 to res i Perform data storage; the universal adder sums the output result of the adder group in the universal computing unit of this level and the value of the register group REG1 of the initial computing unit or the register group REG11 of the previous universal computing unit;
[0010] The summing circuit performs a summing operation on the values stored in the register group REG1 and the register group REG11 in N computing units, where the N computing units include an initial computing unit and N-1 general computing units;
[0011] The adder group in the general computing unit adds two adjacent original filter coefficients a i and a i-1 Perform a signed subtraction operation to obtain the coefficient difference a of the low data bit width i -a i-1 , and use the adder group, combined with the shift operation, to obtain (a i -a i-1 )*x(ni);
[0012] The general computing units have the same circuit structure, and the HDL code corresponding to the hardware logic circuit of the general computing unit is automatically generated by using a scripting language.
[0013] Preferably, the method of automatically generating the HDL code corresponding to the hardware logic circuit of the general computing unit by using the scripting language specifically includes the following steps:
[0014] S10, using MATLAB to pre-generate filter coefficients a0~a N-1 , and quantify it;
[0015] S20, the script language reads the filter coefficient sequence generated in S10 and performs a signed number subtraction operation a i -a i-1, where i∈[1, N-1] and i∈Z, and determine the maximum data bit width of the signed difference sequence and set it to DW;
[0016] S30, combined with a0 in S10 and the signed number a in S20 i -a i-1 As well as the data bit width DW, the script language is used to automatically generate the HDL code corresponding to the hardware logic circuit of the general computing unit. The specific function is: using the addition circuit in conjunction with the shift operation to calculate the expression (1).
[0017] Preferably, the filter in S10 is a low-pass FIR filter with an order of N-1 and a length of N, and its system differential equation is as follows:
[0018]
[0019] Let S1=a(0)x(n)+[a(1)-a(0)]x(n-1)+…+[a(N-2)-a(N-3)]x(n-N+2)+[a(N-1)-a(N-2)]x(n-N+1);
[0020] Let S2=a(0)x(n-N+1)+…+a(N-3)x(x-N+2)+a(N-2)x(n-N+1);
[0021] Therefore, the system difference equation can be simplified to formula (3):
[0022] y(n)=S1+S2 (3).
[0023] Preferably, the expression S1 is completed by the register group REG0, the register group REG00, the multiplier and the adder group in N computing units; the N computing units include an initial computing unit and N-1 general computing units.
[0024] Preferably, the expression S2 is performed by register group REG1 and register group REG11 in N computing units to shift and register the calculation results of the previous stage computing units, and cooperate with a general adder to complete the operation; the N computing units include an initial computing unit and N-1 general computing units.
[0025] The beneficial effects of the present invention are as follows:
[0026] The present invention is a method for realizing a multiplication-optimized low-pass FIR filter. Based on the traditional direct I-type FIR filter, the bit width and structure of the input parameters of the multiplication circuit are optimized. Only a small number of registers need to be added to greatly reduce the area of the multiplication circuit. The larger the bit width of the filter coefficient, the more obvious the optimization effect. The FPGA implementation example of a 15-order low-pass FIR filter (the bit width of the filter coefficient and the input signal data are both 12 bits) verifies that although the present invention increases the DFF resource consumption by about 4.45%, it reduces the LUT resource consumption by 13.12%, achieving the purpose of area optimization through multiplication optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A schematic diagram of the structure after implementation of the multiplication-optimized low-pass FIR filter implementation method;
[0028] Figure 2 A flowchart of the steps for automatically generating HDL codes for general computing units using a scripting language; DETAILED DESCRIPTION
[0029] The technical solution of the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings.
[0030] Example 1
[0031] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0032] On the contrary, the present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention as defined by the claims. Further, in order to make the public have a better understanding of the present invention, some specific details are described in detail in the detailed description of the present invention below. Those skilled in the art can fully understand the present invention without the description of these details.
[0033] First, it is necessary to understand that the multiplication optimized low-pass FIR filter implementation method described in the present invention is used to solve the problem of large multiplication circuit area in the traditional direct I-type FIR filter implementation process. By adopting this implementation method, only a small number of registers need to be added to greatly reduce the area of the multiplication circuit, and the larger the filter coefficient bit width, the more obvious the optimization effect. The FPGA embodiment of the 15-order low-pass FIR filter (the filter coefficient and the input signal data bit width are both 12 bits) verifies that although the present invention increases the DFF resource consumption by about 4.45%, it reduces the LUT resource consumption by 13.12%, achieving the purpose of area optimization through multiplication optimization.
[0034] See also Figure 1 , is a schematic diagram of a method for implementing a multiplication-optimized low-pass FIR filter according to an embodiment of the present invention, considering a passband cutoff frequency f pass =1MHz, the stopband start frequency is f stop =1.5MHz, order 15, length 16, low-pass FIR filter with sampling frequency of 50MHz, including the following three parts:
[0035] Fifteen general computing units 20, an initial computing unit 10 and a summing circuit 30 with similar circuit structures are cascaded in a pipeline form.
[0036] The initial calculation unit 10 has an internal structure of two register groups (REG0 and REG1) and a multiplier 11, wherein the register group REG0 is used to delay the current input signal x(n), and the multiplier 11 is used to perform a product operation on the current input signal x(n) and the first filter coefficient a0, that is, a0*x(n), and store the operation result in the register group REG1;
[0037] The internal structure of the universal computing unit 20 includes two register groups (REG00 and REG11), an adder group 21 and a universal adder 22, wherein the register group REG00 of the i-th universal computing unit 20 is used to delay the output signal of the register group REG00 from the i-1-th universal computing unit 20, and the adder group 21 cooperates with the shift operation and the universal adder 22 to calculate the formula (1):
[0038] res i =(a i -a i-1 )*x(ni)+res i-1 , i∈[1,15] and i∈Z (1)
[0039] res0 is the output data from the register group REG1 in the initial calculation unit 10, a iis the pre-generated i-th filter coefficient, x(n) is the input signal at time n, and it should be noted that res0 can be calculated by formula (2):
[0040] res0=a0*x(n) (2)
[0041] Register group REG11 to res i general adder 22 sums the output result of adder group 21 in the general computing unit of the current level and the value of register group REG1 of the initial computing unit 10 or the register group REG11 of the previous general computing unit 20;
[0042] The summing circuit 30 performs a summing operation on the values stored in the register group REG1 and the register group REG11 in the 16 computing units (including one initial computing unit 10 and 15 general computing units 20);
[0043] The general computing unit 20 has the same circuit structure, and the HDL code corresponding to the hardware logic circuit of the general computing unit 20 is automatically generated by using a scripting language, which specifically includes the following steps:
[0044] S10, using MATLAB to pre-generate the low-pass FIR filter coefficients, and perform a 12-bit quantization operation on them to obtain quantized filter coefficients a0~a 15 =[79, 85, 91, 96, 99, 102, 104, 105, 105...];
[0045] S20, the script language reads the filter coefficient sequence generated in S10 and performs a signed number subtraction operation a i -a i-1 , where i∈[1,15] and i∈Z, and determine the maximum data bit width DW=4 of the signed difference sequence;
[0046] S30, combined with a0 in S10 and the signed number a in S20 i -a i-1 and data bit width DW, and automatically generates HDL code corresponding to the hardware logic circuit of the general computing unit 20 using a scripting language. The specific functions of the code are: using an addition circuit in conjunction with a shift operation to calculate expression (1);
[0047] For this embodiment, the system difference equation is as follows:
[0048]
[0049] Let S1=79*x(n)+[85-79]*x(n-1)+…+[85-91]*x(n-N+2)+[79-85]*x(n-N+1);
[0050] Let S2=79*x(n-N+1)+…+91*x(x-N+2)+85*x(n-N+1);
[0051] Therefore, the system difference equation can be simplified to formula (3):
[0052] y(n)=S1+S2 (3)
[0053] Expression S1 is completed by the register group REG0, register group REG00, multiplier 11 and adder group 21 in 16 computing units (including 15 general computing units 20 and one initial computing unit 10); Expression S2 is completed by the register group REG1 and register group REG11 in the 16 computing units to shift and register the calculation results of the previous stage computing units, and cooperate with the general adder 22 to complete the operation.
[0054] The FPGA embodiment verifies that although the method of the present invention increases the DFF resource consumption by about 4.45%, it reduces the LUT resource consumption by 13.12%, thus achieving the purpose of area optimization through multiplication optimization.
[0055] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for implementing a multiplication-optimized low-pass FIR filter, characterized in that: The multiplication-optimized low-pass FIR filter implemented by the method includes N-1 general computing units with the same circuit structure, an initial computing unit and a summing circuit cascaded in a pipeline form, wherein N-1 is the filter order; The initial calculation unit includes a register group REG0, a register group REG1 and a multiplier, wherein the register group REG0 performs a delay register on the current input signal, and the multiplier performs a product operation on the current input signal and the first filter coefficient, and stores the operation result in the register group REG1; The general calculation unit includes a register group REG00, a register group REG11, an adder group and a general adder, wherein the register group REG00 delays the input signal from the previous stage, and the adder group cooperates with the shift operation and the general adder to calculate formula (1): res i = (a i - a i-1 ) * x(n - i) + res i-1 , i ∈ [1, N - 1] and i ∈ Z (1) res0 is the output data from the register group REG1 in the initial calculation unit, a i is the pre-generated i-th filter coefficient, x(n) is the input signal at time n, and it should be noted that res0 is calculated by formula (2): res0=a0*x(n) (2) Register group REG11 to res i Perform data storage; the universal adder sums the output result of the adder group in the universal computing unit of this level and the value of the register group REG1 of the initial computing unit or the register group REG11 of the previous universal computing unit; The summing circuit performs a summing operation on the values stored in the register group REG1 and the register group REG11 in N computing units, where the N computing units include an initial computing unit and N-1 general computing units; The adder group in the general computing unit adds two adjacent original filter coefficients a i and a i-1 Perform a signed subtraction operation to obtain the coefficient difference a of the low data bit width i -a i-1 , and use the adder group, combined with the shift operation, to obtain (a i -a i-1 )*x(ni); The universal computing units have the same circuit structure, and the HDL code corresponding to the hardware logic circuit of the universal computing unit is automatically generated by using a scripting language; The method of automatically generating the HDL code corresponding to the hardware logic circuit of the general computing unit by using the scripting language specifically includes the following steps: S10, using MATLAB to pre-generate filter coefficients a0~a N-1 , and quantify it; S20, the script language reads the filter coefficient sequence generated in S10 and performs a signed number subtraction operation a i -a i-1 , where i∈[1,N-1] and i∈Z, and determine the maximum data bit width of the signed difference sequence and set it to DW; S30, combined with a0 in S10 and the signed number a in S20 i -a i-1 and data bit width DW, and use the script language to automatically generate the HDL code corresponding to the hardware logic circuit of the general computing unit. The specific functions are: using the addition circuit and the shift operation to calculate the expression (1); The filter in S10 is a low-pass FIR filter with an order of N-1 and a length of N. Its system differential equation is as follows: Let S1=a(0)x(n)+[a(1)-a(0)]x(n-1)+…+[a(N-2)-a(N-3)]x(n-N+2)+[a(N-1)-a(N-2)]x(n-N+1); Let S2=a(0)x(n-N+1)+…+a(N-3)x(x-N+2)+a(N-2)x(n-N+1); Therefore, the system difference equation can be simplified to formula (3): y(n)=S1+S2 (3).
2. A method for implementing a multiplication-optimized low-pass FIR filter according to claim 1, characterized in that: The expression S1 is completed by the register group REG0, the register group REG00, the multiplier and the adder group in N computing units; the N computing units include an initial computing unit and N-1 general computing units.
3. A method for implementing a multiplication-optimized low-pass FIR filter according to claim 1, characterized in that: The expression S2 is that the register groups REG1 and REG11 in the N computing units shift and register the calculation results of the previous computing units, and cooperate with the general adder to complete the operation; the N computing units include an initial computing unit and N-1 general computing units.
Citation Information
Patent Citations
Method and device for implementing bandpass filtering using cascade integration comb filter
CN101136622A
High speed FIR filter realizing device
CN101162895A