Parallel interpolation filtering device
By fusing a parallel interpolation filtering device with double interpolation and semi-band filtering in an FPGA or digital IC, the problems of low working clocks and complex logic in the prior art are solved, and the interpolation filtering effect of high throughput and high clock frequency is achieved.
Patent Information
- Application Number
- CN202510013694.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-06
AI Technical Summary
When implementing high throughput interpolation filtering, the prior art is limited by the low operating clocks of FPGAs and digital ICs, resulting in excessive use of registers in parallel processing and complex combination logic, which affects the clock frequency.
A parallel interpolation filtering device is proposed, which uses concise control logic and fewer registers to achieve high throughput interpolation filtering. The device includes a multiplication unit, an addition unit and a register for buffering the intermediate results of the calculation and providing timing control through the clock module.
This structure can realize large throughput interpolation filtering of data at low frequency, reducing processing delays, reducing register usage, increasing the operating clock frequency of the circuit, and simplifying the control logic.
Smart Images

Figure CN119945382A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the fields of communication, chip and digital signal processing, and in particular relates to a parallel interpolation filtering device. Background Art
[0002] In signal processing, interpolation filters are divided into two parts: interpolation and filtering. The interpolation method is mostly zero insertion, and a signal that meets the adoption rate requirements is obtained after low-pass filtering (mostly half-band filtering). In communication systems, in order to meet real-time requirements, interpolation filters are mostly implemented through FPGAs or digital ICs. However, the internal working clock of FPGAs is mostly lower than 300MHZ, and ICs cannot achieve a higher working clock due to process selection; due to the limitation of the working clock, parallel processing needs to be considered for the implementation of interpolation filtering.
[0003] Existing technical solutions often use multiple filters in parallel to achieve high throughput interpolation filtering. The structure of a single filter is as follows: Figure 1 As shown, the structure of multiple filters in parallel is as follows Figure 2 As shown, two data are input in parallel, and after passing through parallel filters, two data are output in parallel, thereby reducing the operating frequency of the filter to half of the original. The same principle can realize more parallel filtering structures.
[0004] In the patent application CN 112187215 A, a cascaded half-band filter is used to implement parallel filtering of data, and the interpolation value is double interpolation 0 (such as Figure 3 As shown), the half-band filter structure is as follows Figure 4 As shown in , the odd-numbered bits are non-zero coefficients, the even-numbered bits are all zero except for the middle coefficient, and all coefficients are symmetric about the central coefficient. Figure 2 On this basis, the calculation part corresponding to the case where the zero in the data and the half-band filter coefficient is zero is eliminated and simplified, and the symmetry of the filtering system is considered at the same time to obtain the implementation structure of the interpolation filter in the patent application, that is, when performing interpolation filtering, the two terms with the same coefficients are added, and then multiplied with the half-band interpolation filter coefficient, and finally accumulated to obtain the interpolation filtering result.
[0005] In this structure, the multiplication results of multiple channels need to be accumulated. The combinational logic is large. In order to meet the timing requirements, the multiplication results need to be split into multiple steps for accumulation calculation, that is, the operation process is split multiple times using registers, thereby increasing the number of triggers and logic resources. Figure 5 As shown, flip-flops are inserted at the output positions of multiple adders to reduce the size of the combinational logic between the flip-flops, thereby increasing the operating clock.
[0006] The parallel processing of this implementation has a large fan-out due to the large number of taps in the cache register, which affects the clock frequency. In actual implementation, it is considered to split it into multiple low-speed filters to implement high-speed filtering in parallel, such as Figure 6 As shown, in order to meet the requirements of multiple low-speed filter input data, the data needs to be shifted accordingly and then output to each low-speed filter for filtering. At the same time, the parallel input data needs to be converted to serial parallel before being input into each sub-filter, and then the data of each branch is filtered, and finally 4 parallel data outputs are generated, which increases the control logic of the input data. Summary of the invention
[0007] The purpose of the present invention is to propose a parallel interpolation filtering implementation structure with simple control, easy to meet timing requirements, and less register usage. The structure is suitable for application scenarios such as implementing high-throughput interpolation filtering in FPGA and implementing high-throughput interpolation filtering in low-process chips, thereby realizing high-throughput interpolation filtering of data under low-frequency conditions.
[0008] To achieve the above object, the present invention proposes a parallel interpolation filtering device, and adopts the following technical solution:
[0009] A parallel interpolation filter device is a double interpolation filter device, the interpolation value is zero, and mainly includes a parallel interpolation filter module based on a half-band filter and a clock module. The parallel interpolation filter module includes an internal operation unit and a result output unit, and the clock module provides timing control for the parallel interpolation filter module.
[0010] The internal operation unit includes a multiplication unit, an addition unit and a group of registers for caching intermediate calculation results; the multiplication unit is used to multiply each data in the input data sequence with the corresponding non-zero coefficient in the half-band filter; the addition unit performs different accumulation operations with the center coefficient as the demarcation point, and for each coefficient before the center coefficient, the addition unit adds the product of the current coefficient and the corresponding data in the input data sequence to the accumulation result in the previous register indexed based on the number of input data sequences, and for each coefficient after the center coefficient, it adds the accumulation result in the previous register indexed based on the number of output data sequences; the result after the center coefficient is multiplied by the corresponding data in the input data sequence is stored separately in a specific register;
[0011] The result output unit includes a register for caching data results, generates output data of twice the number of input data sequence paths according to the calculation results of the internal calculation unit, and alternately outputs the accumulated result, the center coefficient and the result of multiplying the corresponding data in the input data sequence.
[0012] Furthermore, the data of the double interpolation filter is defined as data, and the numbering starts from 0; the order of the half-band filter is N, the coefficient is h and only the non-zero coefficients are marked, and the coefficients are numbered from 0 to K and There are 2K+1 non-zero coefficients, all of which are symmetrically distributed around h[K]. The register used by the internal computing unit is filter_reg, numbered from 0, denoted as filter_reg[0]. The initial state of the register is 0, and it is used to store intermediate calculation results. The register used by the result output unit is out_reg, numbered from 0, denoted as output_reg[0].
[0013] Furthermore, for 2-way parallel interpolation filtering, one sequence of data is input in each clock cycle, and the internal computing unit simultaneously performs interpolation and filtering processing on each data to generate two output data; the input data sequence is recorded as data[0]; the internal computing unit operates according to the following logical steps:
[0014] filter_reg[0]=data[0]*h[0];
[0015] filter_reg[1]=filter_reg[0]+data[0]*h[1];
[0016] …
[0017] filter_reg[K-1]=filter_reg[K-2]+data[0]*h[K-1];
[0018] filter_reg[K]=data[0]*h[K];
[0019] filter_reg[K+1]=filter_reg[K-1]+data[0]*h[K-1];
[0020] filter_reg[K+2]=filter_reg[K];
[0021] filter_reg[K+3]=filter_reg[K+1]+data[0]*h[K-2];
[0022] filter_reg[K+4]=filter_reg[K+2];
[0023] …
[0024] filter_reg[K+2(K-1)-2]=filter_reg[K+2(K-1)-4];
[0025] filter_reg[K+2(K-1)-1]=filter_reg[K+2(K-1)-3]+data[0]*h[K-(K-1)];
[0026] The output data of the result output unit is:
[0027] output_reg[0]=filter_reg[K+2(K-1)-1]+data[0]*h[0];
[0028] output_reg[1]=filter_reg[K+2(K-1)-2].
[0029] Furthermore, for 4-way parallel interpolation filtering, 2 sequence data are input in each clock cycle, and the internal computing unit simultaneously completes the interpolation and filtering processing to generate 4 output data; the input data sequence is recorded as data[0] and data[1]; the internal computing unit calculates according to the following logical steps:
[0030] filter_reg[0]=data[1]*h[0];
[0031] filter_reg[1]=data[0]*h[0]+data[1]*h[1];
[0032] filter_reg[2]=filter_reg[0]+data[0]*h[1]+data[1]*h[2];
[0033] filter_reg[3]=filter_reg[1]+data[0]*h[2]+data[1]*h[3];
[0034] …
[0035] filter_reg[K-2]=filter_reg[K-4]+data[0]*h[K-3]+data[1]*h[K-2];
[0036] filter_reg[K-1]=filter_reg[K-3]+data[0]*h[K-2]+data[1]*h[K-1];
[0037] filter_reg[K]=data[1]*h[K];
[0038] filter_reg[K+1]=filter_reg[K-1]+data[0]*h[K-1]+data[1]*h[K-1];
[0039] filter_reg[K + 2] = data[0] * h[K];
[0040] filter_reg[K + 3] = filter_reg[K - 1] + data[0] * h[K - 1] + data[1] * h[K - 2];
[0041] filter_reg[K + 4] = filter_reg[K];
[0042] filter_reg[K + 5] = filter_reg[K + 1] + data[0] * h[K - 2] + data[1] * h[K - 3];
[0043] filter_reg[K + 6] = filter_reg[K + 2];
[0044] …
[0045] filter_reg[K + 2(K - 1) - 4] = filter_reg[K + 2(K - 1) - 8];
[0046] filter_reg[K + 2(K - 1) - 3] = filter_reg[K + 2(K - 1) - 7] + data[0] * h[K - (K - 3)] + data[1] * h[K - (K - 2)];
[0047] filter_reg[K + 2(K - 1) - 2] = filter_reg[K + 2(K - 1) - 6];
[0048] filter_reg[K + 2(K - 1) - 1] = filter_reg[K + 2(K - 1) - 5] + data[0] * h[K - (K - 2)] + data[1] * h[K - (K - 1)];
[0049] The result output unit outputs data as:
[0050] output_reg[0] = filter_reg[K + 2(K - 1) - 1] + data[0] * h[0];
[0051] output_reg[1] = filter_reg[K + 2(K - 1) - 2];
[0052] output_reg[2] = filter_reg[K + 2(K - 1) - 3] + data[0] * h[1] + data[1] * h[0];
[0053] output_reg[3]=filter_reg[K+2(K-1)-4].
[0054] Furthermore, for 8-way parallel interpolation filtering, 4 sequence data are input in each clock cycle, and the internal computing unit simultaneously performs interpolation and filtering processing on each data to generate 8 output data; the input data sequence is recorded as data[0], data[1], data[2], data[3]; the internal computing unit calculates according to the following logical steps:
[0055] filter_reg[0]=data[3]*h[0];
[0056] filter_reg[1]=data[2]*h[0]+data[3]*h[1];
[0057] filter_reg[2]=data[1]*h[0]+data[2]*h[1]+data[3]*h[2];
[0058] filter_reg[3]=data[0]*h[0]+data[1]*h[1]+data[2]*h[2]+data[3]*h[3];
[0059] filter_reg[4]=filter_reg[0]+data[0]*h[1]+data[1]*h[2]+data[2]*h[3]+data[3]*h[4];
[0060] filter_reg[5]=filter_reg[1]+data[0]*h[2]+data[1]*h[3]+data[2]*h[4]+data[3]*h[5];
[0061] filter_reg[6]=filter_reg[2]+data[0]*h[3]+data[1]*h[4]+data[2]*h[5]+data[3]*h[6];
[0062] filter_reg[7]=filter_reg[3]+data[0]*h[4]+data[1]*h[5]+data[2]*h[6]+data[3]*h[7];
[0063] …
[0064] filter_reg[K-4]=filter_reg[K-8]+data[0]*h[K-7]+data[1]*h[K-6]+
[0065] data[2]*h[K-5]+data[3]*h[K-4];
[0066] filter_reg[K-3]=filter_reg[K-7]+data[0]*h[K-6]+data[1]*h[K-5]+
[0067] data[2]*h[K-4]+data[3]*h[K-3];
[0068] filter_reg[K-2]=filter_reg[K-6]+data[0]*h[K-5]+data[1]*h[K-4]+
[0069] data[2]*h[K-3]+data[3]*h[K-2];
[0070] filter_reg[K-1]=filter_reg[K-5]+data[0]*h[K-4]+data[1]*h[K-3]+
[0071] data[2]*h[k-2]+data[3]*h[K-1];
[0072] filter_reg[K]=data[3]*h[K];
[0073] filter_reg[K+1]=filter_reg[K-4]+data[0]*h[K-3]+data[1]*h[K-2]+
[0074] data[2]*h[k-1]+data[3]*h[K-1];
[0075] filter_reg[K+2]=data[2]*h[K];
[0076] filter_reg[K+3]=filter_reg[K-3]+data[0]*h[K-2]+data[1]*h[K-1]+
[0077] data[2]*h[k-1]+data[3]*h[K-2];
[0078] filter_reg[K+4]=data[1]*h[K];
[0079] filter_reg[K+5]=filter_reg[K-2]+data[0]*h[K-1]+data[1]*h[K-1]+
[0080] data[2]*h[k-2]+data[3]*h[K-3];
[0081] filter_reg[K+6]=data[0]*h[K];
[0082] filter_reg[K+7]=filter_reg[K-1]+data[0]*h[K-1]+data[1]*h[K-2]+
[0083] data[2]*h[k-3]+data[3]*h[K-4];
[0084] filter_reg[K+8]=filter_reg[K];
[0085] …
[0086] filter_reg[K+2(K-1)-8]=filter_reg[K+2(K-1)-16];
[0087] filter_reg[K+2(K-1)-7]=filter_reg[K+2(K-1)-15]+data[0]*h[K-(K-7)]+data[1]*h[K-(K-6)]+data[2]*h[K-(K-5)]+data[3]*h[K-(K-4)];
[0088] filter_reg[K+2(K-1)-6]=filter_reg[K+2(K-1)-14];
[0089] filter_reg[K+2(K-1)-5]=filter_reg[K+2(K-1)-13]+data[0]*h[K-(K-6)]+data[1]*h[K-(K-5)]+data[2]*h[K-(K-4)]+data[3]*h[K-(K-3)];
[0090] filter_reg[K+2(K-1)-4]=filter_reg[K+2(K-1)-12];
[0091] filter_reg[K + 2(K - 1) - 3] = filter_reg[K + 2(K - 1) - 11] + data[0] * h[K - (K - 5)] + data[1] * h[K - (K - 4)] + data[2] * h[K - (K - 3)] + data[3] * h[K - (K - 2)];
[0092] filter_reg[K + 2(K - 1) - 2] = filter_reg[K + 2(K - 1) - 10];
[0093] filter_reg[K + 2(K - 1) - 1] = filter_reg[K + 2(K - 1) - 9] + data[0] * h[K - (K - 4)] + data[1] * h[K - (K - 3)] + data[2] * h[K - (K - 2)] + data[3] * h[K - (K - 1)];
[0094] The result output unit outputs the data as:
[0095] output_reg[0] = filter_reg[K + 2(K - 1) - 1] + data[0] * h[0];
[0096] output_reg[1] = filter_reg[K + 2(K - 1) - 2];
[0097] output_reg[2] = filter_reg[K + 2(K - 1) - 3] + data[0] * h[1] + data[1] * h[0];
[0098] output_reg[3] = filter_reg[K + 2(K - 1) - 4];
[0099] output_reg[4] = filter_reg[K + 2(K - 1) - 5] + data[0] * h[2] + data[1] * h[1] +
[0100] data[2] * h[0];
[0101] output_reg[5] = filter_reg[K + 2(K - 1) - 6];
[0102] output_reg[6] = filter_reg[K + 2(K - 1) - 7] + data[0] * h[3] + data[1] * h[2] +
[0103] data[2]*h[1]+data[3]*h[0];
[0104] output_reg[7]=filter_reg[K+2(K-1)-8].
[0105] Furthermore, for 8-way parallel interpolation filtering, a register may be inserted into the multiplication unit part to cache the multiplication result, or a register may be inserted into the product result summation operation to cache the summation result, thereby further improving the operating clock frequency.
[0106] Compared with the prior art, the present invention has the following beneficial effects:
[0107] The parallel interpolation filter structure provided by the present invention integrates double interpolation and half-band filtering, and simultaneously completes interpolation and filtering processing for each input data, thereby reducing processing delay; at the same time, the structure eliminates the cache of input data and only caches the results in the calculation process, thereby reducing the use of registers; in addition, there is no large combinational logic between the registers of the parallel interpolation filter structure provided by the present invention, which improves the operating clock of the entire circuit; there is no need to perform serial-to-parallel conversion on the input data, thereby reducing the complexity of the control logic. BRIEF DESCRIPTION OF THE DRAWINGS
[0108] Figure 1 It is a structural diagram of a single filter in the prior art;
[0109] Figure 2 A parallel filter implementation structure in the prior art;
[0110] Figure 3 Schematic diagram of the interpolation method of double interpolation zero;
[0111] Figure 4 Schematic diagram of coefficient arrangement of a 55-order half-band filter;
[0112] Figure 5 A schematic diagram of the insertion position of a trigger in the filter structure in the existing patent application CN112187215A;
[0113] Figure 6 It is a schematic diagram of the parallel filtering input data control method in the existing patent application CN112187215A;
[0114] Figure 7 A schematic diagram of the structure of a 2-way parallel interpolation filter provided by the present invention;
[0115] Figure 8 A schematic diagram of the structure of a 4-way parallel interpolation filter provided by the present invention;
[0116] Fig. 9A schematic diagram of a data input control method for 4-way parallel interpolation filtering provided by the present invention;
[0117] Fig.10 A schematic diagram of a data input control method for 8-way parallel interpolation filtering provided by the present invention;
[0118] Fig.11 This is a schematic diagram of a structure in which a register is further inserted into the 8-way parallel interpolation filter structure provided by the present invention. DETAILED DESCRIPTION
[0119] The technical solution of the present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0120] A parallel interpolation filter device provided by the present invention is a double interpolation filter device, the interpolation value is zero, and mainly includes a parallel interpolation filter module based on a half-band filter and a clock module, the parallel interpolation filter module includes an internal operation unit and a result output unit, and the clock module provides timing control for the parallel interpolation filter module;
[0121] The internal operation unit includes a multiplication unit, an addition unit and a group of registers for caching intermediate calculation results; the multiplication unit is used to multiply each data in the input data sequence with the corresponding non-zero coefficient in the half-band filter; the addition unit performs different accumulation operations with the center coefficient as the demarcation point, and for each coefficient before the center coefficient, the addition unit adds the product of the current coefficient and the corresponding data in the input data sequence to the accumulation result in the previous register indexed based on the number of input data sequences, and for each coefficient after the center coefficient, it adds the accumulation result in the previous register indexed based on the number of output data sequences; the result after the center coefficient is multiplied by the corresponding data in the input data sequence is stored separately in a specific register;
[0122] The result output unit includes a register for caching data results, generates output data of twice the number of input data sequence paths according to the calculation results of the internal calculation unit, and alternately outputs the accumulated result, the center coefficient and the result of multiplying the corresponding data in the input data sequence.
[0123] The data for defining the double interpolation filter is data, and the numbering starts from 0; the filter is a 63-order half-band filter, the coefficient is h and only the non-zero coefficients are marked, the coefficients are numbered from 0 to 16, there are 33 non-zero coefficients, and all coefficients are symmetrically distributed around h
[16] ; the register used by the internal calculation unit is filter_reg, numbered from 0 to 45, with an initial state of 0, and is used to store intermediate calculation results; the register used by the result output unit is out_reg, and the numbering starts from 0.
[0124] In one embodiment, the interpolation filtering device is used to implement 2-way parallel interpolation filtering, and one sequence of data is input in each clock cycle. The internal calculation unit simultaneously performs interpolation and filtering processing on each data to generate 2 output data, that is, each data in the input data sequence is multiplied by the non-zero coefficient in the half-band filter in turn, and the result after the center coefficient is multiplied by the corresponding data in the input data sequence is stored separately, and the result after the remaining data is multiplied by the corresponding non-zero coefficient is accumulated with the cumulative result stored in the previous register (the input data sequence number is indexed before the center coefficient) or the second previous register (the output data number is indexed after the center coefficient) and stored in the current register; the result output unit outputs the result after the corresponding data in the input data sequence is multiplied by the center coefficient in one way, and outputs the result after the remaining data in the input data sequence is multiplied by the corresponding non-zero coefficient in another way. Its logical implementation structure is as follows: Figure 7 As shown:
[0125] The result of data[0]*h[0] is stored in filter_reg[0]; the result of data[0]*h[1] is added to the result stored in filter_reg[0] and stored in filter_reg[1]; the result of data[0]*h[2] is added to the result stored in filter_reg[1] and stored in filter_reg[2]; and so on until the content of filter_reg
[15] register is updated;
[0126] The result of data[0]*h
[16] is stored in filter_reg
[16] ;
[0127] The result of data[0]*h
[15] is added to the result stored in filter_reg
[15] and then stored in filter_reg
[17] . The result of filter_reg
[16] is directly stored in filter_reg
[18] . The result of data[0]*h
[14] is added to the result stored in filter_reg
[17] and then stored in filter_reg
[19] . The result of filter_reg
[18] is directly stored in filter_reg
[20] . The same goes for the update of the contents of the filter_reg
[45] register.
[0128] Finally, filter_reg
[45] +data[0]*h[0] is used as the first output result of the 2-way parallel interpolation filter; filter_reg
[44] is used as the second output result of the 2-way parallel interpolation filter;
[0129] The specific logic operation process of the internal operation unit can be expressed as follows:
[0130] filter_reg[0]=data[0]*h[0];
[0131] filter_reg[1]=filter_reg[0]+data[0]*h[1];
[0132] Filter_reg[2]=filter_reg[1]+data[0]*h[2];
[0133] …
[0134] filter_reg
[15] =filter_reg
[14] +data[0]*h
[15] ;
[0135] filter_reg
[16] =data[0]*h
[16] ;
[0136] filter_reg
[17] =filter_reg
[15] +data[0]*h
[15] ;
[0137] filter_reg
[18] =filter_reg
[16] ;
[0138] filter_reg
[19] =filter_reg
[17] +data[0]*h
[14] ;
[0139] filter_reg
[20] =filter_reg
[18] ;
[0140] …
[0141] filter_reg
[44] =filter_reg
[42] ;
[0142] filter_reg
[45] =filter_reg
[43] +data[0]*h[1];
[0143] The output data of the result output unit is:
[0144] output_reg[0]=filter_reg
[45] +data[0]*h[0];
[0145] output_reg[1]=filter_reg
[44] ;
[0146] Depend on Figure 7 It can be seen that all registers are only used to store the summed results. This structure does not have large combinational logic resources and runs at a high clock frequency.
[0147] In another embodiment, the device is used to implement 4-way parallel interpolation filtering, 2 sequence data are input in each clock cycle, and the internal computing unit performs interpolation and filtering processing on each data at the same time to generate 4 output data. The specific logic implementation process is as follows Figure 8 As shown, the result of data[1]*h[0] is stored in filter_reg[0]; the results of data[0]*h[0] and data[1]*h[1] are added and stored in filter_reg[1]; the results of data[0]*h[1] and data[1]*h[2] are added to the result stored in filter_reg[0] and stored in filter_reg[2]; the results of data[0]*h[2] and data[1]*h[3] are added to the result stored in filter_reg[1] and stored in filter_reg[3]; and so on until the content of filter_reg
[15] register is updated;
[0148] The result of data[1]*h
[16] is stored in filter_reg
[16] ;
[0149] The result of data[0]*h
[15] and data[1]*h
[15] are added to the result stored in filter_reg
[14] and stored in filter_reg
[17] . The result of data[0]*h
[16] is stored in filter_reg
[18] . The result of data[0]*h
[15] and data[1]*h
[14] are added to the result stored in filter_reg
[15] and stored in filter_reg
[19] . The result of filter_reg
[16] is directly stored in filter_reg
[20] . The result of data[0]*h
[14] and data[1]*h
[13] are added to the result stored in filter_reg
[17] and stored in filter_reg
[21] . The result of filter_reg
[18] is directly stored in filter_reg
[22] . And so on until the update of the contents of the filter_reg
[45] register.
[0150] Finally, filter_reg
[45] +data[0]*h[0] is used as the first output result of the 4-way parallel interpolation filter; filter_reg
[44] is used as the second output result of the 4-way parallel interpolation filter; filter_reg
[43] +data[0]*h[1]+data[1]*h[0] is used as the third output result of the 4-way parallel interpolation filter; filter_reg
[42] is used as the fourth output result of the 4-way parallel interpolation filter;
[0151] The entire logical implementation process can be expressed as follows:
[0152] filter_reg[0] = data[1] * h[0];
[0153] filter_reg[1] = data[0] * h[0] + data[1] * h[1];
[0154] filter_reg[2] = filter_reg[0] + data[0] * h[1] + data[1] * h[2];
[0155] filter_reg[3] = filter_reg[1] + data[0] * h[2] + data[1] * h[3];
[0156] …
[0157] filter_reg
[14] = filter_reg
[12] + data[0] * h
[13] + data[1] * h
[14] ;
[0158] filter_reg
[15] = filter_reg
[13] + data[0] * h
[14] + data[1] * h
[15] ;
[0159] filter_reg
[16] = data[1] * h
[16] ;
[0160] filter_reg
[17] = filter_reg
[15] + data[0] * h
[15] + data[1] * h
[15] ;
[0161] filter_reg
[18] = data[0] * h
[16] ;
[0162] filter_reg
[19] = filter_reg
[15] + data[0] * h
[15] + data[1] * h
[14] ;
[0163] filter_reg
[20] = filter_reg
[16] ;
[0164] filter_reg
[21] = filter_reg
[17] + data[0] * h
[14] + data[1] * h
[13] ;
[0165] filter_reg
[22] = filter_reg
[18] ;
[0166] …
[0167] filter_reg
[42] =filter_reg
[38] ;
[0168] filter_reg
[43] =filter_reg
[39] +data[0]*h[3]+data[1]*h[2];
[0169] filter_reg
[44] =filter_reg
[40] ;
[0170] filter_reg
[45] =filter_reg
[41] +data[0]*h[2]+data[1]*h[1];
[0171] The output data of the result output unit is:
[0172] output_reg[0]=filter_reg
[45] +data[0]*h[0];
[0173] output_reg[1]=filter_reg
[44] ;
[0174] output_reg[2]=filter_reg
[43] +data[0]*h[1]+data[1]*h[0];
[0175] output_reg[3]=filter_reg
[42] ;
[0176] In this structure, there is no need to cache multiple input data, only the intermediate calculation results need to be cached, thus reducing the use of registers; since there is no large combinational logic between registers, the operating clock frequency is high; Fig. 9 As shown, due to the fully parallel structure inside the parallel interpolation filter module, the input data does not need to be converted from parallel to serial, and the input control is simple.
[0177] Furthermore, the device is used to implement 8-way parallel interpolation filtering. Four sequence data are input in each clock cycle. The internal calculation unit simultaneously performs interpolation and filtering processing on each data to generate 8 output data. The specific logic implementation process can be expressed as:
[0178] filter_reg[0]=data[3]*h[0];
[0179] filter_reg[1]=data[2]*h[0]+data[3]*h[1];
[0180] filter_reg[2]=data[1]*h[0]+data[2]*h[1]+data[3]*h[2];
[0181] filter_reg[3]=data[0]*h[0]+data[1]*h[1]+data[2]*h[2]+data[3]*h[3];
[0182] filter_reg[4]=filter_reg[0]+data[0]*h[1]+data[1]*h[2]+data[2]*h[3]+data[3]*h[4];
[0183] filter_reg[5]=filter_reg[1]+data[0]*h[2]+data[1]*h[3]+data[2]*h[4]+data[3]*h[5];
[0184] filter_reg[6]=filter_reg[2]+data[0]*h[3]+data[1]*h[4]+data[2]*h[5]+data[3]*h[6];
[0185] filter_reg[7]=filter_reg[3]+data[0]*h[4]+data[1]*h[5]+data[2]*h[6]+data[3]*h[7];
[0186] …
[0187] filter_reg
[12] =filter_reg[8]+data[0]*h[9]+data[1]*h
[10] +data[2]*
[0188] h
[11] +data[3]*h
[12] ;
[0189] filter_reg
[13] =filter_reg[9]+data[0]*h
[10] +data[1]*h
[11] +data[2]
[0190] *h
[12] +data[3]*h
[13] ;
[0191] filter_reg
[14] =filter_reg
[10] +data[0]*h
[11] +data[1]*h
[12] +data[2]*h
[13] +data[3]*h
[14] ;
[0192] filter_reg
[15] =filter_reg
[11] +data[0]*h
[12] +data[1]*h
[13] +data[2]*h
[14] +data[3]*h
[15] ;
[0193] filter_reg
[16] =data[3]*h
[16] ;
[0194] filter_reg
[17] =filter_reg
[12] +data[0]*h
[13] +data[1]*h
[14] +data[2]*h
[15] +data[3]*h
[15] ;
[0195] filter_reg
[18] =data[2]*h
[16] ;
[0196] filter_reg
[19] =filter_reg
[13] +data[0]*h
[14] +data[1]*h
[15] +data[2]*h
[15] +data[3]*h
[14] ;
[0197] filter_reg
[20] =data[1]*h
[16] ;
[0198] filter_reg
[21] =filter_reg
[14] +data[0]*h
[15] +data[1]*h
[15] +data[2]*h
[14] +data[3]*h
[13] ;
[0199] filter_reg
[22] =data[0]*h
[16] ;
[0200] filter_reg
[23] =filter_reg
[15] +data[0]*h
[15] +data[1]*h
[14] +data[2]*h
[13] +data[3]*h
[12] ;
[0201] filter_reg
[24] =filter_reg
[16] ;
[0202] …
[0203] filter_reg
[38] =filter_reg
[30] ;
[0204] filter_reg
[39] = filter_reg
[31] + data[0] * h[7] + data[1] * h[6] + data[2] * h[5] + data[3] * h[4];
[0205] filter_reg
[40] = filter_reg
[32] ;
[0206] filter_reg
[41] = filter_reg
[33] + data[0] * h[6] + data[1] * h[5] + data[2] * h[4] + data[3] * h[3];
[0207] filter_reg
[42] = filter_reg
[34] ;
[0208] filter_reg
[43] = filter_reg
[35] + data[0] * h[5] + data[1] * h[4] + data[2] * h[3] + data[3] * h[2];
[0209] filter_reg
[44] = filter_reg
[36] ;
[0210] filter_reg
[45] = filter_reg
[37] + data[0] * h[4] + data[1] * h[3] + data[2] * h[2] + data[3] * h[1];
[0211] The result output unit is:
[0212] output_reg[0] = filter_reg
[45] + data[0] * h[0];
[0213] output_reg[1] = filter_reg
[44] ;
[0214] output_reg[2] = filter_reg
[43] + data[0] * h[1] + data[1] * h[0];
[0215] output_reg[3] = filter_reg
[42] ;
[0216] output_reg[4] = filter_reg
[41] + data[0] * h[2] + data[1] * h[1] + data[2] * h[0];
[0217] output_reg[5]=filter_reg
[40] ;
[0218] output_reg[6]=filter_reg
[39] +data[0]*h[3]+data[1]*h[2]+data[2]*h[1]+data[3]*h[0];
[0219] output_reg[7]=filter_reg
[38] ;
[0220] The structure and Figure 8 The structure is similar, with 4 data inputs and 8 parallel outputs after interpolation filtering. Fig.10 As shown, there is no need to perform serial-to-parallel conversion on the input data, and the control is simple.
[0221] like Fig.11 As shown, in order to further improve the clock frequency of the 8-way parallel interpolation filter structure, a register can be inserted in the multiplication unit part to store the intermediate results of the multiplication calculation, for example:
[0222] filter_reg[4]=filter_reg[0]+data[0]*h[1]+data[1]*h[2]+data[2]*h[3]+data[3]*h[4];
[0223] After adding registers to the multiplication unit, it is equivalent to:
[0224] Multi0 = data[0]*h[1];
[0225] Multi1 = data[1]*h[2];
[0226] Multi2 = data[2]*h[3];
[0227] Multi3 = data[3]*h[4];
[0228] Further inserting registers into the sum operation is equivalent to:
[0229] Sum=Multi0+Multi1+Multi2+Multi3;
[0230] Finally, solve filter_reg[0]+Sum to get the result in filter_reg[4].
[0231] The above description is only a preferred specific embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made by any person skilled in the art within the technical scope disclosed by the present invention and based on the technical solution and inventive concept of the present invention shall be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A parallel interpolation filter device, which is a double interpolation filter device, the interpolation value is zero, characterized in that: It mainly includes a parallel interpolation filter module based on a half-band filter and a clock module. The parallel interpolation filter module includes an internal operation unit and a result output unit. The clock module provides timing control for the parallel interpolation filter module. The internal operation unit includes a multiplication unit, an addition unit, and a group of registers for caching intermediate calculation results; the multiplication unit is used to multiply each data in the input data sequence with the corresponding non-zero coefficient in the half-band filter; The adding unit performs different accumulation operations with the center coefficient as the dividing point. For each coefficient before the center coefficient, the adding unit adds the product of the current coefficient and the corresponding data in the input data sequence to the accumulation result in the previous register based on the number of input data sequences as the index. For each coefficient after the center coefficient, the adding unit adds the accumulation result in the previous register based on the number of output data sequences as the index. The result of multiplying the center coefficient and the corresponding data in the input data sequence is stored separately in a specific register. The result output unit includes a register for caching data results, generates output data of twice the number of input data sequence paths according to the calculation results of the internal calculation unit, and alternately outputs the accumulated result, the center coefficient and the result of multiplying the corresponding data in the input data sequence.
2. A parallel interpolation filtering device according to claim 1, characterized in that: Define the data of the double interpolation filter as data, numbered from 0; the order of the half-band filter is N, the coefficient is h and only non-zero coefficients are marked, the coefficients are numbered from 0 to K and All coefficients are symmetrically distributed around h[K]; the register used by the internal calculation unit is filter_reg, numbered from 0, filter_reg[0], with an initial state of 0, which is used to store intermediate calculation results; the register used by the result output unit is out_reg, numbered from 0, denoted as output_reg[0].
3. A parallel interpolation filtering device according to claim 2, characterized in that: For 2-way parallel interpolation filtering, one sequence of data is input in each clock cycle, and the internal computing unit simultaneously performs interpolation and filtering processing on each data to generate two output data; the input data sequence is recorded as data[0]; the internal computing unit operates according to the following logical steps: filter_reg[0]=data[0]*h[0]; filter_reg[1]=filter_reg[0]+data[0]*h[1]; … filter_reg[K-1]=filter_reg[K-2]+data[0]*h[K-1]; filter_reg[K]=data[0]*h[K]; filter_reg[K+1]=filter_reg[K-1]+data[0]*h[K-1]; filter_reg[K+2]=filter_reg[K]; filter_reg[K+3]=filter_reg[K+1]+data[0]*h[K-2]; filter_reg[K+4]=filter_reg[K+2]; … filter_reg[K+2(K-1)-2]=filter_reg[K+2(K-1)-4]; filter_reg[K+2(K-1)-1]=filter_reg[K+2(K-1)-3]+data[0]*h[K-(K-1)]; The output data of the result output unit is: output_reg[0]=filter_reg[K+2(K-1)-1]+data[0]*h[0]; output_reg[1]=filter_reg[K+2(K-1)-2].
4. The parallel interpolation filtering device according to claim 2, characterized in that: For 4-way parallel interpolation filtering, 2 sequence data are input in each clock cycle. The internal computing unit simultaneously performs interpolation and filtering on each data to generate 4 output data. The input data sequence is recorded as data[0] and data[1]. The internal computing unit operates according to the following logical steps: filter_reg[0]=data[1]*h[0]; filter_reg[1]=data[0]*h[0]+data[1]*h[1]; filter_reg[2]=filter_reg[0]+data[0]*h[1]+data[1]*h[2]; filter_reg[3]=filter_reg[1]+data[0]*h[2]+data[1]*h[3]; … filter_reg[K-2]=filter_reg[K-4]+data[0]*h[K-3]+data[1]*h[K-2]; filter_reg[K-1]=filter_reg[K-3]+data[0]*h[K-2]+data[1]*h[K-1]; filter_reg[K]=data[1]*h[K]; filter_reg[K+1]=filter_reg[K-1]+data[0]*h[K-1]+data[1]*h[K-1]; filter_reg[K+2]=data[0]*h[K]; filter_reg[K+3]=filter_reg[K-1]+data[0]*h[K-1]+data[1]*h[K-2]; filter_reg[K+4]=filter_reg[K]; filter_reg[K+5]=filter_reg[K+1]+data[0]*h[K-2]+data[1]*h[K-3]; filter_reg[K+6]=filter_reg[K+2]; … filter_reg[K+2(K-1)-4]=filter_reg[K+2(K-1)-8]; filter_reg[K+2(K-1)-3]=filter_reg[K+2(K-1)-7]+data[0]*h[K-(K-3)]+data[1]*h[K-(K-2)]; filter_reg[K+2(K-1)-2]=filter_reg[K+2(K-1)-6]; filter_reg[K+2(K-1)-1]=filter_reg[K+2(K-1)-5]+data[0]*h[K-(K-2)]+data[1]*h[K-(K-1)]; The output data of the result output unit is: output_reg[0]=filter_reg[K+2(K-1)-1]+data[0]*h[0]; output_reg[1]=filter_reg[K+2(K-1)-2]; output_reg[2]=filter_reg[K+2(K-1)-3]+data[0]*h[1]+data[1]*h[0]; output_reg[3]=filter_reg[K+2(K-1)-4].
5. The parallel interpolation filtering device according to claim 2, characterized in that: For 8-way parallel interpolation filtering, 4 sequence data are input in each clock cycle, and the internal computing unit simultaneously performs interpolation and filtering processing on each data to generate 8 output data; the input data sequence is recorded as data[0], data[1], data[2], data[3]; the internal computing unit calculates according to the following logical steps: filter_reg[0]=data[3]*h[0]; filter_reg[1]=data[2]*h[0]+data[3]*h[1]; filter_reg[2]=data[1]*h[0]+data[2]*h[1]+data[3]*h[2]; filter_reg[3]=data[0]*h[0]+data[1]*h[1]+data[2]*h[2]+data[3]*h[3]; filter_reg[4]=filter_reg[0]+data[0]*h[1]+data[1]*h[2]+data[2]*h[3]+data[3]*h[4]; filter_reg[5]=filter_reg[1]+data[0]*h[2]+data[1]*h[3]+data[2]*h[4]+data[3]*h[5]; filter_reg[6]=filter_reg[2]+data[0]*h[3]+data[1]*h[4]+data[2]*h[5]+data[3]*h[6]; filter_reg[7]=filter_reg[3]+data[0]*h[4]+data[1]*h[5]+data[2]*h[6]+data[3]*h[7]; … filter_reg[K-4]=filter_reg[K-8]+data[0]*h[K-7]+data[1]*h[K-6]+data[2]*h[K-5]+data[3]*h[K-4]; filter_reg[K-3]=filter_reg[K-7]+data[0]*h[K-6]+data[1]*h[K-5]+data[2]*h[K-4]+data[3]*h[K-3]; filter_reg[K-2]=filter_reg[K-6]+data[0]*h[K-5]+data[1]*h[K-4]+data[2]*h[K-3]+data[3]*h[K-2]; filter_reg[K-1]=filter_reg[K-5]+data[0]*h[K-4]+data[1]*h[K-3]+data[2]*h[k-2]+data[3]*h[K-1]; filter_reg[K]=data[3]*h[K]; filter_reg[K+1]=filter_reg[K-4]+data[0]*h[K-3]+data[1]*h[K-2]+data[2]*h[k-1]+data[3]*h[K-1]; filter_reg[K+2]=data[2]*h[K]; filter_reg[K+3]=filter_reg[K-3]+data[0]*h[K-2]+data[1]*h[K-1]+data[2]*h[k-1]+data[3]*h[K-2]; filter_reg[K+4]=data[1]*h[K]; filter_reg[K+5]=filter_reg[K-2]+data[0]*h[K-1]+data[1]*h[K-1]+data[2]*h[k-2]+data[3]*h[K-3]; filter_reg[K+6]=data[0]*h[K]; filter_reg[K+7]=filter_reg[K-1]+data[0]*h[K-1]+data[1]*h[K-2]+data[2]*h[k-3]+data[3]*h[K-4]; filter_reg[K+8]=filter_reg[K]; … filter_reg[K+2(K-1)-8]=filter_reg[K+2(K-1)-16]; filter_reg[K+2(K-1)-7]=filter_reg[K+2(K-1)-15]+data[0]*h[K-(K-7)]+data[1]*h[K-(K-6)]+data[2]*h[K-(K-5)]+data[3]*h[K-(K-4)]; filter_reg[K+2(K-1)-6]=filter_reg[K+2(K-1)-14]; filter_reg[K+2(K-1)-5]=filter_reg[K+2(K-1)-13]+data[0]*h[K-(K-6)]+data[1]*h[K-(K-5)]+data[2]*h[K-(K-4)]+data[3]*h[K-(K-3)]; filter_reg[K+2(K-1)-4]=filter_reg[K+2(K-1)-12]; filter_reg[K+2(K-1)-3]=filter_reg[K+2(K-1)-11]+data[0]*h[K-(K-5)]+data[1]*h[K-(K-4)]+data[2]*h[K-(K-3)]+data[3]*h[K-(K-2)]; filter_reg[K+2(K-1)-2]=filter_reg[K+2(K-1)-10]; filter_reg[K+2(K-1)-1]=filter_reg[K+2(K-1)-9]+data[0]*h[K-(K-4)]+data[1]*h[K-(K-3)]+data[2]*h[K-(K-2)]+data[3]*h[K-(K-1)]; The output data of the result output unit is: output_reg[0]=filter_reg[K+2(K-1)-1]+data[0]*h[0]; output_reg[1]=filter_reg[K+2(K-1)-2]; output_reg[2]=filter_reg[K+2(K-1)-3]+data[0]*h[1]+data[1]*h[0]; output_reg[3]=filter_reg[K+2(K-1)-4]; output_reg[4]=filter_reg[K+2(K-1)-5]+data[0]*h[2]+data[1]*h[1]+data[2]*h[0]; output_reg[5]=filter_reg[K+2(K-1)-6]; output_reg[6]=filter_reg[K+2(K-1)-7]+data[0]*h[3]+data[1]*h[2]+data[2]*h[1]+data[3]*h[0]; output_reg[7]=filter_reg[K+2(K-1)-8].
6. The parallel interpolation filtering device according to claim 5, characterized in that: For 8-way parallel interpolation filtering, registers are inserted in the multiplication unit part to cache the multiplication results.
7. The parallel interpolation filtering device according to claim 6, characterized in that: For 8-way parallel interpolation filtering, a register is inserted in the product result summation operation to cache the result of the summation operation.
8. A parallel interpolation filtering device according to any one of claims 2 to 5, characterized in that: The order of the half-band filter is 63, the non-zero coefficients are numbered from 0 to 16, there are 33 non-zero coefficients, and all coefficients are symmetrically distributed around h[16]; the register used by the internal calculation unit is filter_reg, numbered from 0 to 45, with an initial state of 0, which is used to store intermediate calculation results; the register used by the result output unit is out_reg, numbered from 0.
Citation Information
Patent Citations
Cascaded half-band interpolation filter structure
CN112187215A