An Adaptive Filtering Method Based on Multi-Kernel Nystrom Method
By mapping signals to a high-dimensional feature space using the multi-kernel Nystrom method and optimizing weights, the problems of high computational complexity and cumbersome kernel parameter selection in traditional multi-kernel adaptive filters are solved, achieving more efficient nonlinear signal processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2022-09-13
- Publication Date
- 2026-05-26
AI Technical Summary
Traditional multi-core adaptive filters have high computational complexity and cumbersome kernel parameter selection in nonlinear signal processing, while single-core adaptive filters have unstable performance in non-Gaussian environments.
The multi-kernel Nystrom method is adopted to map the input signal to a high-dimensional feature space, use multiple kernel functions for filtering, and optimize the weights through a semi-quadratic optimization method to reduce computational complexity and kernel parameter sensitivity.
It significantly reduces computational complexity and memory consumption, improves filtering accuracy and robustness, and performs exceptionally well in non-Gaussian noise environments.
Smart Images

Figure CN115567038B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of kernel adaptive filtering. Background Technology
[0002] This invention is based on the kernel adaptive filtering method, mainly studying the design and application of the kernel adaptive filtering algorithm. Adaptive filtering algorithms can perform real-time signal processing tasks. Although traditional adaptive filtering algorithms have the characteristics of being lightweight and online, they are still difficult to handle nonlinear signal processing problems. Therefore, the kernel method is introduced to enhance the ability of adaptive filtering to handle nonlinear problems. Considering that data in real-world environments often have nonlinear characteristics, linear adaptive filtering systems have limitations in handling nonlinear problems. To address this, the kernel method is used to nonlinearly map the input signal to a high-dimensional regenerating kernel Hilbert space, and then linear filtering is performed in the feature space, thus developing the kernel adaptive filter (KAF). In the kernel adaptive filtering algorithm, the cost function, kernel function, weight update method, and sparsity strategy determine the performance. Classical kernel adaptive filtering algorithms include the kernel least mean square algorithm and the kernel recursive least square algorithm based on the mean square error (MSE) and least squares (LS) criteria. However, these algorithms, as well as the later proposed kernel conjugate gradient algorithm, are all based on the quadratic loss function (MSE), which can lead to severe instability and performance degradation in non-Gaussian environments. Therefore, kernel adaptive filtering algorithms based on cross-entropy loss function have been proposed to address this problem. Originating from information theory, the cross-entropy loss function captures the higher-order statistical characteristics of signals, effectively resisting interference from non-Gaussian noise. In recent years, a new loss function based on logarithmic functions has been extracted and introduced into kernel adaptive filtering research, including the Cauchy loss function, the Cauchy kernel loss function, and the generalized Cauchy loss function (H. Zhang, B. Yang, L. Wang, and S. Wang, “General cauchy conjugate gradient algorithms based on multiple random fourier features,” IEEE Transactions on Signal Processing, vol. 69, pp. 1859–1873, 2021).
[0003] Among them, the generalized Cauchy function has been shown to have better robustness and filtering performance in impulse noise environments. Furthermore, traditional kernel-adaptive filtering algorithms based on a single kernel function are often very sensitive to the selection of kernel parameters, and the selection of kernel parameters is often tedious and complex. To address this issue, researchers recently proposed a multi-kernel adaptive filtering method (W. Shi, K. Xiong, and S. Wang, “Multikernel adaptive filters under the minimum cauchy kernel loss criterion,” IEEE Access, vol. 7, pp. 120 548–120 558, 2019.), which effectively reduces the algorithm's dependence on kernel parameters and shows a significant improvement in filtering performance compared to single-kernel algorithms.
[0004] In practical adaptive signal processing, traditional multi-core adaptive filtering algorithms, including those mentioned above, inevitably lead to significant computational complexity and memory overhead because the network size of the kernel adaptive filter gradually increases with the iteration process. Summary of the Invention
[0005] Technical problem to be solved by the invention: Inspired by the above research background, this invention proposes a filter based on the multi-core Nystrom method, which can significantly reduce the high computational complexity of traditional multi-core adaptive filters and the kernel parameter selection problem of traditional single-core Nystrom methods, and can effectively improve the performance of adaptive filtering algorithms.
[0006] The technical solution of this invention is: an adaptive filtering method based on the multi-kernel Nystrom method, which can be used in nonlinear adaptive signal processing applications, such as time series prediction, nonlinear regression, nonlinear channel equalization, indoor positioning and other applications;
[0007] Step 1: Define a sequence sample set Let d(i) represent the input signal vector of dimension D at time i to be filtered, d(i) represent the desired filtered signal at time i, and N represent the total length of the training data. The input vector is mapped to a high-dimensional feature space. In the process, feature vectors are generated. Define the kernel matrix in a kernel adaptive filter. Represented as K = Φ T Φ, where Φ is the eigenvector of the reproducing kernel Hilbert space, is composed of the following form:
[0008]
[0009] Step 2: Sample the feature vectors to obtain a subset of length M, denoted as M. Where s(i) is the index value, and M represents the number of sampled data points; then an approximate matrix is constructed. Then the sum matrix is obtained as follows:
[0010]
[0011] Where C=[κ(u,u(s(j)))] N×M , Represented as The pseudoinverse of the sum matrix, κ(u,u(s(j))), represents the value of the kernel function. Eigenvalue decomposition of the sum matrix yields the following form:
[0012] K≈CVΛ -1 V T C T (3)
[0013] Wherein, CVΛ represents The eigenvector matrix, Λ=diag(λ(1),λ(2),...,λ(M)), represents the diagonal matrix of eigenvalues, where λ(j) are the eigenvalues arranged in descending order; the approximation result of the kernel function is:
[0014]
[0015] Where C(i) = [κ(u(i),u(s(1)),…,κ(u(i),u(s(M))]; z(i) represents the transformed input vector, which is:
[0016]
[0017] Step 3: Establish the multi-core adaptive filter as follows:
[0018]
[0019] in κ represents the expansion coefficient. l (u(j),u(i)) represents the value of the l-th kernel function, 1≤l≤n, n>1 represents the number of kernel functions used, and the result is obtained by using the Nestlé approximation with formula (5):
[0020]
[0021] The input Z(i) for constructing the multi-core Nestlé is:
[0022]
[0023] The predicted output f(z(i)) of the kernel adaptive filter based on multi-kernel Nystrom is obtained as follows:
[0024]
[0025] in, It is the weight vector in the feature space of the multi-kernel Nesteron map;
[0026] Step 4: Define the loss function of the multi-kernel Nestlé adaptive filter as follows:
[0027]
[0028] Where e(i) is the prediction error of the predicted output f(z(i)) at time i, θ>0, and p>0 are two fixed constant parameters; when p=2, the prediction result of the kernel adaptive filter is Here, Ω is the weight vector of the reproducing kernel Hilbert space;
[0029] Step 5: Use a semi-quadratic optimization method to transform Equation 10 into a globally convex function, and then minimize Equation 10 to obtain:
[0030]
[0031] in, The auxiliary variable introduced by the semi-square optimization method; Equation 11 is a globally convex function;
[0032] Step 6: Regenerate the kernel Hilbert space using formulas 5 and 8. Transforming into Z(i), the loss function based on multi-kernel Nestlé is:
[0033]
[0034] Formula 12 is a positive definite quadratic function, which can be rewritten in the following form:
[0035]
[0036] Where R is the weighted autocorrelation matrix of Z(i), and b is the weighted correlation vector of Z(i) and d(i). According to the exponential decay method and the conjugate gradient method, R and b have the following forms:
[0037]
[0038] in, λ is the forgetting factor, which is expressed as W according to the iterative optimization form of the weights:
[0039] W(i+1)=W(i)+α(i)p(i) (15)
[0040] Where p(i) is the direction vector:
[0041] p(i+1)=r(i+1)+β(i)p(i) (16)
[0042] The residual vector r(i) is:
[0043] r(i+1)=b(i+1)-R(i+1)W(i+1)
[0044] =λb(i)+ζ(i+1)d(i+1)Z(i+1)-(λR(i)+ζ(i+1)Z(i+1)Z(i+1) T (W(i)+α(i)p(i))
[0045] =λr(i)-α(i)R(i+1)p(i)+Z(i+1)ζ(i+1)e(i+1) (17)
[0046] In the above formula, α(i) and β(i) are step size parameters, which are:
[0047]
[0048]
[0049] Step 7: Calculate formula 13 to obtain the filtered output.
[0050] The multi-kernel Nightstrom method combines multiple kernel functions, including not only Gaussian kernels but also other different kernel functions, to map the original data into multiple independent feature spaces. Compared to the traditional single-kernel Nightstrom method, the multi-kernel Nightstrom method offers superior filtering accuracy and becomes less sensitive to the choice of kernel parameters. Compared to multi-kernel adaptive filtering methods, it significantly reduces computational complexity, time, and memory consumption. This invention then introduces the proposed multi-kernel Nightstrom method into kernel adaptive filters for the first time. Experiments show that this algorithm achieves better filtering accuracy than current state-of-the-art kernel adaptive filtering algorithms with relatively low computational complexity. Attached Figure Description
[0051] Figure 1 The curves show the steady-state MSE and average time consumption of the MNKGCCG algorithm proposed in the MG time series prediction experiment as a function of the number of algorithm dictionary M.
[0052] Figure 2 The parameter settings and learning curves of the MNGCCG algorithm compared with other algorithms in the MG time series prediction experiment are shown.
[0053] Figure 3 In the Lorenz time series prediction experiment, the parameter settings and learning curves of the MNGCCG algorithm were compared with those of other algorithms. Detailed Implementation
[0054] This invention verifies the superiority of the aforementioned multi-core Nestlé algorithm through time series experiments, including prediction experiments on Mackey-Glass (MG) and Lorenz time series under mixed noise conditions. The noise was generated through a mixture model simulation.
[0055] n(i)=(1-c(i))n1(1)+c(i)n2(i), (20)
[0056] Here, c(i) represents a binomial distribution sequence with probability models Pr(c(i)=1)=0.2 and Pr(c(i)=0)=0.8. n1(i) is Gaussian noise with a variance of 0.02, and n2(i) represents impulse noise used to generate larger outliers, simulated by an α-stable distribution with parameters set to P. α = [0.9, 0, 0.1, 0]. Representative algorithms including KRLS, RFFKRLS, KRMC, MRFGCG1, and NKGCCG (single-core form of MNKGCCG) were selected for performance comparison. In the Nestlé method, the first M terms of the training set were sampled. The logarithmic dB mean square error (MSE) was defined as a quantitative metric for filtering error performance as follows:
[0057]
[0058] Here, L represents the number of samples in the test set. This is the predicted value. In the experiment, we sampled the average of 50 Monte Carlo experiments to obtain the final MSE result. The experiment was run on a Matlab 2020b platform (Windows 10, 3.60GHz CPU and 8GB RAM). All kernel adaptive filtering algorithms used the following Gaussian kernel:
[0059]
[0060] Here, h represents the kernel width. As a multi-core algorithm, we use two different kernel widths (h1 and h2) to achieve ideal performance. All algorithms in both sequence prediction experiments underwent multiple experiments to select the most suitable parameters to achieve the highest filtering accuracy.
[0061] Input: Training data
[0062] The parameters of the loss function are θ, p > 0, the forgetting factor λ > 0, and the kernel width {h1,…,h}. n}
[0063] Estimation: Sampling to obtain a subset of data
[0064] Constructing an approximate matrix
[0065] Calculate the eigenvalue matrix Λ l and eigenvector matrix V l
[0066] Calculate the mapping output:
[0067] The first auxiliary variable,
[0068] The correlation matrix R(1) = ζ(1)Z(1)Z(1) T Cross-correlation vector b(1)=ζ(1)d(1)Z(1), weight W(1)=0, residual vector r(1)=b(1), direction vector p(1)=r(1)\\
[0069] Iterate through {i = 1, 2, 3, ..., N}:
[0070] Calculate the mapped input:
[0071]
[0072]
[0073] Error update: e(i+1) = d(i+1) - W(i) T Z(i+1)
[0074] Auxiliary variable update:
[0075]
[0076] Autocorrelation matrix: R(i+1)=λR(i)+ζ(i+1)Z(i+1)Z(i+1) T
[0077] Calculate the step size:
[0078] Update weights: W(i+1) = W(i) + α(i)p(i)
[0079] Calculate the residual vector: r(i+1)=λr(i)-α(i)R(i+1)p(i)+ζ(i+1)Z(i+1)e(i+1)
[0080] Calculate the step size; Update direction vector: p(i+1) = r(i+1) + β(i)p(i)
[0081] Iteration complete.
[0082] A. Validation results of Mackey-Glass time series prediction
[0083] The Mackey-Glass time series can be represented by the following time-delay ordinary differential equation:
[0084]
[0085] The dataset has a sampling period of 6 seconds. We use the data from the first 7 time steps, u(i) = [x(i-7), x(i-6), ..., x(i-1)] T To predict the output d(i) = x(i) at that time. The training set includes 2000 training sets contaminated by the noise mentioned above (20) and another 200 test sets. The Gaussian kernel width h in KRLS, RFFKRLS, KRMC and NKGCCG is set to 1. For multi-core algorithms including MRFGCG1 and the proposed MNKGCCG, the kernel width is set to h1 = 0.5 and h2 = 1.5. The parameters of the MNKGCCG algorithm are set to θ = 20, p = 1.8, λ = 0.999, based on the following Figure 1 The dictionary size for MNKGCCG is set to 100 to achieve a balance between filtering accuracy and computational cost. For fairness, the dictionary size for all kernel adaptive filtering algorithms using sparse methods is set to M=100. All parameters and learning curves are as follows: Figure 2 As shown in the figure, Table 1 lists the detailed simulation results of all algorithms.
[0086] According to Table 1 and Figure 2 We can observe that:
[0087] The MNKGCCG algorithm achieved the lowest steady-state MSE among all algorithms.
[0088] Compared to cross-entropy loss and mean square error kernel adaptive filtering algorithms, it has better robustness. Moreover, in the experiment, the kernel adaptive filter based on mean square error showed serious instability in non-Gaussian noise environment.
[0089] For the sparse kernel adaptive filtering algorithm, the proposed MNKGCCG algorithm achieves significantly better filtering performance in terms of convergence and filtering accuracy, without generating much additional complexity.
[0090] Table 1. Detailed simulation results of the MNGCCG algorithm and other comparison algorithms in the MG time series prediction experiment.
[0091] algorithm Mean squared error (dB) Number of dictionaries M Time elapsed (seconds) KRLS -3.79 2000 45.85 RFFKRLS -0.25 100 0.93 KRMC -29.46 2000 43.58 MRFGCG1 -30.98 100 1.21 NKGCCG -31.43 100 0.79 MNKGCCG -33.83 100 1.36
[0092] B. Validation results of Mackey-Glass time series prediction
[0093] The Lorenz attractor, known for its butterfly shape, is a nonlinear, three-dimensional, deterministic dynamic system. The following differential equation describes how the Lorenz system evolves over time in a complex, non-repeating pattern:
[0094]
[0095] In the simulation experiment, the system parameters are set as μ = 8 / 3, η = 10, and ρ = 28. We sample x as the short-term prediction sample, using the data from the first 6 time steps: u(i) = [x(i-6), x(i-5), ..., x(i-1)] T To predict the current output d(i) = x(i). The training set includes 2000 training sets contaminated with the noise described above (20) and another 200 test sets. The Gaussian kernel width in KRLS, RFFKRLS, KRMC and NKGCCG is set to h = 3. For multi-core algorithms including MRFGCG1 and the proposed MNKGCCG, the kernel width is set to h1 = 2 and h1 = 4 respectively. All parameters and learning curves are as follows Figure 3 As shown. Figure 3 As shown, the proposed MNKGCCG achieves the highest filtering accuracy among all comparison algorithms.
Claims
1. An adaptive filtering method based on the multi-kernel Nystrom method, used for nonlinear adaptive signal processing, comprising: Step 1: Define a sequence sample set Let d(i) represent the input signal vector of dimension D at time i to be filtered, d(i) represent the desired filtered signal at time i, and N represent the total length of the training data. The input vector is mapped to a high-dimensional feature space. In the process, feature vectors are generated. Define the kernel matrix in a kernel adaptive filter. Represented as K = Φ T Φ, where Φ is the eigenvector of the reproducing kernel Hilbert space, is composed of the following form: Step 2: Sample the feature vectors to obtain a subset of length M, denoted as M. Where s(i) is the index value, and M represents the number of sampled data points; then an approximate matrix is constructed. Then the sum matrix is obtained as follows: Where C=[κ(u,u(s(j)))] N×M , Represented as The pseudoinverse of the sum matrix, [κ(u,u(s(j)))], represents the value of the kernel function. Eigenvalue decomposition of the sum matrix yields the following form: K≈CVΛ -1 In T C T (3) Wherein, CVΛ represents The eigenvector matrix, Λ=diag(λ(1),λ(2),...,λ(M)), represents the diagonal matrix of eigenvalues, where λ(j) are the eigenvalues arranged in descending order; the approximation result of the kernel function is: Where C(i) = [κ(u(i),u(s(1)),…,κ(u(i),u(s(M))]; z(i) represents the transformed input vector, which is: Step 3: Establish the multi-core adaptive filter as follows: in κ represents the expansion coefficient. l (u(j),u(i)) represents the value of the l-th kernel function, 1≤l≤n, n>1 represents the number of kernel functions used, and the result is obtained by using the Nestlé approximation with formula (5): The input Z(i) for constructing the multi-core Nestlé is: The predicted output f(z(i)) of the kernel adaptive filter based on multi-kernel Nystrom is obtained as follows: in, It is the weight vector in the feature space of the multi-kernel Nesteron map; Step 4: Define the loss function of the multi-kernel Nestlé adaptive filter as follows: Where e(i) is the prediction error of the predicted output f(z(i)) at time i, θ>0, and p>0 are two fixed constant parameters; when p=2, the prediction result of the kernel adaptive filter is Here, Ω is the weight vector of the reproducing kernel Hilbert space; Step 5: Use a semi-quadratic optimization method to transform Equation 10 into a globally convex function, and then minimize Equation 10 to obtain: in, The auxiliary variable introduced by the semi-square optimization method; Equation 11 is a globally convex function; Step 6: Regenerate the kernel Hilbert space using formulas 5 and 8. Transforming into Z(i), the loss function based on multi-kernel Nestlé is: Formula 12 is a positive definite quadratic function, which can be rewritten in the following form: Where R is the weighted autocorrelation matrix of Z(i), and b is the weighted correlation vector of Z(i) and d(i). According to the exponential decay method and the conjugate gradient method, R and b have the following forms: in, λ is the forgetting factor, which is expressed as W according to the iterative optimization form of the weights: W(i+1)=W(i)+α(i)p(i) (15) Where p(i) is the direction vector: p(i+1)=r(i+1)+β(i)p(i) (16) The residual vector r(i) is: r(i+1)=b(i+1)-R(i+1)W(i+1) =λb(i)+ζ(i+1)d(i+1)Z(i+1)-(λR(i)+ζ(i+1)Z(i+1)Z(i+1) T )(W(i)+α(i)p(i)) =λr(i)-α(i)R(i+1)p(i)+Z(i+1)ζ(i+1)e(i+1) (17) In the above formula, α(i) and β(i) are step size parameters, which are: Step 7: Calculate formula 13 to obtain the filtered output.