A speech signal denoising method based on improved variational mode decomposition and principal component analysis

By combining VMD and PCA, the limitations of traditional speech denoising methods in nonlinear and non-stationary signal processing are overcome, achieving more efficient and accurate speech signal denoising and ensuring signal integrity.

CN113851144BActive Publication Date: 2025-09-30SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111159300.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-30
Publication Date
2025-09-30
Estimated Expiration
2041-09-30

AI Technical Summary

Technical Problem

Existing speech denoising methods have limitations when processing nonlinear and non-stationary signals, resulting in signal distortion or noise residue. In addition, the traditional EMD decomposition method lacks a unified standard and may eliminate valid signals.

Method used

Variational mode decomposition (VMD) combined with principal component analysis (PCA) is adopted. The residual noise after VMD decomposition is eliminated by adding Gaussian white noise. The component category is determined by the correlation coefficient distribution diagram, and PCA is used to reduce noise and reconstruct the signal.

Benefits of technology

The accuracy and efficiency of speech signal decomposition are improved, ensuring that effective signals are not mistakenly removed, the reconstructed signal is more accurate, and noise interference is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113851144B_ABST
    Figure CN113851144B_ABST
Patent Text Reader

Abstract

The present invention relates to a speech signal denoising method based on improved variational modal decomposition and principal component analysis, comprising the following steps: S1: selecting a noisy speech signal as a sample; S2: decomposing the noisy speech signal to obtain K IMF modal components; S3: calculating the correlation coefficient between each IMF modal component and the original noisy speech signal, drawing a correlation coefficient distribution graph, and determining the false components and noise-dominated IMF modal components from the correlation coefficient distribution graph; S4: after removing the false components and the noise-dominated IMF modal components, the remaining IMF modal components are recorded as signal-dominated IMF modal components; S5: removing residual noise from the noise-dominated IMF modal components; and S6: reconstructing the principal component components of the noise-dominated IMF modal components and the signal-dominated IMF modal components to obtain a denoised speech signal. The present invention eliminates the problem of residual noise in the reconstructed signal after VMD decomposition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of signal processing, and in particular relates to a method for denoising a speech signal. Background Art

[0002] Speech signals are inevitably subject to various interferences during the acquisition and transmission process, which will make the acquired speech signals less accurate and not conducive to subsequent analysis. Therefore, speech denoising becomes the most critical step in the speech signal processing process.

[0003] There are many traditional methods for speech denoising. Spectral subtraction-based speech denoising assumes that speech signals are short-term stationary. However, speech signals are inherently nonlinear and nonstationary. Spectral subtraction has certain limitations and can introduce a new type of background noise. Wavelet threshold-based speech denoising relies heavily on the selection of the threshold function. However, hard thresholding can produce oscillations in the reconstructed signal, while soft thresholding can produce distortion. Empirical mode decomposition (EMD), proposed by Huang et al., is a method for processing nonlinear and nonstationary signals. It decomposes the signal into a finite number of intrinsic mode function components (IMFs) and a residual, arranged in descending frequency order. Based on the characteristics of the signal being processed, components that do not conform to the signal characteristics can be removed, while other components that do conform to the signal characteristics can be processed. The remaining processed components are then reconstructed by stacking to obtain the denoised signal. There is no unified standard for selecting IMF modal components obtained by the conventional EMD decomposition method. It is usually assumed that the noise signal is dominant in the high-frequency IMF modal components and is discarded. However, this will cause the effective signal to be eliminated, resulting in distortion of the reconstructed signal. At the same time, the extreme points and envelope lines in the EMD decomposition method cannot be accurately determined, which will produce IMF modal components containing false frequency components. If these components are not eliminated, the reconstructed signal will be inaccurate. Summary of the Invention

[0004] In view of the shortcomings of the existing technology, the present invention proposes an improved speech signal denoising method, which is a technology combining variational mode decomposition (VMD) and principal component analysis (PCA).

[0005] This method eliminates the problem of residual noise in the reconstructed signal after VMD decomposition by adding Gaussian white noise to the original signal. VMD is used to decompose the original speech signal, and the correlation coefficients between each modal component and the original signal are calculated and plotted. Modal components are then classified into three categories: invalid components, signal components, and noise components, using a modal component judgment criterion. The invalid components are directly eliminated, while the signal components are retained. The noise components are then reconstructed with the signal components after subsequent PCA denoising to obtain the final denoised speech signal.

[0006] Explanation of terms:

[0007] 1. VMD decomposition, or variational mode decomposition, is an adaptive, fully non-recursive method for modal variation and signal processing. This technique has the advantage of being able to determine the number of modal decompositions. Its adaptability is reflected in the fact that it determines the number of modal decompositions for a given sequence based on actual conditions. The subsequent search and solution process can adaptively match the optimal center frequency and finite bandwidth of each mode. It can also effectively separate the intrinsic mode components (IMFs) and partition the signal in the frequency domain, thereby obtaining the effective decomposition components of the given signal and ultimately achieving the optimal solution to the variational problem.

[0008] 2. EMD, empirical mode decomposition, is a new adaptive signal time-frequency processing method creatively proposed by Huang E et al. in 1998. It is particularly suitable for the analysis and processing of nonlinear and non-stationary signals.

[0009] The technical solution of the present invention is:

[0010] A speech signal denoising method based on improved variational mode decomposition and principal component analysis comprises the following steps:

[0011] S1: Select a noisy speech signal y(t) as a sample;

[0012] S2: Use the improved VMD method to decompose the noisy speech signal y(t) and obtain K IMF modal components;

[0013] S3: Calculate the correlation coefficient between each IMF modal component and the original noisy speech signal, draw a correlation coefficient distribution graph, determine the false component from the correlation coefficient distribution graph based on the false component judgment principle, and determine the noise-dominated IMF modal component from the correlation coefficient distribution graph based on the noise component judgment principle;

[0014] S4: After removing the false components and noise-dominated IMF modal components, the remaining IMF modal components are recorded as signal-dominated IMF modal components;

[0015] S5: For the noise-dominated IMF modal components, the principal component analysis (PCA) method is adopted to select a certain number of principal component components according to the cumulative contribution rate for reconstruction to remove the residual noise in the noise-dominated IMF modal components;

[0016] S6: Reconstruct the principal component components of the noise-dominated IMF modal components and the signal-dominated IMF modal components after principal component analysis (PCA) to obtain a speech signal with noise removed.

[0017] According to a preferred embodiment of the present invention, the specific implementation process of step S2 includes:

[0018] S2-1: Set VMD decomposition parameters, including the optimal number of decomposition levels and the modal component frequency bandwidth control parameter α;

[0019] S2-2: Construct a constrained variational model, introduce the Lagrangian function, and construct the augmented Lagrangian equation;

[0020] S2-3: Solve the augmented Lagrange equation, initialize the component frequency, and obtain the initial component frequency u^ k 1 , with u^ k 1 The corresponding initial center frequency ω^ k 1 , the initial Lagrange multiplier λ^ k 1 ;

[0021] S2-4: Update the component frequency u^ according to the VMD algorithm formula k , center frequency ω^ k ;

[0022] S2-5: After each update of the component frequency u^ k , center frequency ω^ k Afterwards, update the Lagrange multiplier λ^;

[0023] S2-6: Determine whether the component frequency after iterative update satisfies the convergence equation. If not, continue the iteration while adding Gaussian white noise with gradually decreasing noise intensity, and continue to execute steps S2-S5. If the convergence equation is satisfied, end the iteration and obtain the modal components that complete the VMD decomposition.

[0024] Further preferably, in step S2-1, the method for setting the optimal number of decomposition layers is as follows:

[0025] Perform EMD decomposition on the original noisy speech signal. Assume that the number of decomposition layers is K. After decomposition, K modal components are obtained. Calculate the correlation coefficient between each modal component and the original noisy speech signal. Select the modal component IMF with the largest correlation coefficient. max , calculate its kurtosis and record it as λ, then add 1 to each decomposition level, and record the kurtosis of the modal component with the largest correlation coefficient when the decomposition level is K+1 as λ', and continue to iterate until a λ<λ' appears at some point. At this time, the decomposition level corresponding to λ is the optimal decomposition level. The calculation formula of kurtosis H is shown in formula (I):

[0026]

[0027] In formula (I), IMF i (t) is the i-th modal component, μ i is the mean of the i-th modal component, σi is the standard deviation of the i-th modal component.

[0028] Further preferably, the modal component frequency bandwidth control parameter α is set to 2000.

[0029] Further preferably, in step S2-2, the constrained variational model, i.e., the VMD constrained model, is expressed as shown in formula (II):

[0030]

[0031] In formula (II), δ(t) is the unit impulse function, K is the number of VMD decomposition layers, {u k}={u1,u2,......,u k} is the set of all IMF components, {ω k}={ω1,ω2,......,ω k} is the set of center frequencies of each modal component, and j is an imaginary unit.

[0032] Further preferably, in step S2-3, the augmented Lagrangian equation L is as shown in formula (III):

[0033]

[0034] In formula (III), α is the frequency bandwidth control parameter of the modal component, λ is the Lagrange multiplier, ω k is the center frequency of the kth modal component.

[0035] Further preferably, in step S2-4, the updating formula of the modal component frequency is shown in formula (IV):

[0036]

[0037] In formula (IV), x(ω) is the frequency domain form of the signal x(t), λ^(ω) is the frequency domain form of the Lagrangian operator λ(t), the superscript ^ indicates the conjugate form, and n is the number of iterations;

[0038] The update formula of the center frequency corresponding to the IMF component is shown in formula (V):

[0039]

[0040] In formula (V), u^ k (ω) is the frequency of the kth IMF modal component.

[0041] Further preferably, in step S2-5, the updating formula of the Lagrange multiplier λ is as shown in formula (VI):

[0042]

[0043] In formula (VI), τ is the update parameter of the Lagrange multiplier, τ = 10 -3 .

[0044] Further preferably, in step S2-6, the convergence equation is as shown in formula (VII):

[0045]

[0046] In formula (VII), ε is the convergence criterion tolerance value, and ε is 10 -6 In step S2-6, the decomposed modal components are recorded as IMF1, IMF2, ..., IMF m .

[0047] Further preferably, in step S2-6, the specific method of adding Gaussian white noise with gradually decreasing noise intensity is: adding noise whose amplitude distribution obeys Gaussian distribution and power spectral density distribution obeys uniform distribution to the modal components whose component frequencies do not satisfy the convergence equation, and the noise intensity should follow the principle of gradually decreasing, that is, the noise intensity added later should be lower than the noise intensity added previously.

[0048] According to the preferred embodiment of the present invention, in step S2-1 and step S3, the correlation coefficient ρ xy The calculation formula is shown in formula (VIII):

[0049]

[0050] In formula (VIII), x(i) is the signal whose correlation coefficient is to be calculated, and y(i) is the original signal.

[0051] According to the preferred embodiment of the present invention, in step S3, the false component is determined from the correlation coefficient distribution diagram according to the false component judgment principle, specifically: find the first point with a correlation coefficient less than h from the correlation coefficient distribution diagram, and record the modal component corresponding to this point as IMF h , h refers to the correlation coefficient, which ranges from 0.10 to 0.15. h+1 ~IMF k Recorded as false quantity.

[0052] More preferably, h=0.15.

[0053] According to the preferred embodiment of the present invention, in step S3, the noise-dominated IMF modal component is determined from the correlation coefficient distribution graph according to the noise component judgment principle, specifically: after removing the false component, the correlation coefficient distribution curve is redrawn, and the first turning point on the curve is found and recorded as p. The IMF modal component corresponding to this point is recorded as IMF p , IMF1~IMF p Denoted as the noise-dominated IMF modal component.

[0054] Furthermore, the specific implementation process of step S5 is as follows:

[0055] S5-1: Extract m eigenvalues ​​M from a noise-dominated modal component i , i=1,2,...,m eigenvalue M i The dimension is n{M i1 ,M i2 ,...,M ij}, j = 1, 2, ..., n;

[0056] S5-2: Create a sample matrix A for the eigenvalues mn , that is, by The normalized m×n matrix A mn As a sample matrix, The formula for obtaining is shown in formula (VIII):

[0057]

[0058] In formula (VIII), M ij The standardized value of μ j is the sample mean of the jth component, s j is the sample standard deviation of the jth component,

[0059] S5-3: Calculate the covariance matrix B of the normalized matrix, as shown in formula (IX):

[0060]

[0061] In formula (IX), the covariance matrix B is also called matrix A mn The correlation coefficient matrix of

[0062] S5-4: Calculate the eigenvalue λ of the covariance matrix B and the eigenvector p corresponding to the eigenvalue, and rearrange the eigenvalues ​​in descending order as λ1≥λ1≥...≥λ a , and its corresponding eigenvector is p i , i=1,2,...,a, the eigenvectors are orthogonal to each other, and the eigenvectors form a matrix P=(p1,p2,......,p n );

[0063] S5-5: Let Y = P T B, Y=(y1,y2,...,y n ) T , where y1,y2,...,y i,...are unrelated to each other, and are called y1,y2,...,y n are the 1st, 2nd, ..., i-th... principal component variables respectively;

[0064] S5-6: Select the first p principal component variables and calculate the principal component cumulative contribution rate through their corresponding eigenvalues, as shown in formula (X):

[0065]

[0066] In formula (X), α p is the cumulative contribution rate of the first p principal component variables;

[0067] S5-7: Select the principal component variables with a cumulative contribution rate of more than 85% to reconstruct a new modal component IMF PCA , and reconstructed with the signal-dominated IMF modal component obtained in step S4 to generate a new signal, which is the speech signal with noise removed.

[0068] The reconstruction formula is shown in formula (Ⅺ):

[0069] d(t) = IMF 信号 (t)+IMF PCA (t)(Ⅺ)

[0070] In formula (XI), IMF 信号 (t) is the signal-dominated modal component, IMF PCA (t) is the modal component reconstructed from the principal component variables with a cumulative contribution rate of more than 85%.

[0071] The beneficial effects of the present invention are:

[0072] 1. The present invention uses VMD (Variational Mode Decomposition) method to process nonlinear and non-stationary signals such as speech signals. Compared with other methods, VMD decomposition method is more adaptable and can reduce the complexity and non-stationarity of the signal, which has a significant effect on the processing of speech signals.

[0073] 2. The present invention improves the method for determining the optimal number of decomposition layers. Compared with traditional methods, it improves the accuracy of VMD decomposition, makes the decomposition results more accurate, reduces the decomposition time, and improves the decomposition efficiency.

[0074] 3. The present invention uses a correlation coefficient distribution diagram to determine the noise-dominant component, the signal-dominant component, and the invalid component. Compared with other methods, the present invention will not cause useful components to be mistakenly removed.

[0075] 4. The present invention eliminates the endpoint effect and modal aliasing produced by the VMD decomposition method by adding Gaussian white noise. Compared with other methods, the decomposition of the modal components of the present invention is more accurate.

[0076] 5. The present invention uses the principal component analysis (PCA) method to reduce the noise-dominated components and uses the reduced-noise signal for signal reconstruction. Compared with other methods, the reconstructed signal of the present invention is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 It is a flow chart of VMD decomposition and modal component classification of the present invention.

[0078] Figure 2 4 is a flow chart of the principal component analysis process of the present invention.

[0079] Figure 3 It is the waveform of the original speech signal.

[0080] Figure 4 This is the waveform of the added noise.

[0081] Figure 5 It is the waveform of the speech signal after adding noise.

[0082] Figure 6 It is the waveform of the speech signal after denoising. DETAILED DESCRIPTION

[0083] The present invention will be further described below with reference to the accompanying drawings and specific implementation examples, but is not limited thereto.

[0084] Example 1

[0085] A speech signal denoising method based on improved variational mode decomposition and principal component analysis comprises the following steps:

[0086] S1: Select a noisy speech signal as a sample;

[0087] S2: Use the improved VMD method to decompose the noisy speech signal and obtain K IMF modal components;

[0088] S3: Calculate the correlation coefficient between each IMF modal component and the original noisy speech signal, draw a correlation coefficient distribution graph, determine the false component from the correlation coefficient distribution graph based on the false component judgment principle, and determine the noise-dominated IMF modal component from the correlation coefficient distribution graph based on the noise component judgment principle;

[0089] S4: After removing the false components and noise-dominated IMF modal components, the remaining IMF modal components are recorded as signal-dominated IMF modal components;

[0090] S5: For the noise-dominated IMF modal components, the principal component analysis (PCA) method is adopted to select a certain number of principal component components according to the cumulative contribution rate for reconstruction to remove the residual noise in the noise-dominated IMF modal components;

[0091] S6: Reconstruct the principal component components of the noise-dominated IMF modal components and the signal-dominated IMF modal components after principal component analysis to obtain a speech signal with noise removed.

[0092] like Figure 1 The figure shows a flow chart of the principle of the present invention, wherein x(t) is the collected speech signal containing noise. After VMD decomposition, n modal components are obtained, where n is an integer greater than or equal to 2. The correlation coefficient between each IMF modal component and x(t) is calculated, and a correlation coefficient distribution diagram is drawn. The IMF modal components are classified into two categories: noise-dominated and signal-dominated according to the judgment criteria. The noise-dominated components are further decomposed by the PCA method, and the principal component with the largest contribution rate is reconstructed to obtain the denoised components. Finally, the denoised noise-dominated IMF modal components and the signal-dominated IMF modal components are reconstructed to obtain the denoised speech signal y(t).

[0093] Figure 3 It is the waveform of the original speech signal. Figure 4 This is the waveform of the added noise. Figure 5 It is the waveform of the speech signal after adding noise. Figure 6 It is a waveform diagram of the speech signal after denoising according to the present invention.

[0094] Figure 5 and Figure 6 By comparison, we can see that most of the noise interference has been removed. Figure 3 and Figure 6 By comparison, we can see that the denoised signal has some noise at the initial position, while other positions are close to the original signal, and the denoising effect is better.

[0095] Example 2

[0096] The speech signal denoising method based on improved variational mode decomposition and principal component analysis described in Example 1 is different in that:

[0097] The specific implementation process of step S2 includes:

[0098] like Figure 2 As shown, S2-1: setting VMD decomposition parameters, including the optimal decomposition level and the modal component frequency bandwidth control parameter α;

[0099] In step S2-1, the improvement of the VMD decomposition method in the present invention is to improve the method for determining the optimal number of decomposition levels. The method for setting the optimal number of decomposition levels is as follows:

[0100] Perform EMD decomposition on the original noisy speech signal. Assume that the number of decomposition layers is K. After decomposition, K modal components are obtained. Calculate the correlation coefficient between each modal component and the original noisy speech signal. Select the modal component IMF with the largest correlation coefficient. max , calculate its kurtosis and record it as λ, then add 1 to each decomposition level, and record the kurtosis of the modal component with the largest correlation coefficient when the decomposition level is K+1 as λ', and continue to iterate until a λ<λ' appears at some point. At this time, the decomposition level corresponding to λ is the optimal decomposition level. The calculation formula of kurtosis H is shown in formula (I):

[0101]

[0102] In formula (I), IMF i (t) is the i-th modal component, μ i is the mean of the i-th modal component, σ i is the standard deviation of the ith modal component. Since it is a speech signal, the modal component frequency bandwidth control parameter α is set to 2000.

[0103] S2-2: Construct a constrained variational model, introduce the Lagrangian function, and construct the augmented Lagrangian equation;

[0104] In step S2-2, the VMD decomposition process can be regarded as the construction and solution of the constrained variation problem. The constrained variation model, i.e., the VMD constraint model, is expressed as shown in formula (II):

[0105]

[0106] In formula (II), δ(t) is the unit impulse function, K is the number of VMD decomposition layers, {u k}={u1,u2,......,u k} is the set of all IMF modal components, {ω k}={ω1,ω2,......,ω k} is the set of center frequencies of each modal component, and j is an imaginary unit.

[0107] S2-3: Solve the augmented Lagrange equation, initialize the component frequency, and obtain the initial component frequency u^ k 1 , with u^ k 1 The corresponding initial center frequency ω^ k 1 , the initial Lagrange multiplier λ^ k1 ;

[0108] In step S2-3, the augmented Lagrangian equation L is shown in formula (III):

[0109]

[0110] In formula (III), α is the frequency bandwidth control parameter of the modal component, λ is the Lagrange multiplier, ω k is the center frequency of the kth modal component.

[0111] S2-4: Update the component frequency u^ according to the VMD algorithm formula k , center frequency ω^ k ;

[0112] In step S2-4, the updating formula of the modal component frequency is shown in formula (IV):

[0113]

[0114] In formula (IV), x(ω) is the frequency domain form of the signal x(t), λ^(ω) is the frequency domain form of the Lagrangian operator λ(t), the superscript ^ indicates the conjugate form, and n is the number of iterations;

[0115] The update formula of the center frequency corresponding to the IMF component is shown in formula (V):

[0116]

[0117] In formula (V), u^ k (ω) is the frequency of the kth IMF modal component.

[0118] S2-5: After each update of the component frequency u^ k , center frequency ω^ k Afterwards, update the Lagrange multiplier λ^;

[0119] In step S2-5, the update formula of the Lagrange multiplier λ is shown in formula (VI):

[0120]

[0121] In formula (VI), τ is the update parameter of the Lagrange multiplier, τ = 10 -3 .

[0122] S2-6: Determine whether the component frequency after iterative update satisfies the convergence equation. If not, continue the iteration while adding Gaussian white noise with gradually decreasing noise intensity, and continue to execute steps S2-S5. If the convergence equation is satisfied, end the iteration and obtain the modal components that complete the VMD decomposition.

[0123] In step S2-6, the convergence equation is shown in formula (VII):

[0124]

[0125] In formula (VII), ε is the convergence criterion tolerance value, and ε is 10 -6 In step S2-6, the decomposed modal components are recorded as IMF1, IMF2, ..., IMF m .

[0126] In step S2-6, the specific method of adding Gaussian white noise with gradually decreasing noise intensity is: adding noise whose amplitude distribution obeys Gaussian distribution and power spectral density distribution obeys uniform distribution to the modal components whose component frequencies do not satisfy the convergence equation, and the noise intensity should follow the principle of gradually decreasing, that is, the noise intensity added later should be lower than the noise intensity added previously.

[0127] Example 3

[0128] The speech signal denoising method based on improved variational mode decomposition and principal component analysis described in Example 2 is different in that:

[0129] In step S2-1 and step S3, the correlation coefficient ρ xy The calculation formula is shown in formula (VIII):

[0130]

[0131] In formula (VIII), x(i) is the signal whose correlation coefficient is to be calculated, and y(i) is the original signal.

[0132] In step S3, the false component is determined from the correlation coefficient distribution diagram according to the false component judgment principle, specifically: find the first point with a correlation coefficient less than h from the correlation coefficient distribution diagram, and record the modal component corresponding to this point as IMF h , h refers to the correlation coefficient, h = 0.15, IMF h+1 ~IMF k Recorded as false quantity.

[0133] In step S3, the noise-dominated IMF modal component is determined from the correlation coefficient distribution diagram according to the noise component judgment principle. Specifically, the correlation coefficient distribution curve is redrawn after removing the false component, and the first turning point on the curve is found and recorded as p. The IMF modal component corresponding to this point is recorded as IMF p , IMF1~IMF p Denoted as the noise-dominated IMF modal component.

[0134] Example 4

[0135] The speech signal denoising method based on improved variational mode decomposition and principal component analysis according to embodiment 2 or 3 is different in that:

[0136] The specific implementation process of step S5 is as follows:

[0137] S5-1: Extract m eigenvalues ​​M from a noise-dominated modal component i , i=1,2,...,m eigenvalue M i The dimension is n{M i1 ,M i2 ,...,M ij}, j = 1, 2, ..., n;

[0138] S5-2: Create a sample matrix A for the eigenvalues mn , that is, by The normalized m×n matrix A mn As a sample matrix, The formula for obtaining is shown in formula (VIII):

[0139]

[0140] In formula (VIII), M ij The standardized value of μ j is the sample mean of the jth component, s j is the sample standard deviation of the jth component,

[0141] S5-3: Calculate the covariance matrix B of the normalized matrix, as shown in formula (IX):

[0142]

[0143] In formula (IX), the covariance matrix B is also called matrix A mn The correlation coefficient matrix of

[0144] S5-4: Calculate the eigenvalue λ of the covariance matrix B and the eigenvector p corresponding to the eigenvalue. Use the eig function in Matlab to find the eigenvalue and eigenvector. The specific steps are as follows: enter the matrix A in the command line window. mn , then input [x,y]=eig(A mn ), calculate the two matrices x and y, where each column value of x represents an eigenvector of matrix a, and the diagonal element value of y represents the eigenvalue of matrix a. Rearrange the eigenvalues ​​in descending order as λ1≥λ1≥...≥λ a , and its corresponding eigenvector is p i, i=1,2,...,a, the eigenvectors are orthogonal to each other, and the eigenvectors form a matrix P=(p1,p2,......,p n );

[0145] S5-5: Let Y = P T B, Y=(y1,y2,...,y n ) T , where y1,y2,...,y i ,...are unrelated to each other, and are called y1,y2,...,y n are the 1st, 2nd, ..., i-th... principal component variables respectively;

[0146] S5-6: Select the first p principal component variables and calculate the principal component cumulative contribution rate through their corresponding eigenvalues, as shown in formula (X):

[0147]

[0148] In formula (X), α p is the cumulative contribution rate of the first p principal component variables;

[0149] S5-7: Select the principal component variables with a cumulative contribution rate of more than 85% and reconstruct them with the signal-dominated IMF components obtained in step S4 to generate a new signal. This new signal is the speech signal with noise removed.

[0150] The reconstruction formula is shown in formula (Ⅺ):

[0151] d(t) = IMF 信号 (t)+IMF PCA (t)(Ⅺ)

[0152] In formula (XI), IMF 信号 (t) is the signal-dominated modal component, IMF PCA (t) is the modal component reconstructed from the principal component variables with a cumulative contribution rate of more than 85%.

Claims

1. A speech signal denoising method based on improved variational mode decomposition and principal component analysis, characterized in that: The steps are as follows: S1: Select a noisy speech signal y(t) as a sample; S2: Use the improved VMD method to decompose the noisy speech signal y(t) and obtain K IMF modal components; S3: Calculate the correlation coefficient between each IMF modal component and the original noisy speech signal, draw a correlation coefficient distribution graph, determine the false component from the correlation coefficient distribution graph based on the false component judgment principle, and determine the noise-dominated IMF modal component from the correlation coefficient distribution graph based on the noise component judgment principle; S4: After removing the false components and noise-dominated IMF modal components, the remaining IMF modal components are recorded as signal-dominated IMF modal components; S5: For the noise-dominated IMF modal components, the principal component analysis method is adopted to select the principal component components according to the cumulative contribution rate for reconstruction, so as to remove the residual noise in the noise-dominated IMF modal components; S6: reconstructing the principal component components of the noise-dominated IMF modal components and the signal-dominated IMF modal components after principal component analysis to obtain a speech signal with the noise removed; The specific implementation process of step S2 includes: S2-1: Set VMD decomposition parameters, including the optimal number of decomposition levels and the modal component frequency bandwidth control parameter α; S2-2: Construct a constrained variational model, introduce the Lagrangian function, and construct the augmented Lagrangian equation; S2-3: Solve the augmented Lagrange equation, initialize the frequency of the component, and obtain the initial component frequency u ^ k 1 , with u ^ k 1 The corresponding initial center frequency ω ^ k 1 , the initial Lagrange multiplier λ ^ k 1 ; S2-4: Update component frequency u according to VMD algorithm formula ^ k , center frequency ω ^ k ; S2-5: After each update of the component frequency u ^ k , center frequency ω ^ k Afterwards, update the Lagrange multiplier λ ^ ; S2-6: Determine whether the component frequency after iterative update satisfies the convergence equation. If not, continue the iteration while adding Gaussian white noise with gradually decreasing noise intensity, and continue to execute steps S2-1 to S2-5. If the convergence equation is satisfied, end the iteration and obtain the modal components of the completed VMD decomposition. In step S2-1, the method for setting the optimal number of decomposition levels is as follows: Perform EMD decomposition on the original noisy speech signal. Assume that the number of decomposition layers is K. After decomposition, K modal components are obtained. Calculate the correlation coefficient between each modal component and the original noisy speech signal. Select the modal component IMF with the largest correlation coefficient. max , calculate its kurtosis and record it as λ, then add 1 to each decomposition level, and record the kurtosis of the modal component with the largest correlation coefficient when the decomposition level is K+1 as λ', and continue to iterate until a λ<λ' appears at some point. At this time, the decomposition level corresponding to λ is the optimal decomposition level. The calculation formula of kurtosis H is shown in formula (I): In formula (I), IMF i (t) is the i-th modal component, μ i is the mean of the i-th modal component, σ i is the standard deviation of the i-th modal component.

2. The method according to claim 1, characterized in that Set the modal component frequency bandwidth control parameter α to 2000.

3. The method according to claim 2, characterized in that In step S2-2, the constrained variational model, i.e., the VMD constraint model, is expressed as shown in formula (II): In formula (II), δ(t) is the unit impulse function, K is the number of VMD decomposition layers, {u k }={u1,u2,......,u k } is the set of all IMF components, {ω k }={ω1,ω2,......,ω k } is the set of center frequencies of each modal component, and j is an imaginary unit.

4. The method according to claim 3, characterized in that In step S2-3, the augmented Lagrangian equation L is shown in formula (III): In formula (III), α is the frequency bandwidth control parameter of the modal component, λ is the Lagrange multiplier, ω k is the center frequency of the kth modal component.

5. The method according to claim 1, wherein In step S2-4, the updating formula of the modal component frequency is shown in formula (IV): In formula (Ⅳ), x(ω) is the frequency domain form of signal x(t), λ ^ (ω) is the frequency domain form of the Lagrangian operator λ(t), the superscript ^ indicates the conjugate form, and n is the number of iterations; The update formula of the center frequency corresponding to the IMF component is shown in formula (V): In formula (V), u ^ k (ω) is the frequency of the kth IMF modal component.

6. The method according to claim 5, characterized in that In step S2-5, the update formula of the Lagrange multiplier λ is shown in formula (VI): In formula (VI), τ is the update parameter of the Lagrange multiplier, τ = 10 -3 ; In step S2-6, the convergence equation is shown in formula (VII): In formula (VII), ε is the convergence criterion tolerance value, and ε is 10 -6 In step S2-6, the decomposed modal components are recorded as IMF1, IMF2, ..., IMF m ; In step S2-6, the specific method of adding Gaussian white noise with gradually decreasing noise intensity is: adding noise whose amplitude distribution obeys Gaussian distribution and power spectral density distribution obeys uniform distribution to the modal components whose component frequencies do not satisfy the convergence equation, and the noise intensity should follow the principle of gradually decreasing, that is, the noise intensity added later should be lower than the noise intensity added previously.

7. The method according to claim 1, characterized in that In step S2-1 and step S3, the correlation coefficient ρ xy The calculation formula is shown in formula (VIII): In formula (VIII), x(i) is the signal whose correlation coefficient is to be calculated, and y(i) is the original signal.

8. The method according to claim 1, characterized in that In step S3, the false component is determined from the correlation coefficient distribution diagram according to the false component judgment principle, specifically: find the first point with a correlation coefficient less than h from the correlation coefficient distribution diagram, and record the modal component corresponding to this point as IMF h , h refers to the correlation coefficient, which ranges from 0.10 to 0.

15. h+1 ~IMF k Recorded as false quantity; In step S3, the noise-dominated IMF modal component is determined from the correlation coefficient distribution diagram according to the noise component judgment principle. Specifically, the correlation coefficient distribution curve is redrawn after removing the false component, and the first turning point on the curve is found and recorded as p. The IMF modal component corresponding to this point is recorded as IMF p , IMF1~IMF p Denoted as the noise-dominated IMF modal component.

9. The method according to claim 8, characterized in that h=0.15。 10. The method according to any one of claims 1 to 9, characterized in that: The specific implementation process of step S5 is as follows: S5-1: Extract m eigenvalues ​​M from a noise-dominated modal component i , i=1,2,...,m eigenvalue M i The dimension is n{M i1 ,M i2 ,...,M ij }, j = 1, 2, ..., n; S5-2: Create a sample matrix A for the eigenvalues mn , that is, by The normalized m×n matrix A mn As a sample matrix, The formula for obtaining is shown in formula (VIII): In formula (VIII), M ij The standardized value of μ j is the sample mean of the jth component, s j is the sample standard deviation of the jth component, S5-3: Calculate the covariance matrix B of the standardized matrix, as shown in formula (IX): In formula (IX), the covariance matrix B is also called matrix A mn The correlation coefficient matrix of S5-4: Calculate the eigenvalue λ of the covariance matrix B and the eigenvector p corresponding to the eigenvalue, and rearrange the eigenvalues ​​in descending order as λ1≥λ1≥...≥λ a , and its corresponding eigenvector is p i , i=1,2,...,a, the eigenvectors are orthogonal to each other, and the eigenvectors form a matrix P=(p1,p2,......,p n ); S5-5: Let Y = P T B, Y=(y1,y2,...,y i ,...y n ) T , where y1,y2,...,y i ,...y n are unrelated to each other, and are called y1,y2,...,y i ,...y n They are the 1st, 2nd, ..., ith, ..., and nth principal component variables respectively; S5-6: Select the first p principal component variables and calculate the principal component cumulative contribution rate through their corresponding eigenvalues, as shown in formula (X): In formula (X), α p is the cumulative contribution rate of the first p principal component variables; S5-7: Select the principal component variables with a cumulative contribution rate of more than 85% to reconstruct a new modal component IMF PCA , and reconstructed with the signal-dominated IMF component obtained in step S4 to generate a new signal, which is the speech signal with noise removed.

Citation Information

Patent Citations

  • Mixed noise reduction method for underwater acoustic signals

    CN108387887A

  • Voice enhancement method by combining empirical mode decomposition with wavelet threshold denoising

    CN109785854A