Method and system for optimizing full-band output signal-to-noise ratio filter based on inter-frame correlation

By adopting an optimized full-band output signal-to-noise ratio filter method based on inter-frame correlation in speech noise reduction technology, the problem of insufficient performance of single-channel maximum signal-to-noise ratio filter is solved, and stronger noise reduction capabilities and higher signal fidelity are achieved.

CN120071948APending Publication Date: 2025-05-30SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510219797.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing single-channel maximum signal-to-noise ratio filters lack sufficient performance during speech noise reduction, resulting in a decrease in speech signal quality and increased distortion.

Method used

Using an optimized full-band output signal-to-noise ratio filter method based on inter-frame correlation, by introducing inter-frame correlation in the short-time Fourier transform domain, a noisy voice signal vector is constructed, and the correlation matrix is ​​eigenvalue decomposed, the subspace dominated by the speech signal and the subspace dominated by the noise are extracted, and the maximum signal-to-noise ratio filter is used to process the speech signal, and the subspace dominated by the noise is zeroed.

Benefits of technology

It significantly improves the noise reduction effect and signal fidelity of voice signals, reduces voice distortion, enhances the clarity and intelligibility of voice signals, and adapts to complex and variable noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071948A_ABST
    Figure CN120071948A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice noise reduction, and discloses a method and a system for optimizing a full-band output signal-to-noise ratio filter based on inter-frame correlation, and the method for optimizing the full-band output signal-to-noise ratio filter based on the inter-frame correlation improves the estimation precision of a voice signal by using the inter-frame correlation. Carrying out eigenvalue decomposition on the correlation matrix of the signal, extracting a plurality of dominant subspaces of the voice signal, and processing the voice signal in the subspaces by applying a maximum signal-to-noise ratio filter; the subspaces with dominant noise are zeroed to suppress noise interference to the greatest extent, so that higher noise reduction capability can be shown under different signal-to-noise ratio conditions, meanwhile, distortion of the voice signals is effectively reduced, and the definition and intelligibility of the voice signals are enhanced; besides, the full-band output signal-to-noise ratio filter is optimized, so that the quality of voice signals can be effectively improved in a complex noise environment, and the requirements of a voice enhancement and noise reduction technology on a high-performance algorithm are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of speech noise reduction, and in particular relates to a method and a system for optimizing a full-band output signal-to-noise ratio filter based on inter-frame correlation. Background Art

[0002] Speech signal processing technology has important application value in the field of modern communication and information processing. Speech noise reduction technology, as one of the core links, aims to suppress background noise to the greatest extent while retaining the original quality of the speech signal as much as possible, thereby improving the clarity and intelligibility of the speech. The application scenarios of speech noise reduction are very wide, including hands-free phones, remote video conferencing, speech recognition systems, etc. However, in these scenarios, the complex and changeable background noise environment often significantly reduces the quality of the speech signal and affects the effective transmission of speech information. Therefore, designing an efficient speech noise reduction algorithm has become an important research direction.

[0003] Among the existing speech noise reduction technologies, the Maximum Signal-to-Noise Ratio Filter (MSNRF) has attracted widespread attention due to its superior performance in improving the output signal-to-noise ratio. The algorithm optimizes the filter design to maximize the ratio of the speech component to the noise component in the output signal, thereby achieving efficient noise reduction. Based on its outstanding performance in improving speech quality, the optimization design of the maximum signal-to-noise ratio filter has gradually become an important research topic in the field of speech enhancement.

[0004] At present, the research on the improvement of the maximum signal-to-noise ratio filter mainly focuses on the following two directions: (1) Improved methods based on inter-frame correlation information. By utilizing the continuity and correlation of the speech signal in the time dimension and mining the inter-frame correlation information, the original clean speech signal can be estimated and restored more accurately. This method can not only effectively reduce the speech distortion that may be introduced in the speech denoising process, but also significantly improve the overall quality of the speech signal while ensuring the rationality of the algorithm complexity. (2) Improved strategy for optimizing the full-band output signal-to-noise ratio. From the perspective of full-band signal processing, this strategy performs eigenvalue decomposition on the correlation matrix of the signal to reveal the energy distribution characteristics of the signal in different subspaces. By selecting several subspaces dominated by speech signals to apply the maximum signal-to-noise ratio filter and zeroing the subspaces dominated by noise, the interference of noise on the speech signal can be further weakened. Through this refined subspace processing method, the algorithm can show excellent noise reduction performance under different signal-to-noise ratio conditions and adapt to complex and changeable practical application environments.

[0005] Although the above-mentioned improved methods provide an effective technical approach for the field of speech signal processing, there is still room for further optimization in practical applications. Summary of the Invention

[0006] The object of the present invention is to overcome the above problems, and provide a method and system for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation, aiming to solve the problems of insufficient performance and distortion of speech signals in the process of speech noise reduction by the existing single-channel maximum signal-to-noise ratio filter, and further improve its noise reduction effect and signal fidelity.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] In the first aspect, the present invention provides a method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation, including the following steps:

[0009] Introduce inter-frame correlation in the short-time Fourier transform domain to construct a noisy speech signal vector;

[0010] According to the noisy speech signal vector, construct a full-band output signal-to-noise ratio filter, and calculate the correlation matrix of the speech signal and the noise signal;

[0011] Perform eigenvalue decomposition on the correlation matrix to extract the subspace dominated by the speech signal and the subspace dominated by the noise;

[0012] Apply the maximum signal-to-noise ratio filter in the subspace dominated by the speech signal, and set the subspace dominated by the noise to zero;

[0013] Optimize the full-band output signal-to-noise ratio filter to construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

[0014] A further improvement of the present invention is that when constructing the noisy speech signal vector, continuous time frames are considered, and the noisy speech signal vector is expressed as a speech signal vector and a noise signal vector.

[0015] A further improvement of the present invention is that the noisy speech signal vector can be written as:

[0016] y(k,n) = [Y(k,n) Y(k,n - 1) … Y(k,n - L + 1)] T = x(k,n) + v(k,n)

[0017] Wherein, x(k,n) is the speech signal vector in the short-time Fourier transform domain, v(k,n) is the noise signal vector in the short-time Fourier transform domain, Y is the short-time Fourier transform coefficient of the noisy signal, L is the number of continuous time frames, k is the frequency index, n is the time frame index, and T represents transpose.

[0018] A further improvement of the present invention lies in that when constructing the all-band output signal-to-noise ratio filter, the correlation matrices of the signal and the noise are represented by a block diagonal matrix, and a filter that maximizes the all-band output signal-to-noise ratio is obtained.

[0019] A further improvement of the present invention lies in that the specific method for optimizing the all-band output signal-to-noise ratio filter is as follows: by minimizing the distortion-based mean square error and the mean square error between the original clean signal and the estimated desired signal, several largest eigenvalues corresponding eigenvectors are selected from all possible eigenvalue sets to construct a filter with a length of LK.

[0020] A further improvement of the present invention lies in that the specific method for constructing a filter with a length of LK is as follows:

[0021] From all possible eigenvalue sets {λ 1 (k i ,n), i = 0, 1, …, K - 1}, P largest eigenvalues corresponding eigenvectors are selected to construct a filter with a length of LK. The specific filter form is as follows:

[0022]

[0023] In the formula, h(k i ,n) is a FIR filter with a length of L, where i = 0, 1, … K - 1, α p (k p ,n), p = 0, 1,.., P - 1 are arbitrary complex numbers, and at least one is not equal to 0. is the eigenvector of the matrix , and 0 T is a all-zero vector with a length of L(K - 1);

[0024] In the formula, is the inverse of the matrix D V (n),

[0025] where D X (n) and D V (n) are defined as follows:

[0026] D X (n) = diag[Φ x (k 0 ,n), Φ x (k 1 ,n), …, Φ x (k K-1 ,n)]

[0027] D V (n) = diag[Φ v (k 0, n), Φ v (k 1 , n), …, Φ v (k K-1 , n)]

[0028] Where, Φ x (k i , n) is the correlation matrix of x(k i , n), Φ v (k i , n) is the correlation matrix of v(k i , n);

[0029] The filter can also be written in the following form:

[0030]

[0031] Based on this method, the estimated value of the desired signal can be expressed as:

[0032]

[0033] Where, is the filter coefficient vector, y(k p , n) is the noisy speech signal vector, and H is the conjugate transpose;

[0034] Furthermore, the expression of the distortion-based mean square error filter that is minimized at frequency k p , p = 0, 1…, P - 1 is:

[0035]

[0036] Where, λ 1 (k p , n) is 's maximum eigenvalue, and i L,1 is the first column of the L×L identity matrix I L ;

[0037] The filter expression that minimizes the mean square error between the original clean signal and the estimated desired signal is:

[0038]

[0039] That is, the expression of the full-band distortion-based mean square error filter with length LK is:

[0040]

[0041] Similarly, the filter expression of the full-band that minimizes the mean square error between the original clean signal and the estimated desired signal with length LK can also be obtained as:

[0042]

[0043] In the formula, T represents transpose, and 0 T is a zero vector of length L(K - 1).

[0044] A further improvement of the present invention is that when constructing the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation, it also includes adjusting the parameters of the filter length, forgetting factor, and number of subspaces.

[0045] In a second aspect, the present invention also provides a system for an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation, including the following modules:

[0046] An inter-frame correlation introduction module, configured to introduce inter-frame correlation in the short-time Fourier transform domain to construct a noisy speech signal vector;

[0047] A filter construction module, configured to construct an optimized full-band output signal-to-noise ratio filter according to the noisy speech signal vector and calculate the correlation matrix of the speech signal and the noise signal;

[0048] An eigenvalue decomposition module, configured to perform eigenvalue decomposition on the correlation matrix to extract the subspace dominated by the speech signal and the subspace dominated by the noise;

[0049] A filter application module, configured to apply a maximum signal-to-noise ratio filter in the subspace dominated by the speech signal and perform zeroing processing on the subspace dominated by the noise;

[0050] An optimization module, configured to optimize the optimized full-band output signal-to-noise ratio filter to construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

[0051] In a third aspect, the present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method for the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation described above are implemented.

[0052] In a fourth aspect, the present invention also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation described above are implemented.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] The present invention provides a method for optimizing the all-band output signal-to-noise ratio filter based on inter-frame correlation. By utilizing the inter-frame correlation, the estimation accuracy of the speech signal is improved. At the same time, the eigenvalue decomposition is performed on the correlation matrix of the signal to extract several dominant subspaces of the speech signal. The maximum signal-to-noise ratio filter is applied to process the speech signal in these subspaces; while the subspaces dominated by noise are set to zero to maximize the suppression of noise interference. It can not only exhibit stronger noise reduction ability under different signal-to-noise ratio conditions, but also effectively reduce the distortion of the speech signal, enhance the clarity and intelligibility of the speech signal. In addition, optimizing the all-band output signal-to-noise ratio filter can effectively improve the quality of the speech signal in a complex noise environment, meeting the requirements of speech enhancement and noise reduction technologies for high-performance algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The drawings described herein are for illustrative purposes only and are not intended to limit the scope of the disclosure of the present invention in any way. In addition, the shapes and proportional dimensions of the components in the drawings are only schematic and are used to assist in understanding the present invention, rather than specifically limiting the shapes and proportional dimensions of the components of the present invention.

[0056] Figure 1 It is a flowchart of the method for optimizing the all-band output signal-to-noise ratio filter based on inter-frame correlation of the present invention;

[0057] Figure 2 It is a system flowchart of the optimizing the all-band output signal-to-noise ratio filter based on inter-frame correlation of the present invention;

[0058] FIG. 3(a) shows the influence of p on the output signal-to-noise ratio of filters with different lengths according to the present invention;

[0059] FIG. 3(b) shows the influence of p on the speech distortion coefficient of filters with different lengths according to the present invention;

[0060] FIG. 3(c) shows the influence of p on the PESQ score of filters with different lengths according to the present invention;

[0061] FIG. 4(a) shows the influence of the forgetting factor on the output signal-to-noise ratio of filters with different lengths according to the present invention;

[0062] FIG. 4(b) shows the influence of the forgetting factor on the speech distortion coefficient of filters with different lengths according to the present invention;

[0063] FIG. 4(c) shows the influence of the forgetting factor on the PESQ score of filters with different lengths according to the present invention;

[0064] FIG. 5(a) shows the influence of the filter length on the output signal-to-noise ratio of the filter according to the present invention;

[0065] Figure 5(b) shows the influence of the filter length involved in the present invention on the speech distortion coefficient of the filter;

[0066] Figure 5(c) shows the influence of the filter length involved in the present invention on the PESQ score of the filter;

[0067] Figure 6(a) shows the output signal-to-noise ratio of the optimized full-band output signal-to-noise ratio filter, the maximum signal-to-noise ratio filter, and the Wiener filter based on inter-frame correlation under different noise conditions involved in the present invention;

[0068] Figure 6(b) shows the speech distortion coefficients of the optimized full-band output signal-to-noise ratio filter, the maximum signal-to-noise ratio filter, and the Wiener filter based on inter-frame correlation under different noise conditions involved in the present invention;

[0069] Figure 6(c) shows the PESQ scores of the optimized full-band output signal-to-noise ratio filter, the maximum signal-to-noise ratio filter, and the Wiener filter based on inter-frame correlation under different noise conditions involved in the present invention;

[0070] Figure 7 It is a schematic diagram of an electronic device for the method of an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation. Detailed implementation manners

[0071] The following further describes the present invention in detail with reference to the accompanying drawings:

[0072] As Figure 1 shown, the present invention provides a method for an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation, including the following steps:

[0073] Step S1, in the short-time Fourier transform domain, considering consecutive time frames, represent the noisy speech signal vector as the speech signal vector and the noise signal vector in the short-time Fourier transform domain, and construct the noisy speech signal vector;

[0074] The noisy speech signal vector can be written as:

[0075] y(k,n) = [Y(k,n) Y(k,n - 1) … Y(k,n - L + 1)] T = x(k,n) + v(k,n)

[0076] where x(k,n) is the speech signal vector in the short-time Fourier transform domain, v(k,n) is the noise signal vector in the short-time Fourier transform domain, Y is the short-time Fourier transform coefficient of the noisy signal, L is the number of consecutive time frames, k is the frequency index, n is the time frame index, and T represents the transpose;

[0077] Step S2: According to the noisy speech signal vector, use a block diagonal matrix to represent the correlation matrices of the signal and the noise, construct a full-band output signal-to-noise ratio filter, and calculate the correlation matrices of the speech signal and the noise signal.

[0078] Construct a full-band output signal-to-noise ratio filter with a length of LK The form of the filter can be expressed as:

[0079]

[0080] In the formula, h(k i ,n) is the filter coefficient vector with a length of L, where i = 0, 1, … K-1, and T represents the transpose;

[0081] Then the form of the full-band output signal-to-noise ratio is expressed as:

[0082]

[0083] In the formula, H is the conjugate transpose;

[0084] Among them, D X (n) and D V (n) are defined as follows:

[0085] D X (n) = diag[Φ x (k 0 ,n), Φ x (k 1 ,n), …, Φ x (k K-1 ,n)]

[0086] D V (n) = diag[Φ v (k 0 ,n), Φ v (k 1 ,n), …, Φ v (k K-1 ,n)]

[0087] Furthermore, we get:

[0088]

[0089] In the formula, is the inverse of the matrix D V (n), Φ x (k i ,n) is the correlation matrix of x(k i ,n), and Φ v (k i ,n) is the correlation matrix of v(k i ,n);

[0090] Among them,

[0091] In step S3, perform eigenvalue decomposition on the correlation matrix to extract the subspace dominated by the speech signal and the subspace dominated by the noise;

[0092] In step S4, apply a maximum signal-to-noise ratio filter in the subspace dominated by the speech signal and set the subspace dominated by the noise to zero;

[0093] In step S5, optimize the full-band output signal-to-noise ratio filter by minimizing the distortion-based mean square error and the mean square error between the original clean signal and the estimated desired signal, select several eigenvectors corresponding to the largest eigenvalues from all possible eigenvalue sets, and construct a filter with a length of LK. The specific filter form is:

[0094]

[0095] In the formula, h(k i ,n) is a FIR filter with a length of L, where i = 0, 1, … K-1, α p (k p ,n), p = 0, 1,.., P-1 are arbitrary complex numbers, and at least one is not equal to 0, is the eigenvector of the matrix , 0 T is a zero vector with a length of L(K-1);

[0096] This method enhances the signal at the P most important frequencies and sets the filter coefficients of all the remaining frequency components to zero, thereby reducing speech distortion;

[0097] The filter can also be written in the following form:

[0098]

[0099] Based on this method, the estimated value of the desired signal can be expressed as:

[0100]

[0101] In the formula, is the filter coefficient vector, y(k p ,n) is the vector of the noisy speech signal, and H is the conjugate transpose;

[0102] Furthermore, the expression of the distortion-based mean square error filter that is minimized at frequency k p , p = 0, 1…, P-1 is obtained as:

[0103]

[0104] where λ 1 (k p , n) is the maximum eigenvalue of L,1 the L×L identity matrix I L and i

[0105] The filter expression that minimizes the mean square error between the original clean signal and the estimated desired signal is:

[0106]

[0107] That is, the full-band minimum distortion-based mean square error filter expression of length LK is:

[0108]

[0109] Similarly, the filter expression of length LK that minimizes the mean square error between the original clean signal and the estimated desired signal can also be obtained as:

[0110]

[0111] So far, two optimized full-band output signal-to-noise ratio filters based on inter-frame correlation have been derived, namely the maximum signal-to-noise ratio filter that minimizes the distortion-based mean square error and the maximum signal-to-noise ratio filter that minimizes the mean square error between the original clean signal and the estimated desired signal

[0112] As Figure 2 shown, the present invention also provides a method for an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation, including the following modules:

[0113] An inter-frame correlation introduction module, configured to introduce inter-frame correlation in the short-time Fourier transform domain and construct a noisy speech signal vector;

[0114] A filter construction module, configured to construct a full-band output signal-to-noise ratio filter according to the noisy speech signal vector and calculate the correlation matrix of the speech signal and the noise signal;

[0115] An eigenvalue decomposition module, configured to perform eigenvalue decomposition on the correlation matrix and extract the subspace dominated by the speech signal and the subspace dominated by the noise;

[0116] A filter application module, configured to apply the maximum signal-to-noise ratio filter in the subspace dominated by the speech signal and set the subspace dominated by the noise to zero;

[0117] Optimization module, used to optimize the full-band output signal-to-noise ratio filter and construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

[0118] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0119] In order to further improve the speech denoising effect, based on considering the inter-frame correlation, the present invention proposes an improved strategy for optimizing the full-band output signal-to-noise ratio, aiming to improve the maximum signal-to-noise ratio filter. The core of this strategy is to perform eigenvalue decomposition on the correlation matrix of the signal, and on this basis, select several subspaces dominated by the speech signal to apply the maximum signal-to-noise ratio filter, while setting the remaining subspaces dominated by noise to zero, so as to enhance the denoising ability of the filter under different signal-to-noise ratio conditions and enhance the significance of the speech signal.

[0120] First, construct a full-band output signal-to-noise ratio filter with a length of LK The form of the filter is expressed as:

[0121]

[0122] In the formula, h(k i ,n) is the filter coefficient vector with a length of L, where i = 0, 1,..., K - 1;

[0123] Then the form of the full-band output signal-to-noise ratio is expressed as:

[0124]

[0125] Among them, D X (n) and D V (n) are defined as follows:

[0126] D X (n) = diag[Φ x (k 0 ,n), Φ x (k 1 ,n),..., Φ x (k K-1 ,n)]

[0127] D V (n) = diag[Φ v (k 0 ,n), Φ v (k 1 ,n),..., Φ v (k K-1 ,n)]

[0128] Furthermore, it is obtained that:

[0129]

[0130] Among them,

[0131] In the formula, Φ x (k, n) is the correlation matrix of the speech signal vector in the short-time Fourier transform domain, and Φ v (k, n) is the correlation matrix of the noise signal vector in the short-time Fourier transform domain.

[0132] It can be seen from the above definition that the full-band output signal-to-noise ratio filter is composed of multiple subspace filters. Each subspace filter corresponds to an eigenvector of a maximum eigenvalue of the signal correlation matrix D(k i , n). By sorting the K sub-band eigenvalues, the following inequality can be obtained:

[0133] λ 1 (k 0 , n) ≥ λ 1 (k 1 , n) ≥ … ≥ λ 1 (k K-1 , n)

[0134] In the formula, λ 1 (k i , n) is the maximum eigenvalue of the signal correlation matrix D(k i , n) in the subspace;

[0135] This inequality helps to identify which subspaces should apply the maximum signal-to-noise ratio filter and which subspaces should be set to zero to suppress noise.

[0136] From the above formula, it can be obtained that the maximum eigenvalue of the matrix is λ 1 (k 0 , n). Then, the filter that maximizes the full-band output signal-to-noise ratio is the eigenvector corresponding to the maximum eigenvalue λ (k 1 , n) of the matrix 0 . Therefore, the maximum signal-to-noise ratio filter is:

[0137]

[0138] In the formula, α 0 (k 0 , n) is a non-zero complex number, b 1 (k 0 , n) is the eigenvector corresponding to the maximum eigenvalue. The length of the filter is LK, and 0 T is a zero vector with a length of L(K - 1);

[0139] The maximum signal-to-noise ratio filter can be further expressed as:

[0140]

[0141] Then the relationship between the full-band output signal-to-noise ratio and the input signal-to-noise ratio is:

[0142]

[0143] where, is the sub-band input signal-to-noise ratio at the k-th frequency point.

[0144] Furthermore, the estimated desired signal is obtained as:

[0145]

[0146] In the formula, and are the estimated values of the desired signal at different frequency points;

[0147] To determine the value of α 0 (k 0 ,n), two methods can be adopted. One is to minimize the distortion-based mean square error The other is to minimize the mean square error between the original clean signal and the estimated desired signal

[0148] The formula for the distortion-based mean square error is defined as:

[0149]

[0150] The formula for the mean square error between the original clean signal and the estimated desired signal is defined as:

[0151]

[0152] Taking the method of minimizing the distortion-based mean square error as an example, we can obtain:

[0153]

[0154] In the formula, φ X (k,n) = E[|X(k,n)| 2 is the variance of X(k,n), and i L,1 is the first column of the L×L identity matrix I L .

[0155] It is derived that:

[0156]

[0157] where the superscript (·) *Denotes the complex conjugate.

[0158] Let the equation For α 0 (k 0 ,n) Take the derivative and set the derivative equal to 0 to obtain:

[0159]

[0160] The maximum signal-to-noise ratio filter that minimizes the distortion-based mean square error can be obtained as:

[0161]

[0162] Similarly, by minimizing the mean square error between the original clean signal and the estimated desired signal, another maximum signal-to-noise ratio filter can be obtained:

[0163]

[0164] Although the foregoing method can maximize the output signal-to-noise ratio of the full band, this method will introduce a large distortion to the desired signal because it forces the filter coefficients of all frequency components except k 0 to be set to zero. To solve this problem, the method will be further improved, that is, select the eigenvectors corresponding to the P largest eigenvalues from all possible eigenvalue sets {λ 1 (k i ,n), i = 0, 1,..., K - 1} to construct a filter of length LK. This improvement strategy aims to retain more frequency components, thereby reducing signal distortion.

[0165] The specific filter form is as follows:

[0166]

[0167] In the formula, g(k i ,n) is a FIR filter of length L, where i = 0, 1,..., K - 1, α p (k p ,n), p = 0, 1,.., P - 1 are arbitrary complex numbers, and at least one is not equal to 0, is the eigenvector of the matrix .

[0168] This method enhances the signal at the P most important frequencies while setting the filter coefficients of all the remaining frequency components to zero, thereby reducing speech distortion.

[0169] The filter can also be written in the following form:

[0170]

[0171] Based on this method, the estimated value of the desired signal can be expressed as:

[0172]

[0173] Furthermore, at frequency k p , p = 0, 1…, P - 1, the expression of the distortion-based mean square error filter for minimization is:

[0174]

[0175] The filter expression for minimizing the mean square error between the original clean signal and the estimated desired signal is:

[0176]

[0177] That is, the expression of the full-band distortion-based mean square error filter with length LK for minimization is:

[0178]

[0179] Similarly, the filter expression for minimizing the mean square error between the original clean signal and the estimated desired signal with length LK for the full band can also be obtained as:

[0180]

[0181] So far, two optimized full-band output signal-to-noise ratio filters based on inter-frame correlation have been derived, namely the maximum signal-to-noise ratio filter that minimizes the distortion-based mean square error and the maximum signal-to-noise ratio filter that minimizes the mean square error between the original clean signal and the estimated desired signal When P = K, it corresponds to the traditional maximum signal-to-noise ratio filter, and it corresponds to the traditional Wiener filter.

[0182] Take the maximum signal-to-noise ratio filter that minimizes the distortion-based mean square error For example, the performance of the filter is verified through simulation. The clean speech signals used in the simulation are from the TIMIT database. Different speech signals of 15 male and 15 female speakers are randomly selected, and their sampling rate is reduced from 16 kHz to 8 kHz. Then, under different input signal-to-noise ratio conditions, noise is added to the clean speech signals to generate noisy speech signals. To simulate different background noise environments, two types of noise are introduced in the simulation: Gaussian white noise and car noise, and these noise signals are also sampled at a frequency of 8 kHz. To construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation in the short-time Fourier transform domain, the speech signal is segmented into signal frames with a length of 128 sampling points, and the overlap rate between frames is 75%. To reduce the influence of spectral leakage, a Kaiser window is applied to each signal frame. Then, through the short-time Fourier transform, the signal is transformed into the short-time Fourier transform domain. Subsequently, for each sub-band, the corresponding filter is calculated and applied to the corresponding sub-band to perform noise reduction processing. Finally, through the inverse short-time Fourier transform, the processed signal is transformed back from the short-time Fourier transform domain to the time domain.

[0183] Example 1:

[0184] In this example, the variation of the output signal-to-noise ratio, speech distortion coefficient, and PESQ score of the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation with the parameter p is studied under the Gaussian white noise environment with an input signal-to-noise ratio of 10 dB. The specific values of the filter length L and the corresponding parameter p are shown in Table 1.

[0185] Table 1 p values corresponding to different filter lengths

[0186]

[0187] The simulation results are shown in Figures 3(a), 3(b), and 3(c). The results show that as the parameter p increases, both the output signal-to-noise ratio and the speech distortion coefficient of the filter decrease. This phenomenon indicates that a larger p value makes the filter more conservative in noise reduction processing, and more sub-spaces dominated by speech signals are processed by the filter for noise reduction, while fewer sub-spaces dominated by noise signals are forced to zero. In addition, the PESQ score shows a trend of increasing first and then decreasing. This indicates that by appropriately adjusting the value of p, an effective trade-off can be achieved between improving the output signal-to-noise ratio and controlling speech distortion.

[0188] When the value of p is below 20, the speech distortion coefficient will increase sharply, and the PESQ score of the optimized full-band output SNR filter based on inter-frame correlation is lower than the original PESQ score in the Gaussian white noise environment. This indicates that if the value of p is too low, the filter will overly weaken the subspace dominated by the speech signal, thereby reducing the speech quality. On the contrary, when the value of p is in the range of 35 to 55, the filter can achieve the highest PESQ score, meaning that within this range, the filter reaches the best balance between noise reduction and maintaining speech quality. When the value of p reaches the maximum, the optimized full-band output SNR filter based on inter-frame correlation becomes a traditional maximum SNR filter.

[0189] Example 2:

[0190] In this example, the simulation settings are the same as those of the first simulation, that is, in a Gaussian white noise environment with an input SNR of 10 dB, the output SNR, speech distortion coefficient, and PESQ score of the optimized full-band output SNR filter based on inter-frame correlation are explored as the forgetting factor α changes. In this simulation, the values of the filter length L and the corresponding forgetting factor α are shown in Table 2.

[0191] Table 2 Values of the forgetting factor α corresponding to different filter lengths

[0192]

[0193] Figures 4(a), 4(b), and 4(c) show the simulation results. It can be seen that as the forgetting factor α increases, the output SNR of the filter gradually decreases, while the speech distortion coefficient gradually increases. This indicates that a larger α value makes the filter more dependent on the current observed data and reduces the memory of historical data. Although this helps the filter quickly adapt to environmental changes, it also weakens the ability to capture the long-term trend of the signal, resulting in a decrease in noise suppression effect and an increase in speech distortion. The PESQ score shows a trend of first increasing and then decreasing, indicating that a smaller α value is beneficial for noise suppression but may overly smooth the speech signal. Moderately increasing the α value helps to find a balance between noise suppression and maintaining speech quality and improve the PESQ score. However, an overly large α value will lead to increased speech distortion and a decrease in the PESQ score.

[0194] In addition, the simulation results also show that the larger the filter length, the larger the α value required to obtain the best PESQ score. This indicates that a stronger filter requires a larger α value to balance the relationship between noise suppression and speech fidelity. This means that when designing the filter, an appropriate α value should be selected according to the filter length to ensure the best speech quality at different lengths.

[0195] Example 3:

[0196] In the third example, the noise environment is also set to a Gaussian white noise environment with an input signal-to-noise ratio of 10 dB. This simulation aims to further study the effects of the length L of the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation on the output signal-to-noise ratio, speech distortion coefficient, and PESQ score. During the simulation, the values of the forgetting factor α and the parameter p corresponding to the filter length L are obtained from Table 3.

[0197] Table 3 Values of the forgetting factor α and p corresponding to the filter length

[0198]

[0199] The simulation results are shown in Figures 5(a), 5(b), and 5(c). The results show that as the filter length L increases, the output signal-to-noise ratio shows a gradually increasing trend, which means that a longer filter can more effectively suppress the noise signal in the noisy signal, thus improving the signal-to-noise ratio. However, as the filter length L increases, the speech distortion coefficient also increases correspondingly, and the PESQ score shows a gradually decreasing trend, which indicates that a longer filter will also introduce more speech distortion and affect the speech quality.

[0200] This trade-off relationship emphasizes the importance of comprehensively considering the improvement of the output signal-to-noise ratio and the maintenance of speech quality in filter design. Although a longer filter can provide better noise suppression effect, an overly long filter may damage the quality of the speech signal and lead to a decline in the user's auditory experience. Therefore, in practical applications, it is necessary to select an appropriate filter length according to the specific application scenario and performance requirements.

[0201] Example 4:

[0202] The fourth example studies the performance of the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation in two different noise environments, Gaussian white noise and automotive noise, and under different input signal-to-noise ratio conditions. The simulation not only evaluates the performance of the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation but also compares its performance with that of traditional maximum signal-to-noise ratio filters and Wiener filters, both of which consider inter-frame correlation. Tables 4 and 5 respectively summarize the filter length, the value of the forgetting factor α, and the value of the parameter p when the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation produces the highest PESQ score when the input signal-to-noise ratio takes different values in Gaussian white noise and automotive noise environments.

[0203] Table 4 Parameters of the filter corresponding to the highest PESQ score in Gaussian white noise environment

[0204]

[0205] Table 5 Parameters of the filter corresponding to the highest PESQ score in the automotive noise environment

[0206]

[0207] The simulation results are plotted in Fig. 6(a), Fig. 6(b) and Fig. 6(c), showing the performance comparison of the optimized full-band output SNR filter based on inter-frame correlation, the maximum SNR filter and the Wiener filter under different noise environments and different signal-to-noise ratio conditions.

[0208] From the perspective of output SNR, the output SNR of all filters increases with the increase of input SNR, and this trend is reflected in both Gaussian white noise and automotive noise environments. Moreover, the optimized full-band output SNR filter based on inter-frame correlation always has a better output SNR than the maximum SNR filter and the Wiener filter in these two environments. Especially under low SNR conditions, its noise reduction effect is more significant, which further proves the high efficiency of the designed filter in noise reduction.

[0209] In terms of speech distortion coefficient, in these two noise environments, the speech distortion coefficients of all filters decrease with the increase of input SNR. The optimized full-band output SNR filter based on inter-frame correlation and the maximum SNR filter have similar performance in terms of speech distortion coefficient, while the Wiener filter shows lower speech distortion coefficients in both noise environments, which is its traditional advantage in minimizing distortion.

[0210] In terms of PESQ score, the optimized full-band output SNR filter based on inter-frame correlation performs comprehensively better than the other two filters in the Gaussian white noise environment. This proves its excellent performance in improving the auditory experience. In the automotive noise environment, when the input SNR is lower than 16 dB, the designed filter still maintains its advantage. Although the PESQ score of the Wiener filter is slightly higher when the input SNR is higher than 16 dB, the designed filter still shows good performance, indicating that it can provide high-quality speech output under a wide range of SNR conditions.

[0211] Embodiment 5:

[0212] Please refer to Figure 7 As shown, the present invention also provides an electronic device 100 for a method of an optimized full-band output SNR filter based on inter-frame correlation; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0213] The memory 101 can be used to store the computer program 103. By running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101, the processor 102 implements the steps of the method for the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation described in Embodiment 1. The memory 101 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 may include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices.

[0214] The at least one processor 102 can be a Central Processing Unit (CPU), or can also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 can be a microprocessor or the processor 102 can also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 using various interfaces and lines.

[0215] The memory 101 in the electronic device 100 stores a plurality of instructions to implement the method for the optimized full-band output signal-to-noise ratio filter based on inter-frame correlation. The processor 102 can execute the plurality of instructions to thereby implement:

[0216] Introduce inter-frame correlation in the short-time Fourier transform domain to construct a noisy speech signal vector;

[0217] According to the noisy speech signal vector, construct an optimized full-band output signal-to-noise ratio filter and calculate the correlation matrix of the speech signal and the noise signal;

[0218] Perform eigenvalue decomposition on the correlation matrix to extract the subspace dominated by the speech signal and the subspace dominated by the noise;

[0219] Apply a maximum signal-to-noise ratio filter in the subspace dominated by the speech signal and set to zero the subspace dominated by the noise;

[0220] Optimize the full-band output signal-to-noise ratio filter and construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

[0221] Embodiment 6:

[0222] If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, and read-only memory (ROM, Read-Only Memory).

[0223] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0224] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0225] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0226] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in one or more processes and / or blocks Figure 1 specified in the function.

[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

[0228] Many embodiments and many applications other than the examples provided will be apparent to those skilled in the art from reading the above description. Therefore, the scope of this teaching should not be determined by reference to the above description, but should be determined by reference to the full scope of the foregoing claims and the equivalents of those claims. For the sake of completeness, all articles and references, including patent applications and publications, are incorporated herein by reference. The omission of any aspect of the subject matter disclosed herein in the foregoing claims is not intended to waive that subject matter, nor should it be considered that the applicant has not considered that subject matter to be part of the disclosed inventive subject matter.

[0229] The above content is a further detailed description of the present invention. It cannot be determined that the specific embodiments of the present invention are limited to this. For those of ordinary skill in the technical field of the present invention, without departing from the concept of the present invention, several simple deductions or replacements can still be made, which should all be regarded as belonging to the protection scope determined by the claims submitted by the present invention.

Claims

1. A method for optimizing a full-band output signal-to-noise ratio filter based on inter-frame correlation, characterized in that: The following steps are involved: Introduce inter-frame correlation in the short-time Fourier transform domain to construct the noisy speech signal vector; According to the noisy speech signal vector, a full-band output signal-to-noise ratio filter is constructed to calculate the correlation matrix between the speech signal and the noise signal; Perform eigenvalue decomposition on the correlation matrix to extract the subspace dominated by speech signals and the subspace dominated by noise; Apply the maximum signal-to-noise ratio filter in the subspace dominated by speech signals, and take zero processing for the subspace dominated by noise; Optimize the full-band output signal-to-noise ratio filter and construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

2. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 1, characterized in that: When constructing the noisy speech signal vector, continuous time frames are considered, and the noisy speech signal vector is represented as a speech signal vector and a noise signal vector.

3. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 2, characterized in that: The noisy speech signal vector can be written as: y(k,n)=[Y(k,n)Y(k,n-1)…Y(k,n-L+1)] T =x(k,n)+v(k,n) Where x(k,n) is the speech signal vector in the short-time Fourier transform domain, v(k,n) is the noise signal vector in the short-time Fourier transform domain, Y is the short-time Fourier transform coefficient of the noisy signal, L is the number of consecutive time frames, k is the frequency index, n is the time frame index, and T represents transposition.

4. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 1, characterized in that: When constructing the full-band output signal-to-noise ratio filter, a block diagonal matrix is ​​used to represent the correlation matrix of the signal and the noise, so as to obtain a filter that maximizes the full-band output signal-to-noise ratio.

5. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 1, characterized in that: The specific method for optimizing the full-band output signal-to-noise ratio filter is: by minimizing the mean square error based on distortion and minimizing the mean square error between the original clean signal and the estimated expected signal, several eigenvectors corresponding to the largest eigenvalues ​​are selected from all possible eigenvalue sets to construct a filter with a length of LK.

6. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 5, characterized in that: The specific method of constructing a filter with a length of LK is: From the set of all possible eigenvalues ​​{λ1(k i ,n),i=0,1,…,K-1}, select the eigenvectors corresponding to the P largest eigenvalues ​​to construct a filter with a length of LK. The specific filter form is: In the formula, h(k i ,n) is a FIR filter of length L, where i = 0, 1, ... K-1, α p (k p ,n),p=0,1,..,P-1 is any complex number, and at least one of them is not equal to 0, For the matrix The eigenvector of T is an all-zero vector of length L(K-1); In the formula, is the matrix D V The inverse of (n), Where D X (n) and D V (n) is defined as follows: D X (n)=diag[Φ x (k0,n),Φ x (k1,n),…,Φ x (k K-1 ,n)] D V (n)=diag[Φ v (k0,n),Φ v (k1,n),…,Φ v (k K-1 ,n)] In the formula, Φ x (k i ,n) is x(k i ,n), Φ v (k i ,n) is v(k i ,n) correlation matrix; The filter can also be written as: Based on this approach, the estimated value of the desired signal can be expressed as: In the formula, is the filter coefficient vector, y(k p , n) is the noisy speech signal vector, H is the conjugate transpose; Further, we get p The expression of the distortion-based mean square error filter minimized in ,p=0,1…,P-1 is: In the formula, λ1(k p ,n) is The maximum eigenvalue of L,1 is the L×L identity matrix I L The first column of The filter expression that minimizes the mean square error between the original clean signal and the estimated desired signal is: That is, the expression of the full-band minimum distortion-based mean square error filter with a length of LK is: Similarly, the filter expression for the full-band minimization of the mean square error between the original clean signal and the estimated expected signal with a length of LK can be obtained as follows: In the formula, T represents transpose, 0 T is an all-zero vector of length L(K-1).

7. The method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to claim 1, characterized in that: When constructing an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation, parameters for adjusting the filter length, forgetting factor and number of subspaces are also included.

8. A system for optimizing full-band output signal-to-noise ratio filter based on inter-frame correlation, characterized in that: Includes the following modules: An inter-frame correlation introduction module is used to introduce inter-frame correlation in the short-time Fourier transform domain to construct a noisy speech signal vector; A filter construction module is used to construct a full-band output signal-to-noise ratio filter according to the noisy speech signal vector and calculate the correlation matrix of the speech signal and the noise signal; An eigenvalue decomposition module is used to perform eigenvalue decomposition on the correlation matrix to extract a subspace dominated by speech signals and a subspace dominated by noise; A filter application module, used for applying a maximum signal-to-noise ratio filter in a subspace dominated by speech signals and performing a zeroing process on a subspace dominated by noise; The optimization module is used to optimize the full-band output signal-to-noise ratio filter and construct an optimized full-band output signal-to-noise ratio filter based on inter-frame correlation.

9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for optimizing the full-band output signal-to-noise ratio filter based on inter-frame correlation according to any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for optimizing a full-band output signal-to-noise ratio filter based on inter-frame correlation according to any one of claims 1 to 7 are implemented.