Causal adaptive single-channel speech denoising method and apparatus
By employing a causal adaptive single-channel speech denoising method, which generates a target filter vector to process noise and speech correlation matrices, the problem of poor speech quality and intelligibility in existing technologies is solved, thereby improving speech quality and intelligibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH OF CHINA
- Filing Date
- 2023-07-19
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies, after speech enhancement processing, do not achieve good results in terms of speech quality and intelligibility of the output information.
A causal adaptive single-channel speech denoising method is adopted. By collecting noise information and noisy speech information, noise correlation matrix and speech correlation matrix are generated. Generalized eigenvalue decomposition is performed to generate target filtering vector. This vector is used to process noisy speech information to generate enhanced target output speech information.
It significantly improves speech quality and intelligibility, achieving a balance between the two.
Smart Images

Figure CN116758931B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech signal processing technology, and more specifically, to a causal adaptive single-channel speech denoising method and a causal adaptive single-channel speech denoising device. Background Technology
[0002] Speech enhancement aims to extract clean sound source signals from noisy and reverberant sound signals acquired by acoustic sensors. Its performance metrics primarily include signal-to-noise ratio (SNR) and speech intelligibility. Speech enhancement, or noise reduction (NR), has wide applications in hearing aids, video conferencing, and human-computer interaction, especially in noisy environments, where it is often built as a front-end to improve speech quality and intelligibility. Over the past few decades, numerous noise reduction algorithms have been proposed, which can be categorized into single-channel and multi-channel types.
[0003] In realizing the present invention, the inventors discovered that the related technologies have at least the following problems: after enhancing speech information, the enhanced output information obtained by the related technologies is not effective in terms of speech quality and speech intelligibility. Summary of the Invention
[0004] In view of this, embodiments of the present disclosure provide a causal adaptive single-channel speech denoising method and a causal adaptive single-channel speech denoising device.
[0005] One aspect of this disclosure provides a causal adaptive single-channel speech noise reduction method, including:
[0006] Noise information in the first time period and noisy speech information in the second time period were collected;
[0007] The noise information and the noisy speech information mentioned above are processed to obtain the noise correlation matrix and the speech correlation matrix;
[0008] The noise correlation matrix and speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors;
[0009] Based on the filtering functions of the multiple initial feature vectors and the initial filters mentioned above, a target filtering vector is generated;
[0010] Based on the target filter vector and the noisy speech information described above, enhanced target output speech information is generated.
[0011] According to embodiments of this disclosure, the noise correlation matrix and speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors, including:
[0012] The noise correlation matrix and speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial eigenvalues;
[0013] Each of the above initial feature values is transformed to obtain multiple of the above initial feature vectors.
[0014] According to embodiments of this disclosure, a target filter vector is generated based on a plurality of the aforementioned initial feature vectors and a filter function of the initial filter, including:
[0015] Based on the above filtering function, the above noise correlation matrix, the above speech correlation matrix, and the balance factor, an initial filtering vector is generated;
[0016] Based on the multiple initial eigenvectors described above and the initial eigenvalues corresponding to each of the initial eigenvectors, a vector matrix and a diagonal matrix are generated respectively.
[0017] Based on preset selection rules, multiple initial feature values are filtered to obtain multiple target feature values;
[0018] Based on the aforementioned target eigenvalues, vector matrices, and diagonal matrices, a target filter is generated, wherein the target filter includes the aforementioned target filter vector.
[0019] According to embodiments of this disclosure, multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values, including:
[0020] Based on a preset sorting rule, the above initial feature values are sorted to obtain sorted initial feature values;
[0021] Based on the above-mentioned preset selection rules, multiple target feature values are selected from the sorted initial feature values.
[0022] According to embodiments of this disclosure, a target filter is generated based on a plurality of the aforementioned target feature values, vector matrices, and diagonal matrices, wherein the target filter includes the aforementioned target filter vector, comprising:
[0023] By jointly diagonalizing the above vector matrix and diagonal matrix, we obtain the first diagonalization formula;
[0024] Based on the first diagonalization formula and the noisy covariance matrix described above, the second diagonalization formula is generated.
[0025] Based on the aforementioned target feature values and the aforementioned second diagonalization formula, the aforementioned target filter vector is generated.
[0026] According to embodiments of this disclosure, the target filtering vector is generated based on a plurality of the aforementioned target feature values and the aforementioned second diagonalization formula, including:
[0027] For any correlation matrix in the noise correlation matrix and the speech correlation matrix, the correlation matrix is approximated by using multiple of the above target feature values to obtain an approximate correlation matrix;
[0028] Based on the aforementioned approximate correlation matrix, the aforementioned second diagonalization formula, and the aforementioned initial filter vector, the aforementioned target filter vector is generated, wherein the aforementioned target filter vector includes the aforementioned balance factor.
[0029] According to embodiments of this disclosure, enhanced target output speech information is generated based on the aforementioned target filter vector and the aforementioned noisy speech information, including:
[0030] The target filter vector is conjugate transposed to obtain the transposed vector.
[0031] Perform an inner product operation on the above noisy speech information and the above transpose vector to obtain the initial output speech information;
[0032] The initial output speech information is processed by inverse short-time Fourier transform to obtain the target output speech information.
[0033] According to embodiments of this disclosure, the noise information and the noisy speech information described above are processed to obtain a noise correlation matrix and a speech correlation matrix, including:
[0034] For any of the above noise information and the above noisy speech information, a scalar representation formula is obtained by using frame and frequency index to represent any segment of the above information.
[0035] By performing a column vector transformation on the above scalar representation formula, we obtain the column vector representation formula;
[0036] Based on the column vector representation formulas of multiple segments, the column vector formulas corresponding to the above information are obtained;
[0037] Based on the column vector formulas corresponding to the noise information and the noisy speech information, a noisy speech covariance matrix is generated, wherein the noisy speech covariance matrix includes the noise correlation matrix and the speech correlation matrix.
[0038] Another aspect of this disclosure provides a causal adaptive single-channel speech noise reduction apparatus, comprising:
[0039] A voice acquisition device is used to acquire noise information in a first time period and noisy voice information in a second time period;
[0040] A sound event detection device is used to process the aforementioned noise information and the aforementioned noisy speech information to obtain a noise correlation matrix and a speech correlation matrix;
[0041] Filters are used for:
[0042] The noise correlation matrix and speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors;
[0043] Based on the aforementioned initial feature vectors and the filtering function of the initial filter, a target filter vector is generated; and
[0044] An output device is used to generate enhanced target output speech information based on the target filter vector and the noisy speech information.
[0045] According to embodiments of this disclosure, the causal adaptive single-channel speech noise reduction device further includes:
[0046] A first adjustment device is used to adjust a preset selection rule to filter multiple target feature values from multiple initial feature values using the adjusted preset selection rule, so that the filter generates the target filter vector based on the multiple target feature values, wherein the initial feature values are obtained by generalized eigenvalue decomposition of the noise correlation matrix and the speech correlation matrix; and / or
[0047] The second adjustment device is used to adjust the value of the balance factor so as to generate the target filter vector using the adjusted balance factor.
[0048] According to embodiments of this disclosure, noise information and noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix. Generalized eigenvalue decomposition is then performed on the noise correlation matrix and the speech correlation matrix. The filtering function of the initial filter is adjusted based on the initial eigenvector obtained from the decomposition, thereby obtaining a target filtering vector for speech enhancement. The enhanced target output speech information can be obtained by performing an inner product between the target filtering vector and the noisy speech. Since the target filtering vector is independent of the noisy speech information and depends only on the noise correlation matrix and the speech correlation matrix, the target output speech information achieves a significant improvement in speech quality and speech intelligibility. Attached Figure Description
[0049] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0050] Figure 1 A flowchart illustrating a causal adaptive single-channel speech noise reduction method according to an embodiment of the present disclosure is shown schematically.
[0051] Figure 2The effect of the STFT domain balance factor μ on noise reduction performance is illustrated schematically according to embodiments of the present disclosure;
[0052] Figure 3 The illustration schematically shows the effect of rank r and balance factor μ on noise reduction performance in the STFT domain according to an embodiment of the present disclosure (filter length L = 8);
[0053] Figure 4 This illustration schematically shows the effect of filter length L on noise reduction performance according to an embodiment of the present disclosure;
[0054] Figure 5 This illustration schematically shows the effect of the input signal-to-noise ratio on noise reduction performance according to embodiments of the present disclosure;
[0055] Figure 6 A block diagram of a causal adaptive single-channel speech noise reduction apparatus according to an embodiment of the present disclosure is illustrated schematically. Detailed Implementation
[0056] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0057] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0058] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0059] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (e.g., "a device having at least one of A, B and C" should include, but is not limited to, a device having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B and C, etc.).
[0060] The single-channel Wiener filtering method designs the filter in the short-time frequency domain by minimizing the mean-square error (MSE) between the output signal and the reference target sound source signal. Mathematically, the traditional single-channel Wiener filter relies on the autocorrelation matrix of the noisy speech and the autocorrelation matrix of the noise signal. Although this method yields the smallest MSE between the output signal and the reference sound source signal, the resulting enhanced signal exhibits spectral distortion because it does not control for the distortion of the sound source signal, thus affecting speech intelligibility and listening comfort.
[0061] In practical applications, it is commonly observed that the short-time Fourier transform (STFT) coefficients of speech signals are correlated between adjacent time frames and frequencies, especially when the analysis window length and / or frame overlap rate are small. This can be very useful for tracking the dynamic changes of the target speech source. For example, some literature considers the speech correlation between time frames, some studies the correlation across frequency subbands, and some considers the correlation between time frames and frequencies simultaneously. For each time-frequency point, given a memory size (equal to the FIR filter length), by constructing a vector from the multi-frame STFT coefficients, the time-varying speech correlation matrix and speech correlation vector can be estimated using recursive smoothing techniques, similar to the speech covariance matrix and relative acoustic transfer function in multi-microphone noise reduction. Therefore, multi-frame single-channel noise reduction filters can be designed similarly to multi-channel beamformers.
[0062] On the other hand, environmental noise is a significant factor affecting speech enhancement performance. Utilizing single-channel speech enhancement as a front-end module to improve the quality of the audio signal to be recognized is an important means to improve speech recognition rate. One related technique proposes a multi-frame minimum variance distortionless response (MFMVDR) single-channel denoising FIR filter, similar to a multi-microphone MVDR beamformer, where the speech correlation vector (SCV) is estimated using the normalized first column of the speech correlation matrix. Other related techniques address single-microphone and multi-microphone denoising problems by proposing maximum signal-to-noise ratio (SNR) filters, which highly distort the target speech in both the time and STFT domains, despite maximizing the output SNR. To address this, some researchers have proposed a speech distortion-weighted inter-frame Wiener filter (SDW-IFWF), which is similar in form to a speech distortion-weighted multi-channel Wiener filter.
[0063] However, in practical applications, the speech enhancement methods proposed in the aforementioned related technologies have poor performance in terms of speech quality and speech intelligibility, making it difficult to achieve a balance between the two.
[0064] In view of this, embodiments of the present disclosure provide a causal adaptive single-channel speech denoising method and apparatus. The method includes: acquiring noise information in a first time period and noisy speech information in a second time period; processing the noise information and noisy speech information to obtain a noise correlation matrix and a speech correlation matrix; performing generalized eigenvalue decomposition on the noise correlation matrix and the speech correlation matrix to obtain multiple initial feature vectors; generating a target filter vector based on the multiple initial feature vectors and the filter function of the initial filter; and generating enhanced target output speech information based on the target filter vector and the noisy speech information.
[0065] Figure 1 A flowchart of a causal adaptive single-channel speech denoising method according to an embodiment of the present disclosure is illustrated schematically.
[0066] like Figure 1 As shown, the causal adaptive single-channel speech denoising method includes operations S101 to S105.
[0067] During operation S101, noise information in the first time period and noisy speech information in the second time period are collected.
[0068] In operation S102, noise information and noisy speech information are processed to obtain noise correlation matrix and speech correlation matrix;
[0069] In operation S103, generalized eigenvalue decomposition is performed on the noise correlation matrix and the speech correlation matrix to obtain multiple initial feature vectors.
[0070] In operation S104, the target filter vector is generated based on multiple initial feature vectors and the filter function of the initial filter.
[0071] In operation S105, enhanced target output speech information is generated based on the target filter vector and the noisy speech information.
[0072] According to embodiments of this disclosure, the first time period is earlier than the second time period; for example, the first time period may be a few seconds before the second time period.
[0073] According to embodiments of this disclosure, firstly, noise information of pure noise for several seconds is collected, i.e., the target speech source remains silent during this time, to estimate the initial noise correlation matrix; then, the noise information and noisy speech information are processed, and the noise correlation matrix and speech correlation matrix are updated in real time using the moving average technique in the noise frame and speech frame, respectively.
[0074] According to embodiments of this disclosure, a generalized eigenvalue decomposition is performed on the speech autocorrelation matrix and the noise covariance matrix. Using the generalized eigenvalues and their corresponding eigenvectors, a low-rank approximation can be performed on the speech covariance matrix. Substituting the eigenvalues and eigenvectors into the filtering function of the initial filter yields the target filtering vector of the target filter based on the rank of the covariance matrix, the generalized eigenvalues, and the eigenvectors. By selecting different ranks, this target filter can be transformed into classic single-channel speech denoising filters such as single-channel Wiener filters, maximum signal-to-noise ratio (maxSNR), and multi-frame minimum variance distortionless response filters.
[0075] According to embodiments of this disclosure, the enhanced target output speech information can be obtained by processing the noisy speech information using the target filter vector. For example, the target output speech information can be obtained by performing an inner product operation between the noisy speech vector corresponding to the noisy speech information and the target filter vector, and then performing an inverse short-time Fourier transform.
[0076] According to embodiments of this disclosure, noise information and noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix. Generalized eigenvalue decomposition is then performed on the noise correlation matrix and the speech correlation matrix. The filtering function of the initial filter is adjusted based on the initial eigenvector obtained from the decomposition, thereby obtaining a target filtering vector for speech enhancement. The enhanced target output speech information can be obtained by performing an inner product between the target filtering vector and the noisy speech. Since the target filtering vector is independent of the noisy speech information and depends only on the noise correlation matrix and the speech correlation matrix, the target output speech information achieves a significant improvement in speech quality and speech intelligibility.
[0077] According to embodiments of this disclosure, generalized eigenvalue decomposition is performed on the noise correlation matrix and the speech correlation matrix to obtain multiple initial feature vectors, including:
[0078] Generalized eigenvalue decomposition is performed on the noise correlation matrix and the speech correlation matrix to obtain multiple initial eigenvalues;
[0079] Each initial feature value is transformed to obtain multiple initial feature vectors.
[0080] According to embodiments of this disclosure, for the noise correlation matrix Φ nn and speech correlation matrix Φ xx Generalized eigenvalue decomposition is performed to obtain multiple initial eigenvalues λ, which are arranged in descending order as λ1≥λ2≥…≥λ L Each initial feature value is transformed to obtain multiple initial feature vectors u.
[0081] According to embodiments of this disclosure, a target filter vector is generated based on a plurality of initial feature vectors and a filter function of an initial filter, including:
[0082] The initial filter vector is generated based on the filter function, noise correlation matrix, speech correlation matrix, and balance factor μ.
[0083] Based on multiple initial eigenvectors and the initial eigenvalues corresponding to each initial eigenvector, a vector matrix U and a diagonal matrix Λ are generated respectively.
[0084] Multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values;
[0085] A target filter is generated based on multiple target eigenvalues, a vector matrix U, and a diagonal matrix Λ, wherein the target filter includes a target filter vector.
[0086] According to embodiments of this disclosure, multiple initial feature vectors u are stored in a vector matrix U = [u1, u2, ..., u...]. M Based on the initial eigenvalue λ corresponding to each initial eigenvector u, a diagonal matrix Λ is generated, where the diagonal elements of the diagonal matrix Λ are the initial eigenvalues λ after generalized eigenvalue decomposition.
[0087] According to embodiments of this disclosure, multiple initial feature values λ are filtered based on preset selection rules to obtain multiple target feature values. A target filter is generated based on the multiple target feature values λ, the vector matrix U, and the diagonal matrix Λ, wherein the target filter includes a target filter vector w.
[0088] According to embodiments of this disclosure, multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values, including:
[0089] Multiple initial feature values are sorted according to a preset sorting rule to obtain sorted initial feature values;
[0090] Based on preset selection rules, multiple target feature values are selected from multiple sorted initial feature values.
[0091] According to embodiments of this disclosure, multiple initial feature values λ can be arranged in descending order to obtain sorted initial feature values, namely λ1≥λ2≥…≥λ L .
[0092] According to embodiments of this disclosure, a preset selection rule, for example, refers to selecting the first r initial feature values λ from a plurality of sorted initial feature values as target feature values, thereby obtaining a target feature vector corresponding to each target feature value. The value of r can be adjusted according to actual conditions, for example, it can be 2 or 3.
[0093] According to embodiments of this disclosure, a target filter is generated based on multiple target feature values, a vector matrix U, and a diagonal matrix Λ, wherein the target filter includes a target filter vector, comprising:
[0094] By jointly diagonalizing the vector matrix U and the diagonal matrix Λ, we obtain the first diagonalization formula;
[0095] The second diagonalization formula is generated based on the first diagonalization formula and the noisy covariance matrix.
[0096] The target filtering vector is generated based on multiple target feature values and the second diagonalization formula.
[0097] According to embodiments of this disclosure, based on generalized eigenvalues, the noise correlation matrix Φ nn and speech correlation matrix Φ xx The combined diagonalization can be transformed into the first diagonalization formula as shown in formula (1):
[0098] U H Φ xx U = Λ, U H Φ nn U=I (1)
[0099] Where H represents the conjugate transpose, x represents pure speech information, n represents noise information, I is the identity matrix, and y represents noisy speech information.
[0100] According to embodiments of this disclosure, due to the noisy covariance matrix Φ yy =Φ xx +Φ nn Therefore, the noisy covariance matrix Φ yy It can also be diagonalized to the second diagonalization formula shown in formula (2):
[0101] U H Φ yy U=Λ+I (2)
[0102] According to embodiments of this disclosure, therefore, in practice we can utilize Φ yy and Φ nn The generalized eigenvalue decomposition implements the joint diagonalization operation. Based on Φ xx The diagonalization operation yields formula (3):
[0103]
[0104] Where Q = [q1, q2, ..., q L ] = U -H It can be seen that Φ nn =QQ H ,and Therefore, Q = [q1, q2, ..., qL [Contains a matrix] The left eigenvector.
[0105] According to embodiments of this disclosure, a target filter vector w can be generated based on multiple target feature values and formula (3) derived from the second diagonalization formula.
[0106] According to embodiments of this disclosure, a target filtering vector is generated based on multiple target feature values and a second diagonalization formula, including:
[0107] For any correlation matrix in the noise correlation matrix and the speech correlation matrix, the correlation matrix is approximated by using multiple target feature values to obtain an approximate correlation matrix;
[0108] The target filter vector is generated based on the approximate correlation matrix, the second diagonalization formula, and the initial filter vector, wherein the target filter vector includes a balance factor.
[0109] According to the embodiments of this disclosure, since the speech distortion weighted Wiener filter SDW-SWF is more general, the filter design of SDW-SWF is described below. The design criterion is to minimize the target sound source mean square error plus the weighted residual noise power, i.e., formula (4):
[0110] min w ε[|w0 H xX(k,l)| 2 ]+με[|w0 H n| 2 (4)
[0111] Where w0 represents the initial filter vector, and μ≥0 is the balance factor between speech enhancement performance and speech distortion. Through derivation, the expression for the initial filter vector of this Wiener filter can be obtained as formula (5):
[0112] w0=(Φ xx +μΦ nn ) -1 Φ xx e (5)
[0113] Here, e is the selection vector, with the first element being 1 and all other elements being 0. Clearly, when μ = 0, this filter is equivalent to the classical Wiener filter.
[0114] According to embodiments of this disclosure, this disclosure utilizes the first r target feature values and the corresponding target feature vectors to analyze Φ. xx By approximation, we obtain the approximate correlation matrix shown in formula (6):
[0115]
[0116] The approximate correlation matrix of rank r Substituting the initial filter vector w0 into the original SDW-SWF filter, we can obtain the Wiener filter based on the low-rank approximation, i.e., the target filter vector as shown in equation (7):
[0117]
[0118] Where * denotes complex conjugation, q i1 Represents vector q i The i-th element. It can be seen that when r = L and the balance factor μ = 1, w = e represents a unit filter, and the output signal is equivalent to the input signal, that is, there is no noise reduction operation; when r = 1, SDW-SWF is equivalent to a rank-1 filter, and when μ = 0, it degenerates into a maxSNR filter. In addition, it can be proved that the MFMVDR filter is also a special case of rank-1 SDW-SWF.
[0119] According to embodiments of this disclosure, enhanced target output speech information is generated based on a target filter vector and noisy speech information, including:
[0120] The target filter vector is conjugate transposed to obtain the transpose vector;
[0121] The inner product operation is performed on the noisy speech information and the transpose vector to obtain the initial output speech information;
[0122] The initial output speech information is processed by inverse short-time Fourier transform to obtain the target output speech information.
[0123] According to embodiments of this disclosure, after selecting appropriate rank r and balance factor μ as required, the target filter vector shown in formula (7) is subjected to conjugate transpose processing to obtain the transpose vector w. H The transpose vector w H The inner product operation is performed on the noisy speech signal vector y corresponding to the noisy speech information, that is, beamforming is performed on each frequency point to obtain the initial output speech information in the frequency domain, as shown in formula (8):
[0124]
[0125] After inverse short-time Fourier transform, the speech signal of the target speaker in the time domain can be recovered.
[0126] According to embodiments of this disclosure, the target output speech information of the target speaker in the time domain can be recovered by performing inverse short-time Fourier transform processing on the initial output speech information.
[0127] According to embodiments of this disclosure, noise information and noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix, including:
[0128] For any information in the noise information and the noisy speech information, any segment of the information is represented by the frame and frequency index to obtain the scalar representation formula;
[0129] By performing a column vector transformation on the scalar representation formula, we obtain the column vector representation formula;
[0130] Based on the column vector representation formulas of multiple segments, the column vector formulas corresponding to the information are obtained;
[0131] Based on the column vector formulas corresponding to noise information and noisy speech information, a noisy speech covariance matrix is generated, which includes a noise correlation matrix and a speech correlation matrix.
[0132] According to embodiments of this disclosure, for noisy speech information, l and k represent the frame and frequency index respectively in the short time frequency domain. Then, the noisy speech information can be represented by the scalar representation formula shown in formula (9):
[0133] Y(k,l)=X(k,l)+N(k,l) (9)
[0134] Where X(k,l) and N(k,l) represent the pure sound source component and the noise component (including interfering sound sources, background noise, reverberation, and microphone self-noise, etc.), respectively.
[0135] According to embodiments of this disclosure, by performing column vector transformation on the scalar representation formula, the STFT domain signal model can be written as a column vector representation formula as shown in formula (10).
[0136] y(k,l)=x(k,l)+n(k,l) (10)
[0137] According to an embodiment of this disclosure, for each frequency index k, the STFT coefficients of the L most recently received time frames can be stacked in the vector shown below to obtain the column vector formula as shown in formula (11).
[0138] y(k,l)=[y(k,l),y(k,l-1),...,y(k,l-L+1)] T (11)
[0139] According to embodiments of this disclosure, the method of constructing signal vectors using historical observation data ensures the causal characteristics of the designed filter.
[0140] According to embodiments of this disclosure, the time-frequency index (k, l) is omitted for ease of expression. It is assumed that the target sound source and the noise components are uncorrelated, thus the noisy covariance matrix Φ yy It can be written as the noise correlation matrix Φnn and speech correlation matrix Φ xx The summation form is formula (12).
[0141] Φ yy =ε[yy H ]=ε[xx H ]+ε[nn H ]=Φ xx +Φ nn (12)
[0142] Where, Φ xx Φ represents the speech covariance matrix. nn Let represent the noise covariance matrix, and ε represent the mean operation.
[0143] According to embodiments of this disclosure, in order to more clearly understand the relationship between output signal-to-noise ratio and input signal-to-noise ratio, this disclosure uses a variable rank r to analyze the influence of selecting different rank and different balance factors μ on the output signal-to-noise ratio. The input signal-to-noise ratio and output signal-to-noise ratio of the microphone are defined by formulas (13) and (14), respectively.
[0144]
[0145]
[0146] According to the embodiments of this disclosure, it can be seen from formulas (13) and (14) that the output signal-to-noise ratio decreases with rank r, and the output signal-to-noise ratio of the full-rank SDW-SWF filter is always greater than or equal to the input signal-to-noise ratio, as shown in formula (15):
[0147]
[0148] According to an embodiment of this disclosure, on the other hand, speech distortion (SD) can be calculated using formula (16).
[0149]
[0150] According to the embodiments of this disclosure, it can be seen from formula (16) that the larger the balance factor μ, the more severe the distortion of the output speech signal; the larger the rank of the speech correlation matrix, the smaller the signal distortion. Therefore, this disclosure can adjust the noise reduction level and the output signal quality from two perspectives: the balance factor and the rank of the signal correlation matrix.
[0151] According to embodiments of this disclosure, the causal adaptive single-channel speech denoising method provided in this disclosure is verified using numerical simulation. For this purpose, a clean speech signal is mixed with a Babble noise source to generate a noisy microphone signal, i.e., noisy speech information. The speech source is from the TIMIT database, and the noise signal is from the NoiseX-92 dataset. The sampling frequency is 16kHz. An 8ms Hamming window with a 75% overlap factor is used to segment the noise signal, and the time-domain signal is converted to the STFT domain using a 128-point FFT. The proposed causal adaptive single-channel speech denoising method will be compared with the maximum signal-to-noise ratio (maxSNR) filter and MFMVDR method in related techniques.
[0152] In addition to speech distortion, noise variance, and output signal-to-noise ratio (SNR) mentioned above, performance evaluation can also use short-time objective intelligibility (STOI) and perceptual evaluation of speech quality (PESQ) to assess the speech intelligibility of a device. STOI measures the mutual information between the enhanced signal and the clean reference signal, ranging from 0 to 1. A higher value indicates that the enhanced signal is closer to the clean source signal, and it is widely used to evaluate speech intelligibility. PESQ is one of the most commonly used metrics for evaluating speech quality. The calculation process includes preprocessing, time alignment, perceptual filtering, masking effects, etc., and its value ranges from -0.5 to 4.5. A higher PESQ value indicates better perceived speech quality.
[0153] Figure 2 The effect of the STFT domain balance factor μ on noise reduction performance is illustrated schematically according to embodiments of the present disclosure; Figure 3 The illustration schematically shows the effect of rank r and balance factor μ on noise reduction performance in the STFT domain according to an embodiment of the present disclosure (filter length L = 8); Figure 4 This illustration schematically shows the effect of filter length L on noise reduction performance according to an embodiment of the present disclosure; Figure 5 The illustration schematically shows the effect of the input signal-to-noise ratio on noise reduction performance according to embodiments of the present disclosure.
[0154] According to embodiments of this disclosure, the noise reduction effect is first measured using output signal-to-noise ratio (oSNR), speech distortion, and residual noise power. Figure 2 It can be seen that the output signal-to-noise ratio and speech distortion decrease with increasing rank r and increase with increasing balance factor μ, while the residual noise power shows the opposite trend. Among these, Figure 2 (a) The vertical axis represents the output signal-to-noise ratio. Figure 2(b) The vertical axis represents speech distortion. Figure 2 (c) The vertical axis represents the residual noise power.
[0155] According to embodiments of this disclosure, by Figure 3 It can be clearly observed that the output signal-to-noise ratio (SNR) of SDW-SWF (i.e., the method proposed in this disclosure) is always greater than the input SNR, regardless of the rank r or the balance factor μ. This means that noise reduction can always be achieved using the proposed method. The STOI (Short-Time Objective Intelligibility) obtained by SDW-SWF (i.e., the method proposed in this disclosure) generally increases with increasing rank and decreases with increasing balance factor μ (i.e., the opposite of speech distortion), a trend different from PESQ. SDW-SWF achieves the maximum PESQ (Objective Speech Quality Assessment) when the rank is equal to 3 and the balance factor is greater than 3 and less than 10. Overall, although a larger rank reduces speech signal quality (i.e., output SNR), it may help improve speech intelligibility, which has a higher positive correlation with speech distortion. Compared with the maxSNR filter for highly distorted target speech, the proposed method achieves lower speech distortion and better speech intelligibility performance when the rank is greater than or equal to 2. The method of this disclosure and maxSNR can achieve the same output SNR, which is higher than MFMVDR. Compared to MFMVDR, the obtained speech distortion, STOI, and PESQ are significantly better when the rank is greater than or equal to 2 and the balance factor is between 0 and 1000, although the signal-to-noise ratio may be lower. In the following experiments, a fixed balance factor of 3 was set, which achieved a good balance between noise reduction capabilities (e.g., output signal-to-noise ratio and residual noise variance) and speech quality (e.g., speech distortion SD, STOI, and PESQ).
[0156] According to embodiments of this disclosure, Figure 4 It is evident that filter length has a significant impact on noise reduction performance. Increasing the filter length enhances noise reduction capability, but increases speech distortion, leading to a decrease in both STOI and PESQ. Similarly, for the method proposed in this disclosure, higher rank results in poorer noise reduction capability but better speech quality. Compared to MFMVDR, although SDW-SWF achieves better output signal-to-noise ratio and residual noise power, its speech distortion and intelligibility are both poorer. More importantly, increasing the rank can significantly improve performance. The maxSNR filter severely distorts clean speech, resulting in the worst PESQ and STOI results.
[0157] According to embodiments of this disclosure, Figure 5The performance comparison is presented using input signal-to-noise ratio (SNR) as the metric, with a balance factor set to 3 and a filter length of 8. It can be seen that a higher input SNR leads to a higher output SNR, lower residual noise variance, and higher intelligibility scores for PESQ and STOI, while speech distortion varies slightly with input SNR. Compared to MFMVDR, the method proposed in this disclosure has stronger noise reduction potential because, with iSNR ≥ 2dB, the STOI of MFMVDR becomes smaller than the input STOI, and with iSNR ≥ 15dB, the obtained PESQ becomes smaller than the noisy PESQ. These thresholds for SDW-SWF are much larger than those of MFMVDR. However, for large input SNRs, improving speech intelligibility is more difficult than improving the SNR, as clear speech can easily become distorted. In summary, the STFT domain SDW-SWF method proposed in this disclosure is more robust than other comparative methods, achieving a better balance between speech quality and speech intelligibility.
[0158] In summary, the advantages of the method proposed in this disclosure are as follows: First, in the design of a single-channel Wiener filter, the relationship between the output signal quality and the rank and balance factor of the speech subcorrelation matrix can be analyzed, and the filter form can be flexibly selected by adjusting the rank and balance factor. Second, in the subspace, SDW-SWF can be written as a linear combination of the generalized eigenvectors of the speech and noise covariance matrices. Choosing different ranks for the speech covariance matrix or balance factor will result in some existing single-channel noise reduction filters. Experimental results show that in this method, the larger the balance factor, the stronger the noise reduction capability, but the lower the speech intelligibility; the larger the rank, the lower the noise reduction capability, but the better the speech intelligibility. The proposed SDW-SWF method provides more flexibility to achieve the desired balance between device speech quality and speech intelligibility.
[0159] Figure 6 A block diagram of a causal adaptive single-channel speech noise reduction apparatus according to an embodiment of the present disclosure is illustrated schematically.
[0160] like Figure 6 As shown, the causal adaptive single-channel speech noise reduction device 600 includes:
[0161] The voice acquisition device 610 is used to acquire noise information in a first time period and noisy voice information in a second time period.
[0162] The sound event detection device 620 is used to process noise information and noisy speech information to obtain a noise correlation matrix and a speech correlation matrix;
[0163] Filter 630, used for:
[0164] Generalized eigenvalue decomposition is performed on the noise correlation matrix and the speech correlation matrix to obtain multiple initial feature vectors;
[0165] Based on multiple initial feature vectors and the filtering function of the initial filter, a target filter vector is generated; and
[0166] The output device 640 is used to generate enhanced target output speech information based on the target filter vector and the noisy speech information.
[0167] According to embodiments of this disclosure, the causal adaptive single-channel speech noise reduction device 600 can be installed on recording equipment in hearing aids, human-computer interaction devices, or conference rooms. It collects noisy speech information containing the target sound source through speech acquisition devices such as microphones, and finally outputs target output speech information with better speech quality and speech intelligibility.
[0168] According to embodiments of this disclosure, noise information and noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix. Generalized eigenvalue decomposition is then performed on the noise correlation matrix and the speech correlation matrix. The filtering function of the initial filter is adjusted based on the initial eigenvector obtained from the decomposition, thereby obtaining a target filtering vector for speech enhancement. The enhanced target output speech information can be obtained by performing an inner product between the target filtering vector and the noisy speech. Since the target filtering vector is independent of the noisy speech information and depends only on the noise correlation matrix and the speech correlation matrix, the target output speech information achieves a significant improvement in speech quality and speech intelligibility.
[0169] According to embodiments of this disclosure, the causal adaptive single-channel speech noise reduction device 600 further includes:
[0170] A first adjustment device is used to adjust a preset selection rule to filter multiple target feature values from multiple initial feature values using the adjusted preset selection rule, so that the filter generates a target filter vector based on the multiple target feature values, wherein the initial feature values are obtained by generalized eigenvalue decomposition of the noise correlation matrix and the speech correlation matrix; and / or
[0171] The second adjustment device is used to adjust the value of the balance factor so as to generate the target filter vector using the adjusted balance factor.
[0172] According to embodiments of this disclosure, both the first adjustment device and the second adjustment device can be in the form of a knob or other forms, which can change the number of selected target feature values and the value of the balance factor, thereby changing the target filter vector in real time, and thus obtaining target output speech information with better speech quality and speech intelligibility.
[0173] It should be noted that the causal adaptive single-channel speech denoising device part in the embodiments of this disclosure corresponds to the causal adaptive single-channel speech denoising method part in the embodiments of this disclosure. For a detailed description of the causal adaptive single-channel speech denoising device part, please refer to the causal adaptive single-channel speech denoising method part, which will not be repeated here.
[0174] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A causal adaptive single-channel speech denoising method, comprising: Noise information in the first time period and noisy speech information in the second time period were collected; The noise information and the noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix; The noise correlation matrix and the speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors; Based on the filtering functions of the multiple initial feature vectors and the initial filter, a target filter vector is generated; Based on the target filter vector and the noisy speech information, an enhanced target output speech information is generated; The process of generating a target filter vector based on the filtering functions of the multiple initial feature vectors and the initial filter includes: An initial filter vector is generated based on the filter function, the noise correlation matrix, the speech correlation matrix, and the balance factor. Based on the plurality of initial feature vectors and the initial feature values corresponding to each initial feature vector, a vector matrix and a diagonal matrix are generated respectively; Multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values; A target filter is generated based on multiple target feature values, vector matrices, and diagonal matrices, wherein the target filter includes the target filter vector; Specifically, a target filter is generated based on multiple target feature values, a vector matrix, and a diagonal matrix, wherein the target filter includes the target filter vector, comprising: By jointly diagonalizing the vector matrix and the diagonal matrix, the first diagonalization formula is obtained; Based on the first diagonalization formula and the noisy covariance matrix, a second diagonalization formula is generated. The target filter vector is generated based on the multiple target feature values and the second diagonalization formula; The process of generating the target filter vector based on multiple target feature values and the second diagonalization formula includes: For any correlation matrix in the noise correlation matrix and the speech correlation matrix, the correlation matrix is approximated using multiple target feature values to obtain an approximate correlation matrix; The target filter vector is generated based on the approximate correlation matrix, the second diagonalization formula, and the initial filter vector, wherein the target filter vector includes the balance factor.
2. The method according to claim 1, wherein, The noise correlation matrix and the speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors, including: The noise correlation matrix and the speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial eigenvalues; Each of the initial feature values is transformed to obtain multiple initial feature vectors.
3. The method according to claim 1, wherein, Multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values, including: The initial feature values are sorted according to a preset sorting rule to obtain sorted initial feature values. Based on the preset selection rules, multiple target feature values are selected from the sorted initial feature values.
4. The method according to claim 1, wherein, Based on the target filter vector and the noisy speech information, enhanced target output speech information is generated, including: The target filter vector is conjugate transposed to obtain the transposed vector; Perform an inner product operation on the noisy speech information and the transpose vector to obtain the initial output speech information; The initial output speech information is processed by inverse short-time Fourier transform to obtain the target output speech information.
5. The method according to claim 1, wherein, The noise information and the noisy speech information are processed to obtain a noise correlation matrix and a speech correlation matrix, including: For any of the noise information and the noisy speech information, a scalar representation formula is obtained by representing any segment of the information using frame and frequency indexes. Perform a column vector transformation on the scalar representation formula to obtain the column vector representation formula; Based on the column vector representation formulas of multiple segments, the column vector formula corresponding to the information is obtained; Based on the column vector formulas corresponding to the noise information and the noisy speech information, a noisy speech covariance matrix is generated, wherein the noisy speech covariance matrix includes the noise correlation matrix and the speech correlation matrix.
6. A causal adaptive single-channel speech noise reduction device, comprising: A voice acquisition device is used to acquire noise information in a first time period and noisy voice information in a second time period; A sound event detection device is used to process the noise information and the noisy speech information to obtain a noise correlation matrix and a speech correlation matrix; Filters are used for: The noise correlation matrix and the speech correlation matrix are subjected to generalized eigenvalue decomposition to obtain multiple initial feature vectors; Based on the filtering functions of the multiple initial feature vectors and the initial filter, a target filter vector is generated; as well as An output device is used to generate enhanced target output speech information based on the target filter vector and the noisy speech information; The process of generating a target filter vector based on the filtering functions of the multiple initial feature vectors and the initial filter includes: An initial filter vector is generated based on the filter function, the noise correlation matrix, the speech correlation matrix, and the balance factor. Based on the plurality of initial feature vectors and the initial feature values corresponding to each initial feature vector, a vector matrix and a diagonal matrix are generated respectively; Multiple initial feature values are filtered based on preset selection rules to obtain multiple target feature values; A target filter is generated based on multiple target feature values, vector matrices, and diagonal matrices, wherein the target filter includes the target filter vector; Specifically, a target filter is generated based on multiple target feature values, a vector matrix, and a diagonal matrix, wherein the target filter includes the target filter vector, comprising: By jointly diagonalizing the vector matrix and the diagonal matrix, the first diagonalization formula is obtained; Based on the first diagonalization formula and the noisy covariance matrix, a second diagonalization formula is generated. The target filter vector is generated based on the multiple target feature values and the second diagonalization formula; The process of generating the target filter vector based on multiple target feature values and the second diagonalization formula includes: For any correlation matrix in the noise correlation matrix and the speech correlation matrix, the correlation matrix is approximated using multiple target feature values to obtain an approximate correlation matrix; The target filter vector is generated based on the approximate correlation matrix, the second diagonalization formula, and the initial filter vector, wherein the target filter vector includes the balance factor.
7. The apparatus according to claim 6, further comprising: A first adjustment device is used to adjust a preset selection rule to filter multiple target feature values from multiple initial feature values using the adjusted preset selection rule, such that the filter generates the target filter vector based on the multiple target feature values, wherein the initial feature values are obtained by generalized eigenvalue decomposition of the noise correlation matrix and the speech correlation matrix; and / or The second adjustment device is used to adjust the value of the balance factor so as to generate the target filter vector using the adjusted balance factor.