Multi-channel filtering method and system for optimizing full-band signal-to-noise ratio based on inter-frame correlation

By introducing inter-frame correlation and full-band signal-to-noise ratio optimization into a multi-channel speech system, the problems of temporal discontinuity and speech distortion in multi-channel speech denoising methods under complex environments are solved, achieving stability and speech enhancement effects under complex noise environments.

CN121963761APending Publication Date: 2026-05-01SHAANXI UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610135177.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing multi-channel speech denoising methods are prone to time-domain discontinuity or increased speech distortion in non-stationary noise environments or when speech rate changes rapidly. Furthermore, they lack the ability to optimize the full-band signal-to-noise ratio, making it difficult to balance noise suppression and speech fidelity.

Method used

The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation constructs a multi-channel speech system, uses a microphone array to collect signals and convert them to the time-frequency domain, introduces inter-frame correlation, calculates the correlation matrix and performs eigenvalue decomposition, distinguishes between speech-dominant and noise-dominant subspaces, applies a full-band output signal-to-noise ratio filter for noise reduction and enhancement, and sets the noise-dominant subspace to zero.

Benefits of technology

It improves the stability and robustness of the filter in complex noise environments, enhances the naturalness of speech and overall listening quality, significantly improves speech enhancement effect, and balances noise suppression and speech fidelity performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963761A_ABST
    Figure CN121963761A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-channel filtering method for optimizing a full-band signal-to-noise ratio based on inter-frame correlation, and the method specifically comprises the steps: collecting a band correlation signal in a multi-channel voice system, and converting the band correlation signal to a time-frequency domain; generating a multi-channel vector signal containing inter-frame information; carrying out noise reduction processing on the multi-channel noisy voice signal containing the inter-frame information; calculating a correlation matrix of the multi-channel clean voice signal containing the inter-frame information and the multi-channel noise signal containing the inter-frame information; voice dominant and noise dominant subspaces are distinguished according to the sizes of the characteristic values, the voice dominant subspace with the maximum characteristic value is subjected to noise reduction and enhancement, and other noise dominant subspaces are directly set to zero; and forming a full-band filter with the length of 1 * 10. According to the method, the noise reduction stability, robustness and speech enhancement effect of the filter in a complex noise environment are remarkably improved. The invention further discloses a multi-channel filtering system for optimizing the full-band signal-to-noise ratio based on the inter-frame correlation.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation Technical Field

[0001] This invention belongs to the field of speech noise reduction technology, specifically relating to an optimized full-band signal-to-noise ratio multi-channel filtering method based on inter-frame correlation, and also relating to an optimized full-band signal-to-noise ratio multi-channel filtering system based on inter-frame correlation. Background Technology

[0002] With the rapid development of applications such as voice communication, speech recognition, remote conferencing, and smart terminals, the problem of speech quality in complex acoustic environments is becoming increasingly prominent. In practical applications, speech signals are often affected by background noise during acquisition and transmission, leading to decreased speech intelligibility and poorer listening quality, which severely restricts the performance of speech-related systems.

[0003] To improve the quality of noisy speech, various speech denoising methods have been proposed in existing technologies. Among them, single-channel speech enhancement methods, such as spectral subtraction, Wiener filtering, subspace methods, and maximum signal-to-noise ratio filtering, have been widely studied due to their simple structure and ease of implementation. These methods typically rely on the difference in the statistical characteristics of speech and noise in the time-frequency domain to achieve noise suppression, and can achieve certain results under steady-state noise conditions. However, because single-channel methods cannot utilize spatial information, their denoising performance and stability are often significantly limited in non-stationary noise, strong reverberation, or complex dynamic scenes.

[0004] To overcome the shortcomings of single-channel methods, multi-channel speech denoising technology has gradually become a research hotspot. By introducing multiple microphones, multi-channel methods can utilize the spatial differences between speech and noise to improve the enhancement capability of the target speech. Existing multi-channel speech denoising methods include beamforming, multi-channel Wiener filtering, and maximum signal-to-noise ratio (MSNR) filtering. Among them, the MSNR filtering method achieves a strong noise suppression effect by maximizing the signal-to-noise ratio between the filtered speech signal and the residual noise, and has significant value in theoretical analysis and engineering applications.

[0005] However, existing multi-channel maximum signal-to-noise ratio (MSNR) filtering methods still have certain limitations. On the one hand, these methods typically process each frame independently or frame-by-frame when designing filters, failing to fully utilize the correlation between adjacent time frames of the speech signal. This leads to problems such as temporal discontinuities or increased speech distortion in non-stationary noise environments or when the speech rate changes rapidly. On the other hand, existing methods often focus on enhancing a single frequency band or dominant feature subspace, while directly suppressing or zeroing other frequency bands. Although this can improve the output SNR, it easily causes loss of speech details, affecting the naturalness of the speech and the overall listening quality.

[0006] In addition, some existing multi-channel speech denoising methods only optimize local frequency bands or single subspaces, lacking the ability to optimize the output signal-to-noise ratio from a full-band perspective. In complex acoustic environments, it is difficult to balance noise suppression and speech fidelity.

[0007] Therefore, there is an urgent need for a multi-channel speech denoising method that can simultaneously utilize the inter-frame correlation features of speech signals and optimize the output signal-to-noise ratio across the entire bandwidth, so as to improve the stability, robustness, and speech enhancement effect of the filter in complex noise environments. Summary of the Invention

[0008] The first objective of this invention is to provide an optimized full-band signal-to-noise ratio multi-channel filtering method based on inter-frame correlation, which effectively avoids the loss of speech details and increases the naturalness of speech and overall listening quality; it balances noise suppression effect and speech fidelity performance in complex acoustic environments, and significantly improves the stability, robustness and speech enhancement effect of the filter in denoising in complex noise environments.

[0009] The second objective of this invention is to provide an optimized full-band signal-to-noise ratio multi-channel filtering system based on inter-frame correlation.

[0010] The first technical solution adopted in this invention is a multi-channel filtering method for optimizing full-band signal-to-noise ratio based on inter-frame correlation, comprising the following steps:

[0011] Step S1: Acquire relevant signals in the multi-channel speech system and convert them to time-frequency domain signals; generate multi-channel vector signals containing inter-frame information; perform noise reduction processing on the multi-channel noisy speech signals containing inter-frame information; Step S2: Calculate the correlation matrix between the multi-channel clean speech signal containing inter-frame information and the multi-channel noisy signal containing inter-frame information; Step S3: Distinguish between speech-dominant and noise-dominant subspaces based on the size of the eigenvalues, perform noise reduction and enhancement on the speech-dominant subspace with the largest eigenvalue, and directly set other noise-dominant subspaces to zero; Step S4: Further optimize the strategy in Step S3, select several speech-dominant subspaces with the largest eigenvalues ​​for noise reduction and enhancement, and directly set other noise-dominant subspaces to zero, forming a length of... A full-band filter.

[0012] The invention is further characterized in that: step S1 is specifically implemented according to the following: step S1.1: establish a noisy speech signal model acquired by the microphone array, specifically as follows: first, in a multi-channel speech system, consider the noisy speech signal model containing... An array of microphones is used to acquire the target speech signal, and the multi-channel noisy speech signal is represented as:

[0013] In the formula, Represents linear convolution. Indicates a time index. From the sound source to the first The room impulse response of the channel microphone This represents the clean speech signal at the sound source. It is the first The noise signal of the channel, It is the first The clean speech signal components of the channel; then, assuming the clean speech signal after convolution. With noise signal They are mutually uncorrelated and all are zero-mean, real-value, wideband signals; simultaneously, each channel contains clean voice signals. Maintain coherence among multiple sensors; Step S1.2: Convert the multi-channel noisy speech signal, the multi-channel clean speech signal, and the multi-channel noise signal from the time domain to the time-frequency domain, specifically as follows: Perform a short-time Fourier transform on both sides of equation (1) to obtain:

[0014] In the formula, , and They represent , and In the The first channel, the first The first time frame, the first The short-time Fourier transform coefficients at each frequency point, where the frequency index The range is , The total number of sub-bands of the signal; Step S1.3: Introduce inter-frame correlation, concatenate multiple consecutive time frames into a vector, and construct a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noise signal containing inter-frame information, as follows; The continuous... The time frames are concatenated into a vector form, and the multi-channel noisy speech signal is processed. Vectorization yields:

[0015] In the formula, , and These are, respectively, a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noisy signal containing inter-frame information. , and Then it is the corresponding channel Noisy speech signal, clean speech signal and noisy signal, superscript Indicates transpose; in equation (3), each channel contains noisy speech signal, clean speech signal and noise signal containing inter-frame information in the 1st... The frequency point, the first The multi-frame vector at each time frame is defined as:

[0016] Step S1.4: Introduce inter-frame correlation into the filter and perform noise reduction on the multi-channel noisy speech signal containing inter-frame information, as follows: Construct a length of FIR filter Processing of multi-channel noisy speech signals containing inter-frame information, using filters The vector form is written as:

[0017] In the formula, For the first frame containing inter-frame information Channel filtering; calculating the input signal-to-noise ratio using the first channel as the reference channel; applying filters to the multi-channel noisy speech signal containing inter-frame information. ,have to:

[0018] In the formula, It is a clean speech signal from the first channel containing inter-frame information. The estimate, The filtered desired signal, This is a residual noise signal.

[0019] Step S2 is implemented as follows: Step S2.1: Construct a full-band output signal-to-noise ratio multi-channel filter. Specifically, as follows: Construct a string of length... Full-band output signal-to-noise ratio multi-channel filter Full-band output signal-to-noise ratio multi-channel filter The corresponding vectorized representation is:

[0020] In the formula, It is the first Output signal-to-noise ratio multi-channel filter at each frequency band Step S2.2: Calculate the correlation matrix between the multi-channel clean speech signal containing inter-frame information and the multi-channel noisy signal containing inter-frame information, as follows; determine the multi-channel noisy speech signal containing inter-frame information. The correlation matrix is ​​represented as follows:

[0021] In the formula, For mathematical expectation, and These represent multi-channel clean speech signals containing inter-frame information. and multi-channel noise signals containing inter-frame information The correlation matrix, where the superscript Represents the conjugate transpose operation; based on equations (6) and (8), the first The first time frame, the first The sub-band output signal-to-noise ratio at each frequency point is expressed as:

[0022] In the formula, and They are and The variance; summing the output SNR of all sub-bands yields the full-band output SNR:

[0023] In the formula, and For multi-channels containing inter-frame information Clean speech signals in each frequency band and multi-channel signals containing inter-frame information The correlation matrix of the noise signal at each frequency band; substituting equation (7) into equation (10), the full-band output signal-to-noise ratio is rewritten as:

[0024] In the formula, and There are two A dimensional block diagonal matrix, its specific form is:

[0025] In equations (12) and (13), the blocks on the diagonal and These are the correlation matrices of the clean speech signals and noise signals of each sub-band of the multi-channel system containing inter-frame information, respectively. Step S3 is specifically implemented as follows: Step S3.1: Perform eigenvalue decomposition on the correlation matrices of the clean speech signals and noise signals of the multi-channel system containing inter-frame information, and distinguish the speech-dominated and noise-dominated subspaces according to the magnitude of the eigenvalues; for equation (12) Find the inverse and combine with equation (13) Multiplying together, we get:

[0026] In the formula, For matrix The reverse, for 3D matrix From equation (7), we can obtain the full-band output signal-to-noise ratio filter. It consists of multiple subspace filters Each subspace filter is composed of these elements. All correspond to the signal correlation matrix The eigenvector corresponding to the largest eigenvalue; let Represents a matrix in a subspace The largest eigenvalue, for all The eigenvalues ​​of the sub-bands are arranged in descending order, resulting in the following relationship:

[0027] Step S3.2: During the filtering process, apply a full-band output signal-to-noise ratio filter only to the dominant speech subspace with the largest eigenvalue. To perform noise reduction and enhancement, while directly setting other noise-dominant subspaces to zero, as follows: According to equation (15), the matrix The largest eigenvalue is Therefore, under the condition of maximizing the full-band output signal-to-noise ratio, the full-band output signal-to-noise ratio filter From the matrix Maximum eigenvalue The corresponding eigenvectors constitute the matrix. Maximum eigenvalue The full-band maximum signal-to-noise ratio filter formed by the corresponding eigenvectors The formal representation is as follows:

[0028] In the formula, For non-zero complex numbers, Representation matrix Maximum eigenvalue The corresponding eigenvectors, filters Total length is For length is The all-zero vector; thus, the full-band maximum signal-to-noise ratio filter. Further written in sub-band form:

[0029] In equation (17), the full-band output signal-to-noise ratio and the input signal-to-noise ratio satisfy the following relationship:

[0030] Combining equations (6) and (17), the full-band maximum signal-to-noise ratio filter The expected output signal is estimated as follows:

[0031] Step S3.3: Determine the full-band maximum signal-to-noise ratio filter The coefficients are used to obtain the full-band maximum signal-to-noise ratio filter. The complete structure is as follows: First, to determine the coefficients... The value of is determined using two different criteria: one is to minimize the mean square error based on distortion. Secondly, minimize the mean square error between the clean speech signal and the estimated desired signal. Among them, the mean square error based on distortion Defined as:

[0032] Mean square error between the clean speech signal and the estimated expected signal Defined as:

[0033] Using the method of minimizing the distortion-based mean square error, according to equations (9) and (20), the distortion-based mean square error is... Expanded to:

[0034] In the formula, for identity matrix The first column vector; the full-band maximum signal-to-noise ratio filter shown in equation (16). Substituting the structure into equation (22), we further obtain:

[0035] In the formula, superscript Denotes complex conjugation; where, for Taking the partial derivative and setting it to 0, we get:

[0036] Combining equations (16) and (24), we obtain the maximum signal-to-noise ratio filter that minimizes the mean square error of distortion. for:

[0037] Similarly, another form of maximum signal-to-noise ratio filter is obtained by minimizing the mean square error between the clean speech signal and the estimated desired signal. :

[0038] Step S4 is implemented as follows: Step S4.1: Construct a filter incorporating a full-band optimization strategy, specifically as follows; from the eigenvalue set Before being selected The eigenvectors corresponding to the largest eigenvalues ​​are used to construct an eigenvector of length . Full-band output signal-to-noise ratio filter Among them, the full-band output signal-to-noise ratio filter The format is as follows:

[0039] In the formula, Let be any complex number, and at least one of them is not equal to 0; Step S4.2: Select several speech dominant subspaces with the largest eigenvalues ​​and apply a full-band output signal-to-noise ratio filter. To enhance noise reduction, the filter directly sets other noise-dominant subspaces to zero, as follows; the filter only applies to the front... The frequency corresponding to the largest eigenvalue is enhanced, while the filter coefficients of the remaining sub-bands are set to zero to avoid excessive speech distortion; thus, a full-band output signal-to-noise ratio filter is obtained. The expression in sub-band form is:

[0040] Under the above structure, the estimated signal for each sub-band is written as:

[0041] Based on equations (25) and (26), the frequency of action is obtained respectively. There are two types of optimal filters, one of which is the filter that minimizes the mean square error based on distortion. The expression is:

[0042] Another filter is one that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is:

[0043] Combining the filtering results of each sub-band, a length of [length missing] is formed. A full-band filter.

[0044] In step S4.2, a length of [length missing] is formed. The full-band filter includes a filter that minimizes the mean square error based on distortion. and a filter that minimizes the mean square error between the clean signal and the estimated desired signal. The specific expressions for the two filters are as follows: Minimize the distortion-based mean square error filter The expression is:

[0045] A filter that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is:

[0046] In formulas (32) and (33), when hour, For traditional multi-channel maximum signal-to-noise ratio filters, when hour, This corresponds to a multi-channel Wiener filter.

[0047] The second technical solution adopted in this invention is a multi-channel filtering system for optimizing full-band signal-to-noise ratio based on inter-frame correlation, comprising the following modules: a signal acquisition module for acquiring multi-channel noisy speech signals; a time-frequency transformation module for converting multi-channel speech signals to the time-frequency domain; an inter-frame correlation construction module for constructing multi-frame multi-channel signal vectors; a covariance matrix estimation module for estimating the covariance matrix of the speech signal and the noise signal; an eigenvalue decomposition and subspace selection module for performing eigenvalue decomposition and selecting the speech-dominant subspace; a filter construction module for generating a multi-channel filter for optimizing full-band output signal-to-noise ratio; and a speech enhancement module for outputting the enhanced speech signal.

[0048] The beneficial effects of this invention are as follows: The method of this invention introduces inter-frame correlation in the short-time Fourier transform domain and combines it with the spatial reception differences of multi-microphone arrays to model the joint structure of the speech signal in both time and space dimensions. Based on this, it utilizes the generalized eigenvalue decomposition of the multi-channel speech and noise covariance matrices to divide the multi-channel signal subspace dominated by speech and the interference subspace dominated by noise. Then, it implements maximum output signal-to-noise ratio filtering in the speech-dominated subspace while simultaneously zeroing out the noise subspace. This invention fully utilizes inter-frame correlation, full-band optimization strategies, and the spatial correlation of sound sources on the microphone array. It not only overcomes the temporal discontinuity and speech distortion problems caused by independent processing of single frames but also effectively avoids loss of speech details and increases the naturalness of speech and overall listening quality. Compared with existing multi-channel speech enhancement techniques, this invention balances noise suppression and speech fidelity performance in complex acoustic environments, significantly improving the stability, robustness, and speech enhancement effect of the filter in complex noise environments. Attached Figure Description

[0049] Figure 1 is a flowchart of the multi-channel filtering method for optimizing full-band signal-to-noise ratio based on inter-frame correlation according to the present invention; Figure 2 is a block diagram of the multi-channel filtering system for optimizing full-band signal-to-noise ratio based on inter-frame correlation according to the present invention; Figure 3(a) is a diagram of the method for optimizing full-band signal-to-noise ratio based on inter-frame correlation according to the present invention. The effect of different filter lengths on the output signal-to-noise ratio; Figure 3(b) shows the invention involved. The influence of filters of different lengths on speech distortion coefficients; Figure 3(c) shows the invention involved. The influence of the forgetting factor of this invention on the PESQ score of filters of different lengths; Figure 4(a) shows the influence of the forgetting factor of this invention on the output signal-to-noise ratio of filters of different lengths; Figure 4(b) shows the influence of the forgetting factor of this invention on the speech distortion coefficient of filters of different lengths; Figure 4(c) shows the influence of the forgetting factor of this invention on the PESQ score of filters of different lengths; Figure 5(a) shows the influence of the number of filter channels of this invention on the output signal-to-noise ratio of filters; Figure 5(b) shows the influence of the number of filter channels of this invention on the speech distortion coefficient of filters; Figure 5(c) shows the influence of the number of filter channels of this invention on the PESQ score of filters; Figure 6(a) shows the influence of the filter length of this invention on the output signal-to-noise ratio of filters; Figure 6(b) shows the influence of the filter length of this invention on the speech distortion coefficient of filters; Figure 6(c) shows the influence of the filter length of this invention on the PESQ score of filters; Figure 7(a) shows the influence of the forgetting factor of this invention on the PESQ score of filters based on inter-frame correlation under different noise conditions. The effects of the optimized full-band signal-to-noise ratio (SNR) multi-channel filter on the output SNR of the present invention under different noise conditions are shown in Figure 7(b); Figure 7(c) shows the effects of the optimized full-band SNR multi-channel filter based on inter-frame correlation on the PESQ score under different noise conditions of the present invention; Figure 8(a) shows the effects of the optimized full-band SNR multi-channel filter, maximum SNR filter, and Wiener filter based on inter-frame correlation on the output SNR of the present invention under different noise conditions of the present invention; Figure 8(b) shows the effects of the optimized full-band SNR multi-channel filter, maximum SNR filter, and Wiener filter based on inter-frame correlation on the output SNR of the present invention under different noise conditions of the present invention; Figure 8(c) shows the effects of the optimized full-band SNR multi-channel filter, maximum SNR filter, and Wiener filter based on inter-frame correlation on the PESQ score under different noise conditions of the present invention. Detailed Implementation

[0050] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0051] This invention provides an optimized full-band signal-to-noise ratio multi-channel filtering method based on inter-frame correlation, as shown in Figure 1, including the following steps: Step S1, in a multi-channel speech system, firstly using... An array of microphones acquires noisy speech signals, clean speech signals, and noise signals, and converts them to the time-frequency domain using a short-time Fourier transform. To utilize the inter-frame correlation of the signals, multiple consecutive time frames of each channel are spliced ​​into a vector to form a multi-channel vector signal containing inter-frame information. Based on this, a filter that introduces inter-frame correlation is further used to denoise the multi-channel noisy speech signal containing inter-frame information.

[0052] Step S1 is implemented as follows: Step S1.1: Establish a noisy speech signal model acquired by the microphone array, specifically as follows: First, in a multi-channel speech system, consider the noisy speech signal containing... An array of microphones is used to acquire the target speech signal, and the multi-channel noisy speech signal is represented as:

[0053] In the formula, Represents linear convolution. Indicates a time index. From the sound source to the first The room impulse response of the channel microphone This represents the clean speech signal at the sound source. It is the first The noise signal of the channel, It is the first The clean speech signal components of the channel; then, assuming the clean speech signal after convolution. With noise signal They are mutually uncorrelated and all are zero-mean, real-value, wideband signals; simultaneously, each channel contains clean voice signals. Maintain coherence among multiple sensors; Step S1.2: Convert the multi-channel noisy speech signal, the multi-channel clean speech signal, and the multi-channel noise signal from the time domain to the time-frequency domain, specifically as follows: Perform a short-time Fourier transform on both sides of equation (1) to obtain:

[0054] In the formula, , and They represent , and In the The first channel, the first The first time frame, the first The short-time Fourier transform coefficients at each frequency point, where the frequency index The range is , The total number of sub-bands of the signal; Step S1.3: Introduce inter-frame correlation, concatenate multiple consecutive time frames into a vector, and construct a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noise signal containing inter-frame information, as detailed below; In order to utilize inter-frame related information, the continuous... The time frames are concatenated into a vector form, and the multi-channel noisy speech signal is processed. Vectorization yields:

[0055] In the formula, , and These are, respectively, a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noisy signal containing inter-frame information. , and Then it is the corresponding channel Noisy speech signal, clean speech signal and noisy signal, superscript Indicates transpose; in equation (3), each channel contains noisy speech signal, clean speech signal and noise signal containing inter-frame information in the 1st... The frequency point, the first The multi-frame vector at each time frame is defined as:

[0056] Step S1.4: Introduce inter-frame correlation into the filter and perform noise reduction processing on the multi-channel noisy speech signal containing inter-frame information, as follows; In order to achieve effective recovery of the multi-channel noisy signal containing inter-frame information, a length of... FIR filter Processing of multi-channel noisy speech signals containing inter-frame information, using filters The vector form is written as:

[0057] In the formula, For the first frame containing inter-frame information Channel filters; considering the coherence of the speech signals in each channel of a multi-channel system, the first channel is used as the reference channel for calculating the input signal-to-noise ratio; therefore, the above filters are applied to the noisy multi-channel speech signals containing inter-frame information. ,have to:

[0058] In the formula, It is a clean speech signal from the first channel containing inter-frame information. The estimate, The filtered desired signal, This is a residual noise signal.

[0059] Step S2: Based on the received multi-channel noisy speech signal containing inter-frame information, construct a full-band output signal-to-noise ratio multi-channel filter. And calculate the correlation matrix of the multi-channel clean speech signal containing inter-frame information and the multi-channel noise signal containing inter-frame information; Step S2 is specifically implemented as follows: Step S2.1: Construct a full-band output signal-to-noise ratio multi-channel filter. Specifically, as follows: Construct a string of length... Full-band output signal-to-noise ratio multi-channel filter Full-band output signal-to-noise ratio multi-channel filter The corresponding vectorized representation is:

[0060] In the formula, It is the first Output signal-to-noise ratio multi-channel filter at each frequency band Step S2.2: Calculate the correlation matrix between the multi-channel clean speech signal containing inter-frame information and the multi-channel noisy signal containing inter-frame information, as follows; determine the multi-channel noisy speech signal containing inter-frame information. The correlation matrix is ​​represented as follows:

[0061] In the formula, For mathematical expectation, and These represent multi-channel clean speech signals containing inter-frame information. and multi-channel noise signals containing inter-frame information The correlation matrix, where the superscript Represents the conjugate transpose operation; based on equations (6) and (8), the first The first time frame, the first The sub-band output signal-to-noise ratio at each frequency point is expressed as:

[0062] In the formula, and They are and The variance; summing the output SNR of all sub-bands yields the full-band output SNR:

[0063] In the formula, and For multi-channels containing inter-frame information Clean speech signals in each frequency band and multi-channel signals containing inter-frame information The correlation matrix of the noise signal at each frequency band; substituting equation (7) into equation (10), the full-band output signal-to-noise ratio is rewritten as:

[0064] In the formula, and There are two A dimensional block diagonal matrix, its specific form is:

[0065] In equations (12) and (13), the blocks on the diagonal and The correlation matrices are for each sub-band clean speech signal and each sub-band noise signal of the multi-channel system containing inter-frame information, respectively. Step S3: Perform eigenvalue decomposition on the correlation matrices of the multi-channel clean speech signal and multi-channel noise signal obtained in Step S2. Distinguish between speech-dominant and noise-dominant subspaces based on the magnitude of the eigenvalues. Subspaces with larger eigenvalues ​​correspond to the speech-dominant component, while the remaining subspaces with smaller eigenvalues ​​correspond to the noise-dominant component. During the filtering process, a full-band output signal-to-noise ratio filter is applied only to the speech-dominant subspace with the largest eigenvalue. To perform noise reduction and enhancement, while directly setting other noise-dominant subspaces to zero; Step S3 is specifically implemented as follows: Step S3.1: Perform eigenvalue decomposition on the correlation matrices corresponding to the multi-channel clean speech signal containing inter-frame information and the multi-channel noise signal containing inter-frame information, and distinguish the speech-dominant and noise-dominant subspaces according to the magnitude of the eigenvalues; For equation (12) Find the inverse and combine with equation (13) Multiplying them together gives:

[0066] In the formula, For matrix The reverse, for 3D matrix .

[0067] From equation (7), we can obtain the full-band output signal-to-noise ratio filter. Essentially, it consists of multiple subspace filters Together they form a system in which each subspace filter... All correspond to the signal correlation matrix The eigenvector corresponding to the largest eigenvalue; let Represents a matrix in a subspace The largest eigenvalue, for all The eigenvalues ​​of the sub-bands are arranged in descending order, resulting in the following relationship:

[0068] The ranking result can be used as a basis for subspace selection to determine which subspaces should introduce the maximum signal-to-noise ratio filter. The remaining subspaces can be directly set to zero in the filter design to effectively suppress noise components; Step S3.2: During the filtering process, a full-band output signal-to-noise ratio filter is applied only to the speech-dominant subspace with the largest eigenvalue. To enhance noise reduction, other noise-dominant subspaces are directly set to zero, and this strategy is used to obtain the desired signal estimate, as follows: According to equation (15), the matrix The largest eigenvalue is Therefore, under the condition of maximizing the full-band output signal-to-noise ratio (SNR) of equation (11), the full-band output SNR filter From the matrix Maximum eigenvalue The corresponding eigenvectors constitute the matrix. Maximum eigenvalue The full-band maximum signal-to-noise ratio filter formed by the corresponding eigenvectors The formal representation is as follows:

[0069] In the formula, For non-zero complex numbers, Representation matrix Maximum eigenvalue The corresponding eigenvectors, filters Total length is For length is The all-zero vector; thus, the full-band maximum signal-to-noise ratio filter. Further written in sub-band form:

[0070] In equation (17), the full-band output signal-to-noise ratio and the input signal-to-noise ratio satisfy the following relationship:

[0071] Combining equations (6) and (17), the full-band maximum signal-to-noise ratio filter The expected output signal is estimated as follows:

[0072] Step S3.3: Determine the full-band maximum signal-to-noise ratio filter The coefficients are used to obtain the full-band maximum signal-to-noise ratio filter. The complete structure is as follows: First, to determine the coefficients... The value of is determined using two different criteria: one is to minimize the mean square error based on distortion. Secondly, minimize the mean square error between the clean speech signal and the estimated desired signal. Among them, the mean square error based on distortion Defined as:

[0073] Mean square error between the clean speech signal and the estimated expected signal Defined as:

[0074] Using the method of minimizing the distortion-based mean square error, according to equations (9) and (20), the distortion-based mean square error is... Expanded to:

[0075] In the formula, for identity matrix The first column vector; the full-band maximum signal-to-noise ratio filter shown in equation (16). Substituting the structure into equation (22), we further obtain:

[0076] In the formula, superscript Denotes complex conjugation; where, for Taking the partial derivative and setting it to 0, we get:

[0077] Combining equations (16) and (24), we obtain the maximum signal-to-noise ratio filter that minimizes the mean square error of distortion. for:

[0078] Similarly, another form of maximum signal-to-noise ratio filter is obtained by minimizing the mean square error between the clean speech signal and the estimated desired signal. :

[0079] Step S4: In step S3, the maximum signal-to-noise ratio filter and maximum signal-to-noise ratio filter Only the eigenvalues ​​of the two filters are retained. The corresponding feature vectors are selected, while the filter coefficients of the remaining sub-bands are all set to zero. Although this strategy can significantly improve the output signal-to-noise ratio, it can also easily introduce relatively obvious speech distortion. To alleviate this problem, a full-band optimization strategy is introduced. During the filtering process, several speech dominant subspaces with the largest feature values ​​are selected and a full-band output signal-to-noise ratio filter is applied. This is done to enhance noise reduction while simultaneously setting other noise-dominant subspaces to zero.

[0080] Step S4 is implemented as follows: Step S4.1: Construct a filter incorporating a full-band optimization strategy, specifically as follows; from the eigenvalue set Before being selected The eigenvectors corresponding to the largest eigenvalues ​​are used to construct an eigenvector of length . Full-band output signal-to-noise ratio filter This allows it to further reduce speech distortion while maintaining strong noise suppression capabilities. Among these features is a full-band output signal-to-noise ratio filter. The format is as follows:

[0081] In the formula, Let be any complex number, and at least one of them is not equal to 0; Step S4.2: Select several speech dominant subspaces with the largest eigenvalues ​​and apply a full-band output signal-to-noise ratio filter. To enhance noise reduction, the filter directly sets other noise-dominant subspaces to zero, as detailed below; in this method, the filter only applies to the preceding subspace. The frequency corresponding to the largest eigenvalue is enhanced, while the filter coefficients of the remaining sub-bands are set to zero to avoid excessive speech distortion; thus, a full-band output signal-to-noise ratio filter is obtained. The expression in sub-band form is:

[0082] Under the above structure, the estimated signal for each sub-band is written as:

[0083] Based on equations (25) and (26), the frequency of action is obtained respectively. There are two types of optimal filters, one of which is the filter that minimizes the mean square error based on distortion. The expression is:

[0084] Another filter is one that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is:

[0085] Combining the filtering results of each sub-band, a length of [length missing] is formed. A full-band filter; wherein, the forming length is The full-band filter includes a filter that minimizes the mean square error based on distortion. and a filter that minimizes the mean square error between the clean signal and the estimated desired signal. The specific expressions for the two filters are as follows: Minimize the distortion-based mean square error filter The expression is:

[0086] A filter that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is:

[0087] Based on the above derivation, two filtering schemes with optimal signal-to-noise ratio for multi-channel full-band output are obtained. One of them is a maximum signal-to-noise ratio filter that minimizes the mean square error of distortion. The second is a maximum signal-to-noise ratio filter that minimizes the mean square error between the clean signal and the estimated desired signal. ;when hour, For traditional multi-channel maximum signal-to-noise ratio filters, when hour, This corresponds to the classic multichannel Wiener filter.

[0088] As shown in Figure 2, this invention also provides an optimized full-band signal-to-noise ratio multi-channel filtering system based on inter-frame correlation, comprising the following modules: a signal acquisition module for acquiring multi-channel noisy speech signals; a time-frequency transformation module for converting multi-channel speech signals to the time-frequency domain; an inter-frame correlation construction module for constructing multi-frame multi-channel signal vectors; a covariance matrix estimation module for estimating the covariance matrix of the speech signal and the noise signal; an eigenvalue decomposition and subspace selection module for performing eigenvalue decomposition and selecting the speech-dominant subspace; a filter construction module for generating a multi-channel filter with optimized full-band output signal-to-noise ratio; and a speech enhancement module for outputting the enhanced speech signal.

[0089] The basic environment setup for the experiments in Examples 1-6 is as follows: A filter designed to minimize the mean square error based on distortion. For example, the actual performance of the algorithm is verified. The clean speech signal used in the experiment comes from the TIMIT speech database.

[12] We randomly selected voice samples from 15 men and 15 women, and reduced the sampling rate from 16kHz to 8kHz to unify the signal processing conditions.

[0090] The experiment employed a linear microphone array structure containing 10 omnidirectional microphones. The array was placed parallel to the ground at a height of 1.4m and 0.5m from the north wall. The spatial coordinates of the ten microphones were set as follows: ,in The sound source is simulated by a loudspeaker and placed in... The room impulse response is used to play clean speech signals. Based on the above sound source and microphone positions, the room impulse response from the sound source to each microphone is simulated and generated. The simulation sampling rate is 48kHz, and then downsampled to 8kHz for subsequent processing. The generated room impulse response signal is regarded as the real acoustic channel response and used to generate multi-channel clean speech signals.

[0091] In the signal construction stage, the clean speech signal is convolved with the impulse response corresponding to each channel to obtain a multi-channel clean speech signal, and Gaussian white noise or car noise is superimposed to simulate different noise environments. By adjusting the noise energy ratio, a series of test samples with different input signal-to-noise ratio levels are set up, covering various acoustic environment conditions from low signal-to-noise ratio to high signal-to-noise ratio.

[0092] In the algorithm implementation, to construct an optimized full-band signal-to-noise ratio (SNR) multi-channel filter based on inter-frame correlation in the short-time Fourier transform domain, the speech signal is divided into frames with a length of 128 sampling points and an inter-frame overlap rate of 75%. To reduce the impact of spectral leakage, each frame is weighted using a Kaiser window. Subsequently, the signal is transformed to the frequency domain using a short-time Fourier transform, and the corresponding filters are calculated in each sub-band, optimizing the output SNR across the entire band. Finally, the enhanced speech signal is restored to the time domain using an inverse short-time Fourier transform.

[0093] Example 1 This example studies the room reverberation time under Gaussian white noise with an input signal-to-noise ratio of 10 dB. For 240ms, the output signal-to-noise ratio, speech distortion coefficient, and PESQ score of the optimized full-band signal-to-noise ratio multi-channel filter based on inter-frame correlation are evaluated as a function of the number of subspaces. The variation of the results was investigated. To ensure the generality of the results and avoid potential biases from a single-channel configuration, different configurations of 1, 2, and 4 microphones were used for verification. The filter length was fixed at 2, and the forgetting factor corresponding to the highest PESQ score for each channel was determined. The values ​​are shown in Tables 1, 2 and 3.

[0094] Table 1. Forgetting factor values ​​when the number of channels is 1

[0095] Table 2. Forgetting factor values ​​corresponding to a channel count of 2.

[0096] Table 3. Forgetting factor values ​​corresponding to a channel count of 4

[0097] The results are shown in Figures 3(a), 3(b), and 3(c). With As the signal-to-noise ratio (SNR) of the filter increases, it gradually decreases, while speech distortion initially decreases significantly and then stabilizes. This phenomenon is caused by the smaller [unclear - possibly a specific component or element]. By selecting only a small number of main feature subspaces for filtering, strong noise suppression can be achieved, resulting in a high output signal-to-noise ratio. However, due to the limited feature subspace, the ability to characterize speech structure is insufficient, leading to significant speech distortion. As the filter is increased, it gradually incorporates more speech-related feature subspaces, which significantly enhances the speech reconstruction capability and thus rapidly reduces distortion. However, at the same time, some noise-dominated subspaces are also introduced, thereby weakening the noise suppression effect and causing the output signal-to-noise ratio to decline.

[0098] The changing trend of PESQ scores further confirms the above pattern. With With the improvement, the PESQ score increased rapidly, and when The optimal value is 50. At this value, the filter retains sufficient speech-related information without introducing excessive noise, thus achieving the best trade-off between noise suppression and speech fidelity. As the feature space continues to increase, the improvement in PESQ score gradually saturates, indicating that the additional feature subspace has limited effect on improving speech quality.

[0099] Overall, parameters It plays a crucial balancing role between speech reconstruction and noise suppression capabilities. Smaller Although it can achieve a higher output signal-to-noise ratio, it suffers from significant distortion due to insufficient speech-related features; and with As the distortion increases, the noise suppression capability decreases significantly, but it also weakens. Experiments show that when... A value greater than 50 can simultaneously maintain low distortion and good subjective quality, making it an effective value for achieving the best performance of this filter.

[0100] Example 2: This example studies the performance variation of an optimized full-band signal-to-noise ratio (SNR) multichannel filter based on inter-frame correlation under different forgetting factors. The experimental environment is consistent with the first example, i.e., the input SNR is set to 10dB, the background noise is Gaussian white noise, and the room reverberation time is [not specified]. The time was 240ms. To ensure the generality of the experimental results, 1, 2, and 4 channels were selected for verification, with the filter length kept at 2. The parameters corresponding to the highest PESQ score for each channel were also determined. The values ​​are shown in Table 4.

[0101] Table 4. p-values ​​corresponding to different numbers of channels

[0102] Figures 4(a), 4(b), and 4(c) illustrate the forgetting factor under multi-channel conditions. Impact on algorithm performance. It can be seen that, with... As the number of channels increases, the output signal-to-noise ratio shows a steady-state decreasing trend, while the speech distortion coefficient monotonically increases, and this trend becomes more pronounced with increasing channel count. This indicates that... It plays a crucial role in the performance tuning of filters, and has a relatively large impact. This makes the filter more dependent on historical observation data during recursive updates, resulting in smoother updates to the correlation matrix but slower response. A smaller forgetting factor, on the other hand, depends more on current data, resulting in a faster response but reduced steady-state noise suppression.

[0103] PESQ scores further reveal the changes in speech quality. The pattern of change. Smaller It has certain advantages in noise suppression, but it can lead to overly smooth speech, limiting subjective quality. With... With increased volume, speech details are better preserved, and PESQ reaches its maximum value in the medium range; if As the factor continues to increase, noise leakage and distortion worsen again, resulting in a significant decrease in PESQ. This demonstrates that a moderate forgetting factor can achieve a more ideal trade-off between noise suppression, structure fidelity, and subjective auditory quality.

[0104] Example 3 analyzes the impact of the number of microphones on noise reduction performance. The experimental environment is consistent with the previous two examples, i.e., the input signal-to-noise ratio is set to 10dB, the background noise is Gaussian white noise, and the reverberation time is approximately 240ms. The filter length is fixed. And for different channel numbers Corresponding forgetting factors were set. and subspace parameters The data in the table are the parameter values ​​for each channel when it generates the highest PESQ score. The specific values ​​are shown in Table 5 to ensure the optimal comparison of algorithm performance under each channel configuration.

[0105] Table 5. Forgetting factors corresponding to different numbers of channels and value

[0106] As can be seen from Figures 5(a), 5(b), and 5(c), the number of microphones... It has a significant impact on the performance of multi-channel speech noise reduction. With With the increase of [channels], the overall output signal-to-noise ratio shows a continuous upward trend, indicating that more channels can provide richer spatial information, which is beneficial to enhancing noise suppression capabilities. When When the signal-to-noise ratio is smaller, the improvement is more significant; while... When the signal-to-noise ratio is large, the gain of the output signal-to-noise ratio gradually flattens out, indicating that the performance improvement brought by multiple channels has a certain saturation characteristic.

[0107] At the same time, the voice distortion coefficient increases with... The slight increase in indicates that while enhancing noise reduction capabilities, speech reconstruction errors also increase accordingly. This phenomenon reflects the trade-off between noise suppression and speech fidelity in multi-channel processing; that is, too many channels may introduce additional speech distortion while enhancing noise suppression.

[0108] From the changing trend of the subjective quality indicator PESQ, PESQ is at a relatively small... Significant improvement within the range, and A value of 4 indicates a relatively high level, followed by... The number of channels continues to increase while decreasing slightly. This indicates that an appropriate number of channels can achieve a better balance between noise suppression and speech quality, while too many channels have limited improvement on subjective listening experience and may even affect speech quality due to cumulative distortion. Therefore, in practical applications, it is necessary to combine noise reduction effect and speech quality to select an appropriate channel size.

[0109] Example 4 further investigates the impact of optimizing the full-band signal-to-noise ratio (SNR) multi-channel filter length based on inter-frame correlation on algorithm performance. The experimental environment remains consistent with the previous three examples: input SNR of 10 dB, background noise of Gaussian white noise, and reverberation time of approximately 240 ms. The number of channels is fixed. Optimal forgetting factors were set for different filter lengths. With subspace parameters The data in the table are the parameter values ​​that produce the highest PESQ score for each filter length. The specific values ​​are shown in Table 6.

[0110] Table 6 Forgetting Factors Corresponding to Filter Length and value

[0111] As can be observed from Figures 6(a), 6(b), and 6(c), the filter length It plays a crucial role in the performance of multi-channel speech noise reduction systems. With... As the length of the filter gradually increases, the output signal-to-noise ratio significantly improves, indicating that a longer filter can more fully model the temporal structure of speech and noise, thereby improving noise suppression. This is especially true in smaller... Within the range, the improvement in output signal-to-noise ratio is most significant; when After increasing the filter length to a certain extent, the performance improvement gradually slows down, indicating that the system's dependence on the filter length begins to weaken. Further increases in filter length... The benefits are limited.

[0112] On the other hand, increasing the filter length also has a certain impact on speech distortion. It can be seen that the speech distortion coefficient increases with... The increase in length shows a slow upward trend, indicating that while suppressing noise, the filter's intervention in speech components is also increasing, thus introducing some reconstruction error. This phenomenon suggests that a larger filter length is not necessarily better; its design requires a trade-off between noise suppression capability and speech fidelity.

[0113] The above conclusions can be further verified by combining the subjective quality evaluation index PESQ. When the filter length is small, the PESQ score improves significantly; however, when... As the filter length continues to increase, the PESQ actually decreases, indicating that while an excessively long filter helps reduce noise, the resulting distortion accumulation weakens the improvement in subjective listening experience. Therefore, in practical applications, a suitable filter length should be chosen to achieve a relative balance between noise reduction performance and speech quality.

[0114] Example 5 compares and evaluates the performance of an optimized full-band signal-to-noise ratio (SNR) multi-channel filter based on inter-frame correlation under two typical non-stationary noise environments (car noise and Gaussian white noise), examining performance changes under single-channel and four-channel conditions. Experimental results show the trends in output SNR, speech distortion coefficient, and PESQ. The filter length is kept at 2, and the parameters used for each noise type and channel configuration are as follows. , With forgetting factor The specific values ​​are obtained from the settings given in Table 7. The data in the table are the values ​​of the parameters that produce the highest PESQ score under different noise conditions.

[0115] Table 7. Different noise environments , and The value of

[0116] As shown in Figures 7(a), 7(b), and 7(c), for both types of noise conditions, the output signal-to-noise ratio (SNR) of the filter increases approximately linearly with the increase of the input SNR. Specifically, the output SNR under Gaussian white noise conditions is generally higher than that under automotive noise conditions. This is related to the relatively stable statistical characteristics and uniform energy distribution of Gaussian white noise, which makes it easier for the filter to distinguish between speech and noise components, thus improving the filter's noise reduction performance.

[0117] The changes in speech distortion coefficients further reflect the impact of filters on speech structure under different noise environments. It can be seen that, under both types of noise conditions, the overall distortion level in the Gaussian white noise environment is lower than that in the car noise environment, and it gradually decreases with the increase of the input signal-to-noise ratio, indicating that in a statistically stable noise background, the filter can more effectively recover speech components. In contrast, car noise has stronger non-stationarity, with more dramatic spectral changes, resulting in relatively higher speech distortion.

[0118] From the results of the subjective speech quality index PESQ, the PESQ score under Gaussian white noise is higher than that under car noise throughout the entire input signal-to-noise ratio range, further indicating that speech quality recovery is relatively easier under this noise condition. However, in the car noise environment, although speech quality gradually improves with increasing input signal-to-noise ratio, the overall PESQ level is still constrained by the non-stationary characteristics of the noise.

[0119] Example 6: In this example, the background noise was set to Gaussian white noise. The performance of three types of multi-channel filters based on inter-frame correlation—namely, an optimized full-band SNR multi-channel filter based on inter-frame correlation, a Wiener filter, and a maximum SNR filter—was compared and evaluated in different reverberation environments. The experiments were conducted in... and The experiment was conducted under reverberant conditions, and the input signal-to-noise ratio was extended from 0dB to 20dB to verify the comprehensive performance advantages of the proposed improved filter in speech enhancement tasks. The total number of channels was fixed at 4, and the filter length was set to 2. The parameters used for each filter under different reverberation conditions were... With forgetting factor The specific values ​​are shown in Table 8. The data in the table are the values ​​of the parameters when each filter produces the highest PESQ score.

[0120] Table 8. Different filters , and The value of

[0121] As can be seen from Figures 8(a), 8(b), and 8(c), under all input signal-to-noise ratio (SNR) conditions, the output SNR of the optimized full-band SNR multi-channel filter based on inter-frame correlation is consistently higher than that of the Wiener filter, and generally slightly lower than that of the maximum SNR filter. When the reverberation time increases to... At that time, the output signal-to-noise ratio of the improved filter remained stable as reverberation changed, indicating that the method has good adaptability in strong reverberation environments.

[0122] The speech distortion coefficients further reveal the differences among the three algorithms in terms of speech structure preservation. While the maximum signal-to-noise ratio (MSNR) filter achieves the highest output SNR, it introduces the greatest speech distortion under both reverberation conditions. The Wiener filter exhibits relatively low speech distortion, but its noise suppression capability is limited. In contrast, the improved filter's distortion level consistently falls between the two, showing a more pronounced decreasing trend with increasing input SNR. Combined with its higher output SNR, it can be seen that the improved filter effectively suppresses noise while better controlling the disruption to speech structure, achieving a reasonable trade-off between noise suppression and speech fidelity.

[0123] From the perspective of subjective speech quality metrics (PESQ), the improved filter achieved the highest or near-highest scores across the entire input signal-to-noise ratio (SNR) range, and significantly outperformed the Wiener filter and the maximum SNR filter under both reverberation conditions. Particularly in the low to medium input SNR region, the improved filter significantly enhanced speech intelligibility and naturalness, indicating its superior overall speech enhancement effect in complex noise and reverberation environments.

[0124] This invention overcomes the temporal discontinuity and speech distortion caused by independent processing of single frames by explicitly modeling the time-frequency correlation between adjacent frames in the short-time Fourier transform domain. It also constructs a full-band joint optimization objective function to replace the isolated enhancement strategies for local frequency bands or subspaces in traditional methods, thereby effectively avoiding speech detail loss and increasing speech naturalness and overall listening quality. Comparative experiments in Examples 1, 3, 4, 5, and 6 show that, compared to existing multi-channel maximum signal-to-noise ratio filters and Wiener filters, this invention balances noise suppression and speech fidelity performance in complex acoustic environments, significantly improving the stability, robustness, and speech enhancement effect of the filter in complex noise and reverberation environments.

[0125] It should be emphasized that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit its scope of protection. Any equivalent substitutions or simple deductions made to the method flow, parameter configurations (such as the number of channels, the number of frames, and the subspace dimension) or system architecture within the spirit and scope of the claims of the present invention shall be deemed to fall within the patent protection scope of the present invention.

[0126] The practical application scenarios of this invention are not limited to the examples listed, but can also be extended to any field that relies on multi-microphone arrays for robust speech enhancement, such as remote conferencing, smart homes, hearing aids, and robot voice interaction.

Claims

1. A multi-channel filtering method for optimizing full-band signal-to-noise ratio based on inter-frame correlation, characterized in that, Includes the following steps: Step S1: Acquire relevant signals in the multi-channel speech system and convert them to the time-frequency domain; generate multi-channel vector signals containing inter-frame information; perform noise reduction processing on the multi-channel noisy speech signals containing inter-frame information; Step S2: Calculate the correlation matrix between the multi-channel clean speech signal containing inter-frame information and the multi-channel noisy signal containing inter-frame information; Step S3: Distinguish between speech-dominated and noise-dominated subspaces based on the magnitude of eigenvalues, perform noise reduction and enhancement on the speech-dominated subspace with the largest eigenvalue, and directly set the other noise-dominated subspaces to zero; Step S4: Form a vector signal of length... A full-band filter.

2. The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation according to claim 1, characterized in that, Step S1 is implemented as follows: Step S1.1: Establish a noisy speech signal model acquired by the microphone array, specifically as follows: First, in a multi-channel speech system, consider the noisy speech signal containing... An array of microphones is used to acquire the target speech signal, and the multi-channel noisy speech signal is represented as: In the formula, Represents linear convolution. Indicates a time index. From the sound source to the first The room impulse response of the channel microphone This represents the clean speech signal at the sound source. It is the first The noise signal of the channel, It is the first The clean speech signal components of the channel; then, assuming the clean speech signal after convolution. With noise signal They are mutually uncorrelated and all are zero-mean, real-value, wideband signals; simultaneously, each channel contains clean voice signals. Maintain coherence among multiple sensors; Step S1.2: Convert the multi-channel noisy speech signal, the multi-channel clean speech signal, and the multi-channel noise signal from the time domain to the time-frequency domain, specifically as follows: Perform a short-time Fourier transform on both sides of equation (1) to obtain: In the formula, , and They represent , and In the The first channel, the first The first time frame, the first The short-time Fourier transform coefficients at each frequency point, where the frequency index The range is , The total number of sub-bands of the signal; Step S1.3: Introduce inter-frame correlation, concatenate multiple consecutive time frames into a vector, and construct a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noise signal containing inter-frame information, as follows; The continuous... The time frames are concatenated into a vector form, and the multi-channel noisy speech signal is processed. Vectorization yields: In the formula, 、 and These are, respectively, a multi-channel noisy speech signal containing inter-frame information, a multi-channel clean speech signal containing inter-frame information, and a multi-channel noisy signal containing inter-frame information. 、 and Then it is the corresponding channel Noisy speech signal, clean speech signal and noisy signal, superscript Indicates transpose; in equation (3), each channel contains noisy speech signal, clean speech signal and noise signal containing inter-frame information in the 1st... The frequency point, the first The multi-frame vector at each time frame is defined as: Step S1.4: Introduce inter-frame correlation into the filter and perform noise reduction on the multi-channel noisy speech signal containing inter-frame information, as follows: Construct a length of FIR filter Processing of multi-channel noisy speech signals containing inter-frame information, using filters The vector form is written as: In the formula, For the first frame containing inter-frame information Channel filtering; calculating the input signal-to-noise ratio using the first channel as the reference channel; applying filters to the multi-channel noisy speech signal containing inter-frame information. ,have to: In the formula, It is a clean speech signal from the first channel containing inter-frame information. The estimate, The filtered desired signal, This is a residual noise signal.

3. The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation according to claim 2, characterized in that, Step S2 is implemented as follows: Step S2.1: Construct a full-band output signal-to-noise ratio multi-channel filter. Specifically, as follows: Construct a string of length... Full-band output signal-to-noise ratio multi-channel filter Full-band output signal-to-noise ratio multi-channel filter The corresponding vectorized representation is: In the formula, It is the first Output signal-to-noise ratio multi-channel filter at each frequency band ; Step S2.2: Calculate the correlation matrix between the multi-channel clean speech signal containing inter-frame information and the multi-channel noisy signal containing inter-frame information, as follows; determine the multi-channel noisy speech signal containing inter-frame information. The correlation matrix is ​​represented as follows: In the formula, For mathematical expectation, and These represent multi-channel clean speech signals containing inter-frame information. and multi-channel noise signals containing inter-frame information The correlation matrix, where the superscript This represents the conjugate transpose operation; Based on equations (6) and (8), the first The first time frame, the first The sub-band output signal-to-noise ratio at each frequency point is expressed as: In the formula, and They are and The variance; summing the output SNR of all sub-bands yields the full-band output SNR: In the formula, and For multi-channels containing inter-frame information Clean speech signals in each frequency band and multi-channel signals containing inter-frame information The correlation matrix of the noise signal at each frequency band; substituting equation (7) into equation (10), the full-band output signal-to-noise ratio is rewritten as: In the formula, and There are two A dimensional block diagonal matrix, its specific form is: In equations (12) and (13), the blocks on the diagonal and These are the correlation matrices of the clean speech signals of each sub-band of the multi-channel system containing inter-frame information and the noise signals of each sub-band of the multi-channel system containing inter-frame information, respectively.

4. The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation according to claim 3, characterized in that, Step S3 is implemented as follows: Step S3.1: Perform eigenvalue decomposition on the correlation matrices of the multi-channel clean speech signal containing inter-frame information and the multi-channel noise signal containing inter-frame information, and distinguish the speech-dominated and noise-dominated subspaces according to the magnitude of the eigenvalues; for equation (12) Find the inverse and combine with equation (13) Multiplying together, we get: In the formula, For matrix The reverse, for 3D matrix From equation (7), we can obtain the full-band output signal-to-noise ratio filter. It consists of multiple subspace filters Each subspace filter is composed of these elements. All correspond to the signal correlation matrix The eigenvector corresponding to the largest eigenvalue; let Represents a matrix in a subspace The largest eigenvalue, for all The eigenvalues ​​of the sub-bands are arranged in descending order, resulting in the following relationship: Step S3.2: During the filtering process, apply a full-band output signal-to-noise ratio filter only to the dominant speech subspace with the largest eigenvalue. To perform noise reduction and enhancement, while directly setting other noise-dominant subspaces to zero, as follows: According to equation (15), the matrix The largest eigenvalue is Therefore, under the condition of maximizing the full-band output signal-to-noise ratio, the full-band output signal-to-noise ratio filter From the matrix Maximum eigenvalue The corresponding eigenvectors constitute the matrix. Maximum eigenvalue The full-band maximum signal-to-noise ratio filter formed by the corresponding eigenvectors The formal representation is as follows: In the formula, For non-zero complex numbers, Representation matrix Maximum eigenvalue The corresponding eigenvectors, filters Total length is For length is The all-zero vector; thus, the full-band maximum signal-to-noise ratio filter. Further written in sub-band form: In equation (17), the full-band output signal-to-noise ratio and the input signal-to-noise ratio satisfy the following relationship: Combining equations (6) and (17), the full-band maximum signal-to-noise ratio filter The expected output signal is estimated as follows: Step S3.3: Determine the full-band maximum signal-to-noise ratio filter The coefficients are used to obtain the full-band maximum signal-to-noise ratio filter. The complete structure is as follows: First, to determine the coefficients... The value of is determined using two different criteria: one is to minimize the mean square error based on distortion. Secondly, minimize the mean square error between the clean speech signal and the estimated desired signal. Among them, the mean square error based on distortion Defined as: Mean square error between the clean speech signal and the estimated expected signal Defined as: Using the method of minimizing the distortion-based mean square error, according to equations (9) and (20), the distortion-based mean square error is... Expanded to: In the formula, for identity matrix The first column vector; the full-band maximum signal-to-noise ratio filter shown in equation (16). Substituting the structure into equation (22), we further obtain: In the formula, superscript Denotes complex conjugation; where, for Taking the partial derivative and setting it to 0, we get: Combining equations (16) and (24), we obtain the maximum signal-to-noise ratio filter that minimizes the mean square error of distortion. for: Similarly, another form of maximum signal-to-noise ratio filter is obtained by minimizing the mean square error between the clean speech signal and the estimated desired signal. : 。 5. The multi-channel filtering method for optimizing full-band signal-to-noise ratio based on inter-frame correlation according to claim 4, characterized in that, Step S4 is implemented as follows: Step S4.1: Construct a filter incorporating a full-band optimization strategy, specifically as follows; from the eigenvalue set Before being selected The eigenvectors corresponding to the largest eigenvalues ​​are used to construct an eigenvector of length . Full-band output signal-to-noise ratio filter Among them, the full-band output signal-to-noise ratio filter The format is as follows: In the formula, Let be any complex number, and at least one of them is not equal to 0; Step S4.2: Select several speech dominant subspaces with the largest eigenvalues ​​and apply a full-band output signal-to-noise ratio filter. To enhance noise reduction, the filter directly sets other noise-dominant subspaces to zero, as follows; the filter only applies to the front... The frequency corresponding to the largest eigenvalue is enhanced, while the filter coefficients of the remaining sub-bands are set to zero to avoid excessive speech distortion; thus, a full-band output signal-to-noise ratio filter is obtained. The expression in sub-band form is: Under the above structure, the estimated signal for each sub-band is written as: Based on equations (25) and (26), the frequency of action is obtained respectively. There are two types of optimal filters, one of which is the filter that minimizes the mean square error based on distortion. The expression is: Another filter is one that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is: Combining the filtering results of each sub-band, a length of [length missing] is formed. A full-band filter.

6. The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation according to claim 5, characterized in that, In step S4.2, a length of [length missing] is formed. The full-band filter includes a filter that minimizes the mean square error based on distortion. and a filter that minimizes the mean square error between the clean signal and the estimated desired signal. The specific expressions for the two filters are as follows: Minimize the distortion-based mean square error filter The expression is: A filter that minimizes the mean square error between the clean signal and the estimated desired signal. The expression is: 。 7. The method for optimizing full-band signal-to-noise ratio multi-channel filtering based on inter-frame correlation according to claim 6, characterized in that, In formulas (32) and (33), when hour, For traditional multi-channel maximum signal-to-noise ratio filters, when hour, This corresponds to a multi-channel Wiener filter.

8. A multi-channel filtering system for optimizing full-band signal-to-noise ratio based on inter-frame correlation, characterized in that, It includes the following modules: a signal acquisition module for acquiring multi-channel noisy speech signals; a time-frequency transformation module for converting multi-channel clean speech signals to the time-frequency domain; an inter-frame correlation construction module for constructing multi-frame, multi-channel signal vectors; a covariance matrix estimation module for estimating the covariance matrix between the speech signal and the noise signal; and an eigenvalue decomposition and subspace selection module for performing eigenvalue decomposition and selecting the speech-dominant subspace. The filter building module is used to generate multi-channel filters with optimized full-band output signal-to-noise ratio; The speech enhancement module is used to output the enhanced speech signal.