A method for eliminating echo in intercom system

By evaluating the speaker's position and tone model, the target voice signal is enhanced, the spectral envelope tracking and independent signal source modeling is used, the speech waveform is reconstructed in combination with phase information, and the adaptive filtering is used to suppress noise, which solves the problem of signal separation in multi-channel intercom systems and improves the speech quality and clarity.

CN119342151BActive Publication Date: 2025-08-19GUANGZHOU LIHENGSHENG ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411458368.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-18
Publication Date
2025-08-19
Estimated Expiration
2044-10-18

AI Technical Summary

Technical Problem

In multi-channel intercom systems, the separation of near-end and remote signals is difficult to accurately separate under conditions of highly overlap and dynamic changes, resulting in signal distortion or insufficient separation, affecting call quality and user experience.

Method used

By evaluating the speaker's position and tone model, the target speech signal is enhanced; using spectrum envelope tracking and independent signal source modeling, adaptively distinguishing speech and echo; combining phase information to reconstruct speech waveforms, adaptive filtering is used to suppress noise; judging the dual-speaking state adjustment echo cancellation algorithm, and evaluating propagation characteristics and dereverberation processing.

Benefits of technology

It improves the intelligibility and quality of speech, reduces echoes, enhances the purity and coherence of speech, suppresses background noise, maintains speech integrity, and improves signal-to-noise ratio and clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119342151B_ABST
    Figure CN119342151B_ABST
Patent Text Reader

Abstract

The present application provides a method for canceling echo in an intercom system, comprising: performing residual echo suppression on each of the identified multiple speaker speech segments based on a separated near-end speech evaluation, constructing a binary time-frequency mask, and reducing the residual echo spectrum caused by dynamic changes in echo power; utilizing the phase information of the original mixed signal in combination with the enhanced near-end speech power spectrum to reconstruct the speech waveform of each speaker through an inverse short-time Fourier transform, and employing an overlap-add method to smooth the transition between speech frames; determining whether speech is present simultaneously at the near-end and far-end ends, and if a double-talk state is detected, adjusting the parameters of a nonlinear echo cancellation algorithm to prevent erroneous suppression of the near-end speech; and performing transmission path modeling on the processed near-end speech signal to evaluate the propagation characteristics of the speech signal from the speaker to the microphone, including the time delay and attenuation of the direct sound and early reflections.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for eliminating echo of an intercom system. Background Art

[0002] In multi-channel intercom systems, separating near-end and far-end signals presents a daunting technical challenge. The core of this problem lies in the fact that when multiple users speak simultaneously, the near-end microphone captures both the local speaker's voice and the sound played by the far-end speaker. These two signals overlap significantly in both the time and frequency domains, making it difficult for traditional single-filtering methods to effectively distinguish and separate them. Specifically, the spectral characteristics of the local speaker's voice picked up by the near-end microphone and the sound played by the far-end speaker are very similar. They both contain the fundamental frequency components and harmonic structure of the human voice, making them difficult to distinguish in the frequency domain. Furthermore, the two signals overlap significantly in time, with no clear time interval or demarcation point. This dual overlap in both time and frequency domains makes single-filtering methods based on frequency or time characteristics ineffective. Further complicating matters, in real-world scenarios, factors such as the speaker's position, volume, and timbre are constantly changing, and the far-end signal's transmission path and reverberation characteristics also vary with the environment. These dynamic changes further complicate signal separation. Traditional methods often struggle to adapt to these complex and changing conditions, resulting in signal distortion or inadequate separation. Therefore, how to accurately separate near-end and far-end signals while maintaining their naturalness and integrity under such highly overlapping and dynamically changing conditions has become a key technical issue that needs to be addressed in multi-channel intercom systems. Solving this problem is crucial to improving the system's call quality and user experience. Summary of the Invention

[0003] The present invention provides a method for eliminating echo in an intercom system, which mainly includes:

[0004] Based on the mixed signal collected by the near-end microphone array, the speaker position is estimated, the target speaker's voice is enhanced, the enhanced near-end voice signal is obtained, the timbre characteristics are extracted, a speaker timbre model is established, the reverberation of the far-end reference signal is evaluated, the room impulse response is determined, and the reverberation characteristic parameters of the far-end signal are obtained. Based on the reverberation characteristic parameters and the far-end signal, the near-end voice signal is preliminarily processed to subtract the evaluated echo component;

[0005] The spectral envelope tracking (SEAT) method performs spectral envelope tracking on the preliminarily processed speech signal, calculating the energy distribution at each time-frequency point. Based on the energy dynamic range of the far-end signal leakage echo, the method adaptively distinguishes between speech-dominant and echo-dominant regions, and models the near-end speech and residual echo as independent signal sources. The method then performs a linear transformation on the speech signal by maximizing the negative entropy to evaluate the separation matrix and obtain the separated near-end speech estimate.

[0006] Based on the near-end speech evaluation after separation, the multiple speaker speech segments identified are subjected to residual echo suppression respectively, and a binary time-frequency mask is constructed to reduce the residual echo spectrum caused by the dynamic change of echo power;

[0007] The phase information of the original mixed signal is combined with the enhanced near-end speech power spectrum to reconstruct the speech waveform of each speaker through inverse short-time Fourier transform, and the overlap-add method is used to smooth the transition between speech frames.

[0008] The reconstructed near-end speech signal is post-processed and smoothed using adaptive Wiener filtering. The frequency response of the filter is dynamically adjusted based on the estimated background noise power spectrum to suppress musical noise in two-way communication.

[0009] Determines whether there is simultaneous speech at both the near-end and far-end. If a double-talk state is detected, the parameters of the nonlinear echo cancellation algorithm are adjusted to prevent the near-end speech from being incorrectly suppressed.

[0010] Model the transmission path of the processed near-end speech signal to evaluate the propagation characteristics of the speech signal from the speaker to the microphone, including the delay and attenuation of the direct sound and early reflections;

[0011] According to the transmission path model, the near-end speech signal is subjected to dereverberation processing to suppress the reverberation effect caused by room reflection, thereby obtaining a near-end speech signal with improved clarity. The noise reduction processing is then performed to obtain a clear speech signal with improved signal-to-noise ratio.

[0012] The technical solution provided by the embodiment of the present invention may have the following beneficial effects:

[0013] The present invention provides a method for eliminating echo in an intercom system. By processing a mixed signal, the method can accurately evaluate the position of the speaker, thereby enhancing the speech signal of the target speaker and improving the intelligibility and quality of the speech. By evaluating the impulse response of the room and the reverberation characteristics of the far-end signal, the echo component is effectively subtracted from the near-end speech signal, reducing the echo problem in two-way communication. By adopting spectrum envelope tracking and independent signal source modeling, the optimized separation matrix can further remove residual echo and improve the purity of the speech. By using phase information and the enhanced speech power spectrum, a clear speech waveform is reconstructed through inverse short-time Fourier transform, ensuring a natural transition between speech frames and enhancing the coherence of the speech. The adaptive Wiener filtering technology can dynamically adjust the frequency response, effectively suppress background noise and music noise in two-way communication, and improve the clarity of the speech signal. When a two-way talk state is detected, the echo cancellation algorithm parameters can be intelligently adjusted to prevent the near-end speech from being mistakenly suppressed and maintain the integrity of the speech of both parties.

[0014] Evaluate the propagation characteristics of speech signals and suppress the reverberation effect caused by room reflections through dereverberation processing to further improve the clarity of speech signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The present invention is a flow chart of a method for eliminating echo in an intercom system.

[0016] Figure 2 A schematic diagram of a method for eliminating echo in an intercom system according to the present invention.

[0017] Figure 3 This is another schematic diagram of a method for eliminating echo in an intercom system according to the present invention. DETAILED DESCRIPTION

[0018] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] like Figure 1-3 In this embodiment, a method for eliminating echo in an intercom system may specifically include:

[0020] S101. Based on the mixed signal collected by the near-end microphone array, the speaker position is evaluated, the target speaker's voice is enhanced, the enhanced near-end voice signal is obtained, the timbre features are extracted, a speaker timbre model is established, the reverberation of the far-end reference signal is evaluated, the room impulse response is determined, the reverberation characteristic parameters of the far-end signal are obtained, and the near-end voice signal is preliminarily processed based on the reverberation characteristic parameters and the far-end signal, and the evaluated echo component is subtracted.

[0021] Based on the mixed signal collected by the near-end microphone array, a beamforming algorithm is used to estimate the speaker position. The target speaker's speech is then enhanced using beamforming technology to obtain an enhanced near-end speech signal. The timbre characteristics of the enhanced near-end speech signal are extracted using Mel-frequency cepstral coefficients, and a Gaussian mixture model is established as the speaker's timbre model. Time-frequency analysis is performed on the far-end reference signal to calculate the reverberation time and early decay rate, determine the room impulse response, and obtain the reverberation characteristic parameters of the far-end signal. Based on the reverberation characteristic parameters and the far-end signal, the near-end speech signal is envelope-modulated and filtered. The signal amplitude is adjusted to match the reverberation characteristics of the far-end signal, and the estimated echo component is subtracted to obtain a pre-processed near-end speech signal. Echo cancellation is performed on the pre-processed near-end speech signal using an adaptive filter. The filter coefficients are updated using the minimum mean square error criterion, and double echoes are suppressed through nonlinear processing to obtain the final enhanced near-end speech signal. The final enhanced near-end speech signal is then subjected to timbre verification based on the speaker's timbre model to determine whether the signal enhancement process introduces significant timbre distortion.

[0022] Specifically, the near-end microphone array collects mixed signals, uses a beamforming algorithm to estimate the speaker's position, and enhances the target speaker's speech through beamforming technology to obtain the enhanced near-end speech signal. Mel-frequency cepstral coefficients are used to extract timbre features, and a Gaussian mixture model is constructed as the speaker's timbre model. Time-frequency analysis is performed on the far-end reference signal to calculate the reverberation time and early decay rate, determine the room impulse response, and obtain the reverberation characteristic parameters of the far-end signal. Based on the reverberation characteristic parameters and the far-end signal, the near-end speech signal is envelope-modulated and filtered. The signal amplitude is adjusted to match the reverberation characteristics of the far-end signal, and the estimated echo component is subtracted. A voice activity detection algorithm is used to distinguish speech and non-speech segments from the near-end speech signal. Cepstral subtraction is applied to the speech segments to eliminate residual echo. Specifically, the cepstral of the speech signal is calculated, subtracted from the estimated echo cepstral, and then the signal is reconstructed. Noise suppression is performed on the non-speech segments to obtain the preliminarily processed near-end speech signal. After preliminary processing, the near-end speech signal is further processed using an adaptive filter for echo cancellation. The filter coefficients are updated using the minimum mean square error criterion. Double echoes are suppressed through nonlinear processing, specifically by half-wave or full-wave rectification of the residual echo. This results in an enhanced near-end speech signal. The final enhanced near-end speech signal is then timbre-verified based on the previously established speaker timbre model to ensure that the signal enhancement process does not introduce significant timbre distortion. A near-end microphone array collects the mixed signal. A circular array of eight omnidirectional microphones, spaced 4 cm apart, is used. The beamforming algorithm uses a delay-and-sum method to calculate the energy of sound sources in different directions and identify the direction with the maximum energy as the speaker's location. The beamforming technique uses a minimum variance distortionless response algorithm with a gain of 1 and a suppression coefficient of 0.1. 50 iterations are performed to obtain the enhanced near-end speech signal. To extract timbre features, the signal is pre-emphasized with a coefficient of 0.97, a frame length of 25 milliseconds, and a frame shift of 10 milliseconds. A Hamming window is applied to calculate 13th-order Mel-frequency cepstral coefficients. The Gaussian mixture model consists of 32 components and is trained for 100 iterations using the expectation-maximization algorithm. A short-time Fourier transform (SFT) is performed on the far-end reference signal with a window length of 512 points and a 50% overlap. The energy decay curve is calculated, resulting in a reverberation time T60 of 0.5 seconds and an early decay rate of 10 dB. Envelope modulation filtering uses an exponential decay function with a decay rate matched to T60. Voice activity detection uses an energy threshold with a threshold of -20 dB. For cepstrum subtraction, the logarithmic magnitude spectrum of a 25-ms frame is calculated and discrete cosine transformed to obtain cepstrum coefficients. The estimated echo cepstrum is then subtracted and inversely transformed to reconstruct the signal. Spectral subtraction is used for noise suppression with an oversubtraction factor of 1.5. The adaptive filter uses the normalized minimum mean square error (LMSE) algorithm with a filter length of 1024 and a step size of 0.1. Nonlinear processing uses half-wave rectification with a threshold of -40 dB.The timbre verification calculates the Mel-frequency cepstral coefficients before and after processing, and measures them by cosine similarity. The threshold is set to 0.85. If the value is higher than this, the timbre is considered to be well maintained.

[0023] S102. Perform spectral envelope tracking on the speech signal after preliminary processing, calculate the energy distribution of each time-frequency point, adaptively distinguish between the speech-dominant area and the echo-dominant area based on the energy dynamic range of the far-end signal leakage echo, model the near-end speech and residual echo as independent signal sources, evaluate the separation matrix by maximizing the negative entropy, perform linear transformation on the speech signal, and obtain the separated near-end speech evaluation.

[0024] The speech signal after preliminary processing is subjected to short-time Fourier transform to obtain a time-frequency spectrum of the speech signal; based on the time-frequency spectrum, the spectrum envelope is tracked using the cepstrum analysis method, the energy distribution of each time-frequency point is calculated, and an energy distribution matrix is obtained; according to the energy dynamic range of the far-end signal leakage echo, the threshold is dynamically adjusted based on the sliding average of the far-end signal energy, the area above the threshold in the energy distribution matrix is marked as the speech-dominated area, and the area below the threshold is marked as the echo-dominated area, and a time-frequency masking map is generated; the time-frequency masking map information is applied to the independent component analysis process, and the signal is divided into two sources, speech and echo, according to the time-frequency masking map, to construct a mixing matrix; the near-end speech and residual echo are modeled as independent signal sources, the separation matrix is evaluated using the independent component analysis algorithm, and the separation matrix parameters are optimized by maximizing the negative entropy criterion; the speech signal is linearly transformed, the evaluated separation matrix is multiplied with the initial mixed signal, and the separated signal is reconstructed using a Wiener filter in combination with the time-frequency masking map to obtain a separated near-end speech evaluation signal.

[0025] Specifically, the preliminarily processed speech signal undergoes a short-time Fourier transform (SFT) to obtain a time-frequency spectrum. The spectral envelope is tracked using cepstrum analysis, and the energy distribution at each time-frequency point is calculated to produce an energy distribution matrix. Based on the energy dynamic range of the far-end signal leakage echo, a threshold is dynamically adjusted based on a sliding average of the far-end signal energy. Regions above the threshold in the energy distribution matrix are marked as speech-dominated, while regions below the threshold are marked as echo-dominated. This generates a time-frequency masking map. The time-frequency masking map is then applied to an independent component analysis (ICA) process. Based on the CEF, the signal is separated into two sources: speech and echo, and a 2x2 mixing matrix is constructed. The near-end speech and residual echo are modeled as independent signal sources. An independent component analysis algorithm is used to estimate the separation matrix, and the parameters of the separation matrix are optimized by maximizing the negative entropy criterion. The speech signal is linearly transformed, and the estimated separation matrix is multiplied by the initial mixed signal. The separated signal is reconstructed using a Wiener filter combined with the CEF to obtain the separated near-end speech estimation signal. The preliminarily processed speech signal is subjected to a short-time Fourier transform (SFT) using a Hamming window with a window length of 512 points and a 50% overlap to obtain a time-frequency spectrum. Cepstrum analysis is used to track the spectral envelope. The logarithmic magnitude spectrum is first calculated, followed by a Fourier transform. The first 13 coefficients are taken as cepstral coefficients, and the smoothed spectral envelope is obtained through an inverse transform. The energy distribution of each time-frequency point is calculated to obtain an energy distribution matrix. Based on the dynamic range of the far-end signal leakage echo, the threshold is dynamically adjusted based on a sliding average of the far-end signal energy. The average energy is calculated using a 200ms sliding window, with an initial threshold of -20dB and updated every frame. Regions above the threshold in the energy distribution matrix are marked as speech-dominated, while regions below the threshold are marked as echo-dominated. A binary time-frequency masking map is generated. This information is then applied to independent component analysis to separate the signal into two sources: speech and echo. A 2x2 mixing matrix is then constructed. The near-end speech and residual echo were modeled as independent signal sources. The Fast I CA algorithm was used to evaluate the separation matrix. The separation matrix parameters were optimized by maximizing the negative entropy criterion, with a maximum number of iterations set to 100 and a convergence threshold set to 1e-6. The speech signal was linearly transformed, and the estimated separation matrix was multiplied by the initial mixed signal. The separated signal was reconstructed using a Wiener filter combined with a time-frequency masking map. The Wiener filter's noise evaluation window length was set to 20 frames, and the smoothing factor was set to 0.98. This yielded the separated near-end speech evaluation signal.

[0026] S103 , based on the separated near-end speech evaluation, the identified multiple speaker speech segments are subjected to residual echo suppression respectively, a binary time-frequency mask is constructed, and a residual echo spectrum caused by dynamic changes in echo power is reduced.

[0027] A Gaussian mixture model based on Mel-frequency cepstral coefficients is used to segment speech signals and identify speech segments of multiple speakers, wherein the speech segments include speaker identity and time range information. Residual echo suppression is performed based on the speaker identity information of the speech segments in combination with an adaptive filter and spectral subtraction, and the echo component is subtracted from the power spectral density of the near-end speech signal to obtain an echo-suppressed speech signal. The short-time energy and zero-crossing rate of the echo-suppressed speech signal are calculated, and the speech signal is normalized using the long-term average spectrum of the speech signal. An adaptive threshold is set to determine the speech activity area. If a speech activity area is determined, a binary time-frequency masking matrix is constructed, and the speech activity area is marked as 1 and the non-speech activity area is marked as 0. A Wiener filter is used in combination with the binary time-frequency masking matrix to perform spectrum enhancement, reducing the residual echo spectrum caused by dynamic changes in echo power to obtain an enhanced near-end speech signal. The parameters of the Wiener filter are adjusted by monitoring the short-time changes in echo power to further optimize the enhanced near-end speech signal.

[0028] Specifically, based on the separated near-end speech assessment, a Gaussian mixture model based on Mel-frequency cepstral coefficients is used to segment the speech signal, identifying speech segments from multiple speakers. The time range and speaker identity of each segment are labeled, and overlapping speech is handled to distinguish different speakers using a sound source separation algorithm. For each identified speech segment, residual echo suppression is performed using a combination of adaptive filters and spectral subtraction. The echo component is subtracted from the power spectral density of the near-end speech signal by evaluating the power spectral density of the echo signal. The adaptive filter parameters are personalized based on the speaker identity information. The short-term energy and zero-crossing rate of the speech segments after residual echo suppression are calculated, and the speech signal's long-term average spectrum is used for normalization. An adaptive threshold is set to determine speech activity regions, and a binary time-frequency masking matrix is constructed, marking speech activity regions as 1 and inactivity regions as 0. Using the constructed binary time-frequency masking matrix, the speech signal after residual echo suppression is spectrally weighted. Spectral enhancement is performed using a Wiener filter combined with binary time-frequency masking to reduce the residual echo spectrum caused by dynamic variations in echo power. The filter parameters are adjusted by monitoring short-term variations in echo power to obtain the final enhanced near-end speech signal. Based on the separated near-end speech evaluation, framing is performed with a 25-ms frame length and a 10-ms frame shift to extract 13-dimensional Mel-frequency cepstral coefficient features. Speaker recognition is performed on the speech signal using a Gaussian mixture model with 32 Gaussian mixture components. Speech segments from multiple speakers are identified, and the time range and speaker identity of each segment are labeled. For overlapping speech, independent component analysis is used for source separation, with a maximum number of iterations set to 100 and a convergence threshold set to 1e-6. For each identified speech segment, residual echo suppression is performed using a combination of adaptive filtering and spectral subtraction. The adaptive filter uses a normalized minimum mean square error algorithm with a filter length of 1024, a step size of 0.1, and a spectral subtraction factor of 1.5. The initial coefficients of the adaptive filter are adjusted based on the speaker identity information. The short-term energy and zero-crossing rate of the speech segment are calculated, with a frame length of 20 milliseconds and a 50% overlap ratio. The speech signal's long-term average spectrum of a 500-millisecond window is used for normalization. The initial adaptive energy threshold is set to -20 dB, and the zero-crossing rate threshold is set to 0.3. A binary time-frequency masking matrix is constructed. Using this binary time-frequency masking matrix, a Wiener filter is used for spectral enhancement, with a noise power evaluation window length of 20 frames and a smoothing factor of 0.98. By monitoring short-term changes in echo power, the Wiener filter parameters are updated every 50 milliseconds to reduce the residual echo spectrum caused by dynamic changes in echo power, ultimately resulting in an enhanced near-end speech signal.

[0029] S104 , using the phase information of the original mixed signal in combination with the enhanced near-end speech power spectrum, reconstructing the speech waveform of each speaker through inverse short-time Fourier transform, and smoothing the transition between speech frames using overlap-add method.

[0030] The original mixed signal is subjected to a short-time Fourier transform to obtain phase information of each time-frequency point; an independent component analysis algorithm is applied to the enhanced near-end speech signal based on the phase information to obtain a speech power spectrum of multiple speakers; the original signal phase information is combined with the speech power spectrum, and a complete time-frequency domain representation is constructed through a time-frequency masking method; an inverse short-time Fourier transform is performed on the time-frequency domain representation to obtain a time domain signal; an overlap-add method is used to perform inter-frame smoothing on the time domain signal, and adjacent frames are weightedly superimposed by setting an overlap rate and using a Hanning window function; if the time domain signal contains multi-speaker speech, weighted superposition is performed according to the energy ratio of each speaker to obtain a final reconstructed complete speech waveform.

[0031] Specifically, a short-time Fourier transform (SFT) is performed on the original mixed signal to extract the phase information at each time-frequency point. The enhanced near-end speech signal is then subjected to independent component analysis (ICA) to separate the multi-speaker speech, obtaining the power spectrum of each speaker. The extracted phase information of the original signal is combined with the separated power spectrum, and a time-frequency masking method is used to construct a complete time-frequency domain representation for each speaker. An inverse short-time Fourier transform (IFT) is then performed on the combined time-frequency domain representation to convert the signal from the frequency domain back to the time domain. A dynamic time warping (DTW) algorithm is then used to address the temporal alignment of the different speakers' speech, resulting in a preliminary reconstructed speech waveform for each speaker. The preliminary reconstructed speech waveform is then smoothed across frames using an overlap-add method with a 50% overlap ratio. Adjacent frames are weighted and superimposed using a Hanning window function to eliminate frame boundary artifacts. The multi-speaker reconstructed speech waveforms are then synthesized and weighted according to the energy ratio of each speaker to obtain the final reconstructed complete speech waveform. The original mixed signal is subjected to a short-time Fourier transform (SFT) using a 25ms frame length, a 10ms frame shift, a Hamming window, and a 512-point FFT to extract the phase information at each time-frequency point. The enhanced near-end speech signal is then subjected to the Fast ICA algorithm for multi-speaker speech separation, with a maximum iteration count of 100 and a convergence threshold of 1e-6. Assuming three speakers, the speech power spectrum of each speaker is obtained. The original signal phase information is combined with the separated speech power spectrum, and a binary time-frequency masking method is used with an energy threshold of -40dB to construct a complete time-frequency domain representation for each speaker. A 512-point FFT is performed on the combined time-frequency domain representation to convert the signal from the frequency domain back to the time domain. The dynamic time warping algorithm is used to address the temporal alignment of the different speakers' speech. A window size of 100ms and a step size of 10ms are used to obtain a preliminary reconstructed speech waveform for each speaker. The initially reconstructed speech waveform was smoothed between frames using the overlap-and-add method, with a 50% overlap ratio. Adjacent frames were weighted and superimposed using a 32-millisecond Hanning window to eliminate frame boundary artifacts. The multi-speaker reconstructed speech was then synthesized. The short-term energy of each speaker's speech was calculated, and the average value of the 200-millisecond window was used as the weight for weighted superposition to obtain the final reconstructed complete speech waveform.

[0032] S105 , post-processing the reconstructed near-end speech signal, smoothing the signal through adaptive Wiener filtering, dynamically adjusting the frequency response of the filter based on the evaluated background noise power spectrum, and suppressing music noise in two-way communication.

[0033] The reconstructed near-end speech signal is subjected to a short-time Fourier transform to obtain a time-frequency representation of the signal; based on the time-frequency representation, the power spectrum of the background noise is evaluated using a recursive averaging method to obtain an initial noise power spectrum evaluation value; a decision-guided method is used to adaptively update the prior signal-to-noise ratio, and the prior signal-to-noise ratio is used to calculate the frequency response of the Wiener filter; the Wiener filter frequency response is applied to the amplitude spectrum of the original signal, and the time domain signal is reconstructed by an inverse short-time Fourier transform to obtain a near-end speech signal after preliminary smoothing processing; music noise detection is performed on the signal after preliminary processing, and the stationarity of each frequency band is calculated using a spectral entropy method. If the spectral entropy value is less than a preset threshold, it is determined that music noise exists; if it is determined that music noise exists, the corresponding frequency band is selectively suppressed by adjusting the gain curve of the Wiener filter; by monitoring the change of the short-time signal-to-noise ratio, the Wiener filter coefficient is dynamically updated to implement an adaptive process to obtain a near-end speech signal after suppressing music noise.

[0034] Specifically, the reconstructed near-end speech signal is subjected to a short-time Fourier transform (SFT) to obtain its time-frequency representation. The background noise power spectrum is then estimated using a recursive averaging method to obtain an initial noise power spectrum estimate. Based on the estimated background noise power spectrum, the signal-to-noise ratio (SNR) is calculated for each time-frequency point. The prior SNR is adaptively updated using a decision-directed approach with a smoothing factor set to 0.98. The frequency response of the Wiener filter is then calculated based on the updated SNR. The calculated Wiener filter frequency response is applied to the amplitude spectrum of the original signal, preserving its phase information to obtain an enhanced speech spectrum. The time domain signal is then reconstructed using an inverse short-time Fourier transform (IFT) to obtain a preliminarily smoothed near-end speech signal. Music noise detection is then performed on the preliminarily processed signal. The stationarity of each frequency band is calculated using the spectral entropy method, with a spectral entropy threshold set to 0.6. If music noise is detected, the Wiener filter gain curve is adjusted to selectively suppress the corresponding frequency band. The Wiener filter parameters are dynamically adjusted, and the filter coefficients are updated by monitoring changes in the short-time SNR, implementing an adaptive process. Ultimately, the near-end speech signal with music noise suppressed is obtained. The reconstructed near-end speech signal is subjected to a short-time Fourier transform (SFT) using a 25-ms frame length, a 10-ms frame shift, a Hamming window, and a 512-point FFT to obtain a time-frequency representation of the signal. The power spectrum of the background noise is estimated using a recursive averaging method with a smoothing factor of 0.95, and the noise estimate is updated every 50 frames. The posterior signal-to-noise ratio (SNR) is calculated for each time-frequency point, and the prior SNR is adaptively updated using a decision-directed approach with a smoothing factor of 0.98. The frequency response of the Wiener filter is then calculated. The Wiener filter frequency response is applied to the amplitude spectrum of the original signal, preserving the original phase, and the time domain signal is reconstructed using a 512-point I FFT. The processed signal is then subjected to music noise detection, calculating the spectral entropy using a 32-ms window with a spectral entropy threshold of 0.6. If music noise is detected, the Wiener filter gain is reduced by 3 dB in the corresponding frequency band. The Wiener filter parameters are dynamically adjusted, calculating the short-term SNR every 100 ms. If the SNR changes by more than 5 dB, the filter coefficients are updated and the smoothing factor is adjusted to 0.9. Finally, an overlay-add operation is performed on the enhanced speech signal to produce the near-end speech signal after suppressing the music noise. The latency of the entire process is controlled within 50 milliseconds, making it suitable for real-time two-way communication scenarios.

[0035] S106: Determine whether there are voices at both the near-end and far-end simultaneously. If a double-talk state is detected, adjust the parameters of the nonlinear echo cancellation algorithm to prevent the near-end voice from being incorrectly suppressed.

[0036] The short-term energy of the near-end and far-end signals is calculated. Adaptive thresholding technology is used to dynamically adjust the threshold based on the long-term average energy of the signals to determine whether simultaneous speech activity is present at both the near-end and far-end ends, generating preliminary double-talk detection results. Based on the preliminary double-talk detection results, a cross-correlation analysis method is used to calculate the correlation coefficient between the near-end and far-end signals. If the correlation coefficient exceeds a preset threshold, a double-talk condition is confirmed. Based on the double-talk condition, the suppression factor in the nonlinear echo cancellation algorithm is adjusted within a preset range to prevent false suppression of near-end speech. Spectral envelope similarity between the near-end and far-end signals is compared through spectral analysis. The spectral distance is calculated using the logarithmic spectral distortion metric. A weighted fusion method is used to comprehensively determine the reliability of the double-talk condition by combining time-domain and frequency-domain features. The weighted fusion double-talk detection results are used to dynamically update the parameters of the nonlinear echo cancellation algorithm, including the step size of the adaptive filter and the gain control of the nonlinear processing. The step size of the adaptive filter and the gain control of the nonlinear processing are adjusted in real time based on the double-talk detection results to improve echo cancellation effectiveness.

[0037] Specifically, the system calculates the short-term energy of the near-end and far-end signals. Adaptive thresholding technology dynamically adjusts the threshold based on the long-term average energy of the signals to determine whether simultaneous speech activity is present at both ends, generating preliminary double-talk detection results. Cross-correlation analysis is performed using a 20-millisecond time window to calculate the correlation coefficient between the near-end and far-end signals. If the correlation coefficient exceeds 0.7, a double-talk condition is confirmed. Parallel computing reduces processing delay. Based on the double-talk detection results, the suppression factor in the nonlinear echo cancellation algorithm is adjusted. An exponential decay function is used to reduce the suppression strength during double-talk, with the adjustment range set between 0.5 and 1 to prevent false suppression of near-end speech. Spectral analysis compares the spectral envelope similarity of the near-end and far-end signals. The logarithmic spectral distortion metric is used to calculate spectral distance. A weighted fusion method is used to comprehensively determine the reliability of the double-talk condition, combining time-domain and frequency-domain features. Weights are dynamically adjusted based on the stability of each feature. The fused double-talk detection results are used to dynamically update the parameters of the nonlinear echo cancellation algorithm, including the step size of the adaptive filter and the gain control of the nonlinear processing. The short-time energy of the near-end and far-end signals is calculated, and the Hamming window with a frame length of 25 milliseconds and a frame shift of 10 milliseconds is used for framing, and the logarithmic energy of each frame is calculated. Adaptive threshold technology is used, and the initial threshold is set to -40dB. The threshold is updated every 50 milliseconds based on the average energy of the past 1 second, and the update coefficient is 0.9. If the energy of the near-end and far-end signals exceeds their respective thresholds at the same time, it is preliminarily determined to be a double-talk state. The cross-correlation analysis method is used, and a 20-millisecond rectangular window is used to calculate the normalized cross-correlation coefficient of the near-end signal and the far-end signal. If the maximum correlation coefficient exceeds 0.7, the double-talk state is further confirmed. Through parallel calculation, the cross-correlation results are updated every 10 milliseconds to control the processing delay within 30 milliseconds. According to the double-talk detection results, the suppression factor in the nonlinear echo cancellation algorithm is adjusted, and an exponential decay function is used.

[0038] α(n) = α_min + (α_max - α_min) * exp(-n / τ), where α_min = 0.5, α_max = 1, τ = 50, and n is the number of frames in which double talk persists. The logarithmic spectral distortion (LSD) metric of the near-end and far-end signals is calculated using a 512-point FFT, taking the first 257 frequency bins, and setting the LSD threshold to 5 dB. A weighted fusion method is employed, with weights assigned as 0.3 for short-term energy, 0.4 for cross-correlation, and 0.3 for spectral distortion. Double talk is considered if the fusion score exceeds 0.6. The nonlinear echo cancellation algorithm parameters are dynamically adjusted based on the fusion results. The adaptive filter step size is adjusted between 0.1 and 0.5, and the nonlinear processing gain is controlled between -6 dB and 0 dB. The total latency of the entire process is kept below 50 milliseconds, meeting real-time communication requirements.

[0039] S107. Model the transmission path of the processed near-end speech signal to evaluate the propagation characteristics of the speech signal from the speaker to the microphone, including the delay and attenuation of the direct sound and early reflected sound.

[0040] A short-time Fourier transform is performed on the processed near-end speech signal to obtain a time-frequency representation of the signal, and the envelope and phase information of the signal are calculated to obtain Mel-frequency cepstral coefficients. Based on the Mel-frequency cepstral coefficients, a minimum mean square error algorithm is used to separate the near-end speech signal from direct sound, and the time delay and attenuation coefficient of the direct sound are evaluated by minimizing the mean square error function. The speech signal is decomposed into multiple frequency bands using a uniformly distributed filter bank, and early reflection sound is evaluated for each of the multiple frequency bands. The main reflection path is identified through peak detection and K-means clustering analysis. Based on the time delay and attenuation coefficient of the direct sound and the main reflection path, a convolution-based multipath propagation model is constructed, and the frequency response and phase response are calculated. The mean square error and correlation coefficient of the multipath propagation model are calculated by comparing with a known room impulse response to determine whether the mean square error exceeds a preset threshold. If the mean square error exceeds the preset threshold, the direct sound separation process is returned to and the parameters are re-evaluated until the accuracy requirement is met or the maximum number of iterations is reached.

[0041] Specifically, a short-time Fourier transform is performed on the processed near-end speech signal to obtain its time-frequency representation. By calculating the signal's envelope and phase information, the Mel-frequency cepstral coefficients are extracted as acoustic features for subsequent transmission path modeling and direct sound identification. A minimum mean square error (MMSE) algorithm is used to separate the direct sound from the near-end speech signal. The delay and attenuation coefficients of the direct sound are estimated by minimizing the mean square error function, resulting in a preliminary transmission path model that provides a reference for early reflection assessment. A uniformly distributed filter bank is used to decompose the speech signal into multiple frequency bands. Early reflections are assessed for each frequency band, and the primary reflection paths are identified through peak detection and K-means clustering analysis. The number of cluster centers is set to three, corresponding to the three primary reflection paths. Based on the assessed direct sound and early reflection parameters, a convolution-based multipath propagation model is constructed. The frequency and phase responses are calculated to comprehensively construct a complete transmission path model, including delay and attenuation characteristics. The accuracy of the transmission path model is verified by comparing it with the known room impulse response and calculating the mean square error and correlation coefficient. If the error exceeds a preset threshold, the model returns to the second step and re-evaluates the parameters until the accuracy requirement is met or the maximum number of iterations is reached. The processed near-end speech signal is subjected to a short-time Fourier transform (SFT) using a 25-ms frame length, a 10-ms frame shift, a Hamming window, and a 512-point FFT to obtain a time-frequency representation of the signal. 13th-order Mel-frequency cepstral coefficients are calculated as acoustic features for direct sound identification. A minimum mean square error (MSE) algorithm is used for direct sound separation, with an initial step size of 0.01 and adaptive adjustment to minimize the mean square error function. The direct sound delay and attenuation coefficient are evaluated over 200 iterations. An 8-channel uniformly distributed filter bank is used to decompose the signal into multiple frequency bands with a frequency range of 0-8 kHz. Peak detection is performed on each frequency band with a threshold of 1.5 times the average energy. K-means clustering (K = 3) is used to identify the main reflection paths. A convolution model was constructed to represent multipath propagation. The delays for the direct sound and the three main reflection paths were set to 0ms, 5ms, 10ms, and 15ms, respectively, with corresponding attenuations of 1, 0.8, 0.6, and 0.4. The 512-point frequency response and phase response were calculated to obtain the transmission path model. The model was compared with the pre-measured room impulse response, and the mean square error (MSE) and correlation coefficient were calculated. The MSE thresholds were set to -20dB, and the correlation coefficient threshold to 0.9. If the accuracy requirements were not met, the system would return to the direct sound separation step for re-evaluation, with a maximum of five iterations. The entire process maintained latency within 100 milliseconds, making it suitable for real-time speech enhancement systems.

[0042] S108. Perform dereverberation processing on the near-end speech signal according to the transmission path model to suppress the reverberation effect caused by room reflection to obtain a near-end speech signal with improved clarity, and perform noise reduction processing to obtain a clear speech signal with improved signal-to-noise ratio.

[0043] An inverse filter is constructed based on the transmission path model, and the filter parameters are determined by the minimum mean square error criterion. A time domain convolution operation is performed on the near-end speech signal to obtain a speech signal after preliminary dereverberation processing. The dereverberation speech signal is subjected to a short-time Fourier transform to obtain a frequency domain representation of the speech signal. The noise power spectrum is evaluated using the minimum statistics method, and spectral subtraction is applied in the frequency domain to obtain a denoised speech power spectrum. The denoised speech signal is enhanced using a Wiener filter, and the Wiener filter uses a decision-guided method to evaluate the prior signal-to-noise ratio. The filter parameters are dynamically adjusted according to the evaluated signal-to-noise ratio to obtain an enhanced speech signal. The enhanced speech signal is subjected to an inverse short-time Fourier transform to restore the time domain signal. The inter-frame transition is smoothed by the overlap-add method to obtain a final clear speech signal, and the overlap-add method uses an overlap rate of 50% between the frames. The degree of improvement in speech clarity of the clear speech signal is calculated using the cepstral distance. The improvement in the signal-to-noise ratio of the clear speech signal is evaluated using a segmented signal-to-noise ratio calculation method.

[0044] Specifically, based on the transmission path model, an inverse filter is constructed using the minimum mean square error criterion. Time-domain convolution is performed on the near-end speech signal to suppress the reverberation caused by room reflections, resulting in a preliminarily dereverberated speech signal. The dereverberated speech signal is then subjected to a short-time Fourier transform (SFT), and the noise power spectrum is estimated using the minimum statistic method. Spectral subtraction is then applied in the frequency domain to subtract the noise component from the speech power spectrum, achieving frequency-domain noise reduction. The denoised speech signal is further enhanced using a Wiener filter. A decision-guided method is used to estimate the prior signal-to-noise ratio (SNR). The filter parameters are dynamically adjusted based on the estimated SNR to improve the speech SNR. An inverse short-time Fourier transform is performed on the enhanced speech signal to restore the time-domain signal. Inter-frame transitions are smoothed using the overlap-and-add method with a 50% overlap ratio, resulting in a final, clear speech signal. The improvement in speech clarity is calculated using the cepstral distance, and the improvement in SNR is assessed using the segmented SNR calculation method. The speech signals before and after the processing are compared to quantify the enhancement effect. Based on the transmission path model, an inverse filter was constructed using the minimum mean square error criterion. The filter order was set to 1024, the learning rate was 0.01, and the parameters were optimized over 200 iterations. A time-domain convolution operation was performed on the near-end speech signal with a frame length of 25ms and a frame shift of 10ms to suppress reverberation caused by room reflections. A 512-point short-time Fourier transform was performed on the dereverberated signal, and the noise power spectrum was estimated using the minimum statistic method with a search window length of 1.5 seconds. Spectral subtraction was applied with an oversubtraction factor of 1.2 to achieve frequency-domain noise reduction. Further enhancement was performed using a Wiener filter. A decision-directed approach was used to estimate the prior signal-to-noise ratio (SNR). The smoothing factor α was set to 0.98, and the filter parameters were dynamically adjusted based on the estimated SNR. A 512-point inverse short-time Fourier transform was performed on the enhanced signal, and overlap-add was performed using a Hanning window with a 50% overlap ratio to smooth inter-frame transitions. The cepstral distance between the pre- and post-processing signals was calculated, and the 13th-order Mel-frequency cepstral coefficients were used to evaluate the degree of intelligibility improvement. Segmented signal-to-noise ratios are calculated using non-overlapping 20ms windows, and the average is taken to assess overall signal-to-noise ratio improvement. The entire processing delay is kept within 50ms, making it suitable for real-time voice communication systems. The resulting speech signal, with improved clarity and signal-to-noise ratio, quantifies the enhancement effect.

[0045] The above is only a preferred embodiment of the present invention. It should be pointed out that ordinary technicians in this technical field can make several improvements and supplements without departing from the principles of the present invention. These improvements and supplements should also be regarded as the scope of protection of the present invention.

Claims

1. A method for eliminating echo in an intercom system, characterized in that: The method comprises: Based on the mixed signal collected by the near-end microphone array, the speaker position is estimated, the target speaker's voice is enhanced, the enhanced near-end voice signal is obtained, the timbre characteristics are extracted, a speaker timbre model is established, the reverberation of the far-end reference signal is evaluated, the room impulse response is determined, and the reverberation characteristic parameters of the far-end signal are obtained. Based on the reverberation characteristic parameters and the far-end signal, the enhanced near-end voice signal is preliminarily processed to subtract the evaluated echo component; The spectral envelope tracking (SET) of the preliminarily processed near-end speech signal is performed to calculate the energy distribution at each time-frequency point. Based on the energy dynamic range of the far-end signal leakage echo, the near-end speech and residual echo are adaptively distinguished between speech-dominant and echo-dominant regions, and the near-end speech and residual echo are modeled as independent signal sources. The separation matrix is evaluated by maximizing the negative entropy, and the speech signal is linearly transformed to obtain the separated near-end speech evaluation. Based on the near-end speech evaluation after separation, the multiple speaker speech segments identified are subjected to residual echo suppression respectively, and a binary time-frequency mask is constructed to reduce the residual echo spectrum caused by the dynamic change of echo power; The phase information of the original mixed signal is combined with the enhanced near-end speech power spectrum to reconstruct the speech waveform of each speaker through inverse short-time Fourier transform. The overlap-add method is used to smooth the transition between speech frames to obtain the reconstructed near-end speech signal. The reconstructed near-end speech signal is post-processed and smoothed using adaptive Wiener filtering. The frequency response of the filter is dynamically adjusted based on the estimated background noise power spectrum to suppress musical noise in two-way communication. Determines whether there is simultaneous speech at both the near-end and far-end. If a double-talk state is detected, the parameters of the nonlinear echo cancellation algorithm are adjusted to prevent the near-end speech from being incorrectly suppressed. Model the transmission path of the processed near-end speech signal to evaluate the propagation characteristics of the speech signal from the speaker to the microphone, including the delay and attenuation of the direct sound and early reflections; According to the transmission path model, the near-end speech signal is subjected to dereverberation processing to suppress the reverberation effect caused by room reflection, thereby obtaining a near-end speech signal with improved clarity. The noise reduction processing is then performed to obtain a clear speech signal with improved signal-to-noise ratio.

2. The method according to claim 1, wherein The method includes: evaluating the speaker position based on the mixed signal collected by the near-end microphone array, enhancing the target speaker's voice, obtaining the enhanced near-end voice signal, extracting timbre features, establishing a speaker timbre model, performing reverberation evaluation on the far-end reference signal, determining the room impulse response, obtaining reverberation characteristic parameters of the far-end signal, performing preliminary processing on the enhanced near-end voice signal based on the reverberation characteristic parameters and the far-end signal, and subtracting the evaluated echo component, including: Based on the mixed signal collected by the near-end microphone array, the beamforming algorithm is used to estimate the speaker position, and the target speaker's voice is enhanced through beamforming technology to obtain the enhanced near-end speech signal; extracting the timbre characteristics of the enhanced near-end speech signal using Mel-frequency cepstral coefficients, and establishing a Gaussian mixture model as a speaker timbre model; Perform time-frequency analysis on the far-end reference signal, calculate the reverberation time and early decay rate, determine the room impulse response, and obtain the reverberation characteristic parameters of the far-end signal; performing envelope modulation filtering on the near-end speech signal according to the reverberation characteristic parameters and the far-end signal, adjusting the signal amplitude to match the reverberation characteristic of the far-end signal, and subtracting the evaluated echo component to obtain a preliminarily processed near-end speech signal; Using an adaptive filter to perform echo cancellation on the preliminarily processed near-end speech signal, updating the filter coefficients using a minimum mean square error criterion, suppressing double echoes through nonlinear processing, and obtaining a final enhanced near-end speech signal; The timbre of the final enhanced near-end speech signal is verified according to the speaker timbre model to determine whether the signal enhancement process introduces obvious timbre distortion.

3. The method according to claim 1, wherein The method performs spectral envelope tracking on the preliminarily processed near-end speech signal, calculates the energy distribution of each time-frequency point, adaptively distinguishes the speech-dominant region from the echo-dominant region based on the energy dynamic range of the far-end signal leakage echo, models the near-end speech and residual echo as independent signal sources, and performs a linear transformation on the speech signal by maximizing the negative entropy to evaluate the separation matrix, thereby obtaining a near-end speech evaluation after separation, including: Performing a short-time Fourier transform on the preliminarily processed speech signal to obtain a time-frequency spectrum of the speech signal; According to the time-frequency spectrum, the spectrum envelope is tracked using the cepstrum analysis method, the energy distribution of each time-frequency point is calculated, and an energy distribution matrix is obtained; According to the energy dynamic range of the far-end signal leakage echo, a threshold is dynamically adjusted based on a sliding average of the far-end signal energy, and regions above the threshold in the energy distribution matrix are marked as speech-dominated regions, and regions below the threshold are marked as echo-dominated regions, thereby generating a time-frequency masking map; Applying the time-frequency masking map information to an independent component analysis process, dividing the signal into two sources, speech and echo, according to the time-frequency masking map, and constructing a mixing matrix; Modeling the near-end speech and residual echo as independent signal sources, evaluating the separation matrix using an independent component analysis algorithm, and optimizing the separation matrix parameters by maximizing the negative entropy criterion; The speech signal is linearly transformed, the estimated separation matrix is multiplied by the initial mixed signal, and the separated signal is reconstructed using a Wiener filter in combination with the time-frequency masking graph to obtain a separated near-end speech evaluation signal.

4. The method according to claim 1, wherein The method includes: performing residual echo suppression on the multiple speaker speech segments identified based on the separated near-end speech evaluation, constructing a binary time-frequency mask, and reducing the residual echo spectrum caused by the dynamic change of the echo power, including: Segmenting the speech signal using a Gaussian mixture model based on Mel-frequency cepstral coefficients to identify speech segments of multiple speakers, each containing speaker identity and time range information. performing residual echo suppression based on the speaker identity information of the speech segment in combination with an adaptive filter and spectral subtraction, subtracting the echo component from the power spectral density of the near-end speech signal to obtain an echo-suppressed speech signal; For the speech signal after echo suppression, calculate the short-time energy and zero-crossing rate, use the long-term average spectrum of the speech signal for normalization, set an adaptive threshold, and determine the speech activity area; If a speech activity area is determined, a binary time-frequency masking matrix is constructed, marking the speech activity area as 1 and the non-speech activity area as 0; Using a Wiener filter in combination with the binary time-frequency masking matrix to perform spectrum enhancement, reducing the residual echo spectrum caused by dynamic changes in echo power, and obtaining an enhanced near-end speech signal; The enhanced near-end speech signal is further optimized by adjusting the parameters of the Wiener filter by monitoring the short-term change of the echo power.

5. The method according to claim 1, wherein The method utilizes the phase information of the original mixed signal, combines it with the enhanced near-end speech power spectrum, reconstructs the speech waveform of each speaker through inverse short-time Fourier transform, and uses overlap-add method to smooth the transition between speech frames to obtain the reconstructed near-end speech signal, including: Performing short-time Fourier transform on the original mixed signal to obtain phase information of each time-frequency point; Applying an independent component analysis algorithm to the enhanced near-end speech signal according to the phase information to obtain a speech power spectrum of multiple speakers; Combining the original mixed signal phase information with the speech power spectrum to construct a complete time-frequency domain representation through a time-frequency masking method; Performing an inverse short-time Fourier transform on the time-frequency domain representation to obtain a time-domain signal; The overlap-add method is used to perform inter-frame smoothing on the time domain signal, and adjacent frames are weightedly superimposed by setting an overlap rate and using a Hanning window function; If the time domain signal contains multi-speaker speech, weighted superposition is performed according to the energy ratio of each speaker to obtain a final reconstructed complete speech waveform.

6. The method according to claim 1, wherein The post-processing of the reconstructed near-end speech signal, smoothing the signal by adaptive Wiener filtering, and dynamically adjusting the frequency response of the filter according to the evaluated background noise power spectrum to suppress music noise in two-way communication includes: Performing a short-time Fourier transform on the reconstructed near-end speech signal to obtain a time-frequency representation of the signal; According to the time-frequency representation, the power spectrum of the background noise is evaluated using a recursive averaging method to obtain an initial noise power spectrum evaluation value; Adaptively updating a priori signal-to-noise ratio using a decision-directed approach, wherein the priori signal-to-noise ratio is used to calculate the frequency response of the Wiener filter; Applying the Wiener filter frequency response to the amplitude spectrum of the original mixed signal, reconstructing the time domain signal by inverse short-time Fourier transform, and obtaining a near-end speech signal after preliminary smoothing; Performing music noise detection on the signal after the preliminary processing, calculating the stationarity of each frequency band using a spectral entropy method, and determining that music noise exists if the spectral entropy value is less than a preset threshold; If it is determined that music noise exists, selectively suppressing the corresponding frequency band by adjusting the gain curve of the Wiener filter; By monitoring the short-term signal-to-noise ratio change, the Wiener filter coefficient is dynamically updated to implement an adaptive process, thereby obtaining a near-end speech signal after suppressing music noise.

7. The method according to claim 1, wherein The determining whether there are voices at both the near-end and far-end simultaneously, and if a double-talk state is detected, adjusting the parameters of the nonlinear echo cancellation algorithm to prevent the near-end voice from being incorrectly suppressed, includes: The short-term energy of the near-end and far-end signals is calculated, and adaptive threshold technology is used to dynamically adjust the threshold based on the long-term average energy of the signal to determine whether there is simultaneous speech activity at both the near-end and far-end ends, thereby obtaining preliminary double-talk detection results. Calculating the correlation coefficient between the near-end signal and the far-end signal using a cross-correlation analysis method based on the preliminary double-talk detection result, and further confirming the double-talk state if the correlation coefficient exceeds a preset threshold; According to the double-talk state, adjusting the suppression factor in the nonlinear echo cancellation algorithm, wherein the adjustment range of the suppression factor is set to a preset interval to prevent the near-end speech from being incorrectly suppressed; The spectrum envelope similarity between the near-end and far-end signals is compared through spectrum analysis, and the spectrum distance is calculated using the logarithmic spectrum distortion metric. The reliability of the dual-talk state is comprehensively judged using a weighted fusion method combining time-domain and frequency-domain features. Using the weighted fusion double-talk detection result to dynamically update the parameters of the nonlinear echo cancellation algorithm, including the step size of the adaptive filter and the gain control of the nonlinear processing; The step size of the adaptive filter and the gain control of the nonlinear processing are adjusted in real time according to the double-talk detection result, thereby improving the effect of echo cancellation.

8. The method according to claim 1, wherein The transmission path modeling of the processed near-end speech signal is performed to evaluate the propagation characteristics of the speech signal from the speaker to the microphone, including the time delay and attenuation of the direct sound and early reflection sound, including: Performing a short-time Fourier transform on the processed near-end speech signal to obtain a time-frequency representation of the signal, and obtaining Mel-frequency cepstral coefficients by calculating the envelope and phase information of the signal; Separating the near-end speech signal from the direct sound using a minimum mean square error (MMSE) algorithm based on the Mel-frequency cepstral coefficients, and evaluating the time delay and attenuation coefficient of the direct sound by minimizing the MSE function. Decomposing the speech signal into multiple frequency bands using a uniformly distributed filter bank, performing early reflection sound evaluation on each of the multiple frequency bands, and identifying a main reflection path through peak detection and K-means cluster analysis; Constructing a convolution-based multipath propagation model based on the time delay and attenuation coefficient of the direct sound and the main reflection path, and calculating the frequency response and phase response; Calculating the mean square error and correlation coefficient of the multipath propagation model by comparing with a known room impulse response, and determining whether the mean square error exceeds a preset threshold; If the mean square error exceeds a preset threshold, the process returns to direct sound separation and re-evaluates the parameters until the accuracy requirement is met or the maximum number of iterations is reached.

9. The method according to claim 1, wherein: The method includes performing dereverberation processing on the near-end speech signal according to the transmission path model to suppress the reverberation effect caused by room reflection to obtain a near-end speech signal with improved clarity, and performing noise reduction processing to obtain a clear speech signal with improved signal-to-noise ratio, including: Constructing an inverse filter according to the transmission path model, wherein the inverse filter adopts a minimum mean square error criterion to determine filter parameters; Perform time-domain convolution on the near-end speech signal to obtain a speech signal after preliminary dereverberation processing; Performing a short-time Fourier transform on the dereverberated speech signal to obtain a frequency domain representation of the speech signal; The noise power spectrum is evaluated using the minimum statistics method, and spectral subtraction is applied in the frequency domain to obtain the power spectrum of the denoised speech. Performing enhancement processing on the noise-reduced speech signal using a Wiener filter, wherein the Wiener filter uses a decision-directed method to evaluate a priori signal-to-noise ratio; Dynamically adjusting filter parameters according to the evaluated signal-to-noise ratio to obtain an enhanced speech signal; Performing an inverse short-time Fourier transform on the enhanced speech signal to restore a time domain signal; Smoothing the inter-frame transition by an overlap-add method to obtain a final clear speech signal, wherein the overlap rate between the frames is 50 percent; Calculating the degree of improvement of speech clarity of the clear speech signal using the cepstral distance; A segmented signal-to-noise ratio calculation method is used to evaluate the improvement of the signal-to-noise ratio of the clear speech signal.

Citation Information

Patent Citations

  • Echo cancellation method and device, electronic equipment and storage medium

    CN114171049A

  • Echo cancellation method and device and electronic equipment

    CN114650340A