An echo cancellation method and system based on neural network dual-talker detection
The RNN-based double-talk detection enhances AEC by dynamically adjusting filter coefficients and incorporating auditory masking effects, addressing the robustness and accuracy issues in existing AEC methods, resulting in improved echo suppression.
Patent Information
- Application Number
- CN202210888604.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-27
AI Technical Summary
Among the existing echo cancellation methods, the double-talk detection scheme has poor robustness and low detection accuracy, resulting in unsatisfactory echo cancellation effect, and the linear adaptive filter cannot effectively remove linear and nonlinear residual echoes.
A dual-talk detection method based on neural network is adopted, combined with recurrent neural network (RNN) for vocal detection, the dual-talk detection state results are controlled by a finite state machine, combined with linear adaptive filtering and nonlinear post-processing, and the filter coefficients are optimized using variable step factor and gain function to perform adaptive updates and residual echo suppression.
It improves the robustness and accuracy of double-talk detection, effectively removes linear and nonlinear residual echoes, improves the echo cancellation effect, prevents the adaptive filter from diverging when there is no human voice, and optimizes the fidelity of the near-end signal.
Smart Images

Figure CN115457928B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular, to an echo cancellation method and system based on neural network double-talk detection. Background Art
[0002] The acoustic echo cancelling (AEC) algorithm is one of the commonly used algorithms in the fields of speech signal processing and speech communication, and is widely used in applications such as speech communication and intelligent speech human-computer interaction. In speech communication, it is mainly to solve the problem that any speaker at either end hears their own voice (echo) during a full-duplex call. Using the echo cancellation algorithm, echo cancellation can be performed in advance at the near end and then sent to the far end, so that the far-end speaker will no longer hear their own voice. In the process of intelligent speech human-computer interaction, in order to prevent the music or voice played by the intelligent device itself from interfering with speech recognition, it is also necessary to use the echo cancellation algorithm to remove the sound played by itself in advance to prevent misrecognition and improve the recognition rate.
[0003] In the existing AEC methods, an adaptive linear filter is usually used to estimate the echo signal, and then the echo signal in the communication system is cancelled according to the estimated echo signal. In order to improve the effect of the linear adaptive filter in the AEC method, a double-talk detection module (DTD) is usually added to cooperate with the adaptive linear filter. The double-talk detection module is used to detect the speaking states of both parties in the communication. For example, when both parties in the communication are speaking simultaneously, it is a double-talk state. In the related art, at one end of the communication, whether it is a double-talk state is determined by detecting the local voice signal (i.e., the near-end voice signal) and the voice signal of the other end (i.e., the far-end voice signal).
[0004] However, the problems of poor robustness and low detection accuracy of the existing double-talk detection schemes make the echo cancellation effect unsatisfactory; in addition, only the linear adaptive filter is used to cancel the echo signal, and there are still linear residual echo and non-linear residual echo signals in the obtained signal, which will affect the final echo cancellation effect.
[0005] Therefore, it is necessary to provide an echo cancellation method and system based on neural network double-talk detection to solve the above technical problems. Summary of the Invention
[0006] To solve the above technical problems, an echo cancellation method based on neural network double-talk detection provided by the present invention includes input signal processing, linear adaptive filtering processing, non-linear post-processing, RNN double-talk detection, and output signal processing.
[0007] Specifically, for input signal processing: collect the microphone signal at the near end and the reference signal at the far end, and transmit them in the form of a digital signal stream; store the microphone signal and the reference signal into the input buffer respectively, and the input buffer divides the signals into several data blocks, and the data blocks include the microphone signal data block d l (n) and the reference signal data block x l (n); where l = 1, 2, 3,... represents the data block serial number, n = 0, 1, 2,..., n represents the sample serial number of each data block, and N is the total number of samples in each data block.
[0008] Specifically, for RNN double-talk detection: perform voice detection on the microphone signal and the reference signal through a recurrent neural network (RNN), and use a finite state machine to control and give the double-talk detection status result db_flag(l). Among them, the double-talk detection status result db_flag(l) includes: the far-end single-talk state far_talk_only with only voices at the far end, the near-end single-talk state near_talk_only with only voices at the near end, and the far-end and near-end double-talk state doble_talk with voices at both the far end and the near end; the double-talk detection status result db_flag(l) is used to perform feedback adjustment on linear adaptive filtering processing and nonlinear post-processing.
[0009] Specifically, for linear adaptive filtering processing: receive the microphone signal data block d l (n), the reference signal data block x l (n) and perform point-by-point data processing; the data processing is carried out through the NLMS algorithm and is adaptively adjusted through the double-talk detection status result db_flag(l) to obtain the adaptively updated filter coefficients Calculate the adaptively updated residual signal e (n) through the filter coefficients l (n).
[0010] Specifically, for nonlinear post-processing: used to further remove the linear residual echo and the nonlinear residual echo signals in the residual signal e l (n); obtain the AEC output signal data block out l (n).
[0011] Specifically, for output signal processing: store the AEC output signal data block out l (n) after removing the echo into the output buffer, and perform data merging to obtain a continuous audio data stream for output.
[0012] As a further solution, the linear adaptive filtering processing is carried out through the following steps:
[0013] Receive the microphone signal data block d l(n) and reference signal data block x l (n), initialize the filter coefficient vector
[0014] where denotes the reference signal vector at the n-th point of the l-th reference signal data block; denotes the filter coefficient vector corresponding to the current x l (n); T represents the transpose of the current vector, L is the filter length set during initialization, the initial values of are all set to 0;
[0015] Estimate the echo signal at the n-th point of the current frame through the filter coefficient at the n-th point of the previous frame
[0016]
[0017] Calculate the estimated residual signal e at the n-th point of the current frame: l (n):
[0018]
[0019] Calculate the reference signal energy E l,x (n):
[0020] E l,x (n) = x l (n) T x l (n)
[0021] Calculate the variable step size factor μ l (n);
[0022] Update the estimated echo signal and the auto-correlation function and cross-correlation function of the estimated residual signal e l (n):
[0023]
[0024]
[0025] where r dd (n) is the auto-correlation function of the echo signal ; r de (n) is the cross-correlation function of the echo signal and the residual signal e l (n), where α is the forgetting coefficient, r dd (n) and r de (n) are initially set to 0;
[0026] Perform RNN double-talk detection to obtain the double-talk detection status result db_flag(l);
[0027] According to the double-talk detection status result db_flag(l), the filter coefficient of the nth point in the previous frame Variable step size factor μ l (n), reference signal vector x l (n), residual signal e l (n) and reference signal energy E l,x (n) adaptively update the filter coefficient Perform adaptive update of the filter coefficient;
[0028] Through the adaptively updated filter coefficient Calculate the adaptively updated residual signal e l (n);
[0029] Take the adaptively updated residual signal e l (n) as the linear adaptive filtering output result, and perform point-by-point processing on each point in the data block through the above steps to obtain the residual signal output of the data block: [e l (n), n = 0, 1, 2, …, N].
[0030] As a further solution, the variable step size factor μ l (n) is calculated by the following formula:
[0031]
[0032] where ε is the regularization factor; μ0 is the maximum adaptive step size constant; r dd (n) is the autocorrelation function of the echo signal ; r de (n) is the cross-correlation function of the echo signal and the residual signal e l (n).
[0033] As a further solution, the adaptive update of the filter coefficient is calculated by the following formula:
[0034]
[0035] where db_flag(l - 1) represents the double-talk detection status result corresponding to the double-talk detection status information of the previous frame given by the RNN double-talk detection module; far_talk_only indicates that there is only a far-end human voice speaking; else indicates that when the double-talk detection result is not far_talk_only, stop updating the filter; ε is the regularization factor.
[0036] As a further solution, the non-linear post-processing is carried out through the following steps:
[0037] Perform frequency-domain processing on the residual signal e l (n) of data block l and the estimated echo signal to obtain the complex spectrum S e (l, k) of the residual signal in the frequency-domain sub-band, the complex spectrum of the echo signal The energy spectrum P e (l, k) of the residual signal and the echo signal where k represents the serial number of the discrete sampling points in the frequency domain, k = 0, 1,..., N B -1; N B is the total number of frequency-domain sub-bands;
[0038] Obtain the energy spectrum P l (n) of the remaining echo in the residual signal e res (l, k) through the following formula:
[0039]
[0040] where and are the correlation function values calculated from the last sample point N of the previous data block l-1;
[0041] Perform weighted processing on the complex spectrum S e (l, k) through the gain function G(l, k) to obtain the complex spectrum S o (l, k) of the final output signal:
[0042] S o (l, k) = G(l, k) · S e (l, k)
[0043] where G(l, k) is the gain function and P e (l, k) is the energy spectrum of the residual signal;
[0044] Perform inverse short-time Fourier transform (ISTFT) on the complex spectrum S o (l, k) of the final output signal to obtain the AEC output signal data block out l (n) in the time domain.
[0045] As a further solution, the gain function G(l, k) is obtained through a Wiener filter and is subjected to a maximum suppression amount constraint:
[0046] G(l, k) = max(G wiener (l, k), min_G(l, k))
[0047] min_G(l, k) = linear(max_attenu(l, k))
[0048] where G wiener (l, k) is the Wiener filter output corresponding to the k-th subband of the l-th data block, max_attenu(l, k) represents the maximum amount of residual echo suppression for the k-th subband of the l-th data block, and linear() is a linear function;
[0049] The maximum amount of suppression max_attenu(l, k) is set according to the double-talk detection status result:
[0050] When db_flag(l - 1) == doble_talk, the maximum amount of suppression max_attenu(l, k) is:
[0051] max_attenu(l, k) = db(min_gain(l, k)) - 3
[0052]
[0053] where db() is a function that converts linear gain to db value, ε is a regularization factor, P res (l, k) is the energy spectrum of the residual echo, P e (l, k) is the energy spectrum of the residual signal.
[0054] When db_flag(l - 1) == far_talk_only, the maximum amount of suppression max_attenu(l, k) is set to a preset value, and the preset value is used to increase suppression:
[0055] When db_flag(l - 1) == near_talk_only, no residual echo suppression is performed, and max_attenu(l, k) is set to 0.
[0056] As a further solution, the recurrent neural network RNN learns by extracting MFCC features, and uses Dense layers and GRU layers to implement the estimation of the probability of the presence of human voice, and finally outputs the probability of the presence of human voice.
[0057] As a further solution, the weight coefficients in the recurrent neural network (RNN) are obtained by preprocessing training data, and the training data is obtained through the following steps: pre-record proximal human voice speech data, use far-end human voice single-speaker data for acoustic echo cancellation (AEC) processing to obtain AEC residual data and ambient noise data; mix the proximal human voice speech data with the AEC residual data and the ambient noise data respectively to obtain noisy speech signals; label the proximal human voice speech data to obtain the vocal presence position labels in the noisy speech signals; use the noisy speech signals and the vocal presence position labels as training samples, and pre-train the recurrent neural network (RNN) to obtain the weight coefficients in it.
[0058] As a further solution, the double-talk detection status result db_flag(l) is delayed and output through a delay unit Z -1 for delay output. The delay unit Z -1 delays the data by one unit duration, so that the linear adaptive filtering process and the non-linear post-processing obtain the double-talk detection status result db_flag(l - 1) of the previous frame.
[0059] An echo cancellation system based on neural network double-talk detection runs on a hardware device. The hardware device includes a signal collector, an input buffer, a linear adaptive filtering module, a non-linear post-processing module, an RNN double-talk detection module, and a delay module; and the echo signal is cancelled by using an echo cancellation method based on neural network double-talk detection as described in any one of the above.
[0060] Compared with the related technology, an echo cancellation method based on neural network double-talk detection provided by the present invention has the following beneficial effects:
[0061] 1. In the present invention, the update of the adaptive filter in the linear pre-processing is controlled by using the double-talk detection result; when no far-end human voice signal is detected, the update of the adaptive filter is stopped. This prevents the adaptive filter from diverging due to perturbation and deviating from the stable point when in the double-talk state or when there is only ambient noise at the far end.
[0062] 2. In the present invention, the maximum suppression amount of the echo in the non-linear post-processing is controlled by using the double-talk detection result. When it is detected that the near-end and far-end human ears are speaking simultaneously, the maximum suppression amount required to mask the residual echo is estimated in combination with the auditory masking effect of the human ear. While keeping the distortion of the useful signal at the near end small, the residual echo is also effectively suppressed; when it is detected that there is only a single-speaker signal from the far-end human ear, the suppression of the residual echo signal is enhanced, so that the residual echo can be completely removed.
[0063] 3. The present invention performs far - end and near - end double - talk detection through neural network technology, effectively solving the problems of poor robustness and low detection accuracy in existing double - talk detection schemes; a finite - state machine is used to control the far - end and near - end voice detection states, improving robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 FIG. is a preferred flowchart of an echo cancellation method based on neural network double - talk detection provided by an embodiment of the present invention;
[0065] Figure 2 FIG. is a schematic diagram of the principle of an echo cancellation method based on neural network double - talk detection provided by an embodiment of the present invention;
[0066] Figure 3 FIG. is a preferred schematic diagram of the process of linear adaptive filtering processing provided by an embodiment of the present invention;
[0067] Figure 4 FIG. is a preferred schematic diagram of the process of RNN double - talk detection provided by an embodiment of the present invention;
[0068] Figure 5 FIG. is a preferred schematic diagram of the process of training the RNN double - talk detection model provided by an embodiment of the present invention;
[0069] Figure 6 FIG. is a preferred schematic diagram of the process of constructing training samples provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The present invention will be further described below with reference to the drawings and embodiments.
[0071] As Figure 1 shown in Figure 2 , an echo cancellation method based on neural network double - talk detection provided by this embodiment includes input signal processing, linear adaptive filtering processing, non - linear post - processing, RNN double - talk detection, and output signal processing.
[0072] Specifically, for input signal processing: collect the near - end microphone signal and the far - end reference signal, and transmit them in the form of a digital signal stream; store the microphone signal and the reference signal into the input buffer respectively, and the input buffer divides the signal into several data blocks, and the data block includes the microphone signal data block d l (n) and the reference signal data block x l (n); where l = 1, 2, 3,... represents the data block number, n = 0, 1, 2,..., n represents the sample number of each data block, and N is the total number of sample points of each data block.
[0073] Specifically, as Figure 4As shown in the figure, the RNN double-talk detection: The voice detection is performed on the microphone signal and the reference signal through the Recurrent Neural Network (RNN), and the finite state machine is used to control and give the double-talk detection status result db_flag(l). Among them, the double-talk detection status result db_flag(l) includes: the far-end single-talk state far_talk_only with only the far-end having a voice, the near-end single-talk state near_talk_only with only the near-end having a voice, and the double-talk state doble_talk with both the far-end and the near-end having voices; the double-talk detection status result db_flag(l) is used to perform feedback adjustment on the linear adaptive filtering process and the non-linear post-processing.
[0074] Specifically, for the linear adaptive filtering process: Receive the microphone signal data block d l (n), the reference signal data block x l (n) and perform point-by-point data processing; the data processing is carried out through the NLMS algorithm and is adaptively adjusted through the double-talk detection status result db_flag(l) to obtain the filter coefficients after adaptive update Through the filter coefficients Calculate the residual signal e l (n) after adaptive update.
[0075] Specifically, for the non-linear post-processing: It is used to further remove the linear residual echo and the non-linear residual echo signals in the residual signal e l (n); obtain the AEC output signal data block out l (n).
[0076] Specifically, for the output signal processing: Store the AEC output signal data block out l (n) after removing the echo into the output buffer and perform data merging to obtain a continuous audio data stream for output.
[0077] It should be noted that: As Figure 1 shown, an echo cancellation method based on neural network double-talk detection provided in this embodiment takes the microphone signal and the reference signal as the input signals for the acoustic echo cancellation problem. Here, the microphone signal and the reference signal are considered to be digital signal streams that have undergone analog-to-digital conversion (A / D). In the acoustic echo cancellation problem, the near-end and far-end voice signals often mentioned respectively correspond to the voice signals in the microphone signal and the reference signal.
[0078] The microphone signal and the reference signal will pass through the input buffer to obtain input signal data blocks to be processed in chunks. The two original audio signals will be input into the input buffer of the main processing flow. The input buffer divides the continuous input data stream into equal-length data blocks for subsequent processing. After the chunked microphone and reference audio signals are processed by linear adaptive filtering, the output signal is the microphone signal with linear echo removed. The non-linear post-processing module is used to further remove the linear residual echo and non-linear residual echo signals in the microphone signal. Common methods of the non-linear post-processing module include the correlation-based residual echo estimation method, the method of combining non-linear model modeling similar to Volterra filtering, or the neural network method. Here, for the sake of illustration, the correlation-based residual echo estimation method is taken as an example. The output signal of the non-linear post-processing module is the signal with both linear echo and non-linear residual echo removed.
[0079] In this embodiment, an RNN neural network double-talk detection is also used to improve the stability of the aforementioned linear adaptive filtering module and non-linear post-processing module, thereby improving the overall performance of the echo cancellation system. The input signals of the RNN neural network double-talk detection are the output signal of the non-linear post-processing module and the reference audio data block signal. The RNN neural network double-talk detection module uses two independent recursive neural networks (RNNs) to perform voice detection on the two inputs respectively to obtain voice detection flags (1 indicates that voice speech is detected, and 0 indicates that the current is a noise signal or voice speech is not detected). Then, based on a state machine control, the double-talk detection result is obtained and output. The result of the double-talk detection module will be used by the linear adaptive filtering module and the non-linear post-processing module. The microphone signal blocks after echo cancellation will enter the output buffer to obtain a continuous audio data stream for output again.
[0080] As a further solution, as Figure 3 shown, the linear adaptive filtering process is carried out through the following steps:
[0081] Receive the microphone signal data block d l (n) and the reference signal data block x l (n), and initialize the filter coefficient vector
[0082] where, represents the reference signal vector at the nth point of the lth reference signal data block; represents the filter coefficient vector corresponding to the current x l (n); T represents the transpose of the current vector, L is the filter length set during initialization, the initial values of are all set to 0;
[0083] Through the filter coefficient at the nth point of the previous frame Estimate the echo signal at the n-th point of the current frame
[0084]
[0085] Based on the estimated echo signal Calculate the estimated residual signal e l (n):
[0086]
[0087] Calculate the reference signal energy E l,x (n):
[0088] E l,x (n) = x l (n) T x l (n)
[0089] Calculate the variable step size factor μ l (n);
[0090] It should be noted that: the variable step size factor μ l (n) automatically adjusts the step size according to the residual echo magnitude in the output of the linear filter, which is used to accelerate the filter convergence speed and prevent the filter from being disturbed when the near-end human voice is speaking.
[0091] Update the estimated echo signal And the autocorrelation function and cross-correlation function of the estimated residual signal e l (n):
[0092]
[0093]
[0094] Among them, r dd (n) is the autocorrelation function of the echo signal ; r de (n) is the cross-correlation function of the echo signal and the residual signal e l (n), where α is the forgetting coefficient, r dd (n) and r de (n) are set to 0 for the initial values of the functions;
[0095] Perform RNN double-talk detection to obtain the double-talk detection status result db_flag(l);
[0096] According to the double-talk detection status result db_flag(l), the filter coefficients at the n-th point of the previous frame The variable step size factor μ l(n), reference signal vector x l (n), residual signal e l (n) and reference signal energy E l,x (n) for filter coefficients perform adaptive update of filter coefficients;
[0097] Through the filter coefficients after adaptive update calculate the residual signal e after adaptive update l (n);
[0098] Take the residual signal e after adaptive update l (n) as the output result of linear adaptive filtering, and perform point-by-point processing on each point in the data block through the above steps to obtain the residual signal output of the data block: [e l (n), n = 0, 1, 2, …, N].
[0099] As a further solution, the variable step size factor μ l (n) is calculated by the following formula:
[0100]
[0101] where ε is a regularization factor (to prevent the denominator from being 0); μ0 is the maximum adaptive step size constant; r dd (n) is the autocorrelation function of the echo signal ; r de (n) is the cross-correlation function of the echo signal and the residual signal e l (n).
[0102] As a further solution, the adaptive update of filter coefficients is calculated by the following formula:
[0103]
[0104] where db_flag(l - 1) represents the double-talk detection state result corresponding to the double-talk detection state information of the previous frame given by the RNN double-talk detection module; far_talk_only means only the far-end human voice is speaking; else means to stop updating the filter when the double-talk detection result is not for_talk_omly; ε is a regularization factor.
[0105] As a further solution, the non-linear post-processing estimates the residual echo in the residual based on the correlation principle and combines the human ear auditory masking effect to control the maximum residual echo suppression amount in the double-talk stage, effectively preventing distortion of the proximal human voice signal. At the same time, when there is only the far-end human voice in single talk, a larger echo suppression amount is given to completely remove the residual echo. The non-linear post-processing is performed through the following steps:
[0106] Perform frequency-domain processing on the residual signal e l (n) of data block l and the estimated echo signal to obtain the complex spectrum S of the residual signal in the frequency-domain subband e (l, k), the complex spectrum of the echo signal the energy spectrum P of the residual signal e (l, k) and the echo signal where k represents the serial number of the discrete sampling point in the frequency domain, k = 0, 1,..., N B -1; N B is the total number of frequency-domain subbands;
[0107] Obtain the energy spectrum P of the residual echo in the residual signal e l (n) through the following formula: res (l, k):
[0108]
[0109] where, and are the correlation function values calculated from the last sample point N of the previous data block l - 1;
[0110] Perform weighted processing on the complex spectrum S e (l, k) through the gain function G(l, k) to obtain the complex spectrum S of the final output signal o (l, k):
[0111] S o (l, k) = G(l, k) · S e (l, k)
[0112] where G(l, k) is the gain function and P e (l, k) is the energy spectrum of the residual signal;
[0113] Perform inverse short-time Fourier transform (ISTFT) on the complex spectrum S of the final output signal o (l, k) to obtain the data block out of the AEC output signal in the time domain l (n).
[0114] As a further solution, the gain function G(l, k) is obtained through a Wiener filter and is subject to a maximum suppression amount constraint:
[0115] G(l, k) = max(G wiener (l,), min_G(l, k))
[0116] min_G(l,k) = linear(max_attenu(l,k))
[0117] where G wiener (l,k) is the Wiener filter output corresponding to the k-th subband of the l-th data block, max_attenu(l,k) represents the maximum suppression amount of the k-th subband of the l-th data block for the residual echo, and linear() is a linear function;
[0118] The maximum suppression amount max_attenu(l,k) is set according to the double-talk detection status result:
[0119] According to the human ear auditory masking characteristic, when two sounds with the same frequency appear simultaneously, the sound with larger energy will mask the sound with smaller energy, so the sound perceived by the human ear will be the sound with larger energy. Accordingly, when db_flag(l - 1) == doble_talk, the maximum suppression amount max_attenu(l,k) is:
[0120] max_attenu(l,k) = db(min_gain(l,k)) - 3
[0121]
[0122] where db() is a function that converts linear gain to db value, ε is a regularization factor, P res (l,k) is the energy spectrum of the residual echo, P e (l,k) is the energy spectrum of the residual signal.
[0123] When db_lag(l - 1) == far_talk_only, that is, when only the far-end human voice is speaking alone, it is necessary to increase the suppression of the residual echo, completely remove the residual echo, and obtain a better user experience. Therefore, a larger suppression amount needs to be set. The maximum suppression amount max_attenu(l,k) is set to a preset value, and the preset value is used to increase the suppression. This suppression amount is larger than the suppression amount when db_lag(l - 1) == doble_talk, and its specific value is an empirical value and is preset.
[0124] When db_flag(l - 1) == near_talk_only, no residual echo suppression is performed, and max_attenu(l,k) is set to 0.
[0125] It should be noted that: using the above maximum echo suppression db amount can effectively prevent damage to the proximal speech, and at the same time retain the effective suppression function for the residual echo.
[0126] As a further solution, such as Figure 5As shown, the recursive neural network (RNN) learns by extracting MFCC features, and uses the Dense layer and GRU layer to estimate the probability of the presence of human voice, and finally outputs the probability of the presence of human voice.
[0127] As a further solution, as Figure 6 shown, the weight coefficients in the recursive neural network (RNN) are obtained through preprocessing of training data, and the training data is obtained through the following steps: pre-record proximal human voice speech data, use far-end human voice single-talk data for AEC processing to obtain AEC residual data and environmental noise data; mix the proximal human voice speech data with the AEC residual data and environmental noise data respectively to obtain noisy speech signals; label according to the proximal human voice speech data to obtain the position label of the presence of human voice in the noisy speech signal; use the noisy speech signal and the position label of the presence of human voice as training samples, and pre-train the recursive neural network (RNN) to obtain the weight coefficients in it.
[0128] As a further solution, the double-talk detection status result db_flag(l) is delayed and output through the delay unit Z -1 The delay unit Z -1 delays the data by one unit time length, so that the linear adaptive filtering process and the non-linear post-processing obtain the double-talk detection status result db_flag(l-1) of the previous frame.
[0129] An echo cancellation system based on neural network double-talk detection runs on a hardware device, and the hardware device includes a signal collector, an input buffer, a linear adaptive filtering module, a non-linear post-processing module, an RNN double-talk detection module and a delay module; and the echo signal is cancelled by using an echo cancellation method based on neural network double-talk detection described in any one of the above.
[0130] The above are only embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be included in the patent protection scope of the present invention by the same token.
Claims
1. An echo cancellation method based on neural network dual-talker detection, characterized in that, It includes input signal processing, linear adaptive filtering processing, nonlinear post-processing, RNN double-talk detection, and output signal processing; Input signal processing: Collect the microphone signal at the near end and the reference signal at the far end, and transmit them in the form of a digital signal stream; store the microphone signal and the reference signal into the input buffer respectively, and the input buffer divides the signal into several data blocks, and the data block includes the microphone signal data block d l (n) and the reference signal data block x l (n); where l = 1, 2, 3,... represents the data block serial number, n = 0, 1, 2,..., n represents the sample serial number of each data block, and N is the total number of sample points of each data block; RNN double-talk detection: The voice activity detection is performed on the microphone signal and the reference signal through a Recurrent Neural Network (RNN), and the double-talk detection status result db_flag(l) is given by controlling with a finite state machine. Among them, the double-talk detection status result db_flag(l) includes: the far-end single-talk state far_talk_only with only far-end voice activity, the near-end single-talk state near_talk_only with only near-end voice activity, and the double-end double-talk state doble_talk with both far-end and near-end voice activities; the double-talk detection status result db_flag(l) is used for feedback adjustment of the linear adaptive filtering processing and the nonlinear post-processing; Linear adaptive filtering process: Receive the microphone signal data block d l (n), the reference signal data block x l (n) and perform point-by-point data processing; the data processing is carried out by the NLMS algorithm and adaptively adjusted by the double-talk detection status result db_flag(l) to obtain the adaptively updated filter coefficients Through the filter coefficients Calculate the adaptively updated residual signal e l (n); Nonlinear post-processing: used to further remove the residual signal e l (n) for the linear residual echo and the nonlinear residual echo signals in it; obtain the AEC output signal data block out l (n); Output signal processing: Store the data block out of the AEC output signal after echo removal l (n) into the output buffer, and perform data merging to obtain a continuous audio data stream for output; The nonlinear post-processing is carried out through the following steps: The residual signal e of data block l is processed in the frequency domain through short-time Fourier transform (STFT) l (n) and the estimated echo signal to obtain the complex spectrum S of the residual signal in the frequency-domain subband e (l, k), the complex spectrum of the echo signal the energy spectrum P of the residual signal e (l, k) and the echo signal where k represents the serial number of the discrete sampling points in the frequency domain, k = 0, 1,..., N B - 1; N B is the total number of frequency-domain subbands; The residual signal e is obtained through the following formula l The energy spectrum P of the residual echo in res (l, k): Among them, and is the correlation function value calculated for the last sample point N of the previous data block l-1; The complex spectrum S is weighted by the gain function G(l, k) e to obtain the complex spectrum S of the final output signal o (l, k): S o (l, k) = G(l, k) · S e (l, k) where G(l, k) is the gain function and P e (l, k) is the residual signal energy spectrum; The complex spectrum S of the final output signal o (l, k) is subjected to the inverse short-time Fourier transform (ISTFT) to obtain the AEC output signal data block out l (n) in the time domain.
2. The echo cancellation method based on neural network dual-speaker detection according to claim 1, wherein The linear adaptive filtering processing is carried out through the following steps: Receive the microphone signal data block d l (n) and the reference signal data block x l (n), and initialize the filter coefficient vector Among them, represents the reference signal vector at the n-th point of the l-th reference signal data block; represents the current x l (n) corresponding filter coefficient vector; T represents the transpose of the current vector, L is the filter length set during initialization, The initial values of are all set to 0; Filter coefficients of the n-th point in the previous frame Estimate the echo signal of the n-th point in the current frame Based on the estimated echo signal Calculate the estimated residual signal e at the n-th point of the current frame l (n): Calculate the reference signal energy E l,x(n): E l,x(n)=xl (n) T x l (n) Calculate the variable step size factor μ l (n); Update the estimated echo signal The autocorrelation function and cross-correlation function of the estimated residual signal e l (n): where r dd (n) is the echo signal autocorrelation function; r de (n) is the echo signal cross-correlation function with the residual signal e l (n), where α is the forgetting factor, r dd (n) and r de (n) are initialized to 0; Perform RNN double-talk detection to obtain the double-talk detection status result db_flag(l); According to the double-talk detection status result db_flag(l) and the filter coefficients of the n-th point in the previous frame Variable step size factor μ l (n), reference signal vector x l (n), residual signal e l (n) and reference signal energy E l,x(n)对滤波器系数 Perform adaptive update of the filter coefficients; With the filter coefficients updated adaptively Calculate the residual signal e l (n); Take the adaptively updated residual signal e l (n) as the output result of the linear adaptive filtering. Process each point in the data block point by point through the above steps to obtain the residual signal output of the data block: [e l (n), n = 0, 1, 2,..., N].
3. The echo cancellation method based on neural network double-talk detection according to claim 2, wherein The variable step size factor μ l (n) is calculated by the following formula: where ε is the regularization factor; μ0 is the maximum adaptive step size constant; r dd (n) is the autocorrelation function of the echo signal ; r de (n) is the cross-correlation function of the echo signal and the residual signal e l (n).
4. An echo cancellation method based on neural network double-talk detection according to claim 3, characterized in that, The adaptive update of the filter coefficients is calculated by the following formula: where db_flag(l - 1) represents the double-talk detection status result corresponding to the double-talk detection status information of the previous frame given by the RNN double-talk detection module; far_talk_only represents only far-end voice activity; else represents stopping the update of the filter when the double-talk detection result is not far_talk_only; ε is the regularization factor.
5. The echo cancellation method based on neural network double-talk detection according to claim 1, characterized in that, The gain function G(l, k) is obtained by a Wiener filter with a maximum attenuation constraint: G(l, k) = max(G wiener (l, k), min_G(l, k)) min_G(l, k) = linear(max_attenu(l, k)) Among them, G wiener (l, k) is the Wiener filter output corresponding to the k-th subband of the l-th data block, max_attenu(l, k) represents the maximum suppression amount of the k-th subband of the l-th data block for the residual echo, and linear() is a linear function; The maximum attenuation max_attenu(l, k) is set according to the double-talk detection status result: When db_flag(l - 1) == doble_talk, the maximum attenuation max_attenu(l, k) is: max_attenu(l, k) = db(min_gain(l, k)) - 3 where db() is a function for converting linear gain to db value, ε is a regularization factor, and P res (l, k) is the energy spectrum of the residual echo, and P e (l, k) is the energy spectrum of the residual signal; when db_flag(l - 1) == far_talk_only, the maximum suppression amount max_attenu(l, k) is set to a preset value, and the preset value is used to increase suppression: When db_flag(l - 1) == near_talk_only, no residual echo suppression is performed, and max_attenu(l, k) is set to 0.
6. The echo cancellation method based on neural network double-talk detection according to claim 1, characterized in that The Recurrent Neural Network (RNN) learns by extracting MFCC features, and uses Dense layers and GRU layers to implement the voice presence probability estimation, and finally outputs the voice presence probability.
7. An echo cancellation method based on neural network dual-speaker detection according to claim 6, characterized in that, The weight coefficients in the recursive neural network (RNN) are obtained through preprocessing of training data, and the training data is obtained through the following steps: pre-record proximal human voice data, use distal single-talker human voice data for acoustic echo cancellation (AEC) processing to obtain AEC residual data and ambient noise data; mix the proximal human voice data with the AEC residual data and the ambient noise data respectively to obtain noisy speech signals; label the proximal human voice data to obtain the vocal presence position labels in the noisy speech signals; use the noisy speech signals and the vocal presence position labels as training samples, and pre-train the RNN to obtain the weight coefficients therein.
8. The echo cancellation method based on neural network dual-speaker detection according to claim 1, wherein The two-talk detection status result db_flag(l) passes through the delay unit Z -1 for delayed output. The delay unit Z -1 delays the data by one unit duration, so that the linear adaptive filtering process and the non-linear post-processing obtain the two-talk detection status result db_flag(l-1) of the previous frame.
9. An echo cancellation system based on neural network dual-talker detection, characterized in that, It runs on a hardware device, and the hardware device includes a signal collector, an input buffer, a linear adaptive filtering module, a nonlinear post-processing module, an RNN double-talk detection module, and a delay module; and the cancellation of the echo signal is implemented by using an echo cancellation method based on neural network double-talk detection according to any one of claims 1 to 8.
Citation Information
Patent Citations
Echo cancellation processing method and processing system
CN110838300A