Echo cancellation method, echo cancellation apparatus, electronic device, and storage medium
By estimating and suppressing linear and reverberant echo components in the microphone signal, and utilizing Kalman and Wiener filtering algorithms, the problem of echo affecting speech clarity is solved, achieving fast and accurate echo cancellation and improving call quality.
Patent Information
- Application Number
- CN202311101568.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-29
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2043-08-29
AI Technical Summary
In voice calls, echoes reduce speech clarity and affect call quality. Existing technologies struggle to effectively eliminate echoes, especially reverberant echoes.
By estimating the linear echo component and reverberant echo component in the microphone signal, residual echoes are detected using the Kalman filter algorithm and spectral subtraction, and suppression parameters are determined by combining the Wiener filter algorithm, thus achieving the suppression of echo signals.
With limited computing resources, it can quickly and accurately estimate and eliminate reverberation echoes, improving call quality and speech clarity.
Smart Images

Figure CN119541518B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of communication, and particularly relates to an echo cancellation method, an echo cancellation device, an electronic device and a storage medium. BACKGROUND
[0002] In daily voice communication, voice intelligibility is an important index. For a communication device with a microphone and a loudspeaker, the microphone collects the voice of a near-end user, and the signal of a far-end user is also collected by the microphone after being played by the loudspeaker, thus forming an echo. If the echo is not processed, the signal collected by the microphone will be transmitted to the far end, and the far-end user will hear his / her own echo, which reduces the voice intelligibility and affects the voice communication quality. SUMMARY
[0003] To overcome the problems in the related art, the present disclosure provides an echo cancellation method, an echo cancellation device, an electronic device and a storage medium.
[0004] According to a first aspect of the embodiments of the present disclosure, an echo cancellation method is provided, applied to a terminal with a loudspeaker and a microphone, and the method comprises the following steps.
[0005] According to a loudspeaker signal output by the loudspeaker, a linear echo component in a microphone signal collected by the microphone is estimated to obtain a linear echo signal;
[0006] According to the linear echo signal, the linear echo component in the microphone signal is cancelled to obtain a linear echo cancellation signal;
[0007] According to the spectrum of the microphone signal and the spectrum of the linear echo signal, a residual echo component in the linear echo cancellation signal is estimated, and a first suppression parameter is determined, wherein the first suppression parameter is used to suppress the residual echo component in the linear echo cancellation signal;
[0008] According to the linear echo signal of the previous N frames, a reverberation echo component in the linear echo cancellation signal is estimated to obtain a reverberation echo signal, wherein N is a positive integer;
[0009] According to the linear echo cancellation signal and the reverberation echo signal, a second suppression parameter is determined, wherein the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal;
[0010] According to the first suppression parameter and the second suppression parameter, the linear echo cancellation signal is processed, and an echo cancellation signal is output.
[0011] In some embodiments, the estimating, according to the loudspeaker signal output by the loudspeaker, a linear echo component in the microphone signal collected by the microphone to obtain a linear echo signal comprises:
[0012] detecting whether the loudspeaker signal comprises a speech component;
[0013] in response to detecting that the loudspeaker signal does not comprise a speech component, estimating the linear echo component in the microphone signal by using a Kalman filtering algorithm to obtain the linear echo signal;
[0014] in response to detecting that the loudspeaker signal comprises a speech component, updating a Kalman coefficient and estimating the linear echo component in the microphone signal by using the Kalman filtering algorithm to obtain the linear echo signal.
[0015] In some embodiments, the estimating a residual echo component in the linear echo cancellation signal according to the spectrum of the microphone signal and the spectrum of the linear echo signal comprises:
[0016] detecting the residual echo component in the linear echo cancellation signal by using spectral subtraction according to the spectrum of the microphone signal and the spectrum of the linear echo signal.
[0017] In some embodiments, before estimating a reverberation echo component in the linear echo cancellation signal to obtain a reverberation echo signal, the echo cancellation method further comprises:
[0018] calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter;
[0019] the reverberation echo estimation according to the linear echo signal of the previous N frames to obtain a reverberation echo signal comprises:
[0020] in response to the value of the far-end single-talk flag being greater than or equal to a preset first threshold, determining an echo attenuation coefficient corresponding to each frame in the previous N frames according to the linear echo cancellation signal and the linear echo signal of the previous N frames;
[0021] determining a reverberation attenuation coefficient corresponding to each frame in the previous N frames according to the echo attenuation coefficient and the far-end single-talk flag;
[0022] estimating a reverberation echo component in the linear echo cancellation signal according to the reverberation attenuation coefficient and the linear echo signal of the previous N frames to obtain the reverberation echo signal.
[0023] In some embodiments, the determining a reverberation attenuation coefficient corresponding to each frame in the previous N frames according to the echo attenuation coefficient and the far-end single-talk flag comprises:
[0024] For each frame in the first N frames:
[0025] determine a smoothing factor according to the far-end single-talk flag, the echo decay coefficient corresponding to the frame, and the initial reverb decay coefficient corresponding to the frame;
[0026] smooth the initial reverb decay coefficient according to the smoothing factor and the echo decay coefficient corresponding to the frame, to obtain the reverb decay coefficient corresponding to the frame.
[0027] In some embodiments, before determining the second suppression parameter according to the linear echo-canceled signal and the reverb echo signal, the echo cancellation method further comprises:
[0028] calculate a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter;
[0029] in response to a value of the far-end single-talk flag being greater than or equal to a preset second threshold, estimate a nonlinear echo component in the linear echo-canceled signal according to the far-end single-talk flag, a nonlinear distortion signal corresponding to the linear echo-canceled signal, and the loudspeaker signal, to obtain a nonlinear echo signal, wherein the nonlinear distortion signal is obtained by full-wave rectification of the loudspeaker signal;
[0030] the determining the second suppression parameter according to the linear echo-canceled signal and the reverb echo signal comprises:
[0031] determining the second suppression parameter according to the linear echo-canceled signal, the nonlinear echo signal, and the reverb echo signal.
[0032] In some embodiments, the estimating the nonlinear echo component in the linear echo-canceled signal according to the far-end single-talk flag, the linear echo-canceled signal, and the nonlinear distortion signal corresponding to the loudspeaker signal, to obtain the nonlinear echo signal, comprises:
[0033] calculating a ratio of an amplitude of the linear echo-canceled signal to an amplitude of the nonlinear distortion signal, and transforming the ratio to a logarithmic domain to obtain a logarithmic domain ratio;
[0034] determining a smoothing coefficient according to the far-end single-talk flag;
[0035] smoothing the logarithmic domain ratio according to the smoothing coefficient;
[0036] transforming the smoothed logarithmic domain ratio to an amplitude domain to obtain an amplitude domain ratio;
[0037] obtaining the nonlinear echo signal according to the amplitude domain ratio and the nonlinear distortion signal.
[0038] In some embodiments, the determining the second suppression parameter according to the linear echo cancellation signal, the non-linear echo signal and the reverberation echo signal comprises:
[0039] calculating a post-signal-to-noise ratio according to an energy value of the linear echo cancellation signal, an energy value of the reverberation echo signal and an energy value of the non-linear echo signal;
[0040] determining a prior-signal-to-noise ratio according to the post-signal-to-noise ratio;
[0041] smoothing the prior-signal-to-noise ratio according to the smoothing coefficient, the second suppression parameter corresponding to a previous frame and the post-signal-to-noise ratio corresponding to the previous frame;
[0042] determining the second suppression parameter according to the smoothed prior-signal-to-noise ratio.
[0043] In some embodiments, before determining the second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal, the echo cancellation method comprises:
[0044] calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter;
[0045] determining a smoothing coefficient according to the far-end single-talk flag;
[0046] the determining the second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal comprises:
[0047] calculating a post-signal-to-noise ratio according to an energy value of the linear echo cancellation signal and an energy value of the reverberation echo signal;
[0048] determining a prior-signal-to-noise ratio according to the post-signal-to-noise ratio;
[0049] smoothing the prior-signal-to-noise ratio according to the smoothing coefficient, the second suppression parameter corresponding to a previous frame and the post-signal-to-noise ratio corresponding to the previous frame;
[0050] determining the second suppression parameter according to the smoothed prior-signal-to-noise ratio.
[0051] According to a second aspect of the embodiments of the present disclosure, an echo cancellation device is provided, comprising:
[0052] a linear echo estimation module configured to estimate a linear echo component in a microphone signal collected by a microphone according to a loudspeaker signal output by a loudspeaker to obtain a linear echo signal;
[0053] a linear echo cancellation module configured to cancel the linear echo component in the microphone signal according to the linear echo signal to obtain a linear echo cancellation signal;
[0054] a linear residual echo suppression module configured to estimate a residual echo component in the linear echo cancellation signal according to a spectrum of the microphone signal and a spectrum of the linear echo signal, and determine a first suppression parameter, wherein the first suppression parameter is used to suppress the residual echo component in the linear echo cancellation signal;
[0055] a reverberation echo estimation module configured to estimate a reverberation echo component in the linear echo cancellation signal according to the linear echo signal of previous N frames, to obtain a reverberation echo signal, wherein N is a positive integer;
[0056] an echo suppression module configured to determine a second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal, wherein the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal;
[0057] an echo cancellation module configured to process the linear echo cancellation signal according to the first suppression parameter and the second suppression parameter, to output an echo cancellation signal.
[0058] According to a third aspect of embodiments of the present disclosure, an electronic device is provided, comprising:
[0059] a speaker, a microphone, a processor, and a memory;
[0060] The speaker, the microphone, and the memory are all in communication connection with the processor;
[0061] The memory is configured to store executable instructions;
[0062] The processor is configured to run the executable instructions to implement the steps of the echo cancellation method provided in any of the embodiments of the first aspect of the present disclosure.
[0063] According to a fourth aspect of embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions, and the program instructions are executed by a processor to implement the steps of the echo cancellation method provided in any of the embodiments of the first aspect of the present disclosure.
[0064] The technical scheme provided by the embodiments of the present disclosure can have the following beneficial effects: according to a loudspeaker signal output by a loudspeaker, a linear echo component in a microphone signal collected by a microphone is estimated to obtain a linear echo signal; according to the linear echo signal, the linear echo component in the microphone signal is eliminated to obtain a linear echo cancellation signal; a residual echo component in the linear echo cancellation signal is estimated according to a spectrum of the microphone signal and a spectrum of the linear echo signal, and a first suppression parameter is determined, wherein the first suppression parameter is used to suppress the residual echo component in the linear echo cancellation signal; a reverberation echo component in the linear echo cancellation signal is estimated according to linear echo signals of the previous N frames to obtain a reverberation echo signal, wherein N is a positive integer; a second suppression parameter is determined according to the linear echo cancellation signal and the reverberation echo signal, wherein the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal; and the linear echo cancellation signal is processed according to the first suppression parameter and the second suppression parameter to output an echo cancellation signal. In this way, on the basis of linear echo cancellation, the above-mentioned reverberation echo estimation method can achieve fast and accurate reverberation echo estimation with less occupied computing resources, so as to process a call signal and improve call effect.
[0065] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and are not limiting to the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0066] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure.
[0067] Figure 1 is a flowchart of an echo cancellation method according to an exemplary embodiment.
[0068] Figure 2 is a schematic diagram of a call scene according to an exemplary embodiment.
[0069] Figure 3 is a flowchart of an implementation method of step S110 in the embodiments of the present disclosure.
[0070] Figure 4 is a flowchart of an implementation method of step S130 in the embodiments of the present disclosure.
[0071] Figure 5 is a flowchart of an implementation method of step S150 in the embodiments of the present disclosure.
[0072] Figure 6 is a flowchart of an implementation method of step S152 in the embodiments of the present disclosure.
[0073] Figure 7An implementation method flow chart for step S160 in the embodiments of the present disclosure.
[0074] Figure 8 A flow chart of another echo cancellation method according to an exemplary embodiment is shown.
[0075] Figure 9 An implementation method flow chart for step S241 in the embodiments of the present disclosure.
[0076] Figure 10 An implementation method flow chart for step S265 in the embodiments of the present disclosure.
[0077] Figure 11 A flow chart of yet another echo cancellation method according to an exemplary embodiment is shown.
[0078] Figure 12 A block diagram of an echo cancellation apparatus according to an exemplary embodiment is shown.
[0079] Figure 13 A block diagram of an electronic device 800 according to an exemplary embodiment is shown. DETAILED DESCRIPTION
[0080] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is only exemplary and is not intended to limit the scope, applicability or configuration of the disclosure. Rather, the disclosure is applicable as claimed and drawn to a variety of electronic devices and methods consistent with the following claims.
[0081] It should be noted that all the actions of acquiring signals, information or data in the present disclosure are performed in compliance with the corresponding data protection regulations and policies of the country where the device is located, and with the authorization given by the owner of the corresponding device.
[0082] Figure 1 A flow chart of an echo cancellation method according to an exemplary embodiment is shown. As shown in Figure 1 the method can be applied to a terminal with a speaker and a microphone, and the method comprises the following steps.
[0083] Step S110, estimating a linear echo component in the input microphone signal collected by the microphone according to the speaker signal output by the speaker to obtain a linear echo signal.
[0084] The loudspeaker signal output by the loudspeaker is a far-end signal, which can also be referred to as a loudspeaker reference (REF) signal, and can be played through the loudspeaker. The microphone signal is a near-end signal, which can be a signal collected by the microphone.
[0085] The linear echo component in the microphone signal is estimated according to the far-end signal to obtain a linear echo signal.
[0086] In step S120, the linear echo component in the microphone signal is removed according to the linear echo signal to obtain a linear echo removed signal.
[0087] In some embodiments, the linear echo removed signal aec_out1 is obtained by the following formula
[0088]
[0089] The linear echo removed signal aec_out1 is obtained. Wherein, MIC is the microphone signal, and echo is the linear echo signal.
[0090] In step S130, a residual echo component in the linear echo removed signal is estimated according to the spectrum of the microphone signal and the spectrum of the linear echo signal, and a first suppression parameter is determined.
[0091] The first suppression parameter is used to suppress the residual echo component in the linear echo removed signal.
[0092] Wherein, when removing the estimated linear echo component, the residual echo component can be detected according to the spectrum of the microphone signal and the spectrum of the linear echo signal, and the first suppression parameter is determined according to the detected residual echo component.
[0093] In some embodiments, the determined first suppression parameter includes a suppression value corresponding to each frequency point. For example, for the two possible values 0 and 1 of the suppression value, a value of 1 indicates that the frequency point is not suppressed, and a value of 0 indicates that the frequency point is completely suppressed.
[0094] In step S150, a reverberation echo component in the linear echo removed signal is estimated according to the linear echo signal of the previous N frames to obtain a reverberation echo signal.
[0095] Wherein, N is a positive integer.
[0096] Wherein, the first N frames are the first N frames of the current frame, the reverberation echo estimation is performed on the current frame according to the linear echo signals of the first N frames, and the reverberation echo signal of the current frame is obtained. It can be understood that the first N frames include N frames, and the above-mentioned loudspeaker signal, microphone signal, first suppression parameter and the like can be described as the loudspeaker signal of the current frame, the microphone signal of the current frame, the first suppression parameter of the current frame, and so on. In addition to the signals and parameters that have been explicitly limited to the first N frames or the previous frame, the signals and parameters in the embodiments of the present disclosure are understood as the signals and parameters of the current frame.
[0097] In some embodiments, N is the frame number of the loudspeaker signal after the longest path reflection in the current environment after playing; in some embodiments, N is a preset value calculated according to prior experience and / or algorithm model; in some embodiments, before the reverberation echo estimation is performed on the linear echo signals of the first N frames, it further includes: analyzing the microphone signal or the linear echo signal to determine the value of N.
[0098] Step S160, determining the second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal.
[0099] Wherein, the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal.
[0100] In some embodiments, the second suppression parameter is determined by Wiener filtering algorithm according to the linear echo cancellation signal and the reverberation echo signal.
[0101] Step S170, processing the linear echo cancellation signal according to the first suppression parameter and the second suppression parameter, and outputting the echo cancellation signal.
[0102] Wherein, the linear echo cancellation signal is processed according to the first suppression parameter and the second suppression parameter to eliminate the residual linear echo and reverberation echo in the linear echo cancellation signal, and the echo cancellation signal is output.
[0103] In some embodiments, the echo cancellation signal aec_out is obtained by the following formula:
[0104] aec_out=aec_out1*H(w)*H2
[0105] Wherein, aec_out1 is the linear echo cancellation signal, H(w) is the first suppression parameter, and H2 is the second suppression parameter.
[0106] Figure 2 Fig. 1 is a schematic diagram of a call scenario according to an exemplary embodiment. As shown in Fig. 1, the call scenario includes a loudspeaker 101, a microphone 102, a first suppression parameter H(w) and a second suppression parameter H2. Figure 2As shown, the far-end device K and the near-end device L intercommunicate the phone, and the voice signal A of the first user of the far-end device K is collected and processed by the microphone of the far-end device K, and then transmitted to the near-end device L, and played through the loudspeaker of the near-end device L. The arrow in the figure represents the signal propagation direction.
[0107] As shown, the near-end device L is a wearable device such as a smart watch with less computing resources, and performs duplex voice communication with the far-end device K in a strong reverberation environment such as an elevator or a conference room. Figure 2 As shown, the near-end device L is a wearable device such as a smart watch with less computing resources, and performs duplex voice communication with the far-end device K in a strong reverberation environment such as an elevator or a conference room.
[0108] In the related art, the microphone collected signal of the near-end device L is usually processed by using an adaptive filtering algorithm. For example, with a sampling rate of 16000 and a frame shift of 256, an adaptive filter with an order of 256 is used. As shown, it can remove the direct echo A1 and the reverberation echo A2 within 16 ms, but for the late reverberation echo A3 exceeding 16 ms, it will be transmitted to the far-end device K together with the voice signal B, and the first user at the far end will hear the voice content of himself due to the existence of the reverberation echo A3. Figure 2
[0109] The method comprises the following steps.
[0110] Figure 3 An embodiment of the present disclosure provides a method for echo cancellation, which comprises the following steps. Figure 3
[0111] Step S111, detecting whether the loudspeaker signal comprises a speech component.
[0112] In some embodiments, the loudspeaker signal is subjected to voice activity detection (VAD) to detect whether the loudspeaker signal comprises a speech component.
[0113] For example, in response to detecting that the loudspeaker signal does not comprise a speech component, the attribute value VAD in the detection result is set to 0, and in response to detecting that the loudspeaker signal comprises a speech component, the attribute value VAD in the detection result is set to 1.
[0114] Voice activity detection is a technology for speech processing, which is used to detect whether a speech signal exists.
[0115] Step S112, in response to detecting that the loudspeaker signal does not comprise a speech component, estimating the linear echo component in the microphone signal by using a Kalman filtering algorithm to obtain the linear echo signal.
[0116] Step S113, in response to detecting that the loudspeaker signal includes speech components, updating the Kalman coefficient and estimating the linear echo components in the microphone signal by using the Kalman filtering algorithm to obtain the linear echo signal.
[0117] Wherein, in response to detecting that the loudspeaker signal includes speech components, then updating the Kalman coefficient, estimating the echo path, and then filtering to estimate the linear echo in the microphone signal, in response to detecting that the loudspeaker signal does not include speech components, then directly filtering to estimate the linear echo in the microphone signal.
[0118] Wherein, the Kalman filtering algorithm is an algorithm for optimal estimation of system state by using linear system state equation and observing data of system input and output, and since the observation data includes the influence of noise and interference in the system, the optimal estimation can also be regarded as a filtering process.
[0119] Thus, the linear echo estimation is realized by the Kalman filtering algorithm.
[0120] Figure 4 An implementation method flow chart for step S130 in the embodiments of the present disclosure. As shown in the figure, in step S130, the residual echo components in the linear echo cancellation signal are estimated according to the spectrum of the microphone signal and the spectrum of the linear echo signal, including the following steps. Figure 4
[0121] Step S131, the residual echo components in the linear echo cancellation signal are detected by using the spectrum subtraction according to the spectrum of the microphone signal and the spectrum of the linear echo signal.
[0122] Wherein, the purpose of the spectrum subtraction is to subtract the spectrum of the noise signal from the spectrum of the noisy signal to obtain the pure signal, which can also be realized by the subtraction of the amplitude spectrum and the power spectrum; and correspondingly, the residual echo components in the linear echo cancellation signal can also be detected by using the spectrum subtraction, and the first suppression parameter is determined accordingly.
[0123] In some embodiments, the first suppression parameter H(w) is obtained by the following formula:
[0124]
[0125] The first suppression parameter H(w) is obtained. Wherein, D(w) is the frequency domain signal corresponding to the linear echo signal, and Y(w) is the frequency domain signal corresponding to the microphone signal.
[0126] Thus, the linear residual echo estimation is realized by the spectrum subtraction.
[0127] Figure 5 Figure 1 shows an embodiment of a method flowchart for step S150 in the embodiments of the present disclosure. In step S150, the reverberation echo component in the linear echo cancellation signal is estimated to obtain the reverberation echo signal. Before step S150, the echo cancellation method further comprises: calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter; and Figure 5 As shown in Figure 1, step S150, the reverberation echo estimation is performed on the linear echo signals of the previous N frames to obtain the reverberation echo signal, which comprises the following steps.
[0128] Step S151, in response to the value of the far-end single-talk flag being greater than or equal to a preset first threshold, determining the echo attenuation coefficient corresponding to each frame in the previous N frames according to the linear echo cancellation signal and the linear echo signals of the previous N frames.
[0129] In some embodiments, the echo attenuation coefficient corresponding to each frame in the previous N frames is determined according to the linear echo cancellation signal of the current frame and the linear echo signals of the previous N frames.
[0130] In the far-end single-talk speech state, the microphone signal is obtained by superimposing the reverberation attenuation of the far-end signal of the current frame and the signals collected by the microphone after the reverberation attenuation of the far-end signals of the previous N frames of the current frame, so that the echo attenuation of the previous N frames needs to be considered.
[0131] In some embodiments, the echo attenuation coefficient corresponding to the jth frame of the previous N frames of the current frame is obtained by the following formula:
[0132]
[0133] j wherein aec_out1 is the linear echo cancellation signal, echo_last j is the linear echo signal of the jth frame of the previous N frames of the current frame. The echo attenuation coefficient tmp j characterizes the degree of echo attenuation of the jth frame of the previous N frames of the current frame to the current frame.
[0134] In some embodiments, the first threshold can be 0.9.
[0135] In some embodiments, the far-end single-talk flag represents the probability of the call state being far-end single-talk, and the value of the far-end single-talk flag being greater than or equal to a preset first threshold indicates that the call state is detected as far-end single-talk; in some embodiments, the greater the value of the far-end single-talk flag, the greater the probability of the call state being far-end single-talk; in some embodiments, the value of the far-end single-talk flag ranges from 0 to 1.
[0136] In some embodiments, the echo attenuation coefficient corresponding to the jth frame of the previous N frames of the current frame is obtained by the following formula:
[0137]
[0138] We obtain the remote single-talk flag, pure_echo_flag, where energy ref This represents the energy of one frame in the time domain of the speaker signal REF, where thr is the corresponding energy threshold, and Hw is the energy value. i Let i be the first suppression parameter corresponding to frequency i, i∈[m,n]. For example, m takes the value of 5 and n takes the value of 100.
[0139] Step S152: Determine the reverberation attenuation coefficient for each frame in the first N frames based on the echo attenuation coefficient and the far-end single-talk flag.
[0140] Specifically, the reverberation attenuation coefficient for each of the first N frames is determined based on the far-end single-talk flag and the echo attenuation coefficient corresponding to each frame in the first N frames. Thus, while considering echo attenuation, the reverberation attenuation coefficient is determined based on the call state. For example, in the far-end single-talk state, the reverberation attenuation coefficient is estimated based on the echo attenuation coefficient; in the near-end single-talk and near-end / far-end dual-talk states, the reverberation attenuation coefficient remains unchanged, avoiding interference from near-end speech on echo estimation.
[0141] Figure 6 This is a flowchart illustrating one implementation of step S152 in this disclosure. Figure 6 As shown, in step S152, in the step of determining the reverberation attenuation coefficient corresponding to each frame in the first N frames based on the echo attenuation coefficient and the far-end single-talk flag, the following steps are performed for each frame in the first N frames.
[0142] Step S1521: Determine the smoothing factor based on the remote single-talk flag, the echo attenuation coefficient corresponding to the frame, and the initial reverberation attenuation coefficient corresponding to the frame.
[0143] Step S1522: Smooth the initial reverberation attenuation coefficient according to the smoothing factor and the echo attenuation coefficient corresponding to the frame to obtain the reverberation attenuation coefficient corresponding to the frame.
[0144] The initial reverberation attenuation coefficient is the reverberation attenuation coefficient or the initial value of the reverberation attenuation coefficient previously determined for this frame.
[0145] In some embodiments, the following formula is used:
[0146] w_s j =w_s j *b+tmp j *(1-b)
[0147]
[0148] Obtain the reverberation attenuation coefficient w_s corresponding to the j-th frame preceding the current frame. j Where b is the smoothing factor, tmpj is the echo attenuation coefficient corresponding to the jth frame before the current frame, and pure_echo_flag is a far-end single-talk flag.
[0149] Thus, for each frame in the N frames before the current frame, the reverberation attenuation coefficient corresponding to the frame is determined.
[0150] Step S153, reverberation echo components in the linear echo cancellation signal are estimated according to the reverberation attenuation coefficient and the linear echo signals of the N frames before the current frame, to obtain a reverberation echo signal.
[0151] The reverberation echo signal is obtained by estimating the reverberation echo according to the reverberation attenuation coefficient corresponding to each frame in the N frames before the current frame and the linear echo signals of the N frames before the current frame.
[0152] In some embodiments, the reverberation echo signal reverb_echo is obtained by the following formula:
[0153]
[0154] The reverberation echo signal reverb_echo is obtained by the following formula: j is the echo attenuation coefficient corresponding to the jth frame before the current frame, and pure_echo_flag is a far-end single-talk flag. j is the linear echo signal of the jth frame before the current frame.
[0155] Thus, by the above method, the reverberation echo is quickly and accurately estimated based on the characteristics of the reverberation echo.
[0156] Figure 7 is a flow chart of an implementation method of step S160 in the embodiments of the present disclosure. In step S160, before the second suppression parameter is determined according to the linear echo cancellation signal and the reverberation echo signal, the echo cancellation method further includes: calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter, and determining a smoothing coefficient according to the far-end single-talk flag; as shown in Figure 7 Step S160, determining the second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal, includes the following steps.
[0157] Step S161, calculating a posteriori signal-to-noise ratio according to the energy value of the linear echo cancellation signal and the energy value of the reverberation echo signal.
[0158] Wherein, the optional implementation and related nomenclature of the far-end single-talk flag and the smoothing coefficient can be referred to the optional implementation of the related steps and other related parts in the embodiments involved in the related steps, which will not be described here.
[0159] In some embodiments, the reverberation echo signal reverb_echo is obtained by the following formula:
[0160]
[0161] The postSNR is calculated, where aec_out1 is the linear echo-canceled signal, and reverb_echo is the reverberation echo signal.
[0162] In step S162, the priorSNR is determined according to the postSNR.
[0163] In some embodiments, the priorSNR is determined by the following equation:
[0164] priorSNR = max (postSNR, 0)
[0165] The priorSNR is determined, where postSNR is the postSNR, and max() represents the maximum of two values.
[0166] In step S163, the priorSNR is smoothed according to the smoothing factor, the second suppression parameter corresponding to the previous frame, and the postSNR corresponding to the previous frame.
[0167] In some embodiments, the priorSNR is smoothed by the following equation:
[0168] SNR' = priorSNR * (1 - a) + a * H2_last 2 * SNR_last
[0169] The smoothed priorSNR SNR' is obtained, where priorSNR is the priorSNR, a is the smoothing factor, H2_last is the second suppression parameter corresponding to the previous frame of the current frame, and SNR_last is the postSNR corresponding to the previous frame of the current frame.
[0170] In step S164, the second suppression parameter is determined according to the smoothed priorSNR.
[0171] In some embodiments, the second suppression parameter is determined by the following equation:
[0172]
[0173] The second suppression parameter H2 is determined, where SNR' is the smoothed priorSNR.
[0174] Figure 8 FIG. 2 is a flowchart illustrating another echo cancellation method according to an exemplary embodiment. As shown in FIG. 2, the method includes the following steps. Figure 8
[0175] In step S210, a linear echo estimation is performed on the input microphone signal according to the loudspeaker signal output by the loudspeaker, to obtain a linear echo signal.
[0176] Step S220, linear echo cancellation is performed on the microphone signal according to the linear echo signal, to obtain a linear echo cancelled signal.
[0177] Step S230, whether there is residual echo is detected according to the spectrum of the microphone signal and the spectrum of the linear echo signal, and a first suppression parameter is determined according to the detection result.
[0178] Step S240, a far-end single-talk flag is calculated according to the loudspeaker signal and the first suppression parameter.
[0179] Among them, the optional implementation of the far-end single-talk flag and the related nomenclature can be referred to the optional implementation of the aforementioned related steps and other related parts in the embodiments involved in the related steps, which will not be described here.
[0180] Step S241, in response to the value of the far-end single-talk flag being greater than or equal to a preset second threshold, a nonlinear echo component in the linear echo cancelled signal is estimated according to the far-end single-talk flag, the linear echo cancelled signal and a nonlinear distortion signal corresponding to the loudspeaker signal, to obtain a nonlinear echo signal.
[0181] Among them, the nonlinear distortion signal is obtained by full-wave rectification of the loudspeaker signal.
[0182] Among them, the nonlinear distortion signal is obtained by taking the absolute value of the loudspeaker signal in the time domain and then transforming it to the frequency domain.
[0183] Among them, due to the structure of the device and other reasons, it will cause signal distortion, and the corresponding signal component will cause nonlinear echo. The embodiments of the present disclosure introduce nonlinear signals in the form of nonlinear distortion signals, so that the nonlinear echo can be removed subsequently.
[0184] In some embodiments, the second threshold can be 0.7.
[0185] Step S250, reverberation echo estimation is performed according to the linear echo signals of the previous N frames, to obtain a reverberation echo signal.
[0186] Step S265, a second suppression parameter is determined according to the linear echo cancelled signal, the nonlinear echo signal and the reverberation echo signal.
[0187] Among them, the second suppression parameter is used to suppress the nonlinear echo component and the reverberation echo component in the linear echo cancelled signal.
[0188] Among them, after introducing the nonlinear distortion, the second suppression parameter is determined according to the linear echo cancelled signal, the nonlinear echo signal and the reverberation echo signal, to suppress the nonlinear echo and the reverberation echo.
[0189] Step S270, processing the linear echo cancellation signal according to the first suppression parameter and the second suppression parameter, and outputting an echo cancellation signal.
[0190] Thus, the nonlinear echo estimation is performed by introducing the nonlinear distortion signal, and based on the above reverberation echo estimation, the nonlinear echo and the reverberation echo in the microphone signal are processed on the basis of performing linear echo cancellation.
[0191] Figure 9 An implementation method flow chart for step S241 in the embodiments of the present disclosure is shown in FIG. 7. As shown in FIG. 7, in step S241, the step of estimating the nonlinear echo component in the linear echo cancellation signal according to the far-end single-talk flag, the linear echo cancellation signal and the nonlinear distortion signal corresponding to the loudspeaker signal to obtain the nonlinear echo signal, includes the following steps. Figure 9
[0192] Step S2411, calculating the ratio of the amplitude of the linear echo cancellation signal and the amplitude of the nonlinear distortion signal, and transforming the ratio to the logarithmic domain to obtain the logarithmic domain ratio.
[0193] In some embodiments, the ratio is calculated by the following formula:
[0194]
[0195] The logarithmic domain ratio Hn corresponding to the frequency point i is determined by the following formula: i wherein aec_out1 is the linear echo cancellation signal, ref_abs is the nonlinear distortion signal in the frequency domain, and abs() represents the amplitude of the corresponding frequency domain signal.
[0196] Step S2412, determining the smoothing coefficient according to the far-end single-talk flag.
[0197] In some embodiments, the greater the value of the far-end single-talk flag, the higher the probability that the call state is far-end single-talk, and the smaller the smoothing coefficient.
[0198] Step S2413, smoothing the logarithmic domain ratio according to the smoothing coefficient.
[0199] In some embodiments, the smoothing is performed by the following formula:
[0200] Hn_s = α * Hn_s + (1-α) * Hn
[0201]
[0202] The smoothed logarithmic domain ratio Hn_s is obtained, wherein α is the smoothing coefficient, Hn is the logarithmic domain ratio, and pure_echo_flag is the far-end single-talk flag.
[0203] Step S2414, transform the smoothed log-domain ratio to the amplitude domain to obtain an amplitude-domain ratio.
[0204] In some embodiments, the amplitude-domain ratio map_gain is obtained by the following formula:
[0205]
[0206] The amplitude-domain ratio map_gain is obtained, where Hn_s is the smoothed log-domain ratio.
[0207] Step S2415, obtain the nonlinear echo signal according to the amplitude-domain ratio and the nonlinear distortion signal.
[0208] In some embodiments, the nonlinear echo signal map_echo is obtained by the following formula:
[0209] map_echo = map_gain * ref_abs
[0210] The nonlinear echo signal map_echo is obtained, where map_gain is the amplitude-domain ratio and ref_abs is the nonlinear distortion signal.
[0211] Figure 10 is an embodiment method flowchart of step S265 in the embodiments of the present disclosure. As shown in Figure 10 Step S265, determining the second suppression parameter according to the linear echo cancellation signal, the nonlinear echo signal, and the reverberation echo signal, includes the following steps.
[0212] Step S2651, calculate the post SNR according to the energy value of the linear echo cancellation signal, the energy value of the reverberation echo signal, and the energy value of the nonlinear echo signal.
[0213] In some embodiments, the post SNR is calculated by the following formula:
[0214]
[0215] The post SNR is calculated, where aec_out1 is the linear echo cancellation signal, reverb_echo is the reverberation echo signal, and map_echo is the nonlinear echo signal.
[0216] Step S2652, determine the prior SNR according to the post SNR.
[0217] In some embodiments, the prior SNR is determined by the following formula:
[0218] priorSNR = max (postSNR, 0)
[0219] determine a prior signal-to-noise ratio (SNR) priorSNR, wherein the postSNR is a post signal-to-noise ratio, and max() represents a maximum of two values.
[0220] At step S2653, the prior SNR is smoothed according to the smoothing factor, the second suppression parameter corresponding to the previous frame, and the post SNR corresponding to the previous frame.
[0221] In some embodiments, the smoothing of the prior SNR is performed according to the following equation:
[0222] SNR' = priorSNR * (1 - a) + a * H2_last 2 * SNR_last
[0223] to obtain a smoothed prior SNR SNR', wherein the priorSNR is the prior SNR, a is the smoothing factor, H2_last is the second suppression parameter corresponding to the previous frame of the current frame, and SNR_last is the post SNR corresponding to the previous frame of the current frame.
[0224] At step S2654, the second suppression parameter is determined according to the smoothed prior SNR.
[0225] In some embodiments, the smoothing of the prior SNR is performed according to the following equation:
[0226]
[0227] to determine the second suppression parameter H2, wherein the SNR' is the smoothed prior SNR.
[0228] Thus, the linear echo cancellation signal is processed based on the determined second suppression parameter, and the suppression of the nonlinear echo and the reverberation echo can be simultaneously achieved.
[0229] The echo cancellation method provided by the present disclosure is described below in combination with actual applications.
[0230] Figure 11 is a flowchart of yet another echo cancellation method according to an exemplary embodiment. As shown in Figure 11 The method includes the following steps.
[0231] At step S310, it is detected whether the loudspeaker signal includes a speech component.
[0232] In some embodiments, the speech activity detection is performed on the loudspeaker signal to detect whether the loudspeaker signal includes a speech component.
[0233] In step S310, in response to detecting that the loudspeaker signal does not include a speech component, step S311 is performed, and in response to detecting that the loudspeaker signal includes a speech component, step S312 is performed.
[0234] Step S311, estimate the linear echo component in the microphone signal by using the Kalman filtering algorithm to obtain a linear echo signal.
[0235] Step S312, update the Kalman coefficient, and estimate the linear echo component in the microphone signal by using the Kalman filtering algorithm to obtain a linear echo signal.
[0236] Step S320, eliminate the linear echo component in the microphone signal according to the linear echo signal to obtain a linear echo cancellation signal.
[0237] wherein the linear echo cancellation signal aec_out1 is obtained by the following formula
[0238] aec_out1 = MIC - echo
[0239]
[0240] Step S330, determine the first suppression parameter corresponding to the residual echo component by using the spectral subtraction method according to the spectrum of the microphone signal and the spectrum of the linear echo signal.
[0241] wherein the first suppression parameter H(w) is obtained by the following formula:
[0242]
[0243]
[0244] Step S340, calculate a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter.
[0245] wherein the far-end single-talk flag pure_echo_flag is obtained by the following formula:
[0246]
[0247] ref wherein energy represents the energy of a time-domain frame of the loudspeaker signal REF, thr is a corresponding energy threshold, Hw i is the first suppression parameter corresponding to the frequency point i, i∈[m, n], wherein m is 5 and n is 100.
[0248] Step S341, in response to the value of the far-end single-talk flag being greater than or equal to a preset second threshold, calculate the ratio of the amplitude of the linear echo cancellation signal to the amplitude of the nonlinear distortion signal, and transform the ratio to the logarithmic domain to obtain a logarithmic domain ratio.
[0249] wherein the second threshold is 0.7.
[0250] wherein the following equation is used:
[0251]
[0252] determines the log-domain ratio value Hn corresponding to the frequency point i i wherein ref_abs is the non-linear distortion signal in the frequency domain, and abs() represents the amplitude of the corresponding frequency domain signal.
[0253] Step S342, determines the smoothing coefficient according to the far-end single-talk flag, and smoothes the log-domain ratio value according to the smoothing coefficient.
[0254] In some embodiments, the following equation is used:
[0255] Hn_s = a * Hn_s + (1-a) * Hn
[0256]
[0257] obtains the smoothed log-domain ratio value Hn_s, wherein a is the smoothing coefficient.
[0258] Step S343, transforms the smoothed log-domain ratio value to the amplitude domain to obtain an amplitude-domain ratio value.
[0259] wherein the following equation is used:
[0260]
[0261] obtains the amplitude-domain ratio value map_gain.
[0262] Step S344, obtains a non-linear echo signal according to the amplitude-domain ratio value and the non-linear distortion signal.
[0263] wherein the following equation is used:
[0264] map_echo = map_gain * ref_abs
[0265] obtains the non-linear echo signal map_echo.
[0266] Step S350, in response to the value of the far-end single-talk flag being greater than or equal to a preset first threshold, determines an echo attenuation coefficient corresponding to each frame in the previous N frames according to the linear echo cancellation signal and the linear echo signals of the previous N frames.
[0267] wherein the first threshold is 0.9.
[0268] wherein the following equation is used:
[0269]
[0270] obtain the echo attenuation coefficient tmp corresponding to the jth frame before the current frame j wherein echo_last j is the linear echo signal of the jth frame before the current frame.
[0271] Step S351, determine a smoothing factor according to the far-end single-talk flag, the echo attenuation coefficient corresponding to the frame, and the initial reverberation attenuation coefficient corresponding to the frame, and smooth the initial reverberation attenuation coefficient according to the smoothing factor and the echo attenuation coefficient corresponding to the frame to obtain the reverberation attenuation coefficient corresponding to the frame.
[0272] wherein the following formula is used:
[0273] w_s j = w_s j *b + tmp j *(1-b)
[0274]
[0275] obtain the reverberation attenuation coefficient w_s corresponding to the jth frame before the current frame j wherein b is a smoothing factor.
[0276] Step S352, estimate the reverberation echo component in the linear echo cancellation signal according to the reverberation attenuation coefficient and the linear echo signals of the previous N frames to obtain a reverberation echo signal.
[0277] wherein the following formula is used:
[0278]
[0279] obtain the reverberation echo signal reverb_echo.
[0280] Step S360, calculate the post SNR according to the energy value of the linear echo cancellation signal, the energy value of the reverberation echo signal, and the energy value of the nonlinear echo signal.
[0281] wherein the following formula is used:
[0282]
[0283] calculate the post SNR postSNR.
[0284] Step S361, determine the prior SNR according to the post SNR.
[0285] In some embodiments, the following formula is used:
[0286] priorSNR = max (postSNR, 0)
[0287] determine a prior SNR, wherein max() represents taking the maximum of two values.
[0288] At step S362, the prior SNR is smoothed according to the smoothing coefficient, the second suppression parameter corresponding to the previous frame, and the posterior SNR corresponding to the previous frame.
[0289] wherein the following formula is used:
[0290] SNR' = prior SNR * (1 - a) + a * H2_last 2 * SNR_last
[0291] The smoothed prior SNR SNR' is obtained, wherein H2_last is the second suppression parameter corresponding to the previous frame of the current frame, and SNR_last is the posterior SNR corresponding to the previous frame of the current frame.
[0292] At step S363, the second suppression parameter is determined according to the smoothed prior SNR.
[0293] In some embodiments, the following formula is used:
[0294]
[0295] The second suppression parameter H2 is determined.
[0296] At step S370, the linear echo cancellation signal is processed according to the first suppression parameter and the second suppression parameter, and an echo cancellation signal is output.
[0297] wherein the following formula is used:
[0298] aec_out = aec_out1 * H(w) * H2
[0299] The echo cancellation signal aec_out is obtained.
[0300] Thus, the processing of the microphone signal is completed, the linear echo, the nonlinear echo, and the reverberation echo therein are eliminated, and the quality of the speech signal is improved.
[0301] Figure 12 is a block diagram of an echo cancellation device according to an exemplary embodiment. As shown in Figure 12 The echo cancellation device 70 includes a linear echo estimation module 71, a linear echo cancellation module 72, a linear residual echo suppression module 73, a reverberation echo estimation module 74, an echo suppression module 75, and an echo cancellation module 76.
[0302] The linear echo estimation module 71 is configured to estimate a linear echo component in the input microphone signal collected by the microphone according to the loudspeaker signal output by the loudspeaker, to obtain a linear echo signal.
[0303] The linear echo cancellation module 72 cancels the linear echo component in the microphone signal according to the linear echo signal, to obtain a linear echo cancelled signal.
[0304] The linear residual echo suppression module 73 is configured to estimate a residual echo component in the linear echo cancelled signal according to the spectrum of the microphone signal and the spectrum of the linear echo signal, and determine a first suppression parameter.
[0305] The reverberation echo estimation module 74 estimates a reverberation echo component in the linear echo cancelled signal according to the linear echo signals of the previous N frames, to obtain a reverberation echo signal, where N is a positive integer.
[0306] The echo suppression module 75 is configured to determine a second suppression parameter according to the linear echo cancelled signal and the reverberation echo signal.
[0307] The echo cancellation module 76 is configured to process the linear echo cancelled signal according to the first suppression parameter and the second suppression parameter, to output an echo cancelled signal.
[0308] In some embodiments, the linear echo estimation module 71 is configured to detect whether the loudspeaker signal includes a speech component; in response to detecting that the loudspeaker signal does not include a speech component, estimate the linear echo component in the microphone signal using a Kalman filtering algorithm, to obtain the linear echo signal; and in response to detecting that the loudspeaker signal includes a speech component, update a Kalman coefficient, and estimate the linear echo component in the microphone signal using the Kalman filtering algorithm, to obtain the linear echo signal.
[0309] In some embodiments, the linear residual echo suppression module 73 is configured to detect the residual echo component in the linear echo cancelled signal using spectral subtraction according to the spectrum of the microphone signal and the spectrum of the linear echo signal.
[0310] In some embodiments, the echo cancellation apparatus further comprises a calculation module.
[0311] The calculation module is configured to calculate a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter.
[0312] The reverberation echo estimation module 74 is configured to, in response to the value of the far-end single-talk flag being greater than or equal to a preset first threshold, determine, according to the linear echo-canceled signal and the linear echo signals of the previous N frames, an echo attenuation coefficient corresponding to each frame in the previous N frames; determine, according to the echo attenuation coefficient and the far-end single-talk flag, a reverberation attenuation coefficient corresponding to each frame in the previous N frames; estimate a reverberation echo component in the linear echo-canceled signal according to the reverberation attenuation coefficient and the linear echo signals of the previous N frames, to obtain a reverberation echo signal.
[0313] In some embodiments, the reverberation echo estimation module 74 is configured to, for each frame in the previous N frames: determine a smoothing factor according to the far-end single-talk flag, the echo attenuation coefficient corresponding to the frame, and the initial reverberation attenuation coefficient corresponding to the frame; and smooth the initial reverberation attenuation coefficient according to the smoothing factor and the echo attenuation coefficient corresponding to the frame, to obtain the reverberation attenuation coefficient corresponding to the frame.
[0314] In some embodiments, the echo cancellation device further comprises a calculation module and a nonlinear echo estimation module.
[0315] The calculation module is configured to calculate the far-end single-talk flag according to the loudspeaker signal and the first suppression parameter.
[0316] The nonlinear echo estimation module is configured to, in response to the value of the far-end single-talk flag being greater than or equal to a preset second threshold, estimate a nonlinear echo component in the linear echo-canceled signal according to the far-end single-talk flag, the linear echo-canceled signal, and a nonlinear distortion signal corresponding to the loudspeaker signal, to obtain a nonlinear echo signal, wherein the nonlinear distortion signal is obtained by full-wave rectification of the loudspeaker signal.
[0317] The echo suppression module 75 is configured to determine the second suppression parameter according to the linear echo-canceled signal, the nonlinear echo signal, and the reverberation echo signal.
[0318] In some embodiments, the nonlinear echo estimation module is configured to calculate a ratio of the amplitudes of the linear echo-canceled signal and the nonlinear distortion signal, and transform the ratio to a logarithmic domain to obtain a logarithmic domain ratio; determine a smoothing coefficient according to the far-end single-talk flag; smooth the logarithmic domain ratio according to the smoothing coefficient; transform the smoothed logarithmic domain ratio to an amplitude domain to obtain an amplitude domain ratio; and obtain the nonlinear echo signal according to the amplitude domain ratio and the nonlinear distortion signal.
[0319] In some embodiments, the echo suppression module 75 is configured to calculate a posterior signal-to-echo ratio according to the energy value of the linear echo cancellation signal, the energy value of the reverberation echo signal, and the energy value of the non-linear echo signal; determine a prior signal-to-echo ratio according to the posterior signal-to-echo ratio; smooth the prior signal-to-echo ratio according to the smoothing coefficient, the second suppression parameter corresponding to the previous frame, and the posterior signal-to-echo ratio corresponding to the previous frame; and determine the second suppression parameter according to the smoothed prior signal-to-echo ratio.
[0320] In some embodiments, the echo cancellation apparatus further comprises a calculation module.
[0321] The calculation module is configured to calculate a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter; and determine the smoothing coefficient according to the far-end single-talk flag.
[0322] The echo suppression module 75 is configured to calculate a posterior signal-to-echo ratio according to the energy value of the linear echo cancellation signal and the energy value of the reverberation echo signal; determine a prior signal-to-echo ratio according to the posterior signal-to-echo ratio; smooth the prior signal-to-echo ratio according to the smoothing coefficient, the second suppression parameter corresponding to the previous frame, and the posterior signal-to-echo ratio corresponding to the previous frame; and determine the second suppression parameter according to the smoothed prior signal-to-echo ratio.
[0323] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described in detail here.
[0324] The present disclosure also provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, implement the steps of the echo cancellation method provided by the present disclosure.
[0325] Figure 13 is a block diagram of an electronic device 800 according to an exemplary embodiment. For example, the electronic device 800 can be a wearable device, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0326] Referring to Figure 13 , the electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0327] The processing component 802 generally controls the overall operations of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or a part of steps of the echo cancellation method described above. In addition, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0328] The memory 804 is configured to store various types of data to support operations of the electronic device 800. Examples of these data include instructions to operate any applications or methods on the electronic device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be realized by any type of volatile or nonvolatile memory devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disc, or optical disc.
[0329] The power component 806 provides power to various components of the electronic device 800. The power component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0330] The multimedia component 808 includes a screen to provide an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes the touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a back camera. The front camera and / or the back camera can receive external multimedia data when the electronic device 800 is in an operating mode, such as a shooting mode or a video mode. Each of the front camera and the back camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0331] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive an external audio signal when the electronic device 800 is in an operation mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0332] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0333] The sensor component 814 includes one or more sensors for providing status assessments of various aspects of the electronic device 800. For example, the sensor component 814 can detect an open / closed position of the electronic device 800, relative positioning of components, such as a display and a keypad of the electronic device 800, a change of position of the electronic device 800 or a component of the electronic device 800, presence or absence of user contact with the electronic device 800, orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect presence of a nearby object without any physical touch. The sensor component 814 can further include a light sensor, such as a CMOS or CCD image sensor, for use in an imaging application. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0334] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an example embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcasting management system via a broadcast channel. In an example embodiment, the communication component 816 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technology.
[0335] In exemplary embodiments, the electronic device 800 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, micro-controllers, microprocessors, or other electronic elements for executing the above-described echo cancellation method.
[0336] In exemplary embodiments, a non-transitory computer-readable storage medium including instructions, such as the memory 804 including instructions, is also provided, which can be executed by the processor 820 of the electronic device 800 to complete the above-described echo cancellation method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0337] The above-described apparatus can be an independent electronic device, or can be a part of an independent electronic device, such as an integrated circuit (IC) or a chip in an embodiment. The integrated circuit can be one IC or a collection of multiple ICs. The chip can include, but is not limited to, a GPU (Graphics Processing Unit), a CPU (Central Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a SOC (System on Chip), or a SoC (System on Chip), etc. The above-described integrated circuit or chip can execute executable instructions (or code) to implement the above-described echo cancellation method. The executable instructions can be stored in the integrated circuit or chip, or can be obtained from other apparatuses or devices, such as a processor, a memory, and an interface for communicating with other apparatuses included in the integrated circuit or chip. The executable instructions can be stored in the memory, and when executed by the processor, implement the above-described echo cancellation method. Alternatively, the integrated circuit or chip can receive executable instructions through the interface and transmit the executable instructions to the processor for execution to implement the above-described echo cancellation method.
[0338] In another exemplary embodiment, there is also provided a computer program product comprising a computer program capable of being executed by a programmable apparatus, the computer program having code portions for performing the above-mentioned echo cancellation method when executed by the programmable apparatus.
[0339] Other embodiments of the disclosure will be apparent to those skilled in the art from consideration of the specification and practice of the disclosure. It is intended that the specification and examples be considered as exemplary only, with the true scope and spirit of the disclosure being indicated by the following claims.
[0340] It will be understood that the disclosure is not limited to the precise structures hereinbefore described and illustrated in the drawings, and that various modifications and changes can be made without departing from its scope. The scope of the disclosure is limited only by the claims that follow.
Claims
1. An echo cancellation method, characterized by, Applied to a terminal with a loudspeaker and a microphone, the method comprises: estimating a linear echo component in a microphone signal collected by the microphone according to a loudspeaker signal output by the loudspeaker, to obtain a linear echo signal; eliminating the linear echo component in the microphone signal according to the linear echo signal, to obtain a linear echo cancellation signal; estimating a residual echo component in the linear echo cancellation signal according to a spectrum of the microphone signal and a spectrum of the linear echo signal, and determining a first suppression parameter, wherein the first suppression parameter is used to suppress the residual echo component in the linear echo cancellation signal; estimating a reverberation echo component in the linear echo cancellation signal according to the linear echo signals of the previous N frames, to obtain a reverberation echo signal, wherein N is a positive integer; determining a second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal, wherein the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal; processing the linear echo cancellation signal according to the first suppression parameter and the second suppression parameter, to output an echo cancellation signal.
2. The echo cancellation method of claim 1, wherein, The estimation of the linear echo component in the microphone signal according to the loudspeaker signal output by the loudspeaker, to obtain a linear echo signal, comprises: detecting whether the loudspeaker signal includes a speech component; in response to detecting that the loudspeaker signal does not include a speech component, estimating the linear echo component in the microphone signal by using a Kalman filtering algorithm, to obtain the linear echo signal; in response to detecting that the loudspeaker signal includes a speech component, updating a Kalman coefficient, and estimating the linear echo component in the microphone signal by using the Kalman filtering algorithm, to obtain the linear echo signal.
3. The echo cancellation method of claim 1, wherein, The estimation of the residual echo component in the linear echo cancellation signal according to the spectrum of the microphone signal and the spectrum of the linear echo signal, comprises: detecting the residual echo component in the linear echo cancellation signal by using spectral subtraction according to the spectrum of the microphone signal and the spectrum of the linear echo signal.
4. The echo cancellation method of claim 1, wherein, Before estimating the reverberation echo component in the linear echo cancellation signal to obtain a reverberation echo signal, the echo cancellation method further comprises: calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter; The reverberation echo estimation according to the linear echo signals of the previous N frames to obtain a reverberation echo signal, comprises: in response to the value of the far-end single-talk flag being greater than or equal to a preset first threshold, determining an echo attenuation coefficient corresponding to each of the previous N frames according to the linear echo cancellation signal and the linear echo signals of the previous N frames; determining a reverberation attenuation coefficient corresponding to each of the previous N frames according to the echo attenuation coefficient and the far-end single-talk flag; estimating the reverberation echo component in the linear echo cancellation signal according to the reverberation attenuation coefficient and the linear echo signals of the previous N frames, to obtain the reverberation echo signal.
5. The echo cancellation method of claim 4, wherein, The determination of the reverberation attenuation coefficient corresponding to each of the previous N frames according to the echo attenuation coefficient and the far-end single-talk flag, comprises: For each frame in the first N frames: determine a smoothing factor according to the far-end single-talk flag, the echo decay coefficient corresponding to the frame, and the initial reverb decay coefficient corresponding to the frame; smooth the initial reverb decay coefficient according to the smoothing factor and the echo decay coefficient corresponding to the frame, to obtain the reverb decay coefficient corresponding to the frame.
6. The echo cancellation method of claim 1, wherein, Before determining the second suppression parameter according to the linear echo-canceled signal and the reverberation echo signal, the echo cancellation method further comprises: calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter; in response to the value of the far-end single-talk flag being greater than or equal to a preset second threshold, estimating a nonlinear echo component in the linear echo-canceled signal according to the far-end single-talk flag, the linear echo-canceled signal, and a nonlinear distortion signal corresponding to the loudspeaker signal, to obtain a nonlinear echo signal, wherein the nonlinear distortion signal is obtained by full-wave rectification of the loudspeaker signal; determining the second suppression parameter according to the linear echo-canceled signal, the nonlinear echo signal, and the reverberation echo signal, wherein the second suppression parameter is used to suppress the reverberation echo component and the nonlinear echo component in the linear echo-canceled signal. The estimation of the nonlinear echo component in the linear echo-canceled signal according to the far-end single-talk flag, the linear echo-canceled signal, and a nonlinear distortion signal corresponding to the loudspeaker signal to obtain a nonlinear echo signal comprises:
7. The echo cancellation method of claim 6, wherein, calculating the ratio of the amplitude of the linear echo-canceled signal to the amplitude of the nonlinear distortion signal, and transforming the ratio to the logarithmic domain to obtain a logarithmic domain ratio; determining a smoothing coefficient according to the far-end single-talk flag; smoothing the logarithmic domain ratio according to the smoothing coefficient; transforming the smoothed logarithmic domain ratio to the amplitude domain to obtain an amplitude domain ratio; obtaining the nonlinear echo signal according to the amplitude domain ratio and the nonlinear distortion signal. The determination of the second suppression parameter according to the linear echo-canceled signal, the nonlinear echo signal, and the reverberation echo signal comprises:
8. The echo cancellation method of claim 7, wherein, calculating a post-signal-to-background ratio according to the energy value of the linear echo-canceled signal, the energy value of the reverberation echo signal, and the energy value of the nonlinear echo signal; determining a priori signal-to-background ratio according to the post-signal-to-background ratio; smoothing the priori signal-to-background ratio according to the smoothing coefficient, the second suppression parameter corresponding to the previous frame, and the post-signal-to-background ratio corresponding to the previous frame; determining the second suppression parameter according to the smoothed priori signal-to-background ratio. Before determining the second suppression parameter according to the linear echo-canceled signal and the reverberation echo signal, the echo cancellation method comprises:
9. The echo cancellation method of claim 1, wherein, calculating a far-end single-talk flag according to the loudspeaker signal and the first suppression parameter; determining a smoothing coefficient according to the far-end single-talk flag; the determination of the second suppression parameter according to the linear echo-canceled signal and the reverberation echo signal comprises: calculate a post-signal-to-reverberation ratio according to the energy value of the linear echo cancellation signal and the energy value of the reverberation echo signal; determine a priori signal-to-reverberation ratio according to the post-signal-to-reverberation ratio; smooth the priori signal-to-reverberation ratio according to the smoothing coefficient, the second suppression parameter corresponding to the previous frame and the post-signal-to-reverberation ratio corresponding to the previous frame; determine the second suppression parameter according to the smoothed priori signal-to-reverberation ratio.
10. An echo cancellation device, characterized by comprise: a linear echo estimation module configured to estimate a linear echo component in a microphone signal collected by a microphone according to a loudspeaker signal output by a loudspeaker to obtain a linear echo signal; a linear echo cancellation module configured to cancel the linear echo component in the microphone signal according to the linear echo signal to obtain a linear echo cancellation signal; a linear residual echo suppression module configured to estimate a residual echo component in the linear echo cancellation signal according to a spectrum of the microphone signal and a spectrum of the linear echo signal and determine a first suppression parameter, wherein the first suppression parameter is used to suppress the residual echo component in the linear echo cancellation signal; a reverberation echo estimation module configured to estimate a reverberation echo component in the linear echo cancellation signal according to the linear echo signals of previous N frames to obtain a reverberation echo signal, wherein N is a positive integer; an echo suppression module configured to determine a second suppression parameter according to the linear echo cancellation signal and the reverberation echo signal, wherein the second suppression parameter is used to suppress the reverberation echo component in the linear echo cancellation signal; an echo cancellation module configured to process the linear echo cancellation signal according to the first suppression parameter and the second suppression parameter to output an echo cancellation signal.
11. An electronic device, comprising: comprise: a loudspeaker, a microphone, a processor and a memory; the loudspeaker, the microphone and the memory are all in communication connection with the processor; the memory is used to store processor executable instructions; wherein the processor is configured to run the executable instructions to implement the steps of the echo cancellation method in any one of claims 1-9.
12. A computer-readable storage medium having stored thereon computer program instructions, wherein, The program instructions are executed by the processor to implement the steps of the echo cancellation method in any one of claims 1-9.
Citation Information
Patent Citations
Voice enhancement method and device based on microphone array
CN108447496A
Reverberation inhibiting system and method
CN109712637A