Residual echo estimator, non-transitory computer-readable medium, and application processor
Through the residual echo estimator and frequency domain transformation technology, combined with adaptive filters and suppression gain calculation, the residual echo in the electronic device is effectively eliminated, solving the problem of acoustic echo affecting the recognition of the far-end speaker, and improving voice recognition and call quality.
Patent Information
- Application Number
- CN201911376200.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-01-29
- Filing Date
- 2019-12-27
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2039-12-27
AI Technical Summary
In electronic devices, existing technologies have difficulty in effectively eliminating acoustic echoes, especially residual echoes, which affect the voice and speech recognition of the far-end speaker.
A residual echo estimator is used to estimate the amplitude of the residual echo through time correlation, and a linear echo canceller and a residual echo suppressor are combined to calculate the suppression gain to eliminate the echo using an adaptive filter and frequency domain transformation technology.
Significantly reduces echo return loss, improves the voice recognition effect and speech recognition accuracy of the far-end speaker, and improves the call quality of electronic devices.
Smart Images

Figure CN111489761B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims the benefit of Korean Patent Application No. 10-2019-0011091 filed on January 29, 2019, in the Korean Intellectual Property Office, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] Example embodiments of the inventive concepts described herein relate to a residual echo estimator configured to estimate a residual echo based on a time correlation, a non-transitory computer-readable medium storing a program code configured to estimate a residual echo, and / or an application processor. Background Art
[0004] In electronic devices that include a speaker and a microphone, when the voice of a far-end speaker is output through the speaker, the far-end speaker's voice is input into the microphone. This acoustic echo makes it difficult to recognize the voice of a near-end speaker at the far end. Accordingly, due to the coupling between the adjacent speaker and microphone, it is desirable to eliminate the acoustic echo. Furthermore, with the increasing demand for electronic devices that support voice recognition, effectively eliminating the acoustic echo is becoming even more important. Summary of the Invention
[0005] Example embodiments of the inventive concept provide a residual echo estimator configured to estimate a residual echo based on time correlation, a non-transitory computer-readable medium storing a program code configured to estimate the residual echo, and an application processor.
[0006] According to an example embodiment, a residual echo estimator is configured to estimate a residual echo of a microphone signal. In some example embodiments, the residual echo estimator includes a processing circuit configured to estimate the amplitude of the residual echo at the current frame based on the amplitude of the linear echo of the reference signal at the current frame and the amplitude of the linear echo of the reference signal at the past frame, and to update the weight applied to the amplitude of the linear echo at the current frame and the weight applied to the amplitude of the linear echo at the past frame.
[0007] According to example embodiments, a non-transitory computer-readable medium stores program code executable by a processor. In some example embodiments, the program code is executable by the processor to: estimate the amplitude of a residual echo of a microphone signal at a current frame based on the amplitude of a linear echo of a reference signal at a current frame and the amplitude of a linear echo of the reference signal at a past frame; update weights applied to the amplitude of the linear echo at the current frame and the amplitude of the linear echo at the past frame; and calculate a suppression gain based on the amplitude of the residual echo and the amplitude of an output signal obtained by canceling the linear echo from the microphone signal, wherein the output signal is multiplied by the suppression gain to generate a final output signal.
[0008] According to an example embodiment, an application processor includes: an audio processor; and a non-transitory computer-readable medium configured to store program code executable by the audio processor to: generate an output signal by canceling a linear echo of a reference signal from a microphone signal input to the microphone as the reference signal is output from a speaker, wherein the linear echo is determined based on a transfer path between the speaker and the microphone; and estimate a magnitude of a residual echo at a current frame based on one or more of a magnitude of the linear echo at the current frame and a magnitude of the linear echo at a past frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other aspects and features of the inventive concepts will become apparent by describing in detail example embodiments of the inventive concepts with reference to the attached drawings.
[0010] Figure 1 A communication system according to an example embodiment of the inventive concept is shown.
[0011] Figure 2 An exemplary embodiment according to the inventive concept is shown. Figure 1 AEC system.
[0012] Figure 3A An exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0013] Figure 3B Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0014] Figure 3C Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0015] Figure 3D Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0016] Figure 4 An exemplary embodiment according to the inventive concept is shown. Figure 1 electronic devices.
[0017] Figure 5 Another exemplary embodiment according to the inventive concept is shown. Figure 1 AEC system.
[0018] Figure 6 A flowchart of estimating the amplitude of a residual echo according to an example embodiment of the inventive concept is shown.
[0019] Figure 7An example embodiment according to the inventive concept is shown. Figure 6 Detailed operations associated with operation S110.
[0020] Figure 8 Another exemplary embodiment according to the inventive concept is shown. Figure 6 Detailed operations associated with operation S110.
[0021] Figure 9 An electronic device according to an example embodiment of the inventive concept is shown. DETAILED DESCRIPTION
[0022] Figure 1 A communication system according to an example embodiment of the inventive concept is shown.
[0023] Reference Figure 1 , the communication system 1 may include an electronic device 10 associated with or adjacent to a near-end speaker, and an electronic device 20 associated with or adjacent to a far-end speaker. A far-end signal corresponding to the voice of the far-end speaker (hereinafter referred to as a "reference signal") may be input to the electronic device 20 through a microphone 22. The electronic device 20 may transmit the reference signal to the electronic device 10 based on a wired / wireless communication protocol.
[0024] The electronic device 10 according to an exemplary embodiment of the inventive concept can receive a reference signal from the electronic device 20 and output the reference signal through the speaker 11. The electronic device 10 can receive the ambient sound of the electronic device 10 (e.g., the voice of a near-end speaker or noise) through the microphone 12. However, when the reference signal is output through the speaker 11, an echo (echo signal) relative to the reference signal may be input to the microphone 12 of the electronic device 10. Therefore, the electronic device 10 can cancel or suppress the echo included in the microphone signal to suppress (or alternatively, prevent) the voice of the far-end speaker from being replayed to the far-end speaker through the speaker 21 of the electronic device 20.
[0025] The electronic device 10 may include an application processor 100. The application processor 100 may include an audio processor 110 and a memory 120.
[0026] The electronic device 10 may include an acoustic echo cancellation (AEC) system 1000, which is implemented using a processing circuit including a logic circuit (such as hardware), a hardware / software combination (such as a processor executing software), or a combination thereof. For example, the processing circuit may more specifically include, but is not limited to, a central processing unit (CPU), an arithmetic logic unit (ALU), a digital signal processor, a microprocessor, a field programmable gate array (FPGA), a system on a chip (SoC), a programmable logic unit, a microprocessor, an application-specific integrated circuit (ASIC), etc.
[0027] For example, the application processor 100 may include an acoustic echo cancellation (AEC) system 1000 implemented using hardware components, circuits, or modules to cancel echoes associated with a reference signal. For another example, the application processor 100 may execute or implement the AEC system 1000 that may be executed using software components, functional blocks, or modules.
[0028] In the case where all or part of the AEC system 1000 is implemented by software, the audio processor 110 may execute program code stored in the memory 120 to execute or implement the AEC system 1000. The audio processor 110 may include at least one core (processing unit) that may read and execute commands, algorithms, or functions included in the program code from the memory 120 to implement the AEC system 1000.
[0029] The memory 120 may be a non-transitory computer-readable medium that stores program code for executing the AEC system 1000. The memory 120 may be a random access memory (RAM), a flash memory, a read-only memory (ROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a register, a hard drive, a removable disk, a CD-ROM, or any other type of storage medium. Figure 1 As shown in FIG, memory 120 may be implemented within application processor 100; alternatively, Figure 1 Unlike the illustration, the memory 120 may be a storage medium that is implemented independently of the application processor 100 in the electronic device 10 or is provided outside the electronic device 10 .
[0030] The AEC system 1000 may include a linear echo canceller LEC and a residual echo suppressor RES. The linear echo canceller LEC may receive a reference signal. The linear echo canceller LEC may linearly estimate (calculate) a linear echo X'(κ, m) of the reference signal based on the transfer path between the loudspeaker 11 and the microphone 12. "κ(kappa)" is a frame index associated with time, and "m" is a frequency band. For example, the linear echo canceller LEC may perform a convolution operation on the reference signal and a transfer function, wherein the transfer function is pre-modeled to indicate the transfer path. The linear echo canceller LEC may include an adaptive filter. The linear echo canceller LEC may eliminate the linear echo X'(κ, m) from the microphone signal by using an adaptive filter, and may generate an output signal E(κ, m) as a result of the elimination. The coefficients of the adaptive filter may be continuously updated so that the linear echo X'(κ, m) matches the actual echo of the microphone signal.
[0031] Since the linear echo canceller LEC linearly estimates the linear echo X'(κ,m) of the reference signal, a residual echo may exist in the output signal E(κ,m). For example, the nonlinear distortion of the echo (i.e., the cause of harmonic distortion (HD)) may include the nonlinear characteristics of the speaker 11 itself, vibrations generated by the adjacent speaker 11 and microphone 12, and nonlinear characteristics of components of the electronic device 10.
[0032] The residual echo suppressor RES can suppress residual echo that is not eliminated by the linear echo canceller LEC. The residual echo suppressor RES can generate a suppression gain G(κ,m), can multiply the suppression gain G(κ,m) by the output signal E(κ,m), and can generate an output signal R(κ,m). The residual echo suppressor RES can adaptively update the suppression gain G(κ,m) so that the residual echo is not included in the output signal R(κ,m) to the greatest extent possible. The electronic device 10 can provide the echo-canceled microphone signal to the electronic device 20, and the far-end speaker can hear the microphone signal, which does not include the echo to the greatest extent possible, through the speaker 21.
[0033] Figure 2 An exemplary embodiment according to the inventive concept is shown. Figure 1 AEC system.
[0034] Reference Figure 2 , showing in detail the AEC system 1000 of the electronic device 10, Figure 2 , an exemplary embodiment is shown in which a reference signal is output through a speaker 11 after being amplified by an amplifier (AMP) 13, and an echo is captured by a microphone 12. For brevity and for ease of description, any other components of the electronic device 10 are omitted.
[0035] The AEC system 1000 may include a first Fourier transform (FT) module 1100 and a second Fourier transform (FT) module 1200, a linear echo canceller 1300, a residual echo suppressor 1400, and an inverse Fourier transform (IFT) module 1500. The transform module may also be referred to as a "converter" or "transformer." As discussed above, the AEC system 1000 may be implemented by executing program code for estimating the amplitude of the residual echo, and the executed program code may transform a processor into a special-purpose computer to perform the functions of the components, functions, blocks, or modules of the AEC system 1000.
[0036] The first Fourier transform module 1100 can receive a microphone signal and transform (convert) the microphone signal from the time domain to the frequency domain. For example, the first Fourier transform module 1100 can perform a transform operation based on a modulated overlap transform (MCLT), a Fourier transform (FT), a short-time Fourier transform (STFT), a discrete Fourier transform (DFT), a fast Fourier transform (FFT), etc., and can output a microphone signal Y(κ, m). The second Fourier transform module 1200 can receive a reference signal and transform the reference signal from the time domain to the frequency domain. The second Fourier transform module 1200 can perform a transform operation in a manner similar to the first Fourier transform module 1100 and can output a reference signal X(κ, m).
[0037] The linear echo canceller 1300 may perform a convolution operation on a reference signal and a transfer function that is pre-modeled to indicate a path between the speaker 11 and the microphone 12, and may estimate a linear echo X'(κ,m) of the reference signal. The linear echo canceller 1300 may provide the linear echo X'(κ,m) and an output signal E(κ,m) obtained by canceling the linear echo X'(κ,m) from the microphone signal Y(κ,m) to the residual echo suppressor 1400.
[0038] The residual echo suppressor 1400 may receive a microphone signal Y(κ,m), a linear echo X'(κ,m), an estimated residual echo, and an output signal E(κ,m). The residual echo suppressor 1400 according to an exemplary embodiment of the inventive concept may calculate a suppression gain G(κ,m) for suppressing the residual echo using the microphone signal Y(κ,m), the linear echo X'(κ,m), and the estimated residual echo. The residual echo suppressor 1400 may generate an output signal R(κ,m) by multiplying the suppression gain G(κ,m) by the output signal E(κ,m).
[0039] The inverse Fourier transform module 1500 may receive the output signal R(κ, m) from the residual echo suppressor 1400. The inverse Fourier transform module 1500 may transform the output signal R(κ, m) from the frequency domain to the time domain. For example, the inverse Fourier transform module 1500 may perform a transform operation based on an inverse MCLT, an inverse FT, an inverse STFT, an inverse DFT, an inverse FFT, or the like.
[0040] Figure 3A An exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0041] Reference Figure 3AThe residual echo suppressor 1400a may include a residual echo estimator 1410a, a noise estimator 1420, first to third smoothing calculators 1431 to 1433, a gain calculator 1440, and a multiplier 1450. The residual echo suppressor 1400a may be Figure 2 However, example embodiments are not limited thereto.
[0042] The residual echo estimator 1410a may receive the linear echo X'(κ,m) and the linear echo X'(κ-t,m) at a past frame (or alternatively, in a past frame) stored in the memory 120. Figure 2 Unlike the illustration in FIG, the residual echo estimator 1410a may not receive the microphone signal Y(κ,m). The residual echo estimator 1410a may estimate the residual echo at the current frame by using the amplitude of the linear echo X'(κ,m) at the current frame and the amplitude of the linear echo X'(κ-t,m) at the past frame. For example, the residual echo estimator 1410a can use Equation 1 to calculate the residual echo at the current frame: The amplitude.
[0043] [Equation 1]
[0044]
[0045] in,
[0046] In Equation 1, "i" is the baseband, "M" is the number of subbands, "j" is the harmonic, "H" is the number of harmonics to be considered, "2K+1" as the k (letter) index is the length of the harmonic search window, "m" is the frequency bin index* (frequency bin index), "W R (i, j, k)” is the weight applied to the amplitude of the linear echo X'(κ, m) at the current frame, "T" is the number of past frames to perform estimation by using "t" as an index, and "W T1 (t, m)" is the weight applied to the amplitude of the linear echo X'(κ-t, m) at the past frame. The weight W can be adaptively updated by the residual echo estimator 1410a. R (i,j,k) and W T1 (t,m).
[0047] According to an exemplary embodiment of the inventive concept, the residual echo estimator 1410a may estimate the residual echo at the current frame by using the magnitude of the linear echo X'(κ-t,m) at the past frame and the magnitude of the linear echo X'(κ,m) at the current frame. The time correlation of the linear echo X'(κ,m) can be used to calculate the residual echo Therefore, the echo return loss enhancement (ERLE) of the AEC system 1000 can be improved.
[0048] The amplitude of the linear echo X'(κ-t, m) at the past frame may be stored in the memory 120. As described above, the memory 120 may be any type of storage medium. The memory 120 may be a register, a tightly coupled memory (TCM), or a static random access memory (SRAM) in the application processor 100. The memory 120 may be a dynamic random access memory (DRAM) located outside the application processor 100. The amplitude of the linear echo X'(κ, m) at the current frame may be stored in the memory 120. Both the amplitude of the linear echo X'(κ-t, m) at the past frame and the amplitude of the linear echo X'(κ, m) at the current frame may be used to estimate the residual echo at the next frame. The amplitude.
[0049] The noise estimator 1420 may estimate the magnitude of the noise N(κ,m) by using the magnitude of the output signal E(κ,m) of the linear echo canceller 1300. For example, the noise estimator 1420 may calculate the noise N(κ,m) according to minimum static force.
[0050] The first smoothing calculator 1431 to the third smoothing calculator 1433 may respectively calculate the residual echo. The first to third smoothing calculators 1431 to 1433 can perform smoothing operations on the amplitude of the residual echo, the amplitude of the output signal E(κ, m) of the linear echo canceller 1300, and the amplitude of the noise N(κ, m). The amplitude of the residual echo, the amplitude of the output signal E(κ,m) of the linear echo canceller 1300, and the amplitude of the noise N(κ,m) are recursively averaged. For example, smoothing calculators 1431 to 1433 can calculate the smoothed amplitude of the residual echo, the smoothed amplitude of the output signal of the linear echo canceller 1300, and the smoothed amplitude of the noise using Equations 2 to 4, respectively.
[0051] [Equation 2]
[0052]
[0053] [Equation 3]
[0054]
[0055] [Equation 4]
[0056]
[0057] In Equations 2 to 4, “α1,” “α2,” and “α3” may vary between 0 and 1, and may be the same as or different from each other.
[0058] Gain calculator 1440 can calculate the suppression gain G(κ,m) by using the smoothed amplitude of the residual echo, the smoothed amplitude of the output signal of linear echo canceller 1300, and the smoothed amplitude of the noise. The suppression gain G(κ,m) is then multiplied by the output signal E(κ,m) of linear echo canceller 1300. Gain calculator 1440 can calculate the suppression gain G(κ,m) using various operation techniques. For example, gain calculator 1440 can calculate the suppression gain G(κ,m) using Equation 5 based on spectral subtraction.
[0059] [Equation 5]
[0060]
[0061] In Equation 5, “β” varies between 0 and 1.
[0062] For another example, the suppression gain G(κ,m) may be calculated by Equation 6 based on a Wiener filter.
[0063] [Equation 6]
[0064]
[0065] In Equation 6, It can be calculated by Equation 7.
[0066] [Equation 7]
[0067]
[0068] In Equation 7, “α4” varies between 0 and 1, u(·) is a unit step function, and γ(κ, m) can be calculated by Equation 8.
[0069] [Equation 8]
[0070]
[0071] In Equation 8, “α5” varies between 0 and 1.
[0072] For another example, the gain calculator 1440 may calculate the suppression gain G(κ,m) using Equation 9 based on a minimum mean square error-short time spectrum amplitude estimator (MMSE-STSA).
[0073] [Equation 9]
[0074]
[0075] In Equation 9, Γ(·) is a gamma function, v(κ,m) can be calculated by Equation 10, and M(·; ·; ·) can be a confluent hypergeometric function.
[0076] [Equation 10]
[0077]
[0078] The multiplier 1450 may generate an output signal R(κ,m) by multiplying the suppression gain G(κ,m) and the output signal E(κ,m) of the linear echo canceller 1300 .
[0079] Figure 3B Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0080] Reference Figure 3B The residual echo suppressor 1400b may include a residual echo estimator 1410b, a noise estimator 1420, first to third smoothing calculators 1431 to 1433, a gain calculator 1440, and a multiplier 1450. The residual echo suppressor 1400b may be Figure 2 However, example embodiments are not limited thereto. The description will focus on the differences between the residual echo estimator 1410a and the residual echo estimator 1410b, and will refer to Figure 3A The remaining components 1420 to 1450 are described.
[0081] The residual echo estimator 1410b may receive the linear echo X'(κ,m) and the linear echo X'(κ-t,m) at the past frame stored in the memory 120. Figure 2 Unlike the illustration of FIG, the residual echo estimator 1410b may not receive the microphone signal Y(κ,m). The residual echo estimator 1410b may also receive the residual echo at the past frame stored in the memory 120. The residual echo estimator 1410b can estimate the residual echo by using the magnitude of the linear echo X'(κ,m) at the current frame, the magnitude of the linear echo X'(κ-t,m) at the past frame, and the residual echo at the past frame. The amplitude of the residual echo at the current frame is estimated For example, the residual echo at the current frame It can be calculated by Equation 11.
[0082] [Equation 11]
[0083]
[0084] W T2 (t,m) is the residual echo applied to the past frame The weight W can be adaptively updated by the residual echo estimator 1410b. T2 (t, m). According to another embodiment of the inventive concept, the residual echo estimator 1410b can calculate the residual echo value by using the amplitude of the linear echo X'(κ-t, m) at the past frame, the residual echo value at the past frame, and the residual echo value at the past frame. The amplitude of the linear echo X'(κ,m) at the current frame is used to estimate the residual echo at the current frame. The amplitude of the linear echo X'(κ,m) time correlation and residual echo The time correlation can be used to calculate the residual echo Therefore, the degree of improvement in ERLE by the residual echo estimator 1410b can be greater than that by Figure 3A The residual echo estimator 1410a improves the ERLE to a certain extent.
[0085] Residual echo from past frames The amplitude of can be stored in the memory 120. The residual echo at the current frame The amplitude of can also be stored in the memory 120. The residual echo at the past frame The amplitude of the residual echo at the current frame The amplitude of the linear echo X'(κ-t, m) at the past frame and the amplitude of the linear echo X'(κ, m) at the current frame can be used to estimate the residual echo at the next frame The amplitude.
[0086] Figure 3C Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0087] Reference Figure 3C The residual echo suppressor 1400c may include a residual echo estimator 1410c, a noise estimator 1420, first to third smoothing calculators 1431 to 1433, a gain calculator 1440, and a multiplier 1450. The residual echo suppressor 1400c may be Figure 2 However, example embodiments are not limited thereto. The description will focus on the differences between the residual echo estimator 1410a and the residual echo estimator 1410c, and will refer to Figure 3A The remaining components 1420 to 1450 are described.
[0088] The residual echo estimator 1410c may receive the linear echo X'(κ,m) and the linear echo X'(κ-t,m) at the past frame stored in the memory 120. The residual echo estimator 1410c may also receive the microphone signal Y(κ,m) and the microphone signal Y(κ-t,m) at the past frame stored in the memory 120. The residual echo estimator 1410c may estimate the residual echo at the current frame by using the amplitude of the linear echo X'(κ,m) at the current frame, the amplitude of the linear echo X'(κ-t,m) at the past frame, the amplitude of the microphone signal Y(κ,m) at the current frame, and the amplitude of the microphone signal Y(κ-t,m) at the past frame. For example, the residual echo at the current frame The magnitude can be calculated using Equation 12.
[0089] [Equation 12]
[0090]
[0091] W T3 (t,m) is the weight applied to the amplitude of the microphone signal Y(κ,m) at the current frame and the amplitude of the microphone signal Y(κ-t,m) at the past frame. The weight W can be adaptively updated by the residual echo estimator 1410c. T3 (t, m). According to another embodiment of the inventive concept, the residual echo estimator 1410c can estimate the residual echo at the current frame by using the amplitude of the linear echo X'(κ-t, m) at the past frame, the amplitude of the microphone signal Y(κ, m) at the current frame, and the amplitude of the microphone signal Y(κ-t, m) at the past frame and the amplitude of the linear echo X'(κ, m) at the current frame. The time correlation of the linear echo X'(κ,m) and the microphone signal Y(κ,m) can be used to calculate the residual echo Therefore, the degree of improvement in ERLE by the residual echo estimator 1410c can be greater than that by Figure 3A The residual echo estimator 1410a improves the ERLE to a certain extent.
[0092] The amplitude of the microphone signal Y(κ-t, m) at the past frame may be stored in the memory 120. The amplitude of the microphone signal Y(κ, m) at the current frame may be stored in the memory 120. The amplitude of the microphone signal Y(κ-t, m) at the past frame, the amplitude of the microphone signal Y(κ, m) at the current frame, the amplitude of the linear echo X'(κ-t, m) at the past frame, and the amplitude of the linear echo X'(κ, m) at the current frame may be used to estimate the residual echo at the next frame. The amplitude.
[0093] Figure 3D Another exemplary embodiment according to the inventive concept is shown. Figure 2 Residual echo suppressor.
[0094] Reference Figure 3D The residual echo suppressor 1400d may include a residual echo estimator 1410d, a noise estimator 1420, first to third smoothing calculators 1431 to 1433, a gain calculator 1440, and a multiplier 1450. The residual echo suppressor 1400d may be Figure 2 The description will focus on the difference between the residual echo estimator 1410a and the residual echo estimator 1410d, and will refer to Figure 3A The remaining components 1420 to 1450 are described.
[0095] The residual echo estimator 1410d may receive the linear echo X'(κ, m) and the linear echo X'(κ-t, m) at the past frame stored in the memory 120. The residual echo estimator 1410d may also receive the residual echo X'(κ, m) at the past frame stored in the memory 120. The microphone signal Y(κ,m) at the current frame and the microphone signal Y(κ-t,m) at the past frame stored in the memory 120. The residual echo estimator 1410d can estimate the residual echo by using the amplitude of the linear echo X'(κ,m) at the current frame, the amplitude of the linear echo X'(κ-t,m) at the past frame, and the residual echo at the past frame. The amplitude of the microphone signal Y(κ,m) at the current frame and the amplitude of the microphone signal Y(κ-t,m) at the past frame are used to estimate the residual echo at the current frame. For example, the residual echo at the current frame The magnitude can be calculated from Equation 13.
[0096] [Equation 13]
[0097]
[0098] The weight W can be adaptively updated by the residual echo estimator 1410d. R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t, m). According to another embodiment of the inventive concept, the residual echo estimator 1410d can calculate the residual echo value by using the amplitude of the linear echo X'(κ-t, m) at the past frame, the residual echo value at the past frame, and the residual echo value at the past frame. The amplitude of the microphone signal Y(κ, m) at the current frame and the amplitude of the microphone signal Y(κ-t, m) at the past frame and the amplitude of the linear echo X'(κ, m) at the current frame are used to estimate the residual echo at the current frame. The amplitude of the linear echo X'(κ,m) time correlation, residual echo The time correlation of the microphone signal Y(κ,m) can be used to calculate the residual echo Therefore, the degree of improvement in ERLE by the residual echo estimator 1410d can be greater than that by Figures 3A to 3C The residual echo estimators 1410a to 1410c improve the ERLE to a certain extent.
[0099] The magnitude of the linear echo X'(κ-t, m) at the past frame may be stored in the memory 120. The magnitude of the linear echo X'(κ, m) at the current frame may be stored in the memory 120. The residual echo at the past frame The amplitude of can be stored in the memory 120. The residual echo at the current frame The amplitude of the microphone signal Y(κ-t,m) at the past frame, the amplitude of the microphone signal Y(κ,m) at the current frame, the amplitude of the linear echo X'(κ-t,m) at the past frame, the amplitude of the linear echo X'(κ,m) at the current frame, the amplitude of the residual echo at the past frame The amplitude and residual echo at the current frame The amplitude of the residual echo can be used to estimate the next frame The amplitude.
[0100] In an example embodiment, each of the residual echo estimators 1410a to 1410d may be a hardware accelerator (ie, a processor) implemented by using hardware components, circuits, or modules. The hardware accelerator may refer to Figure 1 The audio processor 110 described herein is implemented in conjunction with the AEC system 1000. The hardware accelerator may perform the Figures 3A to 3D Based on the operations of the residual echo estimators 1410a through 1410d described above, the audio processor 110 may perform the remaining operations of the AEC system 1000. In another example embodiment, each of the residual echo estimators 1410a through 1410d may be implemented by the audio processor 110 by executing a software component.
[0101] Figure 4 An exemplary embodiment according to the inventive concept is shown. Figure 1 Refer to the electronic device. Figure 2 and Figures 3A to 3D To describe Figure 4 .
[0102] Reference Figure 2 、 Figures 3A to 3D and Figure 4 , the electronic device 10 may include an application processor 100 and a memory 200. The application processor 100 may include an audio processor 110 and a memory 120. The AEC system 1000 including the linear echo canceller 1300 and the residual echo suppressor 1400 may be executed in the application processor 100.
[0103] The memory 120 may be a register, TCM, or SRAM provided inside the application processor 100. The memory 200 may be a DRAM provided outside the application processor 100. In addition to the memory 120, Figure 2 and Figures 3A to 3D The amplitude of the linear echo X'(κ,m), the amplitude of the microphone signal Y(κ,m) and the residual echo at the current frame are described The amplitudes may also be stored in memory 200. For example, memory 120 may operate as a cache memory that stores at least a portion of the amplitudes or all of the amplitudes. For another example, the amplitudes may be stored in memory 200. The amplitudes may be transmitted from application processor 100 to memory 200 in units of frames based on a protocol supported by memory 200 (e.g., a double data rate (DDR) interface).
[0104] The amplitude of the linear echo X'(κ,,m) at the current frame, the amplitude of the microphone signal Y(κ,m) and the residual echo stored in the memory 200 The amplitude of the residual echo can be used to estimate the next frame The memory 200 can transmit the stored amplitudes to the application processor 100 as the amplitude of the linear echo X'(κ-t,m) at the past frame, the amplitude of the microphone signal Y(κ-t,m) and the residual echo. The amplitude.
[0105] In addition, in addition to the memory 120, a weight W applied to the amplitude of the linear echo X'(κ,m) at the current frame R (i, j, k) and the weight W applied to the amplitude of the linear echo X'(κ-t, m) at the past frame T1 (t,m), the weight W applied to the amplitude of the microphone signal Y(κ,m) T3 (t,m), and applied to residual echo The magnitude weight W T2 (t, m) may also be stored in the memory 200. The memory 200 may transmit the stored weights to the application processor 100. The weights W may be adaptively updated. R(i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t,m).
[0106] Figure 5 Another exemplary embodiment according to the inventive concept is shown. Figure 1 AEC system. Figure 2 and Figures 3A to 3D To describe Figure 5 .
[0107] Reference Figure 2 、 Figures 3A to 3D and Figure 5 , the AEC system 1000 may include a first Fourier transform module 1100 and a second Fourier transform module 1200 , a linear echo canceller 1300 , a residual echo suppressor 1400 and an inverse Fourier transform module 1500 . Figure 5 The AEC System 1000 can be used with Figure 2 The AEC system 1000 of FIG. 1000 may be implemented or performed similarly. The AEC system 1000 may also include a double talk detector (DTD) 1600 .
[0108] The double talk detector 1600 can detect whether double talk occurs by using a reference signal and a microphone signal. Double talk may indicate a situation where the voice of a far-end speaker and the voice of a near-end speaker are input together into the electronic device 10. For example, the double talk detector 1600 can detect double talk by comparing the amplitude of the reference signal with the amplitude of the microphone signal or based on the correlation between the reference signal and the microphone signal.
[0109] In an exemplary embodiment, when double talk is detected by the double talk detector 1600, the estimation methods of the residual echo estimators 1410b to 1410d may be changed. For example, the residual echoes at the current frame estimated by the residual echo estimators 1410b to 1410d, respectively, may be changed. The magnitudes of can be calculated by equations 14 to 16 respectively.
[0110] [Equation 14]
[0111]
[0112] [Equation 15]
[0113]
[0114] [Equation 16]
[0115]
[0116] In Equations 14 to 16, g(·) is a function associated with double talk. When double talk is detected, the amplitudes of the microphone signals Y(κ-t,,m) in the current frame and the past frame and the residual echo at the past frame are The amplitudes of both may not be used (may be deactivated) and may not be applied to the residual echo at the current frame When double talk is not detected, as described above, the residual echo estimators 1410a to 1410d may estimate the residual echo at the current frame calculated by Equations 1 and 11 to 13. The amplitude.
[0117] Figure 6 FIG. 1 shows a flow chart of estimating the amplitude of residual echo according to an exemplary embodiment of the inventive concept. Figures 1 to 5 To describe Figure 6 In the following description, the residual echo suppressor 1400 may indicate Figures 3A to 3D For any one of the residual echo suppressors 1400a to 1400d, the residual echo estimator of the residual echo suppressor 1400 may indicate Figures 3A to 3D Any one of the residual echo estimators 1410a to 1410d.
[0118] Reference Figures 1 to 6 In operation S110, the residual echo estimator 1410 may selectively use the magnitude of the linear echo X'(κ, m) at the current frame and the magnitude of the linear echo X'(κ-t, m) at the past frame, the residual echo at the past frame, and the residual echo at the past frame. The amplitude of the microphone signal Y(κ,m) at the current frame and the amplitude of the microphone signal Y(κ-t,m) at the past frame are used to estimate the residual echo at the current frame. The amplitude will be referred to Figure 7 and Figure 8 Operation S110 is described more fully.
[0119] In operation S120, the residual echo estimator 1410 may update the weight W selectively used in the estimating operation S110. R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t, m) so that the amplitude of the output signal E(κ, m) of the linear echo canceller 1300 is equal to the residual echo The difference between the magnitudes of is minimized. Equation 17 indicates the above difference.
[0120] [Equation 17]
[0121]
[0122] The residual echo estimator 1410 may update the weight W selectively used in the estimation of operation S110 by using the difference of Equation 17. R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 The residual echo estimator 1410 may use normalized least mean square error (NLMS), recursive least square (RLS) or Kalman filter to optimize the weight W R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t,m). For example, when using NLMS, the weight W R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t,,m) can be optimized as shown in Equation 18. In Equation 18, the weights can be updated in the direction of the arrows, and P1, P2, and P3 can be the linear echo X'(κ,m), the residual echo and the power smoothing of the microphone signal Y(κ,m).
[0123] [Equation 18]
[0124]
[0125]
[0126]
[0127]
[0128] For another example, in the case of using RLS, the weight W R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t, m) may be optimized as expressed in Equation 19 to Equation 21. In Equation 19, the weight may be updated in the direction of the arrow.
[0129] [Equation 19]
[0130] W R (i, j, k)←W R (i,j,k)+k1(κ,m)ξ(κ,m),
[0131] W T1(κ-t, m)←W T1 (κ-t,m)+k2(κ,m)ξ(κ,m), for 1≤t≤T,
[0132] W T2 (κ-t, m)←W T2 (κ-t,m)+k3(κ,m)ξ(κ,m), for 1≤t≤T,
[0133] W T3 (κ-t, m)←W T3 (κ-t,m)+k4(κ,m)ξ(κ,m), for 0≤t≤T
[0134] [Equation 20]
[0135]
[0136]
[0137]
[0138]
[0139] [Equation 21]
[0140]
[0141]
[0142]
[0143]
[0144] In an example embodiment, the residual echo estimator 1410 may implement an artificial neural network that updates the weights W using R (i, j, k), W T1 (t, m), W T2 (t, m) and W T3 (t, m), so that the amplitude of the output signal E(κ, m) of the linear echo canceller 1300 is equal to the residual echo The difference between the magnitudes of is reduced (or alternatively, minimized). The artificial neural network can learn the weights W based on the above adaptive algorithm. R (i, j, k), W T1 (t, m), W T2 (t, m) and W T3(t, m) such that the above difference is minimized. Since the amplitude of the output signal E(κ, m) of the linear echo canceller 1300 indicates the amplitude of the actual residual echo, the ideal value of the difference may be, for example, "0".
[0145] In operation S130, the first to third smoothing calculators 1431 to 1433 may respectively calculate the residual echo. A smoothing operation is performed on the amplitude of , the amplitude of the output signal E(κ, m) of the linear echo canceller 1300 and the amplitude of the noise N(κ, m).
[0146] In operation S140, the gain calculator 1440 may calculate a suppression gain G(κ,m) by using the smoothed magnitude of the residual echo, the smoothed magnitude of the output signal of the linear echo canceller 1300, and the smoothed magnitude of the noise, which is multiplied by the output signal E(κ,m) of the linear echo canceller 1300.
[0147] In operation S150 , the multiplier 1450 may generate an output signal R(κ,m) by multiplying the output signal E(κ,m) of the linear echo canceller 1300 by the suppression gain G(κ,m).
[0148] Figure 7 An example embodiment according to the inventive concept is shown. Figure 6 Detailed operations associated with operation S110.
[0149] Reference Figure 6 and Figure 7 In operation S111a, the residual echo estimator 1410 may select an acoustic echo cancellation (AEC) mode based on user input. For example, the user may represent various objects, such as the electronic device 10, an operating system (OS) supported by the application processor 100, a user of an application, a far-end speaker, and a near-end speaker. For example, the residual echo estimator 1410 may support five AEC modes. Of course, the number of AEC modes is not limited to the above description.
[0150] In case the first mode is selected, in operation S112a, the residual echo estimator 1410 may estimate the residual echo by using only the magnitude of the linear echo X'(κ,m) at the current frame (or alternatively, of the current frame). The first mode may be a default mode in which the time correlation of the linear echo, the time correlation of the residual echo, and the time correlation of the microphone signal may not be used to estimate the residual echo. The amplitude.
[0151] In case the second mode is selected, in operation S113a, the residual echo estimator 1410 may estimate the residual echo by using the amplitude of the linear echo X'(κ-t, m) at the past frame in addition to the amplitude of the linear echo X'(κ, m) at the current frame. The amplitude (refer to Figure 3A Residual echo estimator 1410a). Figure 3A As shown, when the second mode is selected, the time correlation of the linear echo can also be used to estimate the residual echo. Therefore, the degree to which the ERLE of the AEC system 1000 is improved when the second mode is selected may be greater than the degree to which the ERLE of the AEC system 1000 is improved when the first mode is selected.
[0152] In case the third mode is selected, in operation S114a, the residual echo estimator 1410 may further calculate the residual echo value by using the amplitude of the linear echo X'(κ-t, m) at the past frame and the amplitude of the residual echo X'(κ-t, m) at the past frame in addition to the amplitude of the linear echo X'(κ, m) at the current frame. The amplitude of the residual echo is used to estimate The amplitude (refer to Figure 3B Residual echo estimator 1410b). Figure 3B As shown, when the third mode is selected, the time correlation of linear echo and residual echo can also be used. Time correlation to estimate residual echo Therefore, the degree to which the ERLE of the AEC system 1000 is improved when the third mode is selected may be greater than the degree to which the ERLE of the AEC system 1000 is improved when the first mode or the second mode is selected.
[0153] In case the fourth mode is selected, in operation S115a, the residual echo estimator 1410 may estimate the residual echo by using the amplitude of the linear echo X'(κ, m) at the past frame and the amplitude of the microphone signal Y(κ, m) at the current frame and the amplitude of the microphone signal Y(κ-t, m) at the past frame in addition to the amplitude of the linear echo X'(κ, m) at the current frame. The amplitude (refer to Figure 3C Residual echo estimator 1410c). Figure 3C As shown, when the fourth mode is selected, the time correlation of the linear echo and the time correlation of the microphone signal can also be used to estimate the residual echo. Therefore, the degree to which the ERLE of the AEC system 1000 is improved when the fourth mode is selected may be greater than the degree to which the ERLE of the AEC system 1000 is improved when the first mode or the second mode is selected.
[0154] In case the fifth mode is selected, in operation S116a, the residual echo estimator 1410 may use the amplitude of the linear echo X'(κ,m) at the past frame, the amplitude of the residual echo X'(κ-t,m) at the past frame, and the amplitude of the linear echo X'(κ-t,m) at the current frame. The residual echo is estimated by taking the amplitude of the microphone signal Y(κ,m) at the current frame and the amplitude of the microphone signal Y(κ-t,m) at the past frame The amplitude (refer to Figure 3D Residual echo estimator 1410d). Figure 3D As shown, when the fifth mode is selected, the time correlation of the linear echo, the time correlation of the residual echo and the time correlation of the microphone signal can also be used to estimate the residual echo. Therefore, the degree to which the ERLE of the AEC system 1000 is improved when the fifth mode is selected may be greater than the degree to which the ERLE of the AEC system 1000 is improved when the first mode or the fourth mode is selected.
[0155] However, when one of the second to fifth modes is selected, the amount of operation required to update the weight and estimate the amplitude of the residual echo may increase compared to the amount of operation when the first mode is selected. Therefore, the user can appropriately select any of the first to fifth modes in consideration of the ERLE and hardware source of the electronic device 10.
[0156] Figure 8 Another exemplary embodiment according to the inventive concept is shown. Figure 6 Detailed operations associated with operation S110.
[0157] In operation S111b, the residual echo suppressor 1400 may select an AEC mode based on the echo return loss enhancement (ERLE) at the past frame. The residual echo suppressor 1400 may estimate the amplitude of the residual echo at the past frame based on the previously selected mode among the first mode to the fifth mode, and may evaluate the ERLE. The residual echo suppressor 1400 may compare the ERLE with a reference value, and may change the previously selected mode to another mode. For example, when the ERLE does not reach the reference value, the residual echo suppressor 1400 may select any one of the second mode to the fifth mode in which a higher ERLE may be expected. For another example, when the ERLE is significantly greater than the reference value, the residual echo suppressor 1400 may select any one of the second mode to the fifth mode to reduce the amount of operation. Operations S112b to S116b are similar to Figure 7 Operations S112a to S116a are substantially the same.
[0158] Figure 9An electronic device according to an example embodiment of the inventive concept is shown.
[0159] Reference Figure 9 , the electronic device 2000 may be Figure 1 . The electronic device 2000 may be implemented using a data processing device that may use or support an interface proposed by the Mobile Industry Processor Interface (MIPI) Alliance. For example, the electronic device 2000 may be a digital camera, a video camera, a smartphone, a tablet computer, a wearable device (e.g., a smart watch or smart bracelet), a smart speaker, an artificial intelligence based on a voice recognition device, a medical device, a navigation device, a home appliance, etc. The electronic device 2000 may be used for mobile communications, teleconferencing, etc. The electronic device 2000 may include any device that includes both a speaker 2001 and a microphone 2002.
[0160] A speaker 2001 and a microphone 2002 may be included to process sound information of the electronic device 2000. The speaker 2001 and the microphone 2002 may be reference Figure 1 The speaker 11 and the microphone 12 are described.
[0161] The electronic device 2000 can communicate with an external system or an external device (eg, Figure 1 The communication module 2003 can communicate with the electronic device 20). The communication module 2003 can support at least one of the following wireless communication protocols: Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMax), Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Bluetooth, Near Field Communication (NFC), Wireless Fidelity (Wi-Fi), and Radio Frequency Identification (RFID). The communication module 2003 can be embedded in the application processor 2100.
[0162] The electronic device 2000 may include a storage device 2004. The storage device 2004 may be connected to a reference Figure 4 The memory 200 described above corresponds to the auxiliary memory implemented independently of the memory 200, and the data stored in the memory 200 can be backed up to the storage device 2004. Regardless of whether the power is supplied, the storage device 2004 can store or back up the weight W R (i,j,k),W T1 (t,m),W T2 (t,m) and W T3 (t, m) and program code for executing the AEC system 1000. The storage device 2004 may include a non-volatile memory. The storage device 2004 may be a device such as a secure digital (SD) card or an embedded multimedia card (eMMC).
[0163] The display 2005 of the electronic device 2000 can perform interaction with the user under the control of the application processor 2100. For example, the user can activate the above-mentioned AEC system 1000 through the display 2005. Specifically, the user can activate the AEC system 1000 and select each mode ( Figure 7 and Figure 8 The AEC system 1000 may estimate the amplitude of the residual echo based on the first mode. For example, if the user selects the enhanced mode, the AEC system 1000 may estimate the amplitude of the residual echo based on any one of the second mode to the fifth mode. Of course, in addition to the display 2005, the user may select any mode of the AEC system 1000 through any other input / output method of the electronic device 2000.
[0164] The application processor 2100 may control the overall operation of the electronic device 2000. The application processor 2100 may be implemented in the form of a system on a chip (SoC). The application processor 2100 may include an audio processor 2110, a main processor 2120, a memory 2130, and an input / output interface 2140. The audio processor 2110 may be the audio processor 110 described above, and the memory 2130 may be the memory 120 described above.
[0165] The main processor 2120 can control the overall operation of the application processor 2100 independently of the audio processor 2110. The main processor 2120 can load various software supported by the application processor 2100 (for example, application programs, operating systems, and device drivers) onto the memory 2130 and execute them. The main processor 2120 may include one or more central processing units or one or more graphics processing units (i.e., one or more cores).
[0166] Software, programs, or program codes that can be executed by the main processor 2120 can be loaded onto the memory 2130. For example, the programs or program codes may include a kernel, middleware, an application programming interface (API), and application programs AP1 to AP4. At least a portion of the kernel, middleware, and API may be referred to as an "OS." The kernel may manage resources (e.g., memory 2130 or storage 2004) used to perform operations or functions through any other programs (e.g., middleware, API, and application programs AP1 to AP4). The middleware may play an intermediary role for exchanging data between the API or application programs AP1 to AP4 and the kernel. The API may be an interface for application programs AP1 to AP4 to control functions provided by the kernel or middleware.
[0167] The number of applications AP1 to AP4 is not limited to Figure 9 As described above, the program code for executing the AEC system 1000 may be stored in the memory 2130, and the AEC system 1000 may be loaded onto the memory 2130 and executed. For example, any one of the application programs AP1 to AP4 associated with voice communication may select any one of the modes ( Figure 7 and Figure 8 The program code capable of selecting the mode selected by the AEC system 1000 may be the same as or different from the program code for executing the AEC system 1000 .
[0168] The input / output interface 2140 may perform an interface operation for data exchange between the application processor 2100 and any other components of the electronic device 2000. For example, data stored in the memory 2130 may be transferred or backed up through the input / output interface 2140.
[0169] According to an exemplary embodiment of the inventive concept, the time correlation of the linear echo, the time correlation of the residual echo, and the time correlation of the microphone signal can be selectively used to calculate the amplitude of the residual echo. Therefore, the ERLE of the AEC system can also be improved.
[0170] Although example embodiments of the inventive concept have been described with reference to a few example embodiments of the inventive concept, it will be apparent to those skilled in the art that various changes and modifications can be made thereto without departing from the spirit and scope of the inventive concept as set forth in the appended claims.
Claims
1. A residual echo estimator, configured to estimate a residual echo of a microphone signal, the residual echo estimator comprising: processing circuitry configured to: estimating the amplitude of the residual echo at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at the current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, and a weighted sum of an amplitude of a linear echo of the reference signal at a past frame and the weight applied to the amplitude of the linear echo of the reference signal at the past frame, wherein the reference signal is a far-end signal corresponding to the voice of the far-end talker; and A weight applied to the magnitude of the linear echo at a current frame and a weight applied to the magnitude of the linear echo at a past frame are updated.
2. The residual echo estimator according to claim 1, wherein The processing circuit is configured to: The weight applied to the magnitude of the linear echo at a current frame and the weight applied to the magnitude of the linear echo at a past frame are updated based on a difference between an actual magnitude of the residual echo and the magnitude of the residual echo.
3. The residual echo estimator according to claim 1, wherein The amplitude of the residual echo is a value obtained by: multiplying the amplitude of the linear echo at the current frame by a weight associated with the amplitude of the linear echo at the current frame to generate a first value, multiplying the magnitude of the linear echo at the past frame by a weight associated with the magnitude of the linear echo at the past frame to generate a second value, and The first value and the second value are added to generate the value of the amplitude of the residual echo.
4. A residual echo estimator, configured to estimate a residual echo of a microphone signal, the residual echo estimator comprising: processing circuitry configured to: estimating the amplitude of the residual echo at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at a current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, a weighted sum of an amplitude of the linear echo of the reference signal at a past frame and a weight applied to the amplitude of the linear echo of the reference signal at the past frame, a weighted sum of an amplitude of the microphone signal at the current frame and a weight applied to the amplitude of the microphone signal at the current frame, and a weighted sum of an amplitude of the microphone signal at a past frame and a weight applied to the amplitude of the microphone signal at the past frame, wherein the reference signal is a far-end signal corresponding to the voice of a far-end talker; and A weight applied to the amplitude of the linear echo at a current frame, a weight applied to the amplitude of the linear echo at a past frame, a weight applied to the amplitude of the microphone signal at a current frame, and a weight applied to the amplitude of the microphone signal at a past frame are updated.
5. A residual echo estimator, configured to estimate a residual echo of a microphone signal, the residual echo estimator comprising: processing circuitry configured to: estimating the amplitude of the residual echo at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at the current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, a weighted sum of an amplitude of the linear echo of the reference signal at a past frame and a weight applied to the amplitude of the linear echo of the reference signal at the past frame, and a weighted sum of an amplitude of the residual echo at the past frame and the weight applied to the amplitude of the residual echo at the past frame, wherein the reference signal is a far-end signal corresponding to the voice of the far-end talker; and A weight applied to the magnitude of the linear echo at a current frame, a weight applied to the magnitude of the linear echo at a past frame, and a weight applied to the magnitude of the residual echo at a past frame are updated.
6. A non-transitory computer-readable medium comprising program code executable by a processor to: estimating the amplitude of the residual echo of the microphone signal at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at the current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, and a weighted sum of an amplitude of a linear echo of the reference signal at a past frame and the weight applied to the amplitude of the linear echo of the reference signal at the past frame, the reference signal being a far-end signal corresponding to the voice of the far-end talker, updating the weight applied to the amplitude of the linear echo at the current frame and the weight applied to the amplitude of the linear echo at the past frame, and A suppression gain is calculated based on the magnitude of the residual echo and the magnitude of an output signal obtained by canceling the linear echo from the microphone signal, and the output signal is multiplied by the suppression gain to generate a final output signal.
7. The non-transitory computer-readable medium of claim 6, wherein: When executing the program code, the processor is configured to: The amplitude of the linear echo of the reference signal at a past frame is stored in a memory, which is a tightly coupled memory, a static random access memory, or a dynamic random access memory.
8. The non-transitory computer-readable medium of claim 6, wherein: When executing the program code, the processor is configured to: The weight applied to the magnitude of the linear echo at a current frame and the weight applied to the magnitude of the linear echo at a past frame are updated based on a difference between a magnitude of the output signal and a magnitude of the residual echo.
9. The non-transitory computer-readable medium of claim 6, wherein: When executing the program code, the processor is configured to: An artificial neural network is executed that updates the weight applied to the amplitude of the linear echo at a current frame and the weight applied to the amplitude of the linear echo at a past frame so that a difference between the amplitude of the output signal and the amplitude of the residual echo decreases.
10. The non-transitory computer-readable medium of claim 6, wherein: When executing the program code, the processor is configured to: executing a noise estimator that estimates the magnitude of noise by using the magnitude of the output signal, and The suppression gain is further calculated based on the amplitude of the noise.
11. The non-transitory computer-readable medium of claim 6, wherein: When executing the program code, the processor is configured to: executing a smoothing calculator that smoothes the amplitude of the residual echo and the amplitude of the output signal to generate a smoothed amplitude of the residual echo and a smoothed amplitude of the output signal, respectively; and The suppression gain is calculated based on a smoothed amplitude of the residual echo and a smoothed amplitude of the output signal.
12. A non-transitory computer readable medium comprising program code executable by a processor to: estimating the amplitude of the residual echo at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at a current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, a weighted sum of an amplitude of the linear echo of the reference signal at a past frame and a weight applied to the amplitude of the linear echo of the reference signal at the past frame, a weighted sum of an amplitude of a microphone signal at the current frame and a weight applied to the amplitude of the microphone signal at the current frame, and a weighted sum of an amplitude of the microphone signal at a past frame and a weight applied to the amplitude of the microphone signal at the past frame, the reference signal being a far-end signal corresponding to the voice of a far-end talker, updating a weight applied to the amplitude of the linear echo at a current frame, a weight applied to the amplitude of the linear echo at a past frame, a weight applied to the amplitude of the microphone signal at a current frame, and a weight applied to the amplitude of the microphone signal at a past frame, and A suppression gain is calculated based on the magnitude of the residual echo and the magnitude of an output signal obtained by canceling the linear echo from the microphone signal, and the output signal is multiplied by the suppression gain to generate a final output signal.
13. A non-transitory computer readable medium comprising program code executable by a processor to: estimating the amplitude of the residual echo at the current frame based on a weighted sum of an amplitude of a linear echo of a reference signal at the current frame and a weight applied to the amplitude of the linear echo of the reference signal at the current frame, a weighted sum of an amplitude of the linear echo of the reference signal at a past frame and a weight applied to the amplitude of the linear echo of the reference signal at the past frame, and an amplitude of a residual echo at a past frame and a weighted sum of the amplitude of the residual echo at the past frame, wherein the reference signal is a far-end signal corresponding to the voice of the far-end talker. updating a weight applied to the amplitude of the linear echo at a current frame, a weight applied to the amplitude of the linear echo at a past frame, and a weight applied to the amplitude of the residual echo at a past frame, and A suppression gain is calculated based on the magnitude of the residual echo and the magnitude of an output signal obtained by canceling the linear echo from a microphone signal, and the output signal is multiplied by the suppression gain to generate a final output signal.
14. An application processor, comprising: Audio processor; as well as A non-transitory computer-readable medium configured to store program code executable by the audio processor to: generating an output signal by canceling a linear echo of a reference signal input to a microphone as the reference signal is output from a speaker, the linear echo being determined based on a transfer path between the speaker and the microphone, and The magnitude of the residual echo at the current frame is estimated based on at least one of a weighted sum of the magnitude of the linear echo at the current frame and a weighted sum of the magnitude of the linear echo at a past frame and the weighted sum of the magnitude of the linear echo at the past frame.
15. The application processor according to claim 14, wherein: When executing the program code, the audio processor is configured to: in a first mode, estimating the amplitude of the residual echo at the current frame based on a weighted sum of the amplitude of the linear echo at the current frame and the weight applied to the amplitude of the linear echo of the reference signal at the current frame instead of the weighted sum of the amplitude of the linear echo at the past frame and the weight applied to the amplitude of the linear echo of the reference signal at the past frame, In the second mode, the amplitude of the residual echo at the current frame is estimated based on a weighted sum of the amplitude of the linear echo at the current frame and the weight applied to the amplitude of the linear echo of the reference signal at the current frame, and a weighted sum of the amplitude of the linear echo at a past frame and the weight applied to the amplitude of the linear echo of the reference signal at the past frame.
16. The application processor according to claim 14, wherein: The program code is a first program code, and the non-transitory computer-readable medium further stores a second program code, and the application processor further includes: a main processor configured to execute the second program code to determine whether a weighted sum of the amplitude of the linear echo at a past frame and a weight applied to the amplitude of the linear echo of the reference signal at the past frame is used to estimate the amplitude of the residual echo at a current frame.
17. The application processor according to claim 14, wherein: When executing the program code, the audio processor is configured to: calculating a suppression gain based on the amplitude of the residual echo and the amplitude of the output signal, and A final output signal is generated by multiplying the output signal by the suppression gain.
18. The application processor according to claim 14, wherein: When executing the program code, the audio processor is configured to: A weight applied to the amplitude of the linear echo at a current frame and a weight applied to the amplitude of the linear echo at a past frame are adaptively updated so that a difference between the amplitude of the output signal and the amplitude of the residual echo is minimized.
19. The application processor according to claim 14, wherein: When executing the program code, the audio processor is configured to: The amplitude of the linear echo of the reference signal at a past frame is stored in a memory, which is a tightly coupled memory, a static random access memory, or a dynamic random access memory.
20. An application processor, comprising: an audio processor configured to detect the presence of a two-way conversation when executing the program code; as well as A non-transitory computer-readable medium configured to store the program code to: generating an output signal by canceling a linear echo of a reference signal input to a microphone as the reference signal is output from a speaker, the linear echo being determined based on a transfer path between the speaker and the microphone, and In response to the audio processor not detecting the double talk, estimating the amplitude of the residual echo at the current frame based on a weighted sum of the amplitude of the linear echo at the current frame and the weight applied to the amplitude of the linear echo of the reference signal at the current frame, the weighted sum of the amplitude of the linear echo at the past frame and the weight applied to the amplitude of the linear echo of the reference signal at the past frame, the weighted sum of the amplitude of the residual echo at the past frame and the weight applied to the amplitude of the residual echo at the past frame, the weighted sum of the amplitude of the microphone signal at the current frame and the weight applied to the amplitude of the microphone signal at the current frame, and the weighted sum of the amplitude of the microphone signal at the past frame and the weight applied to the amplitude of the microphone signal at the past frame.
Citation Information
Patent Citations
An Apparatus for Heating a Gas
KR1020190011091A
Combined suppression of noise and out-of-location signals
CN103348408A
Acoustic echo mitigating apparatus and method, audio processing apparatus, and voice communication terminal
CN104050971A