Echo cancellation device, echo cancellation method, and storage medium

By using linear and nonlinear echo suppression units in the echo suppression device, the nonlinear echo signal is estimated and suppressed by using the nonlinear echo model, the problem of difficulty in suppressing nonlinear echo signals in the prior art is solved, and a more stable echo suppression effect is achieved.

CN112863532BActive Publication Date: 2025-06-27PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011254315.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-20
Filing Date
2020-11-11
Publication Date
2025-06-27
Estimated Expiration
2040-11-11

AI Technical Summary

Technical Problem

The prior art is difficult to stably suppress the nonlinear echo signal contained in the input signal acquired through the microphone.

Method used

An echo suppression device including a first linear echo suppression unit, a nonlinear echo estimation unit, a nonlinear echo suppression unit, and a second linear echo suppression unit is adopted. The device suppresses the nonlinear echo signal from the input signal by estimating the linear and nonlinear echo signals in the input signal, and suppresses the residual linear echo signal through the second linear echo suppression section.

Benefits of technology

It is possible to stably suppress the nonlinear echo signal contained in the input signal acquired through the microphone, improve the suppression performance of the linear echo signal, and ensure the sound quality during the call.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112863532B_ABST
    Figure CN112863532B_ABST
Patent Text Reader

Abstract

Provided are an echo suppression device, an echo suppression method, and a storage medium. The echo suppression device according to the present invention includes: a nonlinear echo estimator that uses a nonlinear echo model representing a relationship between at least one of a call reception signal output to a speaker and the input signal obtained from a microphone and a nonlinear echo signal to estimate a nonlinear echo signal included in the input signal based on at least one of the call reception signal and the input signal; a nonlinear echo suppressor that suppresses the nonlinear echo signal from the output signal of the echo canceller using the estimated nonlinear echo signal; and an echo suppressor that suppresses a residual linear echo signal not suppressed by the echo canceller from the output signal of the nonlinear echo suppressor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for suppressing a linear echo signal and a non-linear echo signal included in an input signal acquired by a microphone. Background Art

[0002] In a hands-free call system, a video conferencing system, etc., in the case of performing an amplified call using a speaker and a microphone, the voice of a speaker on the call transmitting side is input to the microphone on the call transmitting side, and is transmitted as a call transmission signal via a network line to the device on the call receiving side. The voice amplified from the speaker on the call receiving side is picked up by the microphone on the call receiving side and transmitted via the network line to the device on the call transmitting side. At this time, the voice of the user himself / herself that has passed through the time of the network line and the time of spatial transmission on the call receiving side is reproduced from the speaker on the call transmitting side. In this way, the voice transmitted between the speaker and the microphone on the call receiving side is called an echo and becomes a main factor hindering the call. For this reason, echo suppression techniques such as an echo canceller and an echo suppressor have been proposed.

[0003] For example, in the echo suppression device disclosed in Japanese Patent Laid-Open Publication No. 2017-191992, when reproducing a call reception signal by a speaker, in the case where distortion may occur in the reproduced sound due to an excessive level of the call reception signal, a gain that is larger than the gain used in the case where no distortion occurs is obtained for each frequency, and the gain is multiplied by the value of the received sound signal based on the frequency domain.

[0004] Moreover, for example, in the echo suppression device disclosed in Japanese Patent Laid-Open Publication No. 2010-103875, when the power of the reproduced signal at a certain frequency value is greater than a predetermined threshold, and in the case where the frequency value is a frequency value that is m (m = 2, 3,..., M) times or a frequency value in the vicinity of the frequency value that is m times, a value that makes the gain coefficient corresponding to the frequency value that is m times and the frequency values in its vicinity approach 0 is obtained as the second gain coefficient, and in other cases, the gain coefficient is obtained as the second gain coefficient.

[0005] However, in the above prior art, it is difficult to stably suppress the non-linear echo signal included in the input signal acquired by the microphone, and further improvement is required. Summary of the Invention

[0006] The present invention has been made to solve the above problems, and an object thereof is to provide a technique capable of stably suppressing a non-linear echo signal included in an input signal acquired by a microphone.

[0007] One embodiment of the echo suppression device according to the present invention includes: a first linear echo suppression unit that suppresses a linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by the microphone; a non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal by using a non-linear echo model representing the relationship between at least one of the call reception signal and the input signal and the non-linear echo signal; a non-linear echo suppression unit that suppresses the non-linear echo signal from the output signal of the first linear echo suppression unit by using the non-linear echo signal estimated by the non-linear echo estimation unit; and a second linear echo suppression unit that suppresses the residual linear echo signal from the output signal of the non-linear echo suppression unit by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppression unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 It is a schematic diagram showing a microphone signal, an echo canceller output signal, and an echo suppressor output signal in the case where the input signal does not include a non-linear echo caused by speaker distortion.

[0009] Figure 2 It is a schematic diagram showing a microphone signal, an echo canceller output signal, and an echo suppressor output signal in the case where the input signal includes a non-linear echo caused by speaker distortion.

[0010] Figure 3 It is a schematic diagram showing the configuration of a call device according to the first embodiment of the present invention.

[0011] Figure 4 It is a schematic diagram showing an example of signals output from each unit of the echo suppression device according to the first embodiment.

[0012] Figure 5 It is a flowchart for explaining the operation of the echo suppression device according to the first embodiment of the present invention.

[0013] Figure 6 It is a schematic diagram showing the configuration of a learning device according to the first embodiment of the present invention.

[0014] Figure 7 It is a schematic diagram showing an example of signals output from each unit of the learning device according to the first embodiment.

[0015] Figure 8It is a schematic diagram showing the amplitude spectrum of a call reception signal containing 1 / 3 octave band noise.

[0016] Figure 9 It is a schematic diagram showing Figure 8 the correct value and the estimated value of the amplitude spectrum of the non-linear echo signal contained in the input signal obtained by the microphone when the call reception signal shown is amplified.

[0017] Figure 10 It is a schematic diagram showing the amplitude spectrum of a call reception signal containing a female voice.

[0018] Figure 11 It is a schematic diagram showing Figure 10 the correct value and the estimated value of the amplitude spectrum of the non-linear echo signal contained in the input signal obtained by the microphone when the call reception signal shown is amplified.

[0019] Figure 12 It is a schematic diagram showing the result of frequency analysis of the output signal from an existing echo cancellation device and the output signal from the echo cancellation device of the first embodiment.

[0020] Figure 13 It is a schematic diagram showing the time variation of the amplitude of the input signal containing a male voice and the time variation of the echo cancellation amount (ERLE) with respect to the input signal.

[0021] Figure 14 It is a schematic diagram showing the configuration of the call device of the second embodiment of the present invention.

[0022] Figure 15 It is a schematic diagram showing the configuration of the call device of the third embodiment of the present invention.

[0023] Figure 16 It is a schematic diagram showing the configuration of the call device of the fourth embodiment of the present invention.

[0024] Figure 17 It is a schematic diagram showing the configuration of the call device of the fifth embodiment of the present invention.

[0025] Figure 18 It is a schematic diagram showing the configuration of the call device of the sixth embodiment of the present invention.

[0026] Figure 19 It is a schematic diagram showing the configuration of the call device of the seventh embodiment of the present invention. Detailed Embodiments

[0027] (Basic Knowledge of the Present Invention)

[0028] Echo cancellation is a technique that estimates an echo signal through an adaptive filter and cancels the echo by subtracting the estimated echo signal from the signal received by a microphone. An echo is a superposition of the direct sound and the reflected sound of the sound amplified by a speaker. For this reason, the transmission characteristics between the speaker and the microphone can be represented by an FIR (Finite Impulse Response) filter. The FIR-type adaptive filter learns in a way that approximates the transmission characteristics, and generates an estimated value of the echo, that is, a pseudo echo signal, by convolving the filter coefficients with the received call signal. As learning algorithms for the adaptive filter, there are methods based on the LMS (Least Mean Square) method, the NLMS (Normalized LMS) method, and the ICA (Independent Component Analysis), etc.

[0029] On the other hand, echo suppression is a technique that estimates the power spectrum of an echo in the frequency domain and suppresses the echo by subtracting the estimated power spectrum of the echo from the power spectrum of the signal received by a microphone. An echo suppressor suppresses an echo, for example, by the spectrum subtraction method or the wiener filtering method. Since the learning of the above-mentioned echo cancellation adaptive filter takes time, residual echoes may occur when the power is just turned on and when the echo path changes. Moreover, noise generated by the speaker or the microphone or the transmitted voice signal may cause incorrect learning of the adaptive filter, and estimation errors may occur in the pseudo echo signal, increasing the possibility of residual echoes. For this reason, an echo suppressor is usually used at a later stage of the echo canceller for the purpose of compensating for echo suppression.

[0030] Since existing echo cancellers and existing echo suppressors are used to estimate an echo in a linear model, there is a problem that it is difficult to suppress non-linear echoes mixed with non-linear noise such as speaker distortion. In devices used for laptop computers or portable video conferencing systems, since a large volume of sound is played from a small-diameter speaker, the influence of non-linear echoes caused by speaker distortion is significant, and it may not be possible to have a comfortable conversation.

[0031] Moreover, in the technique disclosed in Japanese Patent Laid-Open Publication No. 2017-191992 mentioned above, it is difficult to suppress non-linear echo signals of frequency components not included in the received call signal such as harmonic distortion.

[0032] Moreover, in the technology disclosed in Japanese Patent Laid-Open Publication No. 2010-103875, it is difficult to suppress broadband distortion components, and it is also difficult to suppress distortion components generated at frequency values other than integer multiples of the frequency value.

[0033] To solve the above problems, an echo cancellation device according to an embodiment of the present invention includes: a first linear echo cancellation unit that suppresses a linear echo signal from the input signal by estimating an amplitude component and a phase component of the linear echo signal included in the input signal acquired by a microphone; a non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of a call reception signal output to a speaker and the input signal, using a non-linear echo model representing a relationship between at least one of the call reception signal and the input signal and the non-linear echo signal; a non-linear echo cancellation unit that suppresses the non-linear echo signal from the output signal of the first linear echo cancellation unit, using the non-linear echo signal estimated by the non-linear echo estimation unit; and a second linear echo cancellation unit that suppresses the residual linear echo signal from the output signal of the non-linear echo cancellation unit by estimating an amplitude component of the residual linear echo signal not suppressed by the first linear echo cancellation unit.

[0034] According to this configuration, a non-linear echo signal included in the input signal acquired by the microphone can be stably suppressed by using a non-linear echo model representing a relationship between at least one of a call reception signal output to a speaker and the input signal and the non-linear echo signal, estimating the non-linear echo signal included in the input signal based on at least one of the call reception signal and the input signal, and suppressing the non-linear echo signal from the output signal of the first linear echo cancellation unit using the estimated non-linear echo signal.

[0035] Moreover, the second linear echo cancellation unit suppresses the residual linear echo signal from the output signal in which the non-linear echo signal has been suppressed. Therefore, the second linear echo cancellation unit can operate stably, and the suppression performance of the linear echo signal can be improved.

[0036] Moreover, in the echo cancellation device, the non-linear echo model may be learned using teacher data, with at least one of the call reception signal and the input signal as an input and the non-linear echo signal as an output, where at least one of the call reception signal and the input signal and the output signal of the second linear echo cancellation unit that suppresses the residual linear echo signal from the output signal of the first linear echo cancellation unit are used as the teacher data, and the first linear echo cancellation unit suppresses the linear echo signal from the input signal.

[0037] According to this configuration, since the first linear echo suppression unit and the second linear echo suppression unit only suppress linear echo signals and do not suppress non-linear echo signals, the signal obtained by suppressing the linear echo signal through the first linear echo suppression unit and the second linear echo suppression unit can be used as a non-linear echo signal for teacher data.

[0038] Moreover, since at least one of the call reception signal and the input signal, and the output signal of the second linear echo suppression unit are used as teacher data to learn the non-linear echo signal, the complex distortion caused by the speaker can be correctly modeled, and the estimation accuracy of the non-linear echo signal can be improved.

[0039] Moreover, in the echo suppression device described above, the non-linear echo model can also be a neural network.

[0040] According to this configuration, the non-linear echo model can be implemented by a neural network.

[0041] Moreover, in the echo suppression device described above, it can also be that the non-linear echo estimation unit uses the non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal according to the call reception signal.

[0042] According to this configuration, since the non-linear echo signal is estimated according to the call reception signal by using the non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal, the non-linear echo signal can be easily estimated according to the call reception signal.

[0043] Moreover, in the echo suppression device described above, it can also be that the non-linear echo estimation unit uses the non-linear echo model representing the relationship between the call reception signal, the input signal and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal according to the call reception signal and the input signal.

[0044] According to this configuration, since the non-linear echo signal is estimated not according to the call reception signal but according to the call reception signal and the input signal, the estimation accuracy of the non-linear echo signal can be improved.

[0045] Moreover, in the echo suppression device described above, it can also be that the non-linear echo estimation unit uses the non-linear echo model representing the relationship between the call reception signal and the output signal of the first linear echo suppression unit and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal according to the call reception signal and the output signal of the first linear echo suppression unit.

[0046] According to this configuration, since the non-linear echo signal is estimated based on the call reception signal and the output signal of the first linear echo canceller instead of based on the call reception signal alone, the estimation accuracy of the non-linear echo signal can be improved.

[0047] Moreover, in the echo canceller described above, it is also possible that the first linear echo canceller includes an adaptive filter that generates an analog linear echo signal representing the component of the call reception signal included in the input signal by convolving a filter coefficient with the call reception signal, and a subtraction unit that subtracts the analog linear echo signal from the input signal. The non-linear echo estimator estimates the non-linear echo signal included in the input signal based on the call reception signal and the analog linear echo signal from the adaptive filter, using the non-linear echo model representing the relationship between the call reception signal, the analog linear echo signal from the adaptive filter, and the non-linear echo signal.

[0048] According to this configuration, since the non-linear echo signal is estimated based on the call reception signal and the analog linear echo signal from the adaptive filter of the first linear echo canceller instead of based on the call reception signal alone, the estimation accuracy of the non-linear echo signal can be improved.

[0049] Moreover, in the echo canceller described above, it is also possible that the non-linear echo estimator estimates the non-linear echo signal included in the input signal based on the input signal, using the non-linear echo model representing the relationship between the input signal and the non-linear echo signal.

[0050] According to this configuration, since the non-linear echo signal is estimated based on the input signal using the non-linear echo model representing the relationship between the input signal and the non-linear echo signal, the non-linear echo signal can be easily estimated based on the call reception signal.

[0051] Moreover, in the echo canceller described above, it may further include a correction unit that calculates a variable gain for minimizing any one of the output signal of the non-linear echo canceller and the output signal of the second linear echo canceller, and corrects the non-linear echo signal estimated by the non-linear echo estimator using the calculated variable gain.

[0052] According to this configuration, since the estimation error of the non-linear echo signal is corrected, the suppression performance of the non-linear echo signal can be improved.

[0053] Another embodiment of the present invention relates to an echo cancellation device, comprising: a first linear echo cancellation unit that suppresses a linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by a microphone; a non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to a speaker and the input signal; a non-linear echo cancellation unit that suppresses the non-linear echo signal from the input signal by using the non-linear echo signal estimated by the non-linear echo estimation unit; and a second linear echo cancellation unit that suppresses the residual linear echo signal by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo cancellation unit.

[0054] According to this configuration, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal output to the speaker and the input signal, and the non-linear echo signal is used to suppress the non-linear echo signal from the input signal. Therefore, the non-linear echo signal included in the input signal acquired by the microphone can be stably suppressed.

[0055] Moreover, the residual linear echo signal is suppressed by the second linear echo cancellation unit. Therefore, the second linear echo cancellation unit can operate stably, and the suppression performance of the linear echo signal can be improved.

[0056] Another embodiment of the present invention relates to an echo cancellation method, in which the first linear echo cancellation unit suppresses the linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by the microphone; the non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal by using a non-linear echo model representing the relationship between at least one of the call reception signal output to the speaker and the input signal and the non-linear echo signal; the non-linear echo cancellation unit suppresses the non-linear echo signal from the output signal of the first linear echo cancellation unit by using the non-linear echo signal estimated by the non-linear echo estimation unit; and the second linear echo cancellation unit suppresses the residual linear echo signal from the output signal of the non-linear echo cancellation unit by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo cancellation unit.

[0057] According to this configuration, using a non-linear echo model that represents the relationship between at least one of the call reception signal output to the speaker and the input signal and the non-linear echo signal, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal and the input signal. Using the estimated non-linear echo signal, the non-linear echo signal is suppressed from the output signal of the first linear echo suppression unit. Therefore, the non-linear echo signal included in the input signal obtained through the microphone can be stably suppressed.

[0058] Moreover, through the second linear echo suppression unit, the residual linear echo signal is suppressed from the output signal from which the non-linear echo signal has been suppressed. Therefore, the second linear echo suppression unit can operate stably, and the suppression performance of the linear echo signal can be improved.

[0059] An echo suppression method according to another embodiment of the present invention causes the first linear echo suppression unit to suppress the linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal obtained through the microphone; causes the non-linear echo estimation unit to estimate the non-linear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal; causes the non-linear echo suppression unit to suppress the non-linear echo signal from the input signal using the non-linear echo signal estimated by the non-linear echo estimation unit; and causes the second linear echo suppression unit to suppress the residual linear echo signal by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppression unit.

[0060] According to this configuration, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal output to the speaker and the input signal, and the non-linear echo signal is suppressed from the input signal using the estimated non-linear echo signal. Therefore, the non-linear echo signal included in the input signal obtained through the microphone can be stably suppressed.

[0061] Moreover, the residual linear echo signal is suppressed by the second linear echo suppression unit. Therefore, the second linear echo suppression unit can operate stably, and the suppression performance of the linear echo signal can be improved.

[0062] Another embodiment of the present invention relates to a storage medium storing an echo cancellation program, which is a non-transitory computer-readable storage medium that causes a computer to function as follows: a first linear echo cancellation unit that suppresses a linear echo signal from the input signal by estimating an amplitude component and a phase component of the linear echo signal included in the input signal acquired by a microphone; a non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of a call reception signal output to a speaker and the input signal, using a non-linear echo model representing a relationship between at least one of the call reception signal and the input signal and the non-linear echo signal; a non-linear echo cancellation unit that suppresses the non-linear echo signal from the output signal of the first linear echo cancellation unit, using the non-linear echo signal estimated by the non-linear echo estimation unit; and a second linear echo cancellation unit that suppresses the residual linear echo signal from the output signal of the non-linear echo cancellation unit by estimating an amplitude component of the residual linear echo signal not suppressed by the first linear echo cancellation unit.

[0063] According to this configuration, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal output to the speaker and the input signal, using a non-linear echo model representing a relationship between at least one of the call reception signal and the input signal and the non-linear echo signal, and the non-linear echo signal is suppressed from the output signal of the first linear echo cancellation unit using the estimated non-linear echo signal. Therefore, the non-linear echo signal included in the input signal acquired by the microphone can be stably suppressed.

[0064] Moreover, the residual linear echo signal is suppressed from the output signal from which the non-linear echo signal has been suppressed by the second linear echo cancellation unit. Therefore, the second linear echo cancellation unit can stably operate, and the suppression performance of the linear echo signal can be improved.

[0065] Another embodiment of the present invention relates to a storage medium storing an echo cancellation program, which is a non-transitory computer-readable storage medium that causes a computer to function as follows: a first linear echo cancellation unit that suppresses a linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by a microphone; a non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to a speaker and the input signal; a non-linear echo cancellation unit that suppresses the non-linear echo signal from the input signal using the non-linear echo signal estimated by the non-linear echo estimation unit; and a second linear echo cancellation unit that suppresses the residual linear echo signal by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo cancellation unit.

[0066] According to this configuration, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal output to the speaker and the input signal, and the non-linear echo signal is suppressed from the input signal using the estimated non-linear echo signal. Therefore, the non-linear echo signal included in the input signal acquired by the microphone can be stably suppressed.

[0067] Moreover, the residual linear echo signal is suppressed by the second linear echo cancellation unit. Therefore, the second linear echo cancellation unit can operate stably, and the suppression performance of the linear echo signal can be improved.

[0068] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In addition, the embodiments described below are all specific examples of the present invention and do not limit the technical protection scope of the present invention.

[0069] (First Embodiment)

[0070] First, the main factors causing non-linear echo will be described.

[0071] Non-linear distortion is a general term for distortion that occurs when the input / output relationship of a system is not a proportional relationship. For example, if a two-tone sine wave of frequencies f1 and f2 is input to a system having an input / output characteristic in which the output amplitude is clipped as the input amplitude increases, non-linear distortion will occur in the amplitude spectrum of the output waveform at frequency components that do not exist in the input signal. Non-linear distortion can be roughly classified into harmonic distortion and intermodulation distortion. Harmonic distortion occurs at frequencies that are integer multiples of the input signal, such as 2f1 and 2f2, and intermodulation distortion occurs at frequencies that are the sum and difference of the input signal, such as f1 + f2 and f2 - f1.

[0072] In an actual system, the nonlinear distortion of the loudspeaker's sound amplification is the main factor of nonlinear echo. In an electrodynamic speaker that is generally widely used, the displacement of the diaphragm increases in the frequency domain near the lowest resonance frequency f0. Moreover, nonlinear distortion is generated due to the nonlinearity of the driving force generated by the voice coil operating outside the range of the magnetic flux formed by the permanent magnet, or due to the mechanical nonlinearity of the support system such as the cone edge or baffle. In addition, for small-diameter speakers, in order to compensate for the reduction of the sound pressure level in the low-frequency range, sometimes the sound pressure near the minimum resonance frequency f0 is increased through preprocessing. In this case, the displacement of the diaphragm increases, becoming the main factor for further generating nonlinear distortion.

[0073] Next, the impact of nonlinear echo on existing echo cancellation technologies will be described. As an existing echo cancellation technology, a system equipped with an echo canceller and an echo suppressor will be described.

[0074] The echo canceller calculates the estimated value of the echo, that is, the analog echo signal, through an adaptive filter, and cancels the echo by subtracting the calculated analog echo signal from the microphone signal. That is, assuming that the call reception signal is x(k), the coefficient of the adaptive filter is w n (k), the number of taps of the adaptive filter is N, and the analog echo y(k) can be expressed by the following formula (1).

[0075]

[0076] The above formula (1) means that the analog echo is represented by the linear sum of the phase and amplitude changes of the call reception signal, and the nonlinear echo cannot be represented regardless of the adaptive algorithm used for coefficient learning.

[0077] Moreover, the echo suppressor is set in the later stage of the echo canceller. The echo suppressor suppresses the residual echo by estimating the power spectrum of the residual echo that cannot be suppressed by the echo canceller. In an echo suppressor based on the generally widely used Wiener filtering method, the acoustic coupling amount A EC (ω) between the short-time spectrum X(ω) of the call reception signal and the short-time spectrum Y E (ω) of the residual echo is estimated, and the Wiener filter G wiener (ω) is calculated based on the following formula (2).

[0078]

[0079] Moreover, as shown in the following formula (3), the echo suppressor obtains the echo-suppressed signal Y wiener (ω) by multiplying the Wiener filter G EC (ω) by the short-time spectrum Y ES (ω) of the residual echo.

[0080] Y ES (ω) = G wiener (ω)Y EC (ω) …… (3)

[0081] That is, the echo suppressor estimates the residual echo based on the acoustic coupling amount A E (ω) estimated for each frequency component and the call reception signal X(ω). For this reason, the echo suppressor cannot estimate frequency components such as non-linear echo that do not exist in the call reception signal.

[0082] To prove the above, the inventors of the present invention conducted an experiment to evaluate the influence of non-linear echo. An existing echo suppression device was used in the evaluation experiment. The existing echo suppression device includes a speaker for amplifying the call reception signal, a microphone, an echo canceller for suppressing the echo signal from the input signal obtained by the microphone, and an echo suppressor for suppressing the echo signal from the output signal from the echo canceller. Moreover, in the evaluation experiment, a 1 / 3 octave band noise with a center frequency of 400 Hz near the lowest resonance frequency f0 of the speaker used for amplification was adopted.

[0083] Figure 1 is a schematic diagram showing the microphone signal, the echo canceller output signal, and the echo suppressor output signal when the input signal does not include non-linear echo caused by speaker distortion. Figure 2 is a schematic diagram showing the microphone signal, the echo canceller output signal, and the echo suppressor output signal when the input signal includes non-linear echo caused by speaker distortion.

[0084] In Figure 1 and Figure 2 , the solid line represents the microphone signal (input signal) output from the microphone, the dashed line represents the echo canceller output signal, and the dotted line represents the echo suppressor output signal. In Figure 1 and Figure 2 , the horizontal axis represents frequency and the vertical axis represents amplitude level.

[0085] In Figure 2 , the second to fourth harmonics of the input signal appear. As described above, it shows the case where non-linear echo cannot be suppressed at all by the existing echo canceller and echo suppressor. In addition, in Figure 1 and Figure 2In the case where the fundamental tone around 400 Hz is concerned, in the absence of non-linear echo, an echo canceller can suppress the echo by about 35 dB, while in the case of including non-linear echo, the suppression amount of the echo canceller deteriorates to about 20 dB. This is because the adaptive filter compulsorily and continuously updates the filter coefficients in order to simulate non-linear echo that cannot be originally expressed, resulting in incorrect learning and thus causing errors in echo estimation.

[0086] The essential problem of the existing echo suppression technology is that non-linear echo cannot be expressed because linear models are used to estimate echo. Here, the echo suppression device according to the first embodiment of the present invention uses a neural network that can approximate any non-linear function to estimate non-linear echo. As methods for introducing the neural network, two methods can be considered. One is a method of estimating the amplitude and phase of non-linear echo and applying them to the echo canceller, and the other is a method of only estimating the amplitude of non-linear echo and applying it to the echo suppressor. Compared with the latter, the former has problems of requiring higher estimation accuracy and increased computational complexity. Therefore, the echo suppression device according to the first embodiment realizes the suppression of non-linear echo in the form of an echo suppressor that can operate with low power consumption, low cost, and less computational complexity.

[0087] Figure 3 It is a schematic diagram showing the configuration of a communication device according to the first embodiment of the present invention. Figure 4 It is a schematic diagram showing an example of signals output from each part of the echo suppression device according to the first embodiment. In addition, the communication device is used in an amplified hands-free communication system, an amplified two-way communication conference system, an intercom system, and the like.

[0088] Figure 3 The communication device shown includes an echo suppression device 1, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22.

[0089] The input terminal 11 outputs a call reception signal x(k) received from a communication device (not shown) of a call recipient to the echo suppression device 1.

[0090] The speaker 12 outputs the input call reception signal x(k) to the outside. Here, when the sound output from the speaker 12 is picked up by the microphone 13, the voice of the speaker of the call recipient is delayed and reproduced from the speaker of the call recipient, thereby generating a so-called acoustic echo. Here, the echo suppression device 1 suppresses the acoustic echo signal included in the input signal x mic (k). At this time, the acoustic echo signal includes a linear echo signal and a non-linear echo signal.

[0091] The microphone 13 is configured within the space where the speaker is located to receive the speaker's voice. The microphone 13 outputs an input signal x mic (k) representing the received voice to the echo cancellation device 1.

[0092] The output terminal 22 outputs an input signal y ES (k) from which the linear echo signal and the non - linear echo signal have been suppressed by the echo cancellation device 1.

[0093] In addition, the input terminal 11 and the output terminal 22 are connected to a communication unit (not shown). The communication unit transmits the input signal y ES (k) to the communication device (not shown) of the call recipient via a network, and receives a call reception signal x(k) from the communication device (not shown) of the call recipient via the network. The network is, for example, the Internet.

[0094] The echo cancellation device 1 includes an echo canceller 14, fast Fourier transform units 15, 16, a non - linear echo model storage unit 17, a non - linear echo estimation unit 18, a non - linear echo suppression unit 19, an echo canceller 20, and an inverse fast Fourier transform unit 21.

[0095] The input terminal 11 outputs the call reception signal x(k) to the speaker 12, the echo canceller 14, and the fast Fourier transform unit 15.

[0096] The echo canceller 14, by estimating the amplitude component and the phase component of the linear echo signal included in the input signal x mic (k) acquired by the microphone 13, suppresses the linear echo signal from the input signal x mic (k). The echo canceller 14 is an example of a first linear echo suppression unit. As Figure 4 shown, the echo canceller 14 only suppresses the linear echo signal included in the input signal x mic (k) output from the microphone 13.

[0097] The echo canceller 14 includes an adaptive filter (not shown) and a subtraction unit.

[0098] The adaptive filter generates an analog echo signal representing the component of the call reception signal included in the input signal x mic (k) acquired by the microphone 13 by convolving the filter coefficients with the call reception signal.

[0099] The subtraction unit calculates the input signal x from the microphone 13 micThe error signal between (k) and the analog echo signal from the adaptive filter, and outputs the calculated error signal to the adaptive filter. The adaptive filter corrects the filter coefficients based on the input error signal, and generates an analog echo signal by convolving the corrected filter coefficients with the call reception signal. The adaptive filter corrects the filter coefficients in such a way as to minimize the error signal using an adaptive algorithm. As the adaptive algorithm, for example, the normalized least mean square method (NLMS (Normalized Least Mean Square) method), the affine projection method, or the recursive least square method (RLS (Recursive Least Square) method) can be adopted.

[0100] Moreover, the subtraction unit subtracts the analog echo signal from the input signal x mic (k) from the microphone 13, and suppresses the linear echo signal from the input signal xmic(k). Moreover, the subtraction unit outputs the input signal y EC (k) with the linear echo signal suppressed to the fast Fourier transform unit 15.

[0101] The fast Fourier transform unit 15 performs the discrete Fourier transform rapidly. The fast Fourier transform unit 15 transforms the input signal y EC (k) in the time domain input to the non-linear echo suppression unit 19 from the echo canceller 14 into the input signal y EC (ω) in the frequency domain. The fast Fourier transform unit 15 outputs the input signal y EC (ω) in the frequency domain with only the linear echo signal suppressed by the echo canceller 14 to the non-linear echo suppression unit 19.

[0102] The fast Fourier transform unit 16 performs the discrete Fourier transform rapidly. The fast Fourier transform unit 16 transforms the call reception signal x(k) in the time domain input to the non-linear echo estimation unit 18 into the call reception signal X(ω) in the frequency domain. The fast Fourier transform unit 16 outputs the call reception signal X(ω) in the frequency domain to the non-linear echo estimation unit 18 and the echo suppressor 20.

[0103] The non-linear echo model storage unit 17 stores in advance a non-linear echo model, which represents the relationship between at least one of the call reception signal output to the speaker 12 and the input signal acquired by the microphone 13 and the non-linear echo signal. In addition, the non-linear echo model storage unit 17 of the first embodiment stores in advance a non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal. The non-linear echo model is, for example, a neural network.

[0104] The non-linear echo model uses teacher data, takes at least one of the call received signal and the input signal as input, and learns with the non-linear echo signal as output. Among them, at least one of the call received signal and the input signal and the output signal of the echo suppressor that suppresses the linear echo signal from the output signal of the echo canceller are used as teacher data, and the echo canceller suppresses the linear echo signal from the input signal. The non-linear echo model of the first embodiment uses teacher data, takes the call received signal as input, and learns with the non-linear echo signal as output. Among them, the output signal of the echo suppressor that suppresses the linear echo signal from the output signal of the echo canceller is used as teacher data, and the echo canceller suppresses the linear echo signal from the call received signal and the input signal.

[0105] The non-linear echo estimation unit 18 estimates the input signal Y mic (k) based on at least one of the call received signal X(ω) output to the speaker 12 and the input signal x EC (ω) and estimates the non-linear echo signal X NN (ω) included in it. Specifically, the non-linear echo estimation unit 18 uses the non-linear echo model to estimate the non-linear echo signal X mic (ω) included in the input signal Y EC (ω) based on at least one of the call received signal X(ω) and the input signal x NN (k). The non-linear echo model represents the relationship between at least one of the call received signal X(ω) output to the speaker 12 and the input signal x mic (k) and the non-linear echo signal. In addition, the non-linear echo estimation unit 18 of the first embodiment uses the non-linear echo model representing the relationship between the call received signal and the non-linear echo signal to estimate the non-linear echo signal X NN (ω) included in the input signal based on the call received signal X(ω).

[0106] The non-linear echo estimation unit 18 reads the non-linear echo model from the non-linear echo model storage unit 17. The non-linear echo estimation unit 18 inputs the call received signal X(ω) output from the fast Fourier transform unit 16 into the non-linear echo model, and obtains the non-linear echo signal X NN (ω) from the non-linear echo model. The non-linear echo estimation unit 18 outputs the non-linear echo signal X NN (ω) estimated using the call received signal X(ω) to the non-linear echo suppression unit 19.

[0107] The non-linear echo suppression unit 19 uses the non-linear echo signal X NN(ω), suppressing the nonlinear echo signal X from the input signal YEC(ω). NN (ω). Specifically, the nonlinear echo suppression unit 19 suppresses the nonlinear echo signal X NN (ω) from the output signal of the echo canceller 14. NN (ω).

[0108] Based on the following formula (4), the nonlinear echo suppression unit 19 calculates the Wiener filter G NN (ω) from the estimated nonlinear echo signal X EC (ω) and the input signal Y NN (ω) from the echo canceller 14.

[0109]

[0110] As shown in the following formula (5), the nonlinear echo suppression unit 19 multiplies the Wiener filter G NN (ω) by the input signal Y EC (ω) to obtain the input signal Y NL-ES (ω) with the nonlinear echo signal suppressed.

[0111] Y NL-ES (ω) = G NN (ω) Y EC (ω)...... (5)

[0112] The nonlinear echo suppression unit 19 outputs the input signal Y NN (ω) with only the nonlinear echo signal X NL-ES (ω) suppressed to the echo suppressor 20.

[0113] The echo suppressor 20 suppresses the residual linear echo signal by estimating the amplitude component of the residual linear echo signal not suppressed by the echo canceller 14. Specifically, the echo suppressor 20 suppresses the residual linear echo signal from the output signal Y NL-ES (ω) of the nonlinear echo suppression unit 19 by estimating the amplitude component of the residual linear echo signal not suppressed by the echo canceller 14. The echo suppressor 20 is an example of a second linear echo suppression unit.

[0114] The echo suppressor 20 suppresses the residual linear echo signal by spectral subtraction or Wiener filtering. The echo suppressor 20 estimates the acoustic coupling amount for each frequency using the space or coherence function of only the echo signal. The echo suppressor 20 uses the estimated acoustic coupling amount and the output signal YNL-ES (ω) and the call reception signal X(ω) to calculate the suppression gain. The echo suppressor 20 suppresses the residual linear echo signal that has not been suppressed by the echo canceller 14 by multiplying the calculated suppression gain by the output signal of the non-linear echo suppression unit 19. The echo suppressor 20 outputs only the input signal Y NL-ES (ω) that has suppressed the residual linear echo signal to the inverse fast Fourier transform unit 21 as the input signal y ES (ω).

[0115] The inverse fast Fourier transform unit 21 performs an inverse discrete Fourier transform quickly. The inverse fast Fourier transform unit 21 converts the input signal y ES (ω) in the frequency domain input from the echo suppressor 20 to the output terminal 22 into the input signal y ES (k) in the time domain. The inverse fast Fourier transform unit 21 outputs the input signal y ES (k) to the output terminal 22.

[0116] Next, the operation of the echo suppression device 1 according to the first embodiment of the present invention will be described.

[0117] Figure 5 is a flowchart for explaining the operation of the echo suppression device according to the first embodiment of the present invention.

[0118] First, in step S1, the echo canceller 14 suppresses the linear echo signal from the input signal xmic(k) by estimating the amplitude component and the phase component of the linear echo signal included in the input signal x mic (k) acquired by the microphone 13.

[0119] Next, in step S2, the non-linear echo estimation unit 18 estimates the non-linear echo signal X NN (ω) included in the input signal based on the non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal, according to the call reception signal X(ω).

[0120] Next, in step S3, the non-linear echo suppression unit 19 suppresses the non-linear echo signal X EC (ω) estimated by the non-linear echo estimation unit 18 from the input signal Y NN (ω) output from the echo canceller 14.

[0121] Next, in step S4, the echo suppressor 20 suppresses the residual linear echo signal from the input signal Y NL-ES (ω) from the non-linear echo suppression unit 19 by estimating the amplitude component of the residual linear echo signal that has not been suppressed by the echo canceller 14. The echo suppressor 20 outputs only the input signal Y NL-ES(ω) suppresses the input signal y of the residual linear echo signal ES (ω) is output to the inverse fast Fourier transform unit 21. The inverse fast Fourier transform unit 21 converts the input signal y in the time domain ES (k) is output to the output terminal 22.

[0122] As described above, using the non-linear echo model, the non-linear echo signal included in the input signal is estimated based on at least one of the call reception signal and the input signal. The non-linear echo model represents the relationship between at least one of the call reception signal output to the speaker 12 and the input signal and the non-linear echo signal. The non-linear echo signal is suppressed from the output signal of the echo canceller 14 using the estimated non-linear echo signal. Therefore, the non-linear echo signal included in the input signal acquired by the microphone 13 can be stably suppressed.

[0123] Moreover, the residual linear echo signal is suppressed from the output signal from which the non-linear echo signal has been suppressed by the echo suppressor 20. Therefore, the operation of the echo suppressor 20 can be stabilized, and the suppression performance of the linear echo signal can be improved.

[0124] Next, the learning method of the non-linear echo model of the first embodiment will be described.

[0125] Figure 6 is a schematic diagram showing the configuration of the learning device of the first embodiment of the present invention. Figure 7 is a schematic diagram showing an example of the signals output from each part of the learning device of the first embodiment.

[0126] Figure 6 The learning device shown includes a non-linear echo model creation device 2, an input terminal 31, a speaker 32, a microphone 33, and an output terminal 39.

[0127] The input terminal 31 outputs the call reception signal x(k) received from the call device (not shown) of the call recipient to the echo suppression device 1.

[0128] The speaker 32 outputs the input call reception signal x(k) to the outside.

[0129] The microphone 33 is arranged in the space where the speaker is located and is used to receive the sound of the speaker. The microphone 33 outputs the input signal x mic (k) to the non-linear echo model creation device 2.

[0130] The output terminal 39 outputs the input signal y ES (k) from which the linear echo signal has been suppressed by the non-linear echo model creation device 2.

[0131] In addition, the configurations of the input terminal 31, the speaker 32, the microphone 33, and the output terminal 39 are the same as those of the input terminal 11, the speaker 12, the microphone 13, and the output terminal 22 in Figure 3 .

[0132] The non-linear echo model creation device 2 includes an echo canceller 34, fast Fourier transform units 35 and 36, an echo suppressor 37, an inverse fast Fourier transform unit 38, a non-linear echo model learning unit 40, and a non-linear echo model storage unit 41.

[0133] The echo canceller 34 estimates the amplitude component and the phase component of the linear echo signal included in the input signal x mic (k) obtained by the microphone 13, and suppresses the linear echo signal from the input signal x mic (k). The configuration of the echo canceller 34 is the same as that of the echo canceller 14 shown in Figure 3 . The echo canceller 34 outputs the input signal y EC (k) with the linear echo signal suppressed to the fast Fourier transform unit 35.

[0134] The fast Fourier transform unit 35 performs a discrete Fourier transform quickly. The fast Fourier transform unit 35 transforms the input signal y EC (k) in the time domain input from the echo canceller 34 to the input signal Y EC (ω) in the frequency domain. The fast Fourier transform unit 35 outputs the input signal Y EC (ω) in the frequency domain with only the linear echo signal suppressed by the echo canceller 34 to the echo suppressor 37.

[0135] The fast Fourier transform unit 36 performs a discrete Fourier transform quickly. The fast Fourier transform unit 36 transforms the call reception signal x(k) in the time domain input to the echo suppressor 37 to the call reception signal X(ω) in the frequency domain. The fast Fourier transform unit 36 outputs the call reception signal X(ω) in the frequency domain to the echo suppressor 37 and the non-linear echo model learning unit 40.

[0136] The echo suppressor 37 estimates the amplitude component of the residual linear echo signal not suppressed by the echo canceller 34, and suppresses the residual linear echo signal from the input signal Y EC (ω). The echo suppressor 37 outputs the input signal y EC (ω) with only the residual linear echo signal suppressed from the input signal Y ES (ω) to the inverse fast Fourier transform unit 21 and the non-linear echo model learning unit 40.

[0137] The inverse fast Fourier transform unit 38 performs an inverse discrete Fourier transform rapidly. The inverse fast Fourier transform unit 38 converts the input signal y ES (ω) in the frequency domain input from the echo suppressor 37 to the input signal y ES (k) in the time domain. The inverse fast Fourier transform unit 38 outputs the input signal y ES (k) to the output terminal 39.

[0138] The non-linear echo model learning unit 40 causes the non-linear echo model to learn using teacher data. The non-linear echo model takes at least one of the call reception signal X(ω) and the input signal x mic (k) as an input and outputs a non-linear echo signal. Among them, at least one of the call reception signal X(ω) and the input signal x mic (k), the output signal Y EC (ω) of the echo suppressor 37 that suppresses the residual linear echo signal from the output signal Y ES (ω) of the echo suppressor 34 is used as teacher data. Among them, the echo suppressor 34 suppresses the linear echo signal from the input signal xmic(k). The non-linear echo model learning unit 40 of the first embodiment causes the non-linear echo model to learn using teacher data. The non-linear echo model takes the call reception signal X(ω) as an input and outputs a non-linear echo signal. Among them, the call reception signal X(ω), the output signal Y EC (ω) of the echo suppressor 37 that suppresses the residual linear echo signal from the output signal Y ES (ω) of the echo suppressor 34 is used as teacher data. Among them, the echo suppressor 34 suppresses the linear echo signal from the input signal xmic(k).

[0139] The non-linear echo model is a neural network that performs prior learning using the amplitude spectrum X(ω) of the call reception signal and the residual echo amplitude spectrum Y ES (ω) of the echo suppressor 34 and the echo suppressor 37 as teacher data. The echo suppressor 34 and the echo suppressor 37 can only suppress the linear echo signal. For this reason, the output signals (residual echo signals) of the echo suppressor 34 and the echo suppressor 37 are approximately equal to the non-linear echo signal. In this way, the non-linear echo model learning unit 40 can model the relationship between the amplitude spectrum of the call reception signal and the amplitude spectrum of the non-linear echo signal.

[0140] In addition, examples of machine learning include supervised learning in which the relationship between input and output is learned using teacher data with labels (output information) assigned to the input information, unsupervised learning in which the structure of data is constructed based only on unlabeled input, semi-supervised learning in which both labeled and unlabeled parts are processed, and reinforcement learning in which actions that maximize rewards are learned through trial and error. Moreover, as specific methods of machine learning, there exist not only neural networks (including deep learning using multi-layer neural networks), but also genetic algorithms, decision trees, Bayesian networks, or support vector machines (SVM), etc. In the machine learning of the non-linear echo model, any one of the specific examples listed above can be used.

[0141] The non-linear echo model learning unit 40 stores the learned non-linear echo model in the non-linear echo model storage unit 41.

[0142] The non-linear echo model storage unit 41 stores the non-linear echo model learned by the non-linear echo model learning unit 40.

[0143] In addition, Figure 3 The echo cancellation device 1 shown can also be provided with a non-linear echo model learning unit 40. In this case, the echo cancellation device 1 can also be provided with a mode switching unit for switching between the learning mode and the echo cancellation mode. When switched to the learning mode by the mode switching unit, the echo canceller 14 outputs the output signal to the echo suppressor 20. The non-linear echo model learning unit 40 can also use the input signal Y ES (ω) and the call reception signal X(ω) as teacher data to make the non-linear echo model learn.

[0144] Moreover, the non-linear echo model learned by the learning device can also be pre-stored in the non-linear echo model storage unit 17 of the echo cancellation device 1. Moreover, the echo cancellation device 1 can also receive the non-linear echo model learned by the learning device and update the non-linear echo model stored in the non-linear echo model storage unit 17.

[0145] Next, the simulation results comparing the echo cancellation amount of the echo cancellation device 1 of the first embodiment and the echo cancellation amount of the existing echo cancellation device will be described.

[0146] First, the neural network (non-linear echo model) used for simulation takes the amplitude spectrum of the short-time Fourier transform as the input / output feature quantity.

[0147] Figures 8 to 11 It is a schematic diagram showing an example of estimating the amplitude spectrum of the non-linear echo signal using a neural network. Figure 8It is a schematic diagram showing the amplitude spectrum of a call reception signal including 1 / 3 octave band noise. Figure 9 It is shown in Figure 8 a schematic diagram showing the correct values and estimated values of the amplitude spectrum of the non-linear echo signal included in the input signal obtained by the microphone when the call reception signal shown is amplified. Figure 10 It is a schematic diagram showing the amplitude spectrum of a call reception signal including a female voice. Figure 11 It is shown in Figure 10 a schematic diagram showing the correct values and estimated values of the amplitude spectrum of the non-linear echo signal included in the input signal obtained by the microphone when the call reception signal shown is amplified.

[0148] In Figures 8 to 11 , the horizontal axis represents frequency and the vertical axis represents amplitude level. In Figure 9 and Figure 11 , the solid line represents the correct value of the non-linear echo signal and the dashed line represents the estimated value of the non-linear echo signal.

[0149] As shown in Figure 9 and Figure 11 , it can be obtained that the neural network can accurately estimate the non-linear echo signal represented by the solid line.

[0150] Next, the simulation results of the echo cancellation device 1 of the first embodiment using the learned neural network and the existing echo cancellation device will be described. Here, the existing echo cancellation device only has an echo canceller and an echo suppressor, and only suppresses the linear echo signal through the echo canceller and the echo suppressor.

[0151] Figure 12 It is a schematic diagram showing the result of frequency analysis of the output signal from the existing echo cancellation device and the output signal from the echo cancellation device of the first embodiment. In addition, in Figure 12 , the horizontal axis represents frequency and the vertical axis represents amplitude level. Moreover, in Figure 12 , the solid line represents the input signal from the microphone 13, the dashed line represents the output signal from the existing echo cancellation device, and the dotted line represents the output signal from the echo cancellation device 1 of the first embodiment. Moreover, the call reception signal is 1 / 3 octave band noise with a center frequency of 315 Hz.

[0152] As shown in Figure 12As shown in FIG. 1 , the echo suppression device 1 of the first embodiment achieves a suppression effect of 15 dB to 20 dB higher than the target value for the nonlinear echo signal, that is, the harmonic distortion. In addition, the echo suppression device 1 of the first embodiment achieves a suppression effect of about 15 dB higher than that of the conventional echo suppression device even for the linear echo signal of 315 Hz. This is because the nonlinear echo suppression unit 19 of the first embodiment suppresses the nonlinear echo signal, so that the estimation of the acoustic coupling amount of the echo suppressor 20 at the later stage can be performed stably.

[0153] Next, the evaluation results of the echo suppression amount of the echo suppression device 1 of the first embodiment and the conventional echo suppression device for an input signal having a complex frequency structure such as human voice are described. In addition, the evaluation index adopts ERLE (Echo Return Loss Enhancement) which means the amount of echo suppression. ERLE can be calculated by the following formula (6).

[0154]

[0155] Figure 13 is a schematic diagram showing the temporal change in the amplitude of an input signal containing a male voice and the temporal change in the echo suppression amount (ERLE) for the input signal. Figure 13 In the upper part of the graph, the horizontal axis represents time and the vertical axis represents amplitude. Figure 13 In the lower part of the graph, the horizontal axis represents time and the vertical axis represents the amount of echo suppression. Figure 13 In the lower half of the figure, the solid line represents the echo suppression amount of the echo suppression device 1 according to the first embodiment, and the dotted line represents the echo suppression amount of the conventional echo suppression device.

[0156] The echo suppression device 1 of the first embodiment achieves a suppression effect that is approximately 10 dB higher than that of the conventional echo suppression device, and thus proves that the echo suppression device 1 of the first embodiment is fully effective even for input signals having a complex frequency structure such as human voice.

[0157] As described above, the echo suppression device 1 of the first embodiment enables comfortable conversation even with a speaker with much distortion, and contributes to the high quality, miniaturization, and low cost of notebook computers, network conference systems, and mobile phones.

[0158] (Second Embodiment)

[0159] The non - linear echo estimation unit 18 of the above - mentioned first embodiment estimates the non - linear echo signal included in the input signal based on the call reception signal by using a non - linear echo model representing the relationship between the call reception signal and the non - linear echo signal. In contrast, the non - linear echo estimation unit of the second embodiment estimates the non - linear echo signal included in the input signal based on the call reception signal and the input signal by using a non - linear echo model representing the relationship between the call reception signal, the input signal, and the non - linear echo signal.

[0160] Figure 14 It is a schematic diagram showing the configuration of the communication device according to the second embodiment of the present invention.

[0161] Figure 14 The communication device shown includes an echo cancellation device 1A, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. In addition, in the second embodiment, the same components as those in the first embodiment are given the same reference numerals, and their descriptions are omitted.

[0162] The echo cancellation device 1A includes an echo canceller 14, fast Fourier transform units 15, 16, 23, a non - linear echo model storage unit 171, a non - linear echo estimation unit 181, a non - linear echo suppression unit 19, an echo canceller 20, and an inverse fast Fourier transform unit 21.

[0163] The microphone 13 outputs the input signal x mic (k) to the echo canceller 14 and outputs it to the non - linear echo estimation unit 181 via the fast Fourier transform unit 23.

[0164] The fast Fourier transform unit 23 performs a discrete Fourier transform quickly. The fast Fourier transform unit 23 transforms the input signal x mic (k) in the time domain input to the non - linear echo estimation unit 181 into the input signal x mic (ω) in the frequency domain. The fast Fourier transform unit 23 outputs the input signal x mic (ω) in the frequency domain to the non - linear echo estimation unit 181.

[0165] The non - linear echo model storage unit 171 pre - stores a non - linear echo model representing the relationship between the call reception signal output to the speaker 12, the input signal acquired by the microphone 13, and the non - linear echo signal. The non - linear echo model is, for example, a neural network.

[0166] The non-linear echo model of the second embodiment uses teacher data to learn with the received call signal and the input signal as inputs and the non-linear echo signal as the output. Among them, the received call signal, the input signal, and the output signal of the echo suppressor that suppresses the residual linear echo signal from the output signal of the echo canceller are used as teacher data. Moreover, the echo canceller suppresses the linear echo signal from the input signal.

[0167] The learning method of the non-linear echo model of the second embodiment inputs the received call signal X(ω) and the input signal x in the frequency domain mic (ω) into Figure 6 the non-linear echo model learning unit 40 shown. Moreover, the non-linear echo model learning unit 40 of the second embodiment uses teacher data to train a non-linear echo model with the received call signal X(ω) and the input signal x mic (ω) as inputs and the non-linear echo signal as the output. Among them, the received call signal X(ω), the input signal x mic (ω), the output signal Y EC of the echo suppressor 37 that suppresses the residual linear echo signal from the output signal of the echo canceller 34 ES (ω) are used as teacher data. Moreover, the echo canceller 34 suppresses the linear echo signal from the input signal x mic (ω).

[0168] The non-linear echo estimation unit 181 uses a non-linear echo model representing the relationship between the received call signal, the input signal, and the non-linear echo signal, and estimates the non-linear echo signal X mic (ω) included in the input signal based on the received call signal X(ω) and the input signal x NN (ω).

[0169] The non-linear echo estimation unit 181 reads the non-linear echo model from the non-linear echo model storage unit 171. The non-linear echo estimation unit 181 inputs the received call signal X(ω) output from the fast Fourier transform unit 16 and the input signal x mic (ω) output from the fast Fourier transform unit 23 into the non-linear echo model, and obtains the non-linear echo signal X NN (ω) from the non-linear echo model. The non-linear echo estimation unit 181 outputs the non-linear echo signal X mic (ω) estimated using the received call signal X(ω) and the input signal x NN (ω) to the non-linear echo suppression unit 19.

[0170] In addition, regarding the operation of the echo suppression device 1A of the second embodiment, only Figure 5The processing of step S2 shown is different. That is, in the second embodiment, the non-linear echo estimation unit 181 uses a non-linear echo model representing the relationship between the call reception signal and the input signal and the non-linear echo signal, and based on the call reception signal X(ω) and the input signal x mic (ω), estimates the non-linear echo signal X NN (ω).

[0171] In the second embodiment, since the non-linear echo signal is estimated based on the call reception signal and the input signal, the estimation accuracy of the non-linear echo signal can be further improved.

[0172] (Third Embodiment)

[0173] The non-linear echo estimation unit 18 of the first embodiment described above uses a non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal based on the call reception signal. In contrast, the non-linear echo estimation unit of the third embodiment uses a non-linear echo model representing the relationship between the call reception signal and the output signal of the echo canceller and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal based on the call reception signal and the output signal of the echo canceller 14.

[0174] Figure 15 is a schematic diagram showing the configuration of the call device according to the third embodiment of the present invention.

[0175] Figure 15 The call device shown includes an echo suppression device 1B, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. In addition, in the third embodiment, the same components as those in the first embodiment are given the same reference numerals and their descriptions are omitted.

[0176] The echo suppression device 1B includes an echo canceller 14, fast Fourier transform units 15, 16, a non-linear echo model storage unit 172, a non-linear echo estimation unit 182, a non-linear echo suppression unit 19, an echo suppressor 20, and an inverse fast Fourier transform unit 21.

[0177] The fast Fourier transform unit 15 outputs the input signal Y EC (ω) in the frequency domain from which only the linear echo signal has been suppressed by the echo canceller 14 to the non-linear echo suppression unit 19 and the non-linear echo estimation unit 182.

[0178] The non-linear echo model storage unit 172 pre-stores a non-linear echo model representing the relationship between the call reception signal output to the speaker 12 and the output signal of the echo canceller and the non-linear echo signal. The non-linear echo model is, for example, a neural network.

[0179] The non-linear echo model of the third embodiment learns by using teacher data, with the received call signal and the output signal of the echo canceller as inputs and the non-linear echo signal as the output. Among them, the received call signal, the output signal of the echo canceller, and the output signal of the echo suppressor that suppresses the residual linear echo signal from the output signal of the echo canceller are used as teacher data, and the echo canceller suppresses the linear echo signal from the input signal.

[0180] The learning method of the non-linear echo model of the third embodiment inputs the received call signal X(ω) and the output signal Y EC (ω) in the frequency domain of the echo canceller 34 into Figure 6 the non-linear echo model learning unit 40 shown. Moreover, the non-linear echo model learning unit 40 of the third embodiment uses teacher data to learn a non-linear echo model that takes the received call signal X(ω) and the output signal Y EC (ω) in the frequency domain of the echo canceller as inputs and the non-linear echo signal as the output. Among them, the received call signal X(ω), the output signal Y EC (ω) in the frequency domain of the echo canceller, and the output signal Y EC (ω) of the echo suppressor 37 that suppresses the residual linear echo signal from the output signal Y EC (ω) in the frequency domain of the echo canceller are used as teacher data, and the echo canceller 34 suppresses the linear echo signal from the input signal xmic(k).

[0181] The non-linear echo estimation unit 182 uses a non-linear echo model representing the relationship between the received call signal, the output signal of the echo canceller, and the non-linear echo signal, and estimates the non-linear echo signal X EC (ω) included in the input signal based on the received call signal X(ω) and the output signal Y NN (ω) in the frequency domain of the echo canceller 14.

[0182] The non-linear echo estimation unit 182 reads the non-linear echo model from the non-linear echo model storage unit 172. The non-linear echo estimation unit 182 inputs the received call signal X(ω) output from the fast Fourier transform unit 16 and the input signal Y EC (ω) output from the fast Fourier transform unit 15 into the non-linear echo model, and obtains the non-linear echo signal X NN (ω) from the non-linear echo model. The non-linear echo estimation unit 182 outputs the non-linear echo signal X EC (ω) estimated by using the received call signal X(ω) and the input signal Y NN (ω) to the non-linear echo suppression unit 19.

[0183] In addition, regarding the operation of the echo suppression device 1B of the third embodiment, only Figure 5 the processing of step S2 shown is different. That is, in the third embodiment, the non-linear echo estimation unit 182 uses a non-linear echo model representing the relationship between the call reception signal and the output signal of the echo canceller 14 and the non-linear echo signal, and based on the call reception signal X(ω) and the output signal Y EC (ω) of the echo canceller 14 in the frequency domain, estimates the non-linear echo signal X NN (ω).

[0184] In the third embodiment, since the non-linear echo signal is estimated based on the call reception signal and the output signal of the echo canceller, the estimation accuracy of the non-linear echo signal can be further improved.

[0185] (Fourth Embodiment)

[0186] The non-linear echo estimation unit 18 of the first embodiment described above uses a non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal based on the call reception signal. Correspondingly, the non-linear echo estimation unit of the fourth embodiment uses a non-linear echo model representing the relationship between the call reception signal and the analog linear echo signal from the adaptive filter of the echo canceller and the non-linear echo signal, and estimates the non-linear echo signal included in the input signal based on the call reception signal and the analog linear echo signal from the adaptive filter of the echo canceller.

[0187] Figure 16 It is a schematic diagram showing the configuration of the call device of the fourth embodiment of the present invention.

[0188] Figure 16 The call device shown includes an echo suppression device 1C, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. In addition, in the fourth embodiment, the same components as those in the first embodiment are given the same reference numerals, and their descriptions are omitted.

[0189] The echo suppression device 1C includes an echo canceller 14, fast Fourier transform units 15, 16, 24, a non-linear echo model storage unit 173, a non-linear echo estimation unit 183, a non-linear echo suppression unit 19, an echo suppressor 20, and an inverse fast Fourier transform unit 21.

[0190] The echo canceller 14 includes an adaptive filter 141 and a subtractor 142. The adaptive filter 141 generates an analog linear echo signal representing the component of the received call signal included in the input signal by convolving the filter coefficients and the received call signal. The subtractor 142 subtracts the analog linear echo signal from the input signal.

[0191] The fast Fourier transform unit 24 performs a discrete Fourier transform quickly. The fast Fourier transform unit 24 transforms the analog linear echo signal in the time domain input to the non-linear echo estimation unit 183 into an analog linear echo signal in the frequency domain. The fast Fourier transform unit 24 outputs the analog linear echo signal in the frequency domain to the non-linear echo estimation unit 183.

[0192] The non-linear echo model storage unit 173 pre-stores a non-linear echo model representing the relationship between the received call signal output to the speaker 12, the analog linear echo signal from the adaptive filter of the echo canceller, and the non-linear echo signal. The non-linear echo model is, for example, a neural network.

[0193] The non-linear echo model of the fourth embodiment learns using teacher data, taking the received call signal and the analog linear echo signal as inputs and the non-linear echo signal as the output. Here, the received call signal, the analog linear echo signal from the adaptive filter of the echo canceller, and the output signal of the echo suppressor that suppresses the residual linear echo signal from the output signal of the echo canceller are used as teacher data, and the echo canceller suppresses the linear echo signal from the input signal.

[0194] In the learning method of the non-linear echo model of the fourth embodiment, the received call signal X(ω) and the analog linear echo signal from the adaptive filter of the echo canceller 34 are input to Figure 6 the non-linear echo model learning unit 40 shown. Moreover, the non-linear echo model learning unit 40 of the fourth embodiment learns a non-linear echo model that takes the received call signal X(ω) and the analog linear echo signal as inputs and the non-linear echo signal as the output using teacher data. Here, the received call signal X(ω), the analog linear echo signal from the adaptive filter of the echo canceller 34, and the output signal Y EC (ω) of the echo suppressor 37 that suppresses the residual linear echo signal from the output signal Y ES (ω) of the echo canceller 34 are used as teacher data. And the echo canceller 34 suppresses the linear echo signal from the input signal xmic(k).

[0195] The non-linear echo estimation unit 183 estimates the non-linear echo signal X included in the input signal based on the call reception signal X(ω) and the analog linear echo signal from the adaptive filter 141 using a non-linear echo model representing the relationship between the call reception signal and the analog linear echo signal and the non-linear echo signal. NN (ω).

[0196] The non-linear echo estimation unit 183 reads the non-linear echo model from the non-linear echo model storage unit 173. The non-linear echo estimation unit 183 obtains the non-linear echo signal X NN NN (ω) from the non-linear echo model by inputting the call reception signal X(ω) output from the fast Fourier transform unit 16 and the analog linear echo signal output from the fast Fourier transform unit 24 into the non-linear echo model. The non-linear echo estimation unit 183 outputs the non-linear echo signal X NN NN (ω) estimated using the call reception signal X(ω) and the analog linear echo signal to the non-linear echo suppression unit 19.

[0197] In addition, regarding the operation of the echo suppression device 1C of the fourth embodiment, only the process of step S2 shown is different. That is, in the fourth embodiment, the non-linear echo estimation unit 183 estimates the non-linear echo signal X NN Figure 5 (ω) based on the call reception signal X(ω) and the analog linear echo signal from the adaptive filter 141 of the echo canceller using a non-linear echo model representing the relationship between the call reception signal and the analog linear echo signal and the non-linear echo signal. NN (ω).

[0198] In the fourth embodiment, since the non-linear echo signal is estimated based on the call reception signal and the analog linear echo signal from the adaptive filter 141 of the echo canceller, the estimation accuracy of the non-linear echo signal can be further improved.

[0199] (Fifth Embodiment)

[0200] The non-linear echo estimation unit 18 of the first embodiment described above estimates the non-linear echo signal included in the input signal based on the call reception signal using a non-linear echo model representing the relationship between the call reception signal and the non-linear echo signal. In contrast, the non-linear echo estimation unit of the fifth embodiment estimates the non-linear echo signal included in the input signal based on the input signal using a non-linear echo model representing the relationship between the input signal and the non-linear echo signal.

[0201] Figure 17 is a schematic diagram showing the configuration of the call device of the fifth embodiment of the present invention.

[0202] Figure 17 The call device shown includes an echo suppression device 1D, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. In addition, in the fifth embodiment, the same components as those in the first and second embodiments are given the same reference numerals, and their descriptions are omitted.

[0203] The echo suppression device 1D includes an echo canceller 14, fast Fourier transform units 15, 16, 23, a non-linear echo model storage unit 174, a non-linear echo estimation unit 184, a non-linear echo suppression unit 19, an echo suppressor 20, and an inverse fast Fourier transform unit 21.

[0204] The microphone 13 outputs an input signal x mic (k) to the echo canceller 14 and outputs it to the non-linear echo estimation unit 184 via the fast Fourier transform unit 23.

[0205] The fast Fourier transform unit 23 performs a discrete Fourier transform quickly. The fast Fourier transform unit 23 transforms the input signal x mic (k) in the time domain input to the non-linear echo estimation unit 184 into an input signal x mic (ω) in the frequency domain. The fast Fourier transform unit 23 outputs the input signal x mic (ω) in the frequency domain to the non-linear echo estimation unit 184.

[0206] The non-linear echo model storage unit 174 pre-stores a non-linear echo model representing the relationship between the input signal acquired by the microphone 13 and the non-linear echo signal. The non-linear echo model is, for example, a neural network.

[0207] The non-linear echo model of the fifth embodiment is learned using teacher data, with the input signal as the input and the non-linear echo signal as the output. Here, the input signal acquired by the microphone and the output signal of the echo suppressor that suppresses the residual linear echo signal from the output signal of the echo canceller are used as the teacher data, and the echo canceller suppresses the linear echo signal from the input signal.

[0208] In the learning method of the non-linear echo model of the fifth embodiment, the input signal x mic (ω) in the frequency domain is input to Figure 6 the non-linear echo model learning unit 40 shown. Moreover, the non-linear echo model learning unit 40 of the fifth embodiment uses teacher data to learn a non-linear echo model with the input signal X mic (ω) as the input and the non-linear echo signal as the output. Here, the input signal x mic (ω) and the output signal Y from the echo canceller 34EC (ω) The output signal Y of the echo suppressor 37 that suppresses the residual linear echo signal EC (ω) is used as the teacher data, and the echo canceller 34 suppresses the linear echo signal from the input signal xmic(k).

[0209] The non-linear echo estimation unit 184 uses a non-linear echo model representing the relationship between the input signal and the non-linear echo signal, and based on the input signal x mic (ω) estimates the non-linear echo signal X included in the input signal NN (ω).

[0210] The non-linear echo estimation unit 184 reads the non-linear echo model from the non-linear echo model storage unit 174. The non-linear echo estimation unit 184 inputs the input signal x mic (ω) output from the fast Fourier transform unit 23 to the non-linear echo model, and obtains the non-linear echo signal X NN (ω) from the non-linear echo model. The non-linear echo estimation unit 184 uses the input signal x mic (ω) to estimate the non-linear echo signal X NN (ω) and outputs it to the non-linear echo suppression unit 19.

[0211] Additionally, regarding the operation of the echo suppression device 1D of the fifth embodiment, only Figure 5 the processing of step S2 shown is different. That is, in the fifth embodiment, the non-linear echo estimation unit 184 uses a non-linear echo model representing the relationship between the input signal and the non-linear echo signal, and based on the input signal x mic (ω) estimates the non-linear echo signal.

[0212] In the fifth embodiment, even only based on the input signal acquired by the microphone 13, the non-linear echo signal can be estimated.

[0213] (Sixth Embodiment)

[0214] In the above-described first embodiment, the non-linear echo signal estimated by the non-linear echo estimation unit 18 is output to the non-linear echo suppression unit 19. Correspondingly, in the sixth embodiment, the output signal of the non-linear echo suppression unit 19 is used to correct the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18.

[0215] Figure 18 It is a schematic diagram showing the configuration of the communication device of the sixth embodiment of the present invention.

[0216] Figure 18The shown communication device is equipped with an echo suppression device 1E, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. Additionally, in the sixth embodiment, the same components as those in the first embodiment are given the same reference numerals, and their descriptions are omitted.

[0217] The echo suppression device 1E includes an echo canceller 14, fast Fourier transform units 15, 16, a non-linear echo model storage unit 17, a non-linear echo estimation unit 18, a non-linear echo suppression unit 19, an echo suppressor 20, an inverse fast Fourier transform unit 21, and a correction unit 25.

[0218] The correction unit 25 calculates a variable gain for minimizing the output signal of the non-linear echo suppression unit 19, and corrects the non-linear echo signal estimated by the non-linear echo estimation unit 18 using the calculated variable gain. At this time, the correction unit 25 calculates the variable gain in such a way that the output signal of the non-linear echo suppression unit 19 approaches 0. Moreover, the correction unit 25 multiplies the non-linear echo signal estimated by the non-linear echo estimation unit 18 by the calculated variable gain. Thereby, the correction unit 25 corrects the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18.

[0219] Additionally, regarding the operation of the echo suppression device 1E in the sixth embodiment, a new process is added between the steps S2 and S3 shown in Figure 5 That is, in the sixth embodiment, after the process of step S2, the correction unit 25 calculates a variable gain in such a way that the output signal of the non-linear echo suppression unit 19 is minimized, and corrects the non-linear echo signal estimated by the non-linear echo estimation unit 18 using the calculated variable gain.

[0220] In the sixth embodiment, since the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18 is corrected using the output signal of the non-linear echo suppression unit 19, the estimation accuracy of the non-linear echo signal can be improved, and the echo suppression performance can be enhanced. In particular, the sixth embodiment is more effective when the non-linear echo model is a fixed value.

[0221] Additionally, the echo suppression devices 1A to 1D of the second to fifth embodiments described above may also be equipped with the correction unit 25 of the sixth embodiment.

[0222] (Seventh Embodiment)

[0223] In the above-described first embodiment, the non-linear echo signal estimated by the non-linear echo estimation unit 18 is output to the non-linear echo suppression unit 19. In contrast, in the seventh embodiment, the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18 is corrected using the output signal of the echo suppressor 20.

[0224] Figure 19 It is a schematic diagram showing the configuration of the communication device according to the seventh embodiment of the present invention.

[0225] Figure 19 The shown communication device includes an echo suppression device 1F, an input terminal 11, a speaker 12, a microphone 13, and an output terminal 22. In addition, in the seventh embodiment, the same components as those in the first embodiment are given the same reference numerals, and their descriptions are omitted.

[0226] The echo suppression device 1F includes an echo canceller 14, fast Fourier transform units 15 and 16, a non-linear echo model storage unit 17, a non-linear echo estimation unit 18, a non-linear echo suppression unit 19, an echo canceller 20, an inverse fast Fourier transform unit 21, and a correction unit 251.

[0227] The correction unit 251 calculates a variable gain that minimizes the output signal of the echo canceller 20, and corrects the non-linear echo signal estimated by the non-linear echo estimation unit 18 using the calculated variable gain. At this time, the correction unit 25 calculates the variable gain so that the output signal of the echo canceller 20 approaches 0. Moreover, the correction unit 251 multiplies the calculated variable gain by the non-linear echo signal estimated by the non-linear echo estimation unit 18. Thus, the correction unit 251 corrects the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18.

[0228] In addition, regarding the operation of the echo suppression device 1F in the seventh embodiment, a new process is added between the steps S2 and S3 shown in Figure 5 That is, in the seventh embodiment, after the process of step S2, the correction unit 251 calculates a variable gain that minimizes the output signal of the echo canceller 20, and corrects the non-linear echo signal estimated by the non-linear echo estimation unit 18 using the calculated variable gain.

[0229] In the seventh embodiment, since the estimation error of the non-linear echo signal estimated by the non-linear echo estimation unit 18 is corrected using the output signal of the echo canceller 20, the estimation accuracy of the non-linear echo signal can be improved, and the echo suppression performance can be improved. In particular, the seventh embodiment is more effective when the non-linear echo model is a fixed value.

[0230] In addition, the echo suppression devices 1A to 1D of the second to fifth embodiments described above may also include the correction unit 251 of the seventh embodiment.

[0231] In addition, in each of the above-described embodiments, each component can be constituted by dedicated hardware or can be implemented by executing a software program suitable for each component. Each component can also be implemented by causing a program execution unit such as a CPU or a processor to read a software program stored in a storage medium such as a hard disk or a semiconductor memory.

[0232] Part or all of the functions of the device according to the embodiment of the present invention can typically be implemented as an integrated circuit LSI (Large Scale Integration). Part or all of these functions can be formed into chips separately or can be formed into a chip including part or all of them. Moreover, the integrated circuit is not limited to LSI, and can also be implemented by a dedicated circuit or a general-purpose processor. An FPGA (Field Programmable Gate Array) that can be programmed after manufacturing the LSI or a reconfigurable processor that can reconstruct the connection or setting of circuit units inside the LSI can also be used.

[0233] Moreover, part or all of the functions of the device according to the embodiment of the present invention can also be implemented by causing a processor such as a CPU to execute a program.

[0234] Moreover, the numbers used above are examples given for specifically illustrating the present invention, and the present invention is not limited to these exemplified numbers.

[0235] Moreover, the order in which the steps shown in the above flowchart are executed is only an example given for specifically illustrating the present invention, and it can also be an order other than the above within the range where the same effect can be obtained. Moreover, a part of the above steps can also be executed simultaneously (in parallel) with other steps.

[0236] The technology related to the present invention has practical value as a technology for suppressing linear echo signals and non-linear echo signals included in an input signal obtained by a microphone because it can stably suppress non-linear echo signals included in the input signal obtained by the microphone.

Claims

1. An echo cancellation device, characterized in that Comprising: A first linear echo suppression unit that suppresses a linear echo signal from the input signal by estimating an amplitude component and a phase component of the linear echo signal included in the input signal acquired by a microphone; A non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to a speaker and the input signal, using a non-linear echo model representing a relationship between at least one of the call reception signal and the input signal and the non-linear echo signal; A non-linear echo suppression unit that suppresses the non-linear echo signal from the output signal of the first linear echo suppression unit, using the non-linear echo signal estimated by the non-linear echo estimation unit; A second linear echo suppression unit that suppresses the residual linear echo signal from the output signal of the non-linear echo suppression unit by estimating an amplitude component of the residual linear echo signal not suppressed by the first linear echo suppression unit.

2. The echo suppression device according to claim 1, wherein: The non-linear echo model is learned using teacher data, with at least one of the call reception signal and the input signal as an input and the non-linear echo signal as an output, where at least one of the call reception signal and the input signal and the output signal of the second linear echo suppression unit that suppresses the residual linear echo signal from the output signal of the first linear echo suppression unit are used as the teacher data, and the first linear echo suppression unit suppresses the linear echo signal from the input signal.

3. The echo suppression device according to claim 1, wherein: The non-linear echo model is a neural network.

4. The echo suppression device according to any one of claims 1 to 3, wherein: The non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on the call reception signal, using the non-linear echo model representing a relationship between the call reception signal and the non-linear echo signal.

5. The echo suppression device according to any one of claims 1 to 3, wherein: The non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on the call reception signal and the input signal, using the non-linear echo model representing a relationship between the call reception signal, the input signal, and the non-linear echo signal.

6. The echo suppression device according to any one of claims 1 to 3, wherein: The non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on the call reception signal and the output signal of the first linear echo suppression unit, using the non-linear echo model representing a relationship between the call reception signal, the output signal of the first linear echo suppression unit, and the non-linear echo signal.

7. The echo suppression device according to any one of claims 1 to 3, wherein: The first linear echo suppression unit includes an adaptive filter and a subtraction unit. The adaptive filter generates an analog linear echo signal representing the component of the call reception signal included in the input signal by convolving filter coefficients with the call reception signal. The subtraction unit subtracts the analog linear echo signal from the input signal. The non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on the call reception signal and the analog linear echo signal from the adaptive filter, using the non-linear echo model representing the relationship between the call reception signal, the analog linear echo signal from the adaptive filter, and the non-linear echo signal.

8. The echo suppression device according to any one of claims 1 to 3, characterized in that The non-linear echo estimation unit estimates the non-linear echo signal included in the input signal based on the input signal, using the non-linear echo model representing the relationship between the input signal and the non-linear echo signal.

9. The echo cancellation device according to any one of claims 1 to 3, characterized in that It further includes: A correction unit that calculates a variable gain for minimizing any one of the output signals of the non-linear echo suppression unit and the output signal of the second linear echo suppression unit, and corrects the non-linear echo signal estimated by the non-linear echo estimation unit using the calculated variable gain.

10. An echo cancellation device, characterized in that It includes: A first linear echo suppression unit that suppresses a linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by a microphone. A non-linear echo estimation unit that estimates the non-linear echo signal included in the input signal based on at least one of the call reception signal output to a speaker and the input signal. A non-linear echo suppression unit that suppresses the non-linear echo signal from the input signal using the non-linear echo signal estimated by the non-linear echo estimation unit. A second linear echo suppression unit that suppresses the residual linear echo signal from the output signal of the non-linear echo suppression unit by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppression unit.

11. An echo cancellation method, characterized in that It includes the following steps: Let the first linear echo suppression unit suppress the linear echo signal from the input signal by estimating the amplitude component and the phase component of the linear echo signal included in the input signal acquired by a microphone. Let the non-linear echo estimation unit estimate the non-linear echo signal included in the input signal based on at least one of the call reception signal output to a speaker and the input signal, using the non-linear echo model representing the relationship between at least one of the call reception signal output to a speaker and the input signal and the non-linear echo signal. Let the non-linear echo suppression unit suppress the non-linear echo signal from the output signal of the first linear echo suppression unit using the non-linear echo signal estimated by the non-linear echo estimation unit. Cause the second linear echo suppressor to suppress the residual linear echo signal from the output signal of the nonlinear echo suppressor by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppressor.

12. An echo cancellation method, characterized in that Including the following steps: Cause the first linear echo suppressor to suppress the linear echo signal from the input signal by estimating the amplitude component and phase component of the linear echo signal included in the input signal acquired by the microphone; Cause the nonlinear echo estimator to estimate the nonlinear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal; Cause the nonlinear echo suppressor to suppress the nonlinear echo signal from the input signal by using the nonlinear echo signal estimated by the nonlinear echo estimator; Cause the second linear echo suppressor to suppress the residual linear echo signal from the output signal of the nonlinear echo suppressor by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppressor.

13. A storage medium is a non-transitory computer-readable storage medium storing an echo cancellation program, characterized in that, Cause the computer to function as the following configuration: First Linear echo suppressor, which suppresses the linear echo signal from the input signal by estimating the amplitude component and phase component of the linear echo signal included in the input signal acquired by the microphone; Nonlinear echo estimator, which estimates the nonlinear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal by using a nonlinear echo model representing the relationship between at least one of them and the nonlinear echo signal; Nonlinear echo suppressor, which suppresses the nonlinear echo signal from the output signal of the first linear echo suppressor by using the nonlinear echo signal estimated by the nonlinear echo estimator; Second linear echo suppressor, which suppresses the residual linear echo signal from the output signal of the nonlinear echo suppressor by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppressor.

14. A storage medium, which is a non-transitory computer-readable storage medium storing an echo cancellation program, is characterized in that Cause the computer to function as the following configuration: First Linear echo suppressor, which suppresses the linear echo signal from the input signal by estimating the amplitude component and phase component of the linear echo signal included in the input signal acquired by the microphone; Nonlinear echo estimator, which estimates the nonlinear echo signal included in the input signal based on at least one of the call reception signal output to the speaker and the input signal; Nonlinear echo suppressor, which suppresses the nonlinear echo signal from the input signal by using the nonlinear echo signal estimated by the nonlinear echo estimator; Second linear echo suppressor, which suppresses the residual linear echo signal from the output signal of the nonlinear echo suppressor by estimating the amplitude component of the residual linear echo signal not suppressed by the first linear echo suppressor.

Citation Information

Patent Citations

  • Echo suppression apparatus, echo suppression method, echo suppression program, and recording medium

    JP2010103875A

  • Echo suppressor, method therefor, program, and recording medium

    JP2017191992A