Echo suppression device, echo suppression method, and echo suppression program

The echo suppression device addresses the challenge of accurately estimating echo suppression amounts by using an estimated echo function to generate an echo suppression mask, thereby enhancing call quality even with large non-linear echo components.

JP7696676B2Active Publication Date: 2025-06-23TRANSTRON INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021054402
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-26
Publication Date
2025-06-23
Estimated Expiration
2041-03-26

AI Technical Summary

Technical Problem

Existing echo canceller devices struggle to accurately estimate echo suppression amounts, especially when non-linear echo components are large, leading to poor call quality.

Method used

An echo suppression device that uses an estimated echo function with variables such as the logarithm of the received signal magnitude, frequency, and total received value to generate an echo suppression mask, allowing for accurate echo suppression even with large non-linear echo components.

Benefits of technology

The device achieves accurate estimation of echo suppression amounts for each frequency, improving call quality by effectively managing echo components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696676000014
    Figure 0007696676000014
  • Figure 0007696676000015
    Figure 0007696676000015
  • Figure 0007696676000016
    Figure 0007696676000016
Patent Text Reader

Abstract

To precisely estimate an echo suppression amount at each frequency even if a nonlinear echo component is large.SOLUTION: An estimation echo function is stored with a logarithm of magnitude at each frequency of a received signal, a frequency of the received signal, a logarithm of a total received value, which is the total sum of received signal magnitude or the total sum of the received signals over an arbitrary frequency range, and a logarithm of an envelope of the total received value as variables. A value of a second received signal (a result of converting the received signal into a frequency domain) is inputted to the function representing an estimation echo and a mask for echo suppression is generated, and echo suppression processing is performed by multiplying a second transmission signal (a result of converting the transmission signal into the frequency domain) by an echo suppression gain calculated on the basis of the mask for echo suppression.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an echo suppression device, an echo suppression method, and an echo suppression program.

Background Art

[0002] Patent Document 1 discloses an echo canceller device used in a voice communication system having a microphone and a speaker. This echo canceller device includes an echo cancellation unit that removes a pseudo echo component from a microphone input signal and outputs a residual signal, a pERL calculation unit that calculates a pERL value indicating a ratio between the microphone input signal and the residual signal, an echo signal based on the echo input from the speaker to the microphone among the microphone input signals, and an ERLE calculation unit that calculates an ERLE value indicating a ratio between the residual echo signal obtained by subtracting the pseudo echo component from the echo signal, a pERL reduction degree calculation unit that calculates a reduction degree indicating a difference between the ERLE value and the pERL value, and a suppression amount calculation unit that calculates a residual echo suppression amount from the formula (K - 1)T / K(T - 1) when the value indicating the reduction degree as a linear value is K and the value indicating the ERLE value as a linear value is T, and a residual echo suppression processing unit that generates an output signal by multiplying the residual signal by the residual echo suppression amount.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Generally, when non - linear echo components generated by reflection, speaker vibration, etc. are large, the estimation of the echo suppression amount often does not function properly. In the echo canceller device described in Patent Document 1, there is a possibility that the echo suppression amount cannot be accurately estimated in a frame with a long reflection time and no signal in the received call.

[0005] The present invention has been made in view of such circumstances, and an object thereof is to provide an echo suppression device, an echo suppression method, and an echo suppression program capable of accurately estimating the amount of echo suppression for each frequency even when the non-linear echo component is large.

Means for Solving the Problems

[0006] In order to solve the above problems, an echo suppression device according to the present invention is, for example, an echo suppression device that suppresses echo generated when a received signal is transmitted through a receiving-side signal path that transmits a signal to a speaker and the sound output from the speaker by the received signal is input to a microphone. The device includes: a storage unit that stores an estimated echo function having as variables the logarithm of the magnitude of the received signal at each frequency, the frequency of the received signal, the logarithm of the sum of the magnitudes of the received signal or the total received value of the received signal in an arbitrary frequency range, and the logarithm of the envelope of the total received value; and a non-linear echo suppression unit that generates an echo suppression mask by inputting the value of a second received signal obtained by converting the received signal into a frequency domain into the function representing the estimated echo, and performs echo suppression processing by multiplying an echo suppression gain calculated based on the echo suppression mask by a second transmission signal obtained by converting a transmission signal transmitted through the transmission-side signal path into a frequency domain.

[0007] According to the echo suppression device of the present invention, an estimated echo function having as variables the logarithm of the magnitude at each frequency of the received signal, the frequency of the received signal, the logarithm of the total received value which is the sum of the magnitudes of the received signal, and the logarithm of the envelope of the total received value is stored. The value of the second received signal (the result of converting the received signal into the frequency domain) is input to the function representing the estimated echo to generate a mask for echo suppression, and echo suppression processing is performed by multiplying the echo suppression gain calculated based on this mask for echo suppression by the second transmission signal (the result of converting the transmission signal into the frequency domain). Thereby, even when the non-linear echo component is large, the amount of echo suppression can be accurately estimated for each frequency. As a result, the call quality can be improved.

[0008] A double talk detection unit that inputs the value of the second received signal to the function representing the estimated echo to generate a mask for double talk detection, and sequentially detects whether or not a speech is input to the microphone based on the second transmission signal and the mask for double talk detection is provided. When a speech is input to the microphone, the non-linear echo suppression unit may make the echo suppression gain smaller than when no speech is input to the microphone. Thereby, when there is a near-end speech and it is considered that the far-end speaker is less likely to feel discomfort from the echo, the suppression of the echo can be weakened, and it is possible to prevent the sound from becoming unnatural due to excessive suppression of the echo.

[0009] The double talk detection unit compares the magnitude of the second transmission signal with the magnitude of the double talk detection mask for each frequency, and determines whether the number of frequencies at which the magnitude of the second transmission signal exceeds the magnitude of the double talk detection mask is less than a first threshold value, whether the sum of the magnitudes of the second transmission signal in the frequency band where the magnitude of the second transmission signal exceeds the magnitude of the double talk detection mask is less than a second threshold value, or whether the sum of the differences between the magnitude of the second transmission signal and the magnitude of the double talk detection mask in the frequency band where the magnitude of the second transmission signal exceeds the magnitude of the double talk detection mask is less than a third threshold value, and may detect that no speech is input to the microphone based on this. Thereby, the presence or absence of near-end speech can be accurately detected.

[0010] A noise estimation unit that estimates a noise component included in the second transmission signal, and a noise suppression unit that multiplies the second transmission signal by a noise suppression gain to suppress the noise signal from the echo cancellation signal, and the non-linear echo suppression unit may obtain the echo suppression mask based on the estimated echo, the noise component, and the noise suppression gain. Thereby, appropriate echo suppression can be performed without being affected by noise.

[0011] A noise estimation unit that estimates a noise component included in the second transmission signal, and a noise suppression unit that multiplies the second transmission signal by a noise suppression gain to suppress the noise signal from the echo cancellation signal, and the double talk detection unit may obtain the double talk detection mask based on the estimated echo, the noise component, and the noise suppression gain. Thereby, false detection due to the influence of noise can be prevented.

[0012] The non-linear echo suppression unit obtains an allowable value indicating the magnitude of the allowable residual echo based on the noise component and the noise suppression gain, and may multiply the second transmission signal by the echo suppression gain such that the magnitude of the echo suppression mask is reduced to the magnitude of the allowable value. Thereby, it is possible to prevent excessive echo suppression.

[0013] When the magnitude of the second transmission signal is greater than the allowable value and less than or equal to the echo suppression mask, the non-linear echo suppression unit may obtain the echo suppression gain based on a value obtained by subtracting the allowable value from the magnitude of the second transmission signal. When the value of the second transmission signal is greater than the allowable value and the echo suppression mask, the echo suppression gain may be obtained based on a value obtained by subtracting the allowable value from the echo suppression mask. Thereby, echo can be appropriately suppressed according to the magnitude of the second transmission signal.

[0014] In the function representing the estimated echo, the coefficients of the respective variables may be obtained based on data obtained by removing outliers from the second training signal. Thereby, it is possible to prevent the size of the echo suppression mask from becoming larger than necessary, and to prevent over-suppression of the echo. In addition, it is possible to prevent the size of the double talk detection mask from becoming larger than necessary, and to accurately detect the presence or absence of near-end speech.

[0015] The function representing the estimated echo has a first function in which the coefficients of the respective variables are obtained based on data obtained by removing outliers from the second training signal, and a second function in which the coefficients of the respective variables are obtained based on the second training signal without removing outliers. The double talk detection mask may be obtained based on the first function, and the echo suppression mask may be obtained based on the second function. Thereby, it is possible to accurately detect the presence or absence of near-end speech while strengthening the suppression of non-linear echo and performing sufficient echo suppression.

[0016] In order to solve the above problems, the echo suppression method according to the present invention is, for example, an echo suppression method for suppressing an echo generated when a received signal is transmitted through a receiving-side signal path for transmitting a signal to a speaker and the sound output from the speaker by the received signal is input to a microphone. The method includes: converting a learning received signal transmitted through the receiving-side signal path into a second learning received signal in the frequency domain; and a second learning signal obtained by converting a learning signal transmitted through a transmitting-side signal path for transmitting a signal input from the microphone when the sound output from the speaker by the learning received signal is input to the microphone, into the frequency domain. An estimated echo calculated based on the above and stored in a storage unit, and obtaining an estimated echo function having as variables the logarithm of the magnitude at each frequency of the received signal, the frequency of the received signal, the logarithm of the total received value which is the sum of the magnitudes of the received signals, and the logarithm of the envelope of the total received value; inputting the value of the second received signal obtained by converting the received signal into the frequency domain into the function representing the estimated echo to generate an echo suppression mask, and multiplying an echo suppression gain calculated based on the echo suppression mask by a second transmitted signal obtained by converting a transmitted signal transmitted through the transmitting-side signal path into the frequency domain, to perform an echo suppression process.

[0017] In order to solve the above problems, the echo suppression program according to the present invention is, for example, an echo suppression program for suppressing an echo generated when a received signal is transmitted through a received signal path that transmits a signal to a speaker and the voice output from the speaker by the received signal is input to a microphone. The computer is configured to convert a second learning received signal obtained by converting a learning received signal transmitted through the received signal path into a frequency domain, and a learning signal transmitted through a transmission signal path that transmits a signal input from the microphone when the sound output from the speaker by the learning received signal is input to the microphone. An estimated echo calculated based on a second learning signal converted into a frequency domain, the logarithm of the magnitude at each frequency of the received signal, the frequency of the received signal, the logarithm of the total received value that is the sum of the magnitudes of the received signal, and the logarithm of the envelope of the total received value. A storage unit that stores an estimated echo function having variables, inputs the value of the second received signal obtained by converting the received signal into a frequency domain into a function representing the estimated echo to generate an echo suppression mask, and calculates an echo suppression gain based on the echo suppression mask. A non-linear echo suppression unit that performs echo suppression processing by multiplying a second transmission signal obtained by converting a transmission signal transmitted through the transmission signal path into a frequency domain. Note that the computer program can be provided by downloading via a network such as the Internet, or can be recorded on various computer-readable recording media such as a CD-ROM and provided.

Effects of the Invention

[0018] According to the present invention, even when the non-linear echo component is large, the echo suppression amount can be accurately estimated for each frequency.

Brief Description of the Drawings

[0019]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0020] Hereinafter, embodiments of the echo suppression device according to the present invention will be described in detail with reference to the drawings. The echo suppression device is a device that suppresses echo generated when an audio signal output from a speaker is input to a microphone in an audio communication system.

[0021] <First Embodiment> FIG. 1 is a diagram schematically showing an audio communication system 100 provided with an echo suppression device 1 according to the first embodiment. The audio communication system 100 mainly includes a terminal 50 having a microphone 51 and a speaker 52, two mobile phones 53, 54, a speaker amplifier 55, and an echo suppression device 1.

[0022] The audio communication system 100 is a system in which a near-end speaker (user A on the near-end side) using the terminal 50 (near-end terminal) makes a voice call with a far-end speaker (user B on the far-end side) using the mobile phone 54 (far-end terminal). The audio signal input via the mobile phone 54 is amplified and output by the speaker 52, and the voice uttered by the user A on the near-end side is collected by the microphone 51 and transmitted to the mobile phone 54, so that the user A can make a hands-free call without holding the mobile phone 53. The mobile phone 53 and the mobile phone 54 are connected by a general telephone line.

[0023] The echo suppression device 1 may be configured as a dedicated board mounted on, for example, a communication terminal or the like (e.g., in-vehicle device, conference system, mobile terminal) within the voice communication system 100. Further, the echo suppression device 1 may be configured, for example, mainly by a computing device such as a CPU (Central Processing Unit) for executing information processing, and a computer system and software (echo suppression program) including storage devices such as a RAM (Random Access Memory) and a ROM (Read Only Memory). The echo suppression program may be pre-stored in an SSD as a storage medium built into a device such as a computer, or in a ROM within a microcomputer having a CPU, and installed on the computer therefrom. Further, the echo suppression program may be temporarily or permanently stored (memorized) in a removable storage medium such as a semiconductor memory, a memory card, an optical disk, a magneto-optical disk, or a magnetic disk.

[0024] FIG. 2 is a diagram showing an outline of the functional blocks of the echo suppression device 1. Functionally, the echo suppression device 1 mainly includes an echo cancellation unit 11, frequency analyzers (FFT units) 12 and 22, a noise estimation unit 13, a noise suppression unit 14, a double talk detection unit 15, a non-linear echo suppression unit 16, a noise superposition unit 17, a restoration unit (IFFT unit) 18, a dynamic range control 21, and a storage unit 23. In FIG. 2, the upper signal path is a transmission-side signal path for transmitting an input signal input from the microphone 51, and the lower signal path is a reception-side signal path for transmitting a signal to the speaker 52. Note that the functional components of the echo suppression device 1 may be further classified into more components according to the processing content, or one component may execute the processing of a plurality of components.

[0025] The echo cancellation unit 11 cancels the echo using, for example, an adaptive filter. The echo cancellation unit 11 updates the filter coefficients according to a given procedure, generates a pseudo echo signal from the signal transmitted through the receiving side signal path, and subtracts the pseudo echo signal from the signal transmitted through the transmitting side signal path to cancel the echo. Since the adaptive filter is already known, the description thereof is omitted.

[0026] In this embodiment, an adaptive filter is applied to the echo cancellation unit 11, but other known echo cancellation techniques can also be applied to the echo cancellation unit 11. Also, although the echo cancellation unit 11 is not essential, by generating a mask using the learning signal with a part of the echo removed, it is possible to more accurately detect the presence of a near-end speech (the speech of user A (see FIG. 1)). Therefore, it is desirable to provide the echo cancellation unit 11.

[0027] When the double talk detection unit 15 (described in detail later) detects the presence of a near-end speech, the dynamic range control 21 amplifies (i.e., compresses) the received signal greater than the threshold value among the input received signals by a predetermined coefficient (the coefficient is a value less than 1) and outputs the result. Note that the dynamic range control 21 may have a gain adjustment unit that automatically changes the gain according to the noise in the environment where the terminal 50 is installed or automatically changes the gain according to the magnitude of the received signal.

[0028] The frequency analyzers (FFT units) 12 and 22 perform a fast Fourier transform (FFT, Fast Fourier Transform) on the signal. The FFT unit 12 performs a fast Fourier transform on the signal transmitted through the transmitting side signal path, here the signal that has passed through the echo cancellation unit 11, and the FFT unit 22 performs a fast Fourier transform on the signal transmitted through the receiving side signal path. The FFT units 12 and 22 convert a signal arranged in time series (time domain) into a signal represented by a set of frequencies (frequency domain). Hereinafter, a signal dependent on time is indicated by...[t], and a signal dependent on frequency is indicated by...[i].

[0029] The noise estimation unit 13 estimates, for each frequency, the power spectrum of the noise component, i.e., the estimated noise signal [i], included in the echo-removed signal [i] that has been input from the microphone 51, passed through the transmission-side signal path, had the echo removed by the echo cancellation unit 11, and been converted to the frequency domain by the FFT unit 12 (hereinafter referred to as the estimated noise power spectrum [i]). The estimated noise power spectrum [i] is output to the noise suppression unit 14, the double-talk detection unit 15, the non-linear echo suppression unit 16, and the noise superposition unit 17.

[0030] The noise suppression unit 14 multiplies the estimated noise power spectrum [i] by a noise suppression gain that is a frequency-dependent signal (hereinafter referred to as the noise suppression gain [i]) to suppress the noise signal from the echo-removed signal [i] and generate the suppressed signal [i]. The noise suppression unit 14 suppresses the noise signal using a known noise suppression method such as the spectral subtraction method or the Wiener filter, and the noise suppression gain [i] is calculated by the noise suppression unit 14 according to the noise suppression method used. The calculated noise suppression gain [i] is output to the double-talk detection unit 15. Note that the noise estimation unit 13 and the noise suppression unit 14 are not essential.

[0031] The storage unit 23 stores the mask generated by the estimated echo calculation unit 24 (see FIG. 3). Hereinafter, the generation of the mask will be described in detail. The mask is generated in advance before the echo suppression device 1 performs the process of suppressing the echo.

[0032] FIG. 3 is a diagram showing an outline of the functional blocks when obtaining the function for calculating the estimated echo in the echo suppression device 1. The echo suppression device 1 functionally includes an estimated echo calculation unit 24. The calculation process of the estimated echo is mainly performed by the estimated echo calculation unit 24.

[0033] The calculation process of the estimated echo will be described in detail. First, after the adaptive filter has been sufficiently trained in the echo cancellation unit 11, in a situation where there is no near-end speech and the background noise is sufficiently small, a learning received signal is transmitted through the receiving-side signal path, and single talk on the far-end side that outputs sound from the speaker 52 using the learning received signal is repeated. Then, the signal transmitted through the transmitting-side signal path during single talk is used as the learning signal. In the echo suppression device 1, the signal from which the echo has been removed by the echo cancellation unit 11 becomes the learning signal.

[0034] The learning signal, which is a time-dependent signal (hereinafter referred to as the learning signal [t]), is input to the FFT unit 12. The FFT unit 12 performs a fast Fourier transform on the learning signal [t] to generate a learning signal that is a frequency-dependent signal (hereinafter referred to as the learning signal [i]), and inputs it to the estimated echo calculation unit 24.

[0035] The learning received signal, which is a time-dependent signal (hereinafter referred to as the learning received signal [t]), is input to the FFT unit 22. The FFT unit 22 performs a fast Fourier transform on the learning received signal [t] to generate a learning received signal that is a frequency-dependent signal (hereinafter referred to as the learning received signal [i]), and inputs it to the estimated echo calculation unit 24.

[0036] The estimated echo calculation unit 24 stores the learning signal [i] and the learning received signal [i] in the storage unit 23. Further, the estimated echo calculation unit 24 calculates the power spectrum for the learning signal [i] and the learning received signal [i] stored in the storage unit 23 for each fixed interval, and obtains a plurality of learning power spectra. Here, the fixed interval is an arbitrarily determined predetermined time domain. The estimated echo calculation unit 24 stores the learning power spectrum in the storage unit 23.

[0037] Note that the power spectrum P[i] is represented by the square of the Fourier spectrum X[i] obtained by fast Fourier transform (see Equation (1)).

[0038] P[i]=|X[i]|^2 2= |X[i]| × |X[i]| ···(1)

[0039] The estimated echo calculation unit 24 creates a plurality of scatter diagrams of the training signal [i] and the training received signal [i] based on the training signal [i], the training received signal [i], and the training power spectrum stored in the storage unit 23.

[0040] FIG. 4 is an example of a scatter diagram of the training signal [i] with respect to the training received signal [i] at a certain time (for example, time t1). (A) is a scatter diagram of the logarithm of the magnitude of the training received signal at each frequency (the power spectrum of the training received signal [t]) and the logarithm of the power spectrum at each frequency of the training signal. (B) is a scatter diagram of the frequency of the training received signal and the logarithm of the power spectrum at each frequency of the training signal. (C) is a scatter diagram of the logarithm of the total received power spectrum (corresponding to the total received value of the present invention), which is the sum of the magnitudes of the training received signals, and the logarithm of the power spectrum at each frequency of the training signal. (D) is a scatter diagram of the logarithm of the envelope of the total received power spectrum and the logarithm of the power spectrum at each frequency of the training signal.

[0041] For example, as shown in FIGS. 4(A) and 4(C), even if the power spectra of the training signals are the same, the power spectra of the training signals, that is, the echoes, are various. Therefore, in the present embodiment, the estimated echo is calculated based not only on the power spectrum of the training signal but also on a plurality of scatter diagrams with different horizontal axes.

[0042] Here, the power spectrum at each frequency of the training signal means the power spectrum of the echo caused by the training received signal. Also, the total received power spectrum is the same as the sum of the power spectra at each frequency of the training signal, that is, the sum of the power spectra of the training received signal [t] before passing through the FFT unit 22, and is represented by the following mathematical formula (2).

[0043]

Equation

[0044] Note that the total received power spectrum may be the sum of the power spectra at each frequency within an arbitrary frequency range of the learning signal. The total received power spectrum at this time is represented by the following mathematical formula (3). Here, A is 0 or greater, and B is less than the maximum frequency (A > 0, B < F_MAX). [Number]

[0045] When the double talk detection unit 15 performs speech detection (to be described in detail later), the accuracy may be better when using the sum of the power spectra in an arbitrary frequency range of the learning signal (mathematical formula (3)) than when using the sum of the power spectra at all frequencies of the learning signal (mathematical formula (2)). Therefore, in such a case, it is desirable for the estimated echo calculation unit 24 to obtain the total received power spectrum using mathematical formula (3).

[0046] Note that the scatter diagram shown in FIG. 4 is an example, and the scatter diagram will be different depending on the situation of sound reflection, the arrangement of the speaker 52 and the microphone 51, the shape of the speaker 52, the presence or absence of the echo cancellation unit 11, and the like.

[0047] As shown in FIG. 4, a certain relationship holds between the logarithm and frequency information of the learning received signal and the learning signal, that is, the power spectrum of the echo. In the present embodiment, sufficient learning signals [i] and learning received signals [i] are acquired in advance, and the estimated echo amount is obtained based on the certain relationship between them.

[0048] Specifically, the estimated echo calculation unit 24 calculates an estimated echo function using the following mathematical formula (4). The estimated echo function (estimated echo power spectrum [i]) is a signal that depends on frequency, and is represented by a function with the logarithm of the magnitude at each frequency of the learning received signal, the frequency of the learning received signal, the logarithm of the total received power spectrum of the learning received signal, and the logarithm of the envelope of the total received value of the learning received signal as variables.

[0049] Estimated echo power spectrum [i] = α × Received power spectrum [i] + β × Frequency + γ × Total received power spectrum + δ × Envelope of total received power spectrum... (4)

[0050] The calculation of the estimated echo function will be described in detail with reference to FIGS. 5 to 8. The estimated echo calculation unit 24 sequentially calculates the coefficient α of the received power spectrum [i], the coefficient β of the frequency, the coefficient γ of the total received power spectrum, and the coefficient δ of the envelope of the total received power spectrum. The coefficients α, β, γ, and δ of each variable are obtained based on the data obtained by removing outliers from the learning signal [i].

[0051] FIG. 5 is a scatter diagram of the logarithm of the power spectrum at each frequency of the learning received signal (hereinafter referred to as the received power spectrum) and the logarithm of the power spectrum at each frequency of the learning signal (hereinafter referred to as the transmitted power spectrum). In FIG. 5, the measured data is plotted, and α is indicated by a line.

[0052] α indicates the relationship between the logarithm of the received power spectrum and the maximum value of the logarithm of the transmitted power spectrum. α is obtained based on the result of removing outliers from the scatter diagram of the logarithm of the received power spectrum and the logarithm of the transmitted power spectrum. α is represented by a linear function (without conditional branches) or a non-linear function (with conditional branches).

[0053] As shown in FIG. 5, when the received power spectrum is large, the transmitted power spectrum (i.e., the echo) does not necessarily become large. Instead, when the received power spectrum becomes larger than a certain level, the echo becomes smaller. This is because of the characteristics of the speaker 52 (there is a region where sound cannot be emitted) and the fact that the echo cancellation unit 11 is provided upstream of the FFT unit 12. In the example shown in FIG. 5, α is represented by equations (5) and (6). Thus, α is a non-linear function.

[0054] When the logarithm of the received power spectrum < -1 α = 0.5 × Logarithm of received power spectrum - 0.5... (5)

[0055] When the logarithm of the received power spectrum ≥ -1 α = -1.0 × logarithm of the received power spectrum - 2.0 ··· (6)

[0056] In addition, when the echo cancellation unit 11 is not provided, compared with the example shown in FIG. 5, the peak of the line indicating α shifts to the right, and the slope of the descending line after the peak becomes smaller, but α remains a non-linear function (with conditional branching).

[0057] When α is calculated, the estimated echo calculation unit 24 calculates β. FIG. 6 is a scatter diagram of the logarithm of the frequency of the learning received signal (hereinafter referred to as the received frequency) and the logarithm of the transmitted power spectrum. In FIG. 6, the result of subtracting the α component from the measured data is plotted, and β is shown by a line.

[0058] β shows the relationship between the received frequency and the maximum value of the logarithm of the transmitted power spectrum. β is obtained based on the result of removing outliers from the scatter diagram of the received frequency and the logarithm of the transmitted power spectrum. β is represented by a linear function or a non-linear function.

[0059] Since the speaker 52 has the characteristic that it is difficult to emit low frequencies and high frequencies, in FIG. 6, the echo is small for low frequencies and high frequencies. Also, when the terminal 50 is provided inside the vehicle, as shown in FIG. 6, there is a dip where the echo becomes small due to the influence of the intermediate environment (reflection, etc.) near 1 kHz. Therefore, β is a non-linear function.

[0060] When β is calculated, the estimated echo calculation unit 24 calculates γ. FIG. 7 is a scatter diagram of the logarithm of the total received power spectrum of the learning received signal and the logarithm of the transmitted power spectrum. In FIG. 7, the result of subtracting the α component and the β component from the measured data is plotted, and γ is shown by a line.

[0061] For example, when the speaker 52 outputs sounds of 100 Hz and 110 Hz, a sound of 105 Hz may also sound from the speaker 52 in addition to the 100 Hz and 110 Hz. Therefore, in order to refer to information on whether sounds other than the frequencies to be sounded are sounding, in the present embodiment, a term using the logarithm of the total received power spectrum as a variable is added to the estimated echo function (mathematical formula (4)).

[0062] γ indicates the relationship between the logarithm of the total received power spectrum and the maximum value of the logarithm of the transmitted power spectrum. γ is obtained based on the result of removing outliers from the scatter diagram of the received frequency and the logarithm of the transmitted power spectrum. γ is represented by a linear function or a non-linear function. In the example shown in FIG. 7, γ is a non-linear function.

[0063] When γ is calculated, the estimation echo calculation unit 24 calculates δ. FIG. 8 is a scatter diagram of the logarithm of the envelope of the total received power spectrum of the learning received signal and the logarithm of the transmitted power spectrum. In FIG. 8, the result of subtracting the α component, β component, and γ component from the measured data is plotted, and δ is indicated by a line.

[0064] Since the reflection of sound in the vehicle interior, the vibration of the speaker 52, etc. are output as sounds from the speaker 52, an echo may exist even without the learning received signal. Therefore, it is necessary to estimate the echo by referring not only to the current total received power spectrum but also to the learning signals for a certain period immediately before. For this reason, in the present embodiment, a term using the logarithm of the envelope of the total received power spectrum as a variable is added to the estimated echo function (mathematical formula (4)).

[0065] The envelope A is the maximum value in a certain period immediately before, and is gradually calculated as in the following mathematical formula (7) using the time constant B and the total received power spectrum C. In the present embodiment, the time constant B is set to 0.5 to 1.

[0066] If(A<C): A=C Else: A=B×A+(1-B)×C ···(7)

[0067] δ represents the relationship between the logarithm of the envelope of the total received power spectrum and the maximum value of the logarithm of the transmitted power spectrum. δ is obtained based on the result of removing outliers from the scatter diagram of the received frequency and the logarithm of the transmitted power spectrum. δ is represented by a linear function or a non-linear function. In the example shown in FIG. 8, δ is a linear function.

[0068] When the estimated echo function (the function representing the estimated echo power spectrum [i]) is calculated in this way, the estimated echo calculation unit 24 stores the estimated echo function in the storage unit 23.

[0069] Return to the description of FIG. 2. In the description of FIG. 2, the input signal input from the microphone 51 includes the sound output from the speaker 52 and its echo by the received signal transmitted through the received signal path, the noise input to the microphone 51, and the sound (proximal speech) input to the microphone 51 by the speech of the user A on the proximal side (see FIG. 1).

[0070] The double talk detection unit 15 sequentially detects whether it is in a double talk state based on the received signal [i] obtained by converting the received signal [t] transmitted through the received signal path into a frequency-dependent signal by the FFT unit 22, the transmitted signal [i] (here, the suppressed signal after passing through the echo removal unit 11, the FFT unit 12, and the noise suppression unit 14) through which the input signal is input from the microphone 51 and transmitted through the transmitted signal path, and the double talk detection mask.

[0071] Note that the double talk state is a state in which there is proximal speech and distal speech, and the single talk state is a state in which there is only proximal speech or only distal speech. This embodiment is characterized in that the double talk detection unit 15 detects the presence or absence of proximal speech, and the method of detecting the presence or absence of distal speech is not limited. For example, the double talk detection unit 15 may detect that there is distal speech when the envelope of the total received power spectrum is greater than a threshold value.

[0072] Next, a method for the double-talk detection unit 15 to detect the presence or absence of near-end speech will be described. The received signal [i] and the transmitted signal [i] are sequentially input to the double-talk detection unit 15. When the received signal [i] and the transmitted signal [i] are input (a sample point is acquired), the double-talk detection unit 15 generates a double-talk detection mask based on the estimated echo power spectrum [i] stored in the storage unit 23, and detects whether it is in a double-talk state. Also, the double-talk detection unit 15 performs the process of detecting whether it is in a double-talk state every time a sample point is acquired.

[0073] First, the double-talk detection mask will be described. The double-talk detection unit 15 calculates a double-talk detection mask based on the estimated echo power spectrum [i], the estimated noise power spectrum [i], and the noise suppression gain [i]. Specifically, as shown in Equation (8), the double-talk detection mask is obtained by adding a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the estimated echo power spectrum [i]. Since the double-talk detection mask is a frequency-dependent signal, it is hereinafter referred to as the double-talk detection mask [i].

[0074] Double-talk detection mask [i] = Estimated echo power spectrum [i] + Estimated noise power spectrum [i] × Noise suppression gain [i] ··· (8)

[0075] In Equation (8), the estimated echo power spectrum [i] is obtained by inputting the value of the received signal [i] into a function representing the estimated echo (Equation (4)). The estimated noise power spectrum [i] is obtained by the noise estimation unit 13, and the noise suppression gain [i] is stored in the storage unit 23.

[0076] Next, the process of detecting whether it is in a double-talk state will be described with reference to FIG. 9. FIG. 9 is a diagram showing a state of comparing a suppressed signal of one frame at a certain time with a double-talk detection mask. In FIG. 9, each plot is a suppressed signal, and the line is a double-talk detection mask. Also, the horizontal axis of FIG. 9 is the frequency of the suppressed signal, and the vertical axis is the logarithm of the power spectrum of the suppressed signal.

[0077] The double-talk detection unit 15 compares the suppressed signal with the double-talk detection mask for each frequency to detect whether it is in a double-talk state. As methods for detecting whether it is in a double-talk state, there are the following three patterns A, B, and C. Patterns A, B, and C are methods for determining whether, when the plot in FIG. 9 exceeds the double-talk detection mask, it is due to a near-end speech or an outlier.

[0078] <Pattern A> The double-talk detection unit 15 compares the magnitude of the suppressed signal with the magnitude of the double-talk detection mask for each frequency, and counts the number of frequencies at which the magnitude of the suppressed signal exceeds the magnitude of the double-talk detection mask (hereinafter referred to as the excess number). In other words, in the scatter diagram shown in FIG. 9, the number of plots above the double-talk detection mask is counted. The double-talk detection unit 15 determines whether the excess number is equal to or less than a threshold I (corresponding to the first threshold) prepared in advance. Note that the threshold I can be set to any value.

[0079] <Pattern B> The double-talk detection unit 15 compares the magnitude of the suppressed signal with the magnitude of the double-talk detection mask for each frequency, and calculates the sum of the magnitudes of the suppressed signals at the frequencies at which the magnitude of the suppressed signal exceeds the magnitude of the double-talk detection mask. In other words, in the scatter diagram shown in FIG. 9, the sum of the values of the plots above the double-talk detection mask (see the two-dot chain line in FIG. 9) is obtained.

[0080] For example, the sum of the magnitudes of the suppressed signals is a value obtained by subtracting a constant (e.g., -7) from the logarithmic value of the power spectrum of the suppressed signals. Since the logarithm of the power spectrum of the suppressed signals can take negative values, subtracting a negative value makes it a positive value. Also, for example, the sum of the magnitudes of the suppressed signals may be the sum of the power spectra of the suppressed signals. Since the power spectrum of the suppressed signals is positive without taking the logarithm, it is only necessary to simply calculate the sum.

[0081] The double talk detection unit 15 determines whether the sum of the magnitudes of the suppressed signals is less than or equal to a threshold II (corresponding to the second threshold) prepared in advance. Note that the threshold II can be set to any value.

[0082] <Pattern C> The double talk detection unit 15 compares the magnitude of the suppressed signal with the magnitude of the double talk detection mask for each frequency, and calculates the sum of the differences between the magnitude of the suppressed signal (here, the logarithm of the power spectrum of the suppressed signal) at frequencies where the magnitude of the suppressed signal exceeds the magnitude of the double talk detection mask and the magnitude of the double talk detection mask. In other words, in the scatter diagram shown in FIG. 9, the sum of the differences (see the dotted line in FIG. 9) between the magnitude of the plot above the double talk detection mask and the magnitude of the double talk detection mask is obtained.

[0083] The double talk detection unit 15 determines whether the sum of the differences between the magnitude of the suppressed signal and the magnitude of the double talk detection mask is less than or equal to a threshold III (corresponding to the third threshold) prepared in advance. Note that the threshold III can be set to any value.

[0084] The double talk detection unit 15 detects whether the value calculated by any of the methods of patterns A to C is greater than or equal to a threshold (threshold I, II, or III). Then, when the frames in which the calculated value is greater than or equal to the threshold continue for a predetermined number (e.g., 2 frames) or more, it is determined that there is near-end speech.

[0085] For example, when the calculated value is equal to or greater than the threshold value, the double talk detection unit 15 increments (counts up) the value of the counter by 1, and when the calculated value is less than the threshold value, it decrements (counts down) the value of the counter by 1 or sets the counter to 0. Then, when the value of the counter becomes equal to or greater than the threshold value (for example, 2), the double talk detection unit 15 determines that there is a near-end speech.

[0086] Pattern C has the largest amount of calculation, but when it exceeds the double talk detection mask, it can most accurately determine whether it is due to a near-end speech or an outlier.

[0087] Note that the double talk detection unit 15 does not necessarily need to detect whether it is in a double talk state when, for example, transitioning from a state of only near-end speech to a state of only far-end speech, or when transitioning from a double talk state to a state of only near-end speech, only far-end speech, or no near-end and far-end speech. In particular, when transitioning from a double talk state to a state with no near-end and far-end speech, there is a high possibility that the echo still remains, and when transitioning from a double talk state to a state with no far-end speech, there is a high possibility of a near-end speech. Therefore, it is not necessary to detect whether it is in a double talk state for a predetermined time after the transition.

[0088] Returning to the description of FIG. 2. The non-linear echo suppression unit 16 receives an input signal from the microphone 51 and performs a process of suppressing non-linear echo (hereinafter referred to as non-linear echo suppression process) on the transmission signal [i] transmitted through the transmission side signal path (here, the signal to be suppressed after passing through the echo cancellation unit 11, the FFT unit 12, and the noise suppression unit 14). In the present embodiment, the non-linear echo suppression unit 16 performs the non-linear echo suppression process by multiplying the transmission signal [i] by an echo suppression gain calculated based on an echo suppression mask generated based on the estimated echo. Further, the non-linear echo suppression unit 16 sets the echo suppression gain to different values based on the detection result in the double talk detection unit 15.

[0089] The received signal [i], the transmitted signal [i], and the detection result in the double-talk detection unit 15 are sequentially input to the non-linear echo suppression unit 16. When the transmitted signal [i] is input (a sample point is acquired), the non-linear echo suppression unit 16 generates an echo suppression mask based on the estimated echo function stored in the storage unit 23 and performs non-linear echo suppression processing.

[0090] The non-linear echo suppression unit 16 calculates an echo suppression mask based on the estimated echo power spectrum [i], the estimated noise power spectrum [i], and the noise suppression gain [i]. Specifically, as shown in Equation (9), the echo suppression mask is obtained by adding a term obtained by multiplying the estimated echo power spectrum [i] by the estimated noise power spectrum [i] and the noise suppression gain [i]. Since the echo suppression mask is a frequency-dependent signal, it is hereinafter referred to as the echo suppression mask [i].

[0091] Echo suppression mask [i]=Estimated echo power spectrum [i]+Estimated noise power spectrum [i]×Noise suppression gain [i] ···(9)

[0092] In the case of Equation (9) as well, similar to the case of Equation (8), the estimated echo power spectrum [i] is obtained by inputting the value of the received signal [i] into Equation (4), the estimated noise power spectrum [i] is obtained by the noise estimation unit 13, and the noise suppression gain [i] is stored in the storage unit 23.

[0093] FIG. 10 is a diagram showing a state of comparing a suppressed signal of one frame at a certain time with an echo suppression mask. In FIG. 10, each plot is the suppressed signal, the solid line is the echo suppression mask, and the dotted line is the allowable value. Also, the horizontal axis of FIG. 10 is the frequency of the suppressed signal, and the vertical axis is the logarithm of the power spectrum of the suppressed signal.

[0094] The non-linear echo suppression unit 16 performs echo suppression processing on each plot so as to reduce the size of the echo suppression mask to the size of the allowable value. Hereinafter, the echo suppression processing will be described in detail.

[0095] First, the allowable value will be described. The allowable value indicates the magnitude of the residual echo allowed for the transmission signal [i], and is obtained based on the estimated noise power spectrum [i] and the noise suppression gain [i] as shown in Equation (10). Since the allowable value is a frequency-dependent signal, it is hereinafter referred to as the allowable value [i].

[0096] Allowable value [i] = Estimated noise power spectrum [i] × Noise suppression gain [i] + L ··· (10)

[0097] L is a constant. Note that L may be changed based on the magnitude of the estimated noise power spectrum [i] and the detection result in the double-talk detection unit 15.

[0098] FIG. 11 is a graph showing an example of the allowable value [i]. When the estimated noise power spectrum [i] is large, the allowable value becomes large, and when the estimated noise power spectrum [i] is small, the allowable value becomes small.

[0099] Returning to the description of FIG. 10. The allowable value [i] in FIG. 10 is the allowable value [i] when the estimated noise power spectrum [i] in FIG. 11 is small. The non-linear echo suppression unit 16 calculates a basic gain G based on the following Equation (11). Since the gain G is a frequency-dependent signal, it is hereinafter referred to as G [i].

[0100]

Equation

[0101] Note that Equation (9) takes the input signal as X (Z = log 10 Re(X) × Re(X) + Im(X) × Im(X), where Z is the logarithm of the power spectrum of the input signal, Re is the real part, and Im is the imaginary part), and the target signal as Y (Re(Y) = Re(X) × G, Im(Y) = Im(X) × G), and is calculated based on the following Equations (12) to (15).

[0102]

Equation

Number

Number

Number

[0103] The non - linear echo suppression unit 16 generates an echo suppression mask [i] and an allowable value [i] for each frame. Then, for each frame, the non - linear echo suppression unit 16 compares the magnitude of the transmission signal [i] with the magnitude of the echo suppression mask [i] and the magnitude of the transmission signal [i] with the magnitude of the allowable value [i]. And for each frame, the non - linear echo suppression unit 16 calculates echo suppression gains G1 to G5 based on the comparison result and the detection result in the double - talk detection unit 15. The echo suppression gains G1 to G5 are obtained as shown in the following mathematical formulas (16) to (20) using the basic gain G obtained by the mathematical formula (11). Note that Z in the mathematical formulas (16) to (20) is the logarithm of the power spectrum of the transmission signal [i] (the magnitude of the transmission signal [i]), which is the value of the vertical axis of each plot in FIG. 10.

[0104] Z ≤ allowable value: G1 = 1.0 ··· (16)

Number

Number

Number

Number

[0105] As shown in the mathematical formula (16), when Z is less than or equal to the allowable value (the shaded part I in FIG. 10), the non - linear echo suppression unit 16 sets the echo suppression gain G1 to 1 and does not perform echo suppression.

[0106] As shown in Expressions (17) and (18), when Z is greater than the allowable value and less than or equal to the size of the echo suppression mask (the shaded part II in FIG. 10), the echo suppression gains G2 and G3 are obtained based on the value obtained by subtracting the allowable value from the size of the transmission signal (Z - allowable value). In other words, when Z is greater than the allowable value and less than or equal to the size of the echo suppression mask, the non-linear echo suppression unit 16 performs echo suppression so as to reduce the size of the transmission signal to the allowable value.

[0107] When there is a near-end speech, the non-linear echo suppression unit 16 obtains the echo suppression gain G3 by multiplying the value obtained by subtracting the allowable value from the size of the transmission signal by a constant W1. The constant W1 is an arbitrary number between 0 and 1. In other words, when there is a near-end speech, the non-linear echo suppression unit 16 weakens the echo suppression. Note that when W1 is 1, the echo suppression gain G2 and the echo suppression gain G3 coincide.

[0108] As shown in Expressions (19) and (20), when Z is greater than the allowable value and the size of the echo suppression mask (the non-shaded part III in FIG. 10), the echo suppression gains G4 and G5 are obtained based on the value obtained by subtracting the allowable value from the size of the echo suppression mask (echo suppression mask - allowable value). In other words, when Z is greater than the allowable value and the echo suppression mask, the non-linear echo suppression unit 16 performs echo suppression so as to reduce the size of the echo suppression mask to the allowable value.

[0109] When there is a near-end speech, the non-linear echo suppression unit 16 obtains the echo suppression gain G5 by multiplying the value obtained by subtracting the allowable value from the echo suppression mask by a constant W2. The constant W2 is an arbitrary number between 0 and 1. In other words, when there is a near-end speech, the non-linear echo suppression unit 16 weakens the echo suppression. Note that when W2 is 1, the echo suppression gain G4 and the echo suppression gain G5 coincide. Note that the value of W2 may be the same as or different from the value of W1.

[0110] In each frame, the non-linear echo suppression unit 16 performs non-linear echo suppression processing for each measurement point using the obtained echo suppression gains G1 to G5.

[0111] Returning to the description of FIG. 2. The noise superposition unit 17 generates comfort noise based on the estimated noise signal estimated by the noise estimation unit 13, and superimposes the comfort noise on the transmission signal after the echo suppression processing is performed by the non-linear echo suppression unit 16.

[0112] The IFFT unit 18 performs an inverse FFT (IFFT, Inverse FFT) on the input signal that has passed through the noise superposition unit 17.

[0113] FIG. 12 is a flowchart showing the flow of the process in which the echo suppression device 1 sequentially reduces echoes. This process is continuously performed at predetermined time intervals while the received signal and the input signal are input to the echo suppression device 1.

[0114] First, the echo removal unit 11 removes the echo from the input signal (step S11). The noise estimation unit 13 estimates the estimated noise signal included in the echo removal signal, and the noise suppression unit 14 suppresses the noise signal from the echo removal signal based on the estimated noise signal to generate a signal to be suppressed (step S12).

[0115] The double-talk detection unit 15 calculates the power spectra of the signal to be suppressed and the received signal (step S13), acquires the estimated echo power spectrum [i] from the storage unit 23, generates a double-talk detection mask based on the acquired estimated echo and the power spectrum calculated in step S13 (step S14), and detects the presence or absence of near-end speech using the double-talk detection mask generated in step S14 (step S15).

[0116] Next, the non-linear echo suppression unit 16 acquires the estimated echo power spectrum [i] from the storage unit 23, generates an echo suppression mask based on the acquired estimated echo and the power spectrum calculated in step S13 (step S16), and performs echo suppression processing on the signal to be suppressed using the presence or absence of the near-end speech detected in step S15 and the echo suppression mask generated in step S16 (step S17).

[0117] Next, the noise superposition unit 17 generates comfort noise based on the estimated noise signal estimated by the noise estimation unit 13, and superimposes the comfort noise on the transmission signal after the echo suppression processing is performed in step S17 (step S18). Finally, the IFFT unit 18 returns the transmission signal after the noise is superimposed to the time-axis signal (step S19).

[0118] According to the present embodiment, since the non-linear echo suppression processing is performed using the echo suppression mask generated by inputting the value of the received signal to the estimated echo function (estimated echo power spectrum [i]), even when the non-linear echo component is large, the echo suppression amount can be accurately estimated.

[0119] Further, according to the present embodiment, since the presence or absence of the near-end speech is detected using the double-talk detection mask generated by inputting the value of the received signal to the function representing the estimated echo power spectrum [i], the presence or absence of the near-end speech can be accurately detected. In particular, by detecting the presence or absence of the near-end speech by calculating the sum of the differences between the magnitude of the signal to be suppressed and the magnitude of the double-talk detection mask at frequencies where the signal to be suppressed exceeds the value of the double-talk detection mask (pattern C), it is possible to accurately detect whether the data is near-end speech or an outlier when the input signal is larger than the double-talk detection mask.

[0120] Also, according to the present embodiment, by adding a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the mathematical formula (8) for obtaining the double talk detection mask, it is possible to accurately detect the presence or absence of near-end speech. For example, there is a risk that the value of the transmission signal may become larger than the double talk detection mask due to the influence of noise rather than near-end speech. On the other hand, by adding a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the mathematical formula (8) for obtaining the double talk detection mask, false detection due to the influence of noise can be prevented.

[0121] Also, according to the present embodiment, by adding a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the mathematical formula (9) for obtaining the echo suppression mask, appropriate non-linear echo suppression processing can be performed.

[0122] Also, according to the present embodiment, an allowable value, which is a value of allowable residual echo based on the noise component estimated by the noise estimation unit 13 and the noise suppression gain used by the noise suppression unit 14, is obtained, and a non-linear echo suppression process is performed using an echo suppression gain obtained based on the difference between the echo suppression mask and the allowable value. Therefore, it is possible to prevent over-suppressing the echo more than necessary. For example, it is not necessary to make the magnitude after the non-linear echo suppression process smaller than the value of the transmission signal [i] when there is no speech at the near end and the far end, and the demerit that the sound becomes unnatural due to over-increasing the echo suppression gain in the non-linear echo suppression process becomes larger. Therefore, in the non-linear echo suppression process, it is desirable to adjust the echo suppression gain so that the magnitude of the processed signal does not become smaller than the allowable value obtained based on the noise component. In particular, when the value of the transmission signal [i] is larger than the allowable value and below the echo suppression mask (the shaded portion II in FIG. 10), the echo suppression gains G2 and G3 are obtained based on the value obtained by subtracting the allowable value from the transmission signal (Z - allowable value), and when the transmission signal [i] is larger than the allowable value and the echo suppression mask (the non-shaded portion III in FIG. 10), the echo suppression gains G4 and G5 are obtained based on the value obtained by subtracting the allowable value from the echo suppression mask (echo suppression mask - allowable value), thereby appropriately suppressing the echo.

[0123] Also, according to the present embodiment, when Z is larger than the allowable value and below the magnitude of the echo suppression mask, echo suppression is performed so as to reduce the magnitude of the transmission signal to the allowable value, and when Z is larger than the allowable value and the echo suppression mask, echo suppression is performed so as to reduce the magnitude of the echo suppression mask to the allowable value, thereby appropriately suppressing the echo according to the magnitude of Z.

[0124] Also, according to the present embodiment, by making the echo suppression gains G3 and G5 when there is a near-end speech smaller than the echo suppression gains G2 and G4 when there is no near-end speech, it is possible to prevent excessive echo suppression. Generally, when there is a near-end speech, the speaker tends not to care about the echo. Therefore, when there is a near-end speech, the echo suppression can be weakened, and it is possible to prevent the sound from becoming unnatural due to excessive echo suppression.

[0125] Also, according to the present embodiment, since each coefficient α, β, γ, δ of the estimated echo power spectrum [i] is obtained based on the data obtained by removing outliers from the learning signal [i], it is possible to prevent the size of the double-talk detection mask from becoming larger than necessary, and accurately detect the presence or absence of near-end speech. For example, if each coefficient of the estimated echo power spectrum [i] is obtained while including outliers, there is a possibility that the value of the transmission signal [i] does not exceed the double-talk detection mask when the voice of the near-end speaker is small. On the other hand, by obtaining each coefficient α, β, γ, δ of the estimated echo power spectrum [i] based on the data obtained by removing outliers from the learning signal [i], it is possible to detect the presence of near-end speech even when the voice of the near-end speaker is small. Also, since each coefficient α, β, γ, δ of the estimated echo power spectrum [i] is obtained based on the data obtained by removing outliers from the learning signal [i], it is possible to prevent the size of the echo suppression mask from becoming larger than necessary, and prevent excessive echo suppression.

[0126] Note that in the present embodiment, in the non-linear echo suppression unit 16, the detection result in the double-talk detection unit 15 is used to make the echo suppression gain smaller when there is a near-end speech than when there is no near-end speech. However, the double-talk detection unit 15 is not essential, and the non-linear echo suppression unit 16 may not perform processing using the detection result in the double-talk detection unit 15. For example, the non-linear echo suppression unit 16 may perform non-linear echo suppression processing using the echo suppression gains G1, G2, and G5 obtained by the mathematical formulas (15), (16), and (18).

[0127] Also, in the present embodiment, the coefficients α, β, γ, and δ of the estimated echo power spectrum [i] are obtained based on the data obtained by removing outliers from the learning signal [i], and the double talk detection mask and the echo suppression mask are obtained using these coefficients. However, the estimated echo power spectrum [i] serving as the basis for the double talk detection mask and the estimated echo power spectrum [i] serving as the basis for the echo suppression mask may be different.

[0128] For example, the estimated echo calculation unit 24 generates a first estimated echo function (first estimated echo power spectrum [i]) in which the coefficients of each variable are obtained based on the data obtained by removing outliers from the learning signal [i], and a second estimated echo function (second estimated echo power spectrum [i]) in which the coefficients of each variable are obtained based on the learning received speech signal [i] without removing outliers. The storage unit 23 stores the first estimated echo power spectrum [i] and the second estimated echo power spectrum [i] as the estimated echo power spectrum [i]. Then, the double talk detection unit 15 obtains the double talk detection mask based on the first estimated echo power spectrum [i], and the non-linear echo suppression unit 16 obtains the echo suppression mask based on the second estimated echo power spectrum [i]. Thereby, it is possible to accurately detect the presence or absence of near-end speech while strongly suppressing non-linear echo and performing sufficient echo suppression.

[0129] Also, in the present embodiment, the noise estimation unit 13 and the noise suppression unit 14 are provided. In the mathematical expressions (8) for obtaining the double talk detection mask [i] and (9) for obtaining the echo suppression mask [i], a term obtained by multiplying the estimated echo power spectrum [i] by the estimated noise power spectrum [i] and the noise suppression gain [i] is added. However, the noise estimation unit 13 and the noise suppression unit 14 are not essential, and it is not essential to add a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the mathematical expressions (8) and (9). However, in order to accurately detect near-end speech and perform appropriate echo suppression, it is desirable to add a term obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] to the mathematical expressions (8) and (9).

[0130] Also, in the present embodiment, in the non-linear echo suppression process, an allowable value is obtained based on the estimated noise power spectrum [i] and the noise suppression gain [i], and an echo suppression gain is obtained such that the size of the echo suppression mask [i] is reduced to the size of the allowable value. However, it is not necessary to use the allowable value in the non-linear echo suppression process. For example, the non-linear echo suppression unit 16 may perform non-linear echo suppression processing using an echo suppression gain that reduces the size of the echo suppression mask [i] to 0 or an arbitrary value. However, in order to prevent the sound from becoming unnatural due to excessive echo suppression, it is desirable to perform non-linear echo suppression processing so that the size of the echo suppression mask [i] is reduced to the size of the allowable value.

[0131] Also, in the present embodiment, the allowable value [i] is a signal that depends on the frequency. However, the allowable value may be a constant that does not depend on the frequency. For example, the average value of the allowable value [i] may be used as an allowable value (constant) that does not depend on the frequency, and G [i] may be obtained using the allowable value (constant).

[0132] Also, in the present embodiment, the estimated echo calculation unit 24 is provided in the echo suppression device 1. However, the estimated echo calculation unit 24 may be provided in an arithmetic device different from the echo suppression device 1. For example, the estimated echo calculation unit 24 may acquire the learning signal [i] and the learning received signal [i] via a storage medium or network (not shown), and store the generated estimated echo power spectrum [i] in the storage unit 23 via a storage medium or network (not shown).

[0133] Also, in the present embodiment, the estimated echo power spectrum [i] is obtained using the scatter diagram (Figs. 5 to 8) of the learning signal [i] with respect to the learning received signal [i] at a certain time. However, since the data with the logarithm of the transmission power spectrum below a certain value (for example, -5) in each scatter diagram does not affect the calculation of the estimated echo power spectrum [i], the estimated echo power spectrum [i] may be obtained using the data obtained by deleting the data with the logarithm of the transmission power spectrum below a certain value. This can reduce the amount of data and the amount of calculation.

[0134] In addition, in this embodiment, the estimated echo power spectrum [i] is obtained using the scatter diagram (Figs. 5 to 8) of the training signal [i] with respect to the training received signal [i] at a certain time. However, the method for obtaining the estimated echo power spectrum [i] from the training received signal [i] and the training signal [i] is not limited to this. For example, the estimated echo power spectrum [i] may be obtained using known statistical methods or deep learning.

[0135] Also, in this embodiment, the power spectrum is used, but an amplitude spectrum may be used instead of the power spectrum. When using the amplitude spectrum, the magnitude of the signal of the present invention may be the absolute value of the amplitude of the signal as the magnitude of the signal. The total received amplitude spectrum corresponding to the total received value of the present invention may be the sum of the absolute values of the amplitude spectra at each frequency of the training signal, as shown in Equation (21). Also, the total received amplitude spectrum may be the sum of the amplitude spectra at each frequency in an arbitrary frequency range of the training signal, as shown in Equation (22) (A > 0, B < F_MAX).

Number

Number

[0136] In addition, in this embodiment, the echo cancellation unit 11 is provided in the front stage of the FFT unit 12. However, the echo cancellation unit 11 may be provided in the rear stage of the FFT unit 12 or may be provided in the rear stage of the noise suppression unit 14. Also, the noise superposition unit 17 is provided in the rear stage of the non-linear echo suppression unit 16. However, the noise superposition unit 17 may be provided in the rear stage of the restoration unit (IFFT unit) 18.

[0137] Also, in this embodiment, the noise suppression unit 14 was provided in front of the non-linear echo suppression unit 16, but the noise suppression unit 14 may be provided behind the non-linear echo suppression unit 16. In this case, the terms obtained by multiplying the estimated noise power spectrum [i] and the noise suppression gain [i] in the mathematical formulas (8) and (9) are unnecessary.

[0138] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and design changes and the like within the scope not departing from the gist of the present invention are also included. In particular, in the embodiment, the generation of the basic mask, the generation and selection of the optimal mask, the detection of the double talk state, etc. were performed based on the power spectrum represented by the square of the amplitude, but these processes may be performed based on the absolute value of the amplitude.

Explanation of Reference Numerals

[0139] 1: Echo suppression device 11: Echo removal unit 12, 22: FFT unit 13: Noise estimation unit 14: Noise suppression unit 15: Double talk detection unit 16: Non-linear echo suppression unit 17: Noise superposition unit 18: IFFT unit 21: Dynamic range control 23: Storage unit 24: Estimated echo calculation unit 50: Terminal 51: Microphone 52: Speaker 53, 54: Mobile phone 55: Speaker amplifier 100: Voice communication system

Claims

1. An echo suppression device that suppresses an echo generated when a received signal is transmitted through a receiving-side signal path that transmits a signal to a speaker and the sound output from the speaker by the received signal is input to a microphone, a second learning received signal obtained by converting a learning received signal transmitted through the receiving-side signal path into a frequency domain, and a signal input from the microphone when the sound output from the speaker by the learning received signal is input to the microphone An estimated echo function calculated based on a second learning signal obtained by converting a learning signal transmitted through a transmitting-side signal path into a frequency domain, the logarithm of a received power spectrum that is a power spectrum at each frequency of the learning received signal, and the logarithm of the transmitting power spectrum that is the power spectrum at each frequency of the learning signal A storage unit that stores an estimated echo function having as variables α indicating the relationship with the maximum value, β indicating the relationship between the frequency of the learning received signal and the result of subtracting the α component from the maximum value of the logarithm of the transmitting power spectrum, and the logarithm of the power spectrum of the total received value that is the sum of the magnitudes of the learning received signals or the sum of the learning received signals in an arbitrary frequency range, and γ indicating the relationship with the result of subtracting the α component and the β component from the maximum value of the logarithm of the transmitting power spectrum, and δ indicating the relationship with the result of subtracting the α component, the β component, and the γ component from the logarithm of the envelope of the total received value, a non-linear echo suppression unit that generates an echo suppression mask by inputting the value of the second received signal obtained by converting the received signal into a frequency domain into the estimated echo function, and multiplies the echo suppression gain calculated based on the echo suppression mask by a second transmitted signal obtained by converting a transmitted signal transmitted through the transmitting-side signal path into a frequency domain to perform echo suppression processing, comprising The estimated echo function is a function represented by the sum of the product of the α and the power spectrum of the received signal, the product of the β and the frequency, the product of the γ and the total received power spectrum which is the power spectrum of the sum of the magnitudes of the received signal or the sum of the received signals in an arbitrary frequency range, and the product of the δ and the envelope of the total received power spectrum. An echo suppression device characterized by this.

2. It has a double talk detection unit that inputs the value of the second received signal into the estimated echo function to generate a mask for double talk detection, and sequentially detects whether speech is input to the microphone based on the second transmission signal and the mask for double talk detection. When speech is input to the microphone, the non-linear echo suppression unit makes the echo suppression gain smaller than when speech is not input to the microphone. The echo suppression device according to claim 1, characterized by this.

3. The double talk detection unit compares the magnitude of the second transmission signal and the magnitude of the mask for double talk detection for each frequency, and determines whether the number of frequencies at which the magnitude of the second transmission signal exceeds the magnitude of the mask for double talk detection is smaller than a first threshold, whether the sum of the magnitudes of the second transmission signal in the frequency band where the magnitude of the second transmission signal exceeds the magnitude of the mask for double talk detection is smaller than a second threshold, or whether the sum of the differences between the magnitude of the second transmission signal and the magnitude of the mask for double talk detection in the frequency band where the magnitude of the second transmission signal exceeds the magnitude of the mask for double talk detection is smaller than a third threshold, and detects that no speech is input to the microphone based on this. The echo suppression device according to claim 2, characterized by this.

4. A noise estimation unit that estimates the noise component included in the second transmission signal. It includes a noise suppression unit that multiplies the second transmission signal by a noise suppression gain to suppress the noise signal from the echo cancellation signal. The non-linear echo suppression unit obtains the echo suppression mask based on the estimated echo function, the noise component, and the noise suppression gain. The echo suppression device according to any one of claims 1 to 3, characterized in that.

5. A noise estimation unit that estimates a noise component included in the second transmission signal, A noise suppression unit that multiplies the second transmission signal by a noise suppression gain to suppress a noise signal from an echo removal signal, and is provided with. The double talk detection unit obtains the double talk detection mask based on the estimated echo function, the noise component, and the noise suppression gain. The echo suppression device according to claim 2 or 3, characterized in that.

6. The non-linear echo suppression unit obtains an allowable value indicating the magnitude of an allowable residual echo based on the noise component and the noise suppression gain, and multiplies the second transmission signal by the echo suppression gain such that the magnitude of the echo suppression mask is reduced to the magnitude of the allowable value. The echo suppression device according to claim 4 or 5, characterized in that.

7. When the magnitude of the second transmission signal is greater than the allowable value and less than or equal to the echo suppression mask, the non-linear echo suppression unit obtains the echo suppression gain based on a value obtained by subtracting the allowable value from the magnitude of the second transmission signal. When the value of the second transmission signal is greater than the allowable value and the echo suppression mask, the non-linear echo suppression unit obtains the echo suppression gain based on a value obtained by subtracting the allowable value from the echo suppression mask. The echo suppression device according to claim 6, characterized in that.

8. In the estimated echo function, the coefficient of each variable is obtained based on data obtained by removing outliers from the second learning signal. The echo suppression device according to any one of claims 1 to 7, characterized in that.

9. The estimated echo function has a first function in which coefficients of each variable are obtained based on data obtained by removing outliers from the second learning signal, and a second function in which coefficients of each variable are obtained based on the second learning signal without removing outliers. The double talk detection mask is obtained based on the first function. The echo suppression mask is obtained based on the second function. The echo suppression device according to claim 2, 3, or 5, characterized by the above.

10. An echo suppression method for suppressing an echo generated when a received signal is transmitted through a receiving-side signal path that transmits a signal to a speaker and the sound output from the speaker by the received signal is input to a microphone, A second learning received signal obtained by converting a learning received signal transmitted through the received signal path into the frequency domain, and a second learning signal obtained by converting a learning signal transmitted through a transmission signal path that transmits a signal input from the microphone when the sound output from the speaker by the learning received signal is input to the microphone into the frequency domain. An estimated echo function calculated based on the above and stored in the storage unit, which is the logarithm of the received power spectrum, which is the power spectrum at each frequency of the learning received signal, and the logarithm of the transmission power spectrum, which is the power spectrum at each frequency of the learning signal. An α indicating the relationship with the maximum value, a β indicating the relationship between the frequency of the learning received signal and the result obtained by subtracting the α component from the maximum value of the logarithm of the transmission power spectrum, and the logarithm of the power spectrum of the total received value, which is the sum of the magnitudes of the learning received signals or the sum of the learning received signals in an arbitrary frequency range. A γ indicating the relationship with the result obtained by subtracting the α component and the β component from the maximum value of the logarithm of the transmission power spectrum, and a δ indicating the relationship between the logarithm of the envelope of the total received value and the result obtained by subtracting the α component, the β component, and the γ component from the maximum value of the logarithm of the transmission power spectrum are used as variables, and the product of the α and the power spectrum of the received signal, the product of the β and the frequency, and the γ and the total received power spectrum, which is the power spectrum of the sum of the magnitudes of the received signals or the sum of the received signals in an arbitrary frequency range. A step of obtaining an estimated echo function represented by the sum of the product of the δ and the envelope of the total received power spectrum; Inputting the value of the second received signal obtained by converting the received signal into the frequency domain into the estimated echo function to generate a mask for echo suppression, and multiplying the echo suppression gain calculated based on the mask for echo suppression by the second transmission signal obtained by converting the transmission signal transmitted through the transmission signal path into the frequency domain to perform echo suppression processing; An echo suppression method characterized by including the above.

11. An echo suppression program that suppresses an echo generated when a received signal is transmitted through a receiving-side signal path that transmits a signal to a speaker and the sound output from the speaker by the received signal is input to a microphone, a computer, a second learning received signal obtained by converting a learning received signal transmitted through the receiving-side signal path into a frequency domain, and a second learning signal obtained by converting a learning signal transmitted through a transmitting-side signal path that transmits a signal input from the microphone when the sound output from the speaker by the learning received signal is input to the microphone into a frequency domain. An estimated echo function calculated based on the logarithm of the magnitude at each frequency of the received signal, the frequency of the received signal, the logarithm of the received power spectrum which is the power spectrum at each frequency of the learning received signal, and α indicating the relationship with the maximum value of the logarithm of the power spectrum at each frequency of the learning signal, β indicating the relationship between the frequency of the learning received signal and the result of subtracting the α component from the maximum value of the logarithm of the transmitting power spectrum, γ indicating the relationship between the logarithm of the power spectrum of the total received value which is the sum of the magnitudes of the learning received signals or the sum of the learning received signals in an arbitrary frequency range and the result of subtracting the α component and the β component from the maximum value of the logarithm of the transmitting power spectrum, and δ indicating the relationship between the logarithm of the envelope of the total received value and the result of subtracting the α component, the β component, and the γ component from the maximum value of the logarithm of the transmitting power spectrum as variables, and the product of α and the power spectrum of the received signal, the product of β and the frequency, the product of γ and the total received power spectrum which is the power spectrum of the sum of the magnitudes of the received signals or the sum of the received signals in an arbitrary frequency range, and the sum of the product of δ and the envelope of the total received power spectrum. A storage unit that stores an estimated echo function that is a function represented by, Input the value of the second received signal obtained by converting the received signal into the frequency domain into the estimated echo function to generate a mask for echo suppression, and perform echo suppression processing by multiplying the echo suppression gain calculated based on the mask for echo suppression by the second transmitted signal obtained by converting the transmitted signal transmitted through the transmitting-side signal path into the frequency domain. A non-linear echo suppression unit, An echo suppression program characterized by functioning as

Citation Information

Patent Citations

  • Digital disk record reproducing device

    JP1986080689A

  • Echo canceller and microphone device

    JP2007053512A

  • Telecommunication apparatus

    JP2009094802A

  • Echo suppression

    JP2016503263A

  • Echo suppression device, echo suppression method, and echo suppression program

    WO2019220951A1