Echo suppression device, echo suppression method, and echo suppression program

The echo suppression device addresses erroneous learning in adaptive filters by generating delayed reference signals and correcting filter coefficients, achieving efficient echo removal with reduced processing load and cost.

JP7756052B2Active Publication Date: 2025-10-17TRANSTRON INC
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022109997
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-07-07
Publication Date
2025-10-17
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

Existing echo suppression technologies face issues with ensuring sufficient echo removal due to erroneous learning of adaptive filters, leading to inadequate suppression performance.

Method used

The echo suppression device employs a reference signal adjustment unit to generate delayed reference signals, a filter generation unit to obtain convergence values of adaptive filters, and a filter correction unit to compare and correct the adaptive filter coefficients, ensuring appropriate echo removal by setting certain coefficients to zero or reducing their magnitude based on convergence values and delay time changes.

Benefits of technology

This approach effectively removes echoes while reducing processing load on the arithmetic device, allowing for low-cost configuration and robust echo cancellation even in changing environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007756052000002
    Figure 0007756052000002
  • Figure 0007756052000003
    Figure 0007756052000003
  • Figure 0007756052000004
    Figure 0007756052000004
Patent Text Reader

Abstract

To enable echoes to be eliminated using an appropriate adaptive filter that was not subjected to incorrect training.SOLUTION: Provided is an echo suppression device that eliminates echoes from an input signal having been sound-collected by a microphone of a terminal that has a speaker and the microphone. The device comprises: a reference signal adjustment unit that adds a plurality of mutually different delay times to a reference signal that propagates through a receiver-side signal path that transmits signals to the speaker and generates a plurality of delayed reference signals; a filter generation unit that obtains the convergence values of the plurality of adaptive filters on the basis of each of the plurality of delayed reference signals; a filter correction unit that compares the plurality of convergence values and reduces a first filter coefficient to be smaller than the convergence values, the first filter coefficient that is a filter coefficient concerning an element that does not change with a change of delay time; and an echo rejection unit that applies, to the input signal, the adaptive filter that was corrected by the filter correction unit, so as to eliminate linear echoes.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an echo suppression device, an echo suppression method, and an echo suppression program. [Background technology]

[0002] Patent Document 1 discloses an echo canceller device that generates first and second echo replica signals by convolving a received signal with a first filter coefficient or a second filter coefficient, and outputs a signal in the call frequency band of the first echo replica signal as an echo canceller signal. In this device, the second filter coefficient is corrected so as to minimize a signal obtained by subtracting the second echo replica signal from a signal of an arbitrary measurement frequency in the input sound signal, and the filter coefficient is set by determining a first filter coefficient that is paired in advance with the second filter coefficient. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2010-11265 Summary of the Invention [Problem to be solved by the invention]

[0004] However, in the invention described in Patent Document 1, if an erroneous solution occurs in learning the coefficients of the adaptive filter, a sufficient amount of echo suppression may not be ensured.

[0005] The present invention has been made in view of the above circumstances, and has as its object to provide an echo suppression device, an echo suppression method, and an echo suppression program that are capable of performing echo removal using an appropriate adaptive filter that has not undergone erroneous learning. [Means for solving the problem]

[0006] In order to solve the above problems, the echo suppression device of the present invention is, for example, an echo suppression device that removes echo from an input signal picked up by the microphone of a terminal having a speaker and a microphone, and is characterized by comprising: a reference signal adjustment unit that generates a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiving side signal path that transmits a signal to the speaker; a filter generation unit that obtains convergence values ​​of a plurality of adaptive filters based on each of the plurality of delayed reference signals; a filter correction unit that compares the plurality of convergence values ​​and corrects the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change in accordance with changes in the delay time, is smaller than the convergence value; and an echo removal unit that removes linear echo by applying the adaptive filter corrected by the filter correction unit to the input signal.

[0007] An echo suppression method according to another aspect of the present invention is characterized by including, for example, a reference signal adjustment step of generating a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiving side signal path that transmits a signal to a speaker of a terminal having a speaker and a microphone; a filter generation step of obtaining convergence values ​​of a plurality of adaptive filters based on each of the plurality of delayed reference signals; a filter modification step of comparing the plurality of convergence values ​​and modifying the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change with changes in the delay time, is smaller than the convergence value; and an echo removal step of removing linear echo by applying the adaptive filter modified in the filter modification step to an input signal picked up by the microphone.

[0008] An echo suppression program according to another aspect of the present invention is characterized in that it causes a computer to function as, for example, a reference signal adjustment unit that generates a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiving side signal path that transmits a signal to a speaker of a terminal having a speaker and a microphone; a filter generation unit that obtains convergence values ​​of a plurality of adaptive filters based on each of the plurality of delayed reference signals; a filter correction unit that compares the plurality of convergence values ​​and corrects the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change in accordance with changes in the delay time, is smaller than the convergence value; and an echo removal unit that removes linear echo by applying the adaptive filter corrected by the filter correction unit to an input signal picked up by the microphone.

[0009] According to several aspects of the present invention, convergence values ​​of several adaptive filters are obtained based on several delayed reference signals generated by applying several different delay times to a reference signal transmitted through a receiving side signal path that transmits a signal to a speaker, and as a result of comparing the several convergence values, the first filter coefficient, which is the filter coefficient for an element that does not change with changes in delay time, is made smaller than the convergence value, and linear echo is removed from the input signal using the modified adaptive filter, so that echo removal can be performed using an appropriate adaptive filter that has not been erroneously learned.

[0010] The filter correction unit may set the first filter coefficient to 0. This reduces the calculation load and the processing load of the arithmetic device. Furthermore, since the processing load of the arithmetic device can be reduced, the echo suppression device can be configured at low cost.

[0011] The filter correction unit may gradually decrease the first filter coefficient from the convergence value, thereby preventing noise or excessive removal caused by a sudden change in the shape of the adaptive filter.

[0012] The filter correction unit may set a filter coefficient of an element that does not change according to a change in the delay time when the convergence value is equal to or greater than a threshold as the first filter coefficient, thereby reducing the filter coefficient of an element that has a large effect on removing linear echoes.

[0013] The filter correction unit may calculate the threshold value based on the convergence value, thereby automatically determining an appropriate threshold value according to the state of the reference signal.

[0014] The reference signal adjuster may continuously generate the delayed reference signals while acquiring the reference signal, the filter generator may continuously generate the linear filters, and the filter corrector may continuously correct the adaptive filter, thereby enabling appropriate echo cancellation even when an echo path changes during use.

[0015] The system may further include a double-talk detection unit that detects whether the system is in a single-talk state or a double-talk state, and the filter correction unit may set the first filter coefficient to a value smaller than the convergence value if the system does not detect the double-talk state. In this way, when double-talk is occurring, there is a lot of disturbance and the effect of the processing may be small. With this configuration, the calculation load can be reduced by stopping the processing that has little effect.

[0016] The reference signal adjustment unit may generate the delayed reference signal only once before the echo cancellation unit cancels the linear echo, which reduces the calculation load and ultimately allows the echo suppression device 1 to be configured at low cost. [Effects of the Invention]

[0017] According to the present invention, echoes can be effectively removed while reducing the processing load on the arithmetic unit. [Brief explanation of the drawings]

[0018] [Figure 1]1 is a diagram schematically illustrating a voice communication system 100 provided with an echo suppression device 1. FIG. [Figure 2] 1 is a block diagram showing a schematic configuration of an echo suppression device 1. FIG. [Figure 3] 1 is a flowchart showing the flow of processing performed by the echo suppressing device 1. [Figure 4] 10 is a graph showing a process of comparing convergence values ​​of a plurality of adaptive filters with each other. [Figure 5] 1 is a block diagram showing a schematic configuration of an echo suppression device 1A. [Figure 6] 2 is a block diagram showing a schematic configuration of an echo suppression device 2. FIG. [Figure 7] 10 is a flowchart showing the flow of processing performed by the echo suppressing device 2. [Figure 8] FIG. 2 is a block diagram showing a schematic configuration of an echo suppression device 3. [Figure 9] 10 is a flowchart showing the flow of processing performed by the echo suppressor 3. DETAILED DESCRIPTION OF THE INVENTION

[0019] An echo suppression device according to an embodiment of the present invention is described below in detail with reference to the drawings. The echo suppression device is a device that suppresses echoes that occur during calls in a voice communication system, and is used in products that incorporate speakers and microphones, such as headsets for telephone and video conferences, in-vehicle communication devices, and intercoms.

[0020] First Embodiment 1 is a diagram schematically illustrating an audio communication system 100 provided with an echo suppression device 1 according to the first embodiment. The audio communication system 100 mainly includes a terminal 50 (for example, an in-vehicle device, a conference system, or a mobile terminal) having a microphone 51 and a speaker 52, two communication devices 53 and 54, a speaker amplifier 55, and the echo suppression device 1.

[0021] The voice communication system 100 is a system in which a user (user A on the near-end side) using a terminal 50 (near-end terminal) performs voice communication with a user (user B on the far-end side) using a communication device 54 (far-end terminal). An audio signal input via the communication device 54 is amplified and output by a speaker 52, and the voice emitted by the user on the near-end side is collected by a microphone 51 and transmitted to the communication device 54, thereby enabling user A to make a voice call (hands-free call) without holding the communication device 53. The communication devices 53 and 54 are connected by a general telephone line, and can communicate with each other.

[0022] The echo suppressing device 1 is provided on a transmitting side signal path that transmits an input signal input from a microphone 51 from a terminal 50 to a communication device 53, and removes echo from the input signal.

[0023] The echo suppression device 1 may be constructed as a dedicated board mounted on, for example, a terminal 50 in the voice communication system 100. Alternatively, the echo suppression device 1 may be configured by, for example, computer hardware and software (echo suppression program). The echo suppression program may be stored in advance in a storage medium such as a hard disk drive (HDD) built into a computer or a ROM in a microcomputer having a CPU, and then installed into the computer from there. Alternatively, the echo suppression program may be temporarily or permanently stored (memorized) in a removable storage medium such as a semiconductor memory, a memory card, an optical disk, a magneto-optical disk, or a magnetic disk.

[0024] 2 is a block diagram showing a schematic configuration of the echo suppression device 1. The echo suppression device 1 is connected between a microphone 51 and a signal input terminal 531 on the transmitting side of a communication device 53. In FIG. 2, the upper signal path is a transmitting side signal path that transmits an input signal input from the microphone 51, and the lower signal path is a receiving side signal path that transmits a signal to a speaker 52.

[0025] An input signal picked up by microphone 51 and an audio signal received by communication device 53 are input to echo suppression device 1. Echo suppression device 1 removes the echo from the input signal based on a reference signal, which is an audio signal received by communication device 53 and transmitted through a receiving-side signal path, and outputs the result to signal input terminal 531 on the transmitting side.

[0026] The echo suppression device 1 mainly includes a reference signal adjustment unit 10, a linear echo suppression unit 20, and a nonlinear echo suppression unit 30. The echo suppression device 1 may further include a general configuration related to noise canceling.

[0027] The reference signal adjustment unit 10 is a functional unit that acquires a reference signal input from the signal output terminal 532 on the receiving side and generates a delayed reference signal based on the reference signal. The reference signal adjustment unit 10 generates a plurality of delayed reference signals by applying a plurality of mutually different delay times.

[0028] The linear echo suppression unit 20 is a functional unit that generates an adaptive filter using a reference signal and suppresses linear echoes in an input signal. The adaptive filter used by the linear echo suppression unit 20 has linear filter characteristics. The linear echo suppression unit 20 mainly includes a filter generation unit 21, a filter correction unit 22, and an echo removal unit 23.

[0029] The filter generation unit 21 is a functional unit that obtains a convergence value of the adaptive filter based on the delayed reference signal generated by the reference signal adjustment unit 10. The filter generation unit 21 is configured to be able to execute a learning algorithm using linear processing. The learning algorithm using linear processing is, for example, NLMS or LMS, but is not limited to these, and any known algorithm can be applied.

[0030] The filter generation unit 21 configures the adaptive filter as a non-recursive (FIR, Finite Impulse Response) filter, for example, as shown in Equation (1). In Equation (1), x(k) is the input signal, y(k) is the output signal, and h(l) is the filter coefficient of the multiplier. Furthermore, l is the index of the coefficient, and N is the filter length.

number

[0031] The filter generation unit 21 acquires a plurality of delayed reference signals and obtains convergence values ​​of a plurality of adaptive filters based on the delayed reference signals. The convergence values ​​of the adaptive filters include the convergence values ​​of the filter coefficients of each index of the non-recursive filter. The filter generation unit 21 stores the convergence values ​​of the adaptive filters in an appropriate storage unit.

[0032] In addition, the filter generation unit 21 can determine that the adaptive filter has converged, for example, when the variation in the convergence value falls within a certain range, when the convergence value falls within a certain range, or when a certain amount of time has passed since the start of processing.

[0033] The filter correction unit 22 is a functional unit that corrects the adaptive filter generated by the filter generation unit 21. The filter correction unit 22 corrects an adaptive filter that has converged using a delayed reference signal to which a predetermined delay time has been applied, based on an adaptive filter that has converged using a delayed reference signal to which another delay time has been applied. The predetermined delay time is, for example, a delay time expected depending on the time it takes for sound to be emitted from the speaker 52 and picked up by the microphone 51 in the voice communication system 100. The filter correction unit 22 will be described in detail later.

[0034] The echo removal unit 23 is a functional unit that removes linear echo from the input signal picked up by the microphone 51. The echo removal unit 23 removes linear echo using an adaptive filter that has been modified by the filter modification unit 22. The processing performed by the echo removal unit 23 is already known, so a description thereof will be omitted. The signal output from the echo removal unit 23 is input to the nonlinear echo suppression unit 30.

[0035] The nonlinear echo suppressor 30 is a functional unit that executes a learning algorithm using nonlinear processing. Any known algorithm can be applied as the learning algorithm using nonlinear processing. The signal output from the nonlinear echo suppressor 30 is output to a signal input terminal 531 on the transmitting side and transmitted to a communication device 54 owned by user B via a communication device 53. Note that the nonlinear echo suppressor 30 is not essential.

[0036] 3 is a flowchart showing the flow of processing performed by the echo suppression device 1. First, the echo suppression device 1 performs a learning process (step S1). The learning process may be performed in advance, or may be performed when the echo suppression device 1 or a predetermined function of the echo suppression device 1 is activated. In other words, the echo suppression device 1 needs to perform the learning process (step S1) only once before the linear echo removal process (step S7) that removes linear echo by applying an adaptive filter to an input signal.

[0037] The learning process (step S1) includes steps S2 to S6. In the learning process (step S1), first, the reference signal adjuster 10 acquires a reference signal from the signal output terminal 532 on the receiving side (step S2), and delays the reference signal by times t1, t2, ... tn to generate a plurality of delayed reference signals (step S3). Next, the filter generator 21 obtains a convergence value of the adaptive filter trained based on each delayed reference signal (step S4).

[0038] Steps S3-1 to S3-n that make up step S3 may be performed simultaneously or sequentially. Step S4 is made up of steps S4-1 to S4-n, and is performed following steps S3-1 to S3-n, respectively.

[0039] Next, the filter correction unit 22 compares the convergence values ​​of the multiple adaptive filters generated in step S4 (step S5), and corrects the adaptive filters based on the comparison results (step S6).

[0040] Fig. 4 is a graph showing how the filter correction unit 22 compares the convergence values ​​of multiple adaptive filters generated by the filter generation unit 21. The horizontal axis of Fig. 4 represents the index of the filter coefficient, and the vertical axis represents the magnitude of the convergence value. Fig. 4 shows values ​​for delay times tn (n is a natural number) of 160, 170, 180, 190, and 200 (unit: samples). Note that the sampling frequency used to create the graph shown in Fig. 4 was 16 kHz, and the sampling interval for 160 samples, for example, was 10 ms. However, the sampling frequency is not limited to this.

[0041] In graph L160 of an adaptive filter generated based on a delayed reference signal with a delay time of 160 ms, a positive peak appears at index 9, and a negative peak (peak P160) appears at index 40. Peak P160 is an element that changes with changes in delay time tn, and moves horizontally as the delay time tn changes. For example, peak P160 moves to peak P170 (index 30) in graph L170 of an adaptive filter generated based on a delayed reference signal with a delay time of 170 ms, and moves to peak P180 (index 20) in graph L180 of an adaptive filter generated based on a delayed reference signal with a delay time of 180 ms. That is, in the example shown in FIG. 4, the longer the delay time applied to the reference signal, the smaller the index of the filter coefficient at which the peak appears. In this way, peaks that change with delay time indicate that an appropriate acoustic path has been learned.

[0042] The determination of whether or not a peak has moved can be made by applying an appropriate known technique, such as determining whether or not a peak has moved by pattern matching.

[0043] On the other hand, in graph L160, the peak at index 9 does not shift horizontally even when the delay time tn is changed, and no change occurs according to the delay time tn. If an incorrect local optimum solution is reached, a peak is likely to occur at the same filter coefficient regardless of the given delay time. Furthermore, even if a signal other than the audio signal from the receiver's signal output terminal 532 is mixed into the reference signal, the signal is transmitted without being affected by delay, resulting in a peak at the same filter coefficient regardless of the given delay time. Signal contamination can occur, for example, when a signal or generated noise is introduced on the CPU board, when the reference signal has high autocorrelation, or when a signal that changes significantly is introduced via wiring immediately after the speaker amplifier 55. Signal contamination can also occur due to vibrations in the echo suppression device 1 itself.

[0044] Therefore, the filter correction unit 22 corrects the adaptive filter for the element (index I1) that does not change according to the delay time, and reduces the magnitude of the filter coefficient below the convergence value (reduces the learning update width). In this embodiment, the filter correction unit 22 sets the filter coefficient of index I1 to 0 (stops learning). This removes the influence of inappropriate learning from the adaptive filter used for echo cancellation, and makes it possible to improve the adaptive filter to an appropriate one. Note that setting the filter coefficient to 0 (stopping learning) is included in the form of reducing the magnitude of the filter coefficient (reducing the learning update width).

[0045] The filter correction unit 22 determines whether to extract an element (index) based on a threshold value. In the example shown in Fig. 4, the filter correction unit 22 extracts index 9, whose convergence value is equal to or greater than the threshold value, as an element that does not change depending on the delay time tn, and does not extract other indexes.

[0046] For example, the filter correction unit 22 may use any value as the threshold value. For example, the user may set the threshold value in advance based on the design values ​​or experimental values ​​of the microphone 51 and speaker 52 to be used. This configuration eliminates the need for a process to calculate the threshold value, thereby simplifying the process.

[0047] Furthermore, for example, the filter correction unit 22 may calculate the threshold value based on the magnitude of the convergence value. For example, the filter correction unit 22 may calculate the magnitude of the convergence value of each index for an arbitrary delay time, and set the threshold value to a value obtained by subtracting a predetermined value or a predetermined percentage from the maximum value of the calculated magnitude. For example, the filter correction unit 22 may calculate one or more elements that change according to the delay time by pattern matching or the like, calculate the maximum value of the convergence value of the index that changes according to the delay time, and set the threshold value to a value obtained by subtracting a predetermined value or a predetermined percentage from the calculated maximum value. With this configuration, it is possible to automatically determine an appropriate threshold value according to the status of the reference signal.

[0048] If the magnitude of the filter coefficient of an index that does not change depending on the delay time is equal to or greater than the threshold value thus determined or set, the filter correction unit 22 reduces the magnitude of the filter coefficient, thereby reducing the filter coefficient of an index that has a large effect on removing linear echoes.

[0049] In addition, the filter correction unit 22 may reduce the filter coefficients of index I1 and the indexes within a predetermined range before and after it, taking into consideration that peaks appearing in the opposite directions to index I1 appear at the indexes to the left and right of index I1, where a peak that does not change depending on the delay time occurs, in the graph of the filter coefficients.

[0050] Returning to the explanation of Fig. 3, following the learning process (step S1), the echo canceller 23 collects an input signal from the microphone 51 and applies the adaptive filter corrected in step S6 to the input signal to cancel linear echoes (step S7). The signal from which the linear echo has been cancelled has nonlinear echoes suppressed by the nonlinear echo suppressor 30 and is output to the communication device 54 (step S8).

[0051] Steps S7 and S8 are performed continuously while the echo suppression device 1 is operating. That is, the echo suppression device 1 constantly removes linear echoes by using the adaptive filter corrected once by the filter correction unit 22 in step S6.

[0052] According to this embodiment, multiple adaptive filters trained using multiple reference signals with different delay times are compared with each other, and the magnitude of the filter coefficients of elements (indexes) that do not change depending on the delay time is set to 0 (learning is stopped), thereby making it possible to perform echo cancellation using an appropriate adaptive filter that has not been erroneously trained. Furthermore, by stopping learning of indices that do not change depending on the delay time, the calculation load can be reduced, and the processing load of the arithmetic device can be alleviated. Furthermore, because the processing load of the arithmetic device can be reduced, the echo suppression device 1 can be configured at low cost.

[0053] For example, in order to reduce the processing load of the arithmetic device, a method of setting appropriate initial values ​​for the adaptive filter coefficients is conceivable. Possible methods of setting appropriate initial values ​​for the adaptive filter coefficients include, for example, a method of starting learning by setting each filter coefficient to a state of 0 as the initial value (Method 1), and a method of estimating an index and the magnitude (peak value) of the filter coefficient at that index based on prior physical information, i.e., the configuration of the devices constituting the voice communication system, and setting the coefficients of other indexes to 0 (Method 2). In Method 2, the index of the filter coefficient can be obtained by summing the contributions estimated for the speaker amplifier 55, the physical configuration between the speaker 52 and the microphone 51, and appropriate devices required from the microphone 51 to the echo suppression device 1, and subtracting the contribution from the reference signal adjustment unit 10. The peak value can be obtained by multiplying the contributions of the speaker amplifier 55, the speaker 52, the physical configuration between the speaker 52 and the microphone 51, and appropriate devices required from the microphone 51 to the echo suppression device 1.

[0054] However, with Method 1, it is not possible to eliminate the risk of falling into an erroneous local optimum solution. Furthermore, with Method 2, it is difficult to accurately calculate the physical characteristics of the device using digital values, and as with Method 1, it is not possible to eliminate the risk of falling into a local optimum solution due to errors that occur. In contrast, the echo suppression device 1 performs echo cancellation using an appropriate adaptive filter that has not undergone erroneous learning, so it has high echo cancellation accuracy and can determine an appropriate adaptive filter with a light processing load.

[0055] Furthermore, according to this embodiment, the learning process (steps S1 to S6) is performed in advance or only once when a predetermined function of the echo suppression device 1 is activated, thereby reducing the processing load on the arithmetic device and enabling echo removal using an adaptive filter that has not undergone erroneous learning.

[0056] In this embodiment, the filter correction unit 22 sets the filter coefficient of the element that does not change according to the delay time (index I1 in FIG. 4) to 0, but setting the filter coefficient to 0 is not essential; it is sufficient to reduce the magnitude of the filter coefficient (reduce the learning update width). For example, the filter correction unit 22 may reduce the magnitude of the filter coefficient of index I1 and hold a value other than 0 (for example, 0.5). When the magnitude of the filter coefficient of index I1 is reduced, learning is also performed little by little for index I1, so appropriate learning can be performed even when the echo path changes due to a change in the environment of the space in which the microphone 51 and speaker 52 are located, for example.

[0057] Furthermore, in this embodiment, the linear echo removal process (step S7) is performed using an adaptive filter in which the filter coefficient of index I1 is set to 0 in the learning process (step S5). However, the linear echo removal process (step S7) may also be performed while gradually decreasing the filter coefficient of index I1 of the adaptive filter. For example, in the example shown in FIG. 4, the magnitude of the filter coefficient of index I1 is calculated to be approximately 2. Therefore, the filter correction unit 22 may set the filter coefficient of index I1 to 2 initially (t=0) and continuously correct the adaptive filter over an arbitrary time period (t=0 to t1) so that the filter coefficient reaches a final value (e.g., 0) after an arbitrary time period has elapsed (t=t1). In this case, the final value may be 0 or a small value other than 0 (e.g., 0.5). This makes it possible to prevent noise or excessive removal caused by abrupt changes in the adaptive filter.

[0058] Furthermore, in this embodiment, the filter correction unit 22 extracts an index (index I1 in FIG. 4) whose convergence value is equal to or greater than a threshold value from among the indexes that do not change according to the delay time, and reduces the filter coefficient of the extracted index. However, the filter coefficient of an index that does not change according to the delay time regardless of the threshold value may also be reduced. In this case, it is desirable to set the filter coefficient to an appropriate value such as 0.5 rather than 0. This makes it possible to converge to an appropriate adaptive filter even if unnecessary reduction has been performed.

[0059] <Modification of the first embodiment> The echo suppression device 1A according to the modification differs from the echo suppression device 1 according to the first embodiment in that it has a plurality of reference signal adjustment units that generate delayed reference signals, and a plurality of filter generation units that generate one filter from each delayed reference signal. In the following description, the same components as those in the first embodiment are denoted by the same reference symbols, and description thereof will be omitted.

[0060] 5 is a block diagram showing a schematic configuration of an echo suppression device 1A according to a modification of the first embodiment. The echo suppression device 1A mainly includes a reference signal adjustment unit 10A, a linear echo suppression unit 20A, and a nonlinear echo suppression unit 30.

[0061] The reference signal adjuster 10A has a plurality of reference signal adjusters 10-1...10-n (n is a natural number). The reference signal adjusters 10-1...10-n are functional units that acquire reference signals and generate delayed reference signals based on the reference signals. Different delay times are stored in the reference signal adjusters 10-1...10-n beforehand. The reference signal adjusters 10-1...10-n each generate and output a delayed reference signal based on the stored delay time.

[0062] The linear echo suppressor 20A is a functional unit that generates an adaptive filter using a reference signal and suppresses linear echoes in an input signal, and is mainly provided with a filter generator 21A, a filter corrector 22A, and an echo remover .

[0063] The filter generating unit 21A has a plurality of filter generating units 21-1...21-n (n is a natural number). The number of reference signal adjusting units 10-1...10-n and the number of filter generating units 21-1...21-n are the same.

[0064] The filter generators 21-1...21-n are functional units that generate adaptive filters based on the delayed reference signals output from the reference signal adjusters 10-1...10-n, respectively. The process of generating adaptive filters by the filter generators 21-1...21-n is similar to that of the filter generator 21, and therefore a description thereof will be omitted.

[0065] The filter correction unit 22A is a functional unit that corrects an adaptive filter generated by any filter generation unit 21 (for example, filter generation unit 21-1) among the filter generation units 21-1...21-n based on the adaptive filters generated by the filter generation units 21-1...21-n. The optional filter generation unit 21 generates an adaptive filter based on a delayed reference signal to which a predetermined delay time (for example, the delay time expected due to the time it takes for sound to be emitted from the speaker 52 and picked up by the microphone 51 in the voice communication system 100) has been given.

[0066] The filter correction unit 22A compares the convergence values ​​of the multiple adaptive filters generated by the filter generation units 21-1...21-n, and corrects the adaptive filters for elements (indexes) that do not change depending on the delay time, thereby reducing the magnitude of the filter coefficients. The process by which the filter correction unit 22A corrects the adaptive filters is the same as that of the filter correction unit 22, and therefore a description thereof will be omitted.

[0067] The echo removal unit 23 is a functional unit that removes echo from the input signal collected by the microphone 51 using the adaptive filter modified by the filter modification unit 22A.

[0068] According to the echo suppression device 1A of this modification, the magnitudes of the filter coefficients generated by the filter generation units 21-1...21-n are compared, so that it is possible to reduce the amount of memory used. For example, in the echo suppression device 1, the filter generation unit 21 sequentially generates adaptive filters with different delay times, so that the filter correction unit 22 needs a functional unit for storing each of the adaptive filters (magnitudes of the filter coefficients) generated by the filter generation unit 21, but in the echo suppression device 1A, the filter generation units 21-1...21-n each generate different adaptive filters, so that when comparing multiple adaptive filters, a functional unit for storing multiple adaptive filters is not needed, and it is possible to reduce the amount of memory used.

[0069] <Second embodiment> The echo suppression device 1 according to the first embodiment of the present invention performs the learning process (step S1) only once (for example, in advance or when the echo suppression device 1 is started), but the echo suppression device may perform the learning process continuously.

[0070] An echo suppression device 2 according to the second embodiment of the present invention is configured to continuously perform a learning process (step S1), i.e., correct an adaptive filter. The following description of the echo suppression device 2 will focus on the differences from the first embodiment. In the following description, the same components as those in the first embodiment will be denoted by the same reference numerals, and description thereof will be omitted.

[0071] 6 is a block diagram showing a schematic configuration of the echo suppressor 2. The echo suppressor 2 mainly includes a reference signal adjuster 10B, a linear echo suppressor 20B, and a nonlinear echo suppressor 30.

[0072] The reference signal adjuster 10B is a functional unit that continuously acquires reference signals and continuously generates delayed reference signals based on the reference signals. The reference signal adjuster 10B continuously generates delayed reference signals while acquiring the reference signals. The processing performed by the reference signal adjuster 10B is similar to that performed by the reference signal adjuster 10, except that the processing is performed continuously.

[0073] The linear echo suppressor 20B is a functional unit that generates an adaptive filter using a reference signal and performs processing to suppress linear echoes in an input signal. The linear echo suppressor 20B mainly includes a filter generator 21B, a filter corrector 22B, and an echo remover 23A.

[0074] The filter generation unit 21B is a functional unit that continuously obtains a convergence value of the adaptive filter based on the delayed reference signal continuously generated by the reference signal adjustment unit 10B. The processing performed by the filter generation unit 21B is similar to that of the filter generation unit 21, except that the processing is continuously performed.

[0075] The filter correction unit 22B is a functional unit that continuously corrects the adaptive filters that are continuously generated by the filter generation unit 21 B. The processing performed by the filter correction unit 22B is similar to that performed by the filter correction unit 22, except that the processing is performed continuously.

[0076] The echo removal unit 23A is a functional unit that removes echo from the input signal collected by the microphone 51. Adaptive filters corrected by the filter correction unit 22B are continuously input to the echo removal unit 23A, and the echo is removed using the continuously input adaptive filters. The process performed by the echo removal unit 23A is different from the process performed by the echo removal unit 23 in the adaptive filters used, but is otherwise similar. The signal output from the echo removal unit 23A is input to the nonlinear echo suppression unit 30.

[0077] 7 is a flowchart showing the flow of processing performed by the echo suppression device 2. While the echo suppression device 2 is operating and acquiring a reference signal, for example, during a call, a learning process (step S1), an echo removal process (step S7), and a signal output process (step S8) are performed continuously.

[0078] According to this embodiment, while the echo suppression device 2 is operating and acquiring a reference signal, the learning process, i.e., the process of correcting the adaptive filter, is continuously performed, and the echo is removed using the continuously corrected adaptive filter, so that the echo can be removed appropriately even if the echo path changes during use.

[0079] <Third embodiment> The echo suppressing device 2 according to the second embodiment of the present invention continuously performs the learning process (correction of the adaptive filter), but may stop correcting the adaptive filter under certain conditions.

[0080] An echo suppression device 3 according to the third embodiment of the present invention is configured to stop modifying the adaptive filter in the event of double talk. The following description of the echo suppression device 3 will focus on the differences from the first and second embodiments. In the following description, the same components as those in the first and second embodiments will be assigned the same reference numerals, and description thereof will be omitted.

[0081] 8 is a block diagram showing a schematic configuration of the echo suppressor 3. The echo suppressor 2 mainly includes a reference signal adjuster 10C, a linear echo suppressor 20C, a nonlinear echo suppressor 30, and a double-talk detector .

[0082] The double-talk detection unit 40 is a functional unit that detects whether the voice signal input to the echo suppression device 3 is in a single-talk state or a double-talk state. Here, single-talk refers to a state in which either user A or user B is emitting voice and a signal is transmitted to either the transmitting-side signal path or the receiving-side signal path (near-end speech or far-end speech). Double-talk refers to a state in which both user A and user B are emitting voice and a signal is transmitted simultaneously to the transmitting-side signal path and the receiving-side signal path (near-end speech and far-end speech).

[0083] The double-talk detection unit 40 sequentially compares the power spectrum value of the reference signal with the power spectrum value of the input signal for each frequency band and detects whether or not a double-talk state is occurring based on the comparison result. For example, the double-talk detection unit 40 holds a frequency mask that acquires the maximum value of the power spectrum value of the signal transmitted through the transmitting-side signal path during one-sided speech (single talk) on the far-end side, in which only sound output from the speaker 52 is input to the microphone 51. The double-talk detection unit 40 compares the power spectrum value of the input signal picked up by the microphone 51 with the value of the frequency mask for each frequency band, and if the number of frequency bands in which the value of the input signal exceeds the value of the frequency mask is equal to or greater than a certain value, it detects that sound is being input from the microphone 51 and that a signal is being transmitted through the transmitting-side signal path (there is near-end speech). Furthermore, for example, the double-talk detection unit 40 compares the power spectrum value of the reference signal with the value of the frequency mask for each frequency band, and if the number of frequency bands in which the value of the reference signal exceeds the value of the frequency mask is equal to or greater than a certain value, it detects that a signal is being transmitted through the receiving-side signal path (there is far-end speech).

[0084] However, the double-talk detection unit 40 may use various other known methods to detect whether the state is single-talk or double-talk.

[0085] The detection result by the double-talk detector 40 is input to the reference signal adjuster 10C. The reference signal adjuster 10C is a functional unit that continuously acquires reference signals when a double-talk state is not occurring and continuously generates delayed reference signals based on the reference signals. The process by which the reference signal adjuster 10C continuously generates delayed reference signals is the same as that of the reference signal adjuster 10B, and therefore a description thereof will be omitted. The delayed reference signals generated by the reference signal adjuster 10C are input to the linear echo suppressor 20C.

[0086] The linear echo suppressor 20C is a functional unit that generates an adaptive filter using a reference signal and performs processing to suppress linear echoes in the input signal. The linear echo suppressor 20B mainly includes a filter generator 21C, a filter corrector 22C, and an echo remover 23B.

[0087] The filter generation unit 21C is a functional unit that obtains convergence values ​​of a plurality of adaptive filters based on a delayed reference signal when the delayed reference signal is generated by the reference signal adjustment unit 10C. The process by which the filter generation unit 21C continuously obtains convergence values ​​of the adaptive filters is the same as that of the filter generation unit 21B, and therefore a description thereof will be omitted.

[0088] The filter correction unit 22C is a functional unit that corrects the adaptive filter generated by the filter generation unit 21B when multiple adaptive filters are generated by the filter generation unit 21C. The process by which the filter correction unit 22C corrects the adaptive filter is the same as that of the filter correction unit 22B, and therefore a description thereof will be omitted.

[0089] The echo removal unit 23B is a functional unit that removes echo from the input signal collected by the microphone 51. When an adaptive filter corrected by the filter correction unit 22B is input to the echo removal unit 23B, the echo is removed using the input adaptive filter. When an adaptive filter corrected by the filter correction unit 22B is not input, the echo removal unit 23B removes echo using the adaptive filter input last (immediately before) from the filter correction unit 22B. For example, the echo removal unit 23B may have a functional unit that stores the adaptive filter corrected by the filter correction unit 22B, and the echo removal unit 23B may update the adaptive filter stored in the functional unit every time an adaptive filter is input from the filter correction unit 22B, and remove echo from the input signal using the adaptive filter stored in the functional unit. The processing of the echo removal unit 23B is already known, so a description thereof will be omitted. The signal output from the echo removal unit 23B is input to the nonlinear echo suppression unit 30.

[0090] 9 is a flowchart showing the flow of processing performed by the echo suppression device 3. While the echo suppression device 3 is operating and acquiring a reference signal, for example, during a call, the double-talk detection unit 40 detects whether the state is single talk or double talk (step S10).

[0091] If the double-talk detection unit 40 does not detect a double-talk state (No in step S10), learning processing (step S11) and echo cancellation processing (step S12) are performed. The processing in step S11 is the same as step S1, so its explanation will be omitted. In step S12, echo cancellation is performed using the adaptive filter corrected in the immediately preceding step S6.

[0092] If the double-talk detection unit 40 detects that a double-talk state is occurring (Yes in step S10), the learning process (step S11) is not performed, and the process proceeds to the echo cancellation process (step S12). In the echo cancellation process (step S12), the echo is canceled using the adaptive filter corrected in the immediately preceding step S6. The process of canceling the echo in the echo cancellation process (step S12) is the same as in step S7, and therefore will not be described here.

[0093] Following the echo removal process (step S12), a signal output process (step S13) is performed. The process of step S13 is the same as step S8, so a description thereof will be omitted. Thereafter, the process returns to step S10, and the process shown in FIG. 9 is repeated.

[0094] According to this embodiment, when the double-talk detector detects double-talk, i.e., when there are many disturbances and it seems that appropriate correction processing cannot be performed, the correction processing of the adaptive filter can be interrupted. Furthermore, by interrupting the correction processing of the adaptive filter when there are many disturbances, the calculation load can be reduced.

[0095] In this embodiment, the detection results of the double-talk detection unit 40 are input to the reference signal adjustment unit 10C, but the detection results of the double-talk detection unit 40 may also be input to the filter generation unit 21C or the filter correction unit 22C. For example, when the detection results of the double-talk detection unit 40 are input to the filter correction unit 22C, multiple delayed reference signals and convergence values ​​of the adaptive filter are found even in the case of double talk, and correction of the adaptive filter is stopped.

[0096] The above describes an embodiment of the present invention in detail with reference to the drawings, but the specific configuration is not limited to this embodiment, and design changes and the like are also included within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0097] 1, 1A, 2, 3: Echo suppressor 10, 10A, 10B, 10C, 10-1, 10-n: Reference signal adjustment section 20, 20A, 20B, 20C: Linear echo suppression section 21, 21A, 21B, 21C, 21-1, 21-n: filter generation unit 22, 22A, 22B, 22C: Filter correction section 23, 23A, 23B: Echo cancellation section 30: Nonlinear echo suppression unit 40: Double talk detector 50: Terminal 51: Microphone 52: Speaker 53: Communication equipment 54: Communication equipment 55: Speaker amplifier 100: Voice communication system 531: Signal input terminal 532: Signal output terminal

Claims

1. An echo suppression device for removing echo from an input signal picked up by a microphone of a terminal having a speaker and a microphone, a reference signal adjustment unit that generates a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiver side signal path that transmits a signal to the speaker; a filter generation unit for obtaining convergence values ​​of a plurality of adaptive filters based on the plurality of delayed reference signals; a filter correction unit that compares a plurality of the convergence values ​​and corrects the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change according to a change in the delay time, is smaller than the convergence value; an echo canceller that applies the adaptive filter corrected by the filter corrector to the input signal to cancel a linear echo; An echo suppression device comprising:

2. The filter correction unit sets the first filter coefficient to zero.

2. The echo suppression device according to claim 1.

3. The filter correction unit gradually decreases the first filter coefficient from the convergence value.

3. The echo suppression device according to claim 1 or 2.

4. The filter correction unit sets a filter coefficient, among elements that do not change according to a change in the delay time, when the convergence value is equal to or greater than a threshold value as the first filter coefficient.

3. The echo suppression device according to claim 1 or 2.

5. The filter correction unit calculates the threshold value based on the convergence value.

5. The echo suppression device according to claim 4.

6. the reference signal adjustment unit continuously generates a plurality of the delayed reference signals while acquiring the reference signal; the filter generation unit obtains the plurality of convergence values ​​after the reference signal adjustment unit generates the delayed reference signal; When the plurality of convergence values ​​are obtained, the filter correction unit compares the plurality of convergence values ​​and corrects the adaptive filter.

3. The echo suppression device according to claim 1 or 2.

7. A double talk detector is provided to detect whether the device is in a single talk state or a double talk state. The reference signal adjustment unit generates the plurality of delayed reference signals when the double talk state is not detected.

7. The echo suppression device according to claim 6.

8. The reference signal adjustment unit generates the delayed reference signal only once before the echo canceller cancels the linear echo.

3. The echo suppression device according to claim 1 or 2.

9. a reference signal adjustment step of generating a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiver side signal path that transmits a signal to the speaker of a terminal having a speaker and a microphone; a filter generation step of obtaining convergence values ​​of a plurality of adaptive filters based on the plurality of delayed reference signals; a filter correction step of comparing a plurality of the convergence values ​​and correcting the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change according to a change in the delay time, is smaller than the convergence value; an echo removal step of removing a linear echo by applying the adaptive filter corrected in the filter correction step to the input signal picked up by the microphone; 10. An echo suppression method comprising:

10. Computer, a reference signal adjustment unit that generates a plurality of delayed reference signals by applying a plurality of different delay times to a reference signal transmitted through a receiver signal path that transmits a signal to a speaker of a terminal having a speaker and a microphone; a filter generation unit that obtains convergence values ​​of a plurality of adaptive filters based on the plurality of delayed reference signals; a filter correction unit that compares a plurality of the convergence values ​​and corrects the adaptive filter so that a first filter coefficient, which is a filter coefficient for an element that does not change according to a change in the delay time, is smaller than the convergence value; an echo canceller that applies the adaptive filter corrected by the filter corrector to the input signal picked up by the microphone to cancel a linear echo; An echo suppression program characterized by causing the program to function as follows.

Citation Information

Patent Citations

  • Method for adaptive control of digital echo canceller in telecommunication system

    JP1995273691A

  • Method and system for eliminating echo for multiplex channel

    JP2000196507A

  • Speech processor and echo removing method

    JP2009033549A

  • Echo canceller apparatus and its method

    JP2010011265A

  • Echo canceller

    JP2011061449A