Echo suppression device, echo suppression method, and echo suppression program product
By generating or selecting the optimal masking signal corresponding to the size of the received signal in the call-side signal path, the problem of improper echo suppression when the sound is small is solved, and effective echo suppression and stable call are achieved in the case of less sound is achieved.
Patent Information
- Application Number
- CN202180013053.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-13
- Filing Date
- 2021-04-07
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-04-07
AI Technical Summary
The prior art cannot effectively detect vocalization and appropriately suppress echoes when the vocalization is small, resulting in the possibility of voice loss of proximal speakers.
By providing a mask signal storage unit, a mask signal selection unit and a dual-end call detection unit in the signal path on the sending side, an optimal mask signal corresponding to the size of the received signal is generated or selected, and an echo suppression process is performed when a microphone input is not sounded.
Even when the sound is small, it can accurately detect sound and appropriately suppress echoes to ensure stable call quality.
Smart Images

Figure CN115053460B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an echo suppression device, an echo suppression method and an echo suppression program. Background Art
[0002] Patent document 1 discloses an echo suppression device that performs the following processing: a mask signal based on the power spectrum of a learning signal transmitted in a receiver-side signal path and a value of the power spectrum of an input signal input from a microphone are compared in each frequency band to detect whether a double-ended call state is present. If it is detected that a signal is not being transmitted in a transmitter-side signal path but is being transmitted in a receiver-side signal path, an echo suppressor is used to suppress the echo of the input signal.
[0003] Prior art literature
[0004] Patent Literature
[0005] Patent Document 1: Japanese Patent Application Publication No. 2018-201147 Summary of the Invention
[0006] -Problems to be solved by the invention-
[0007] However, in the call signal processing device described in Patent Document 1, a masking signal is generated assuming that the signal in the receiver-side signal path is larger. Therefore, when the user on the microphone side (near-end speaker) speaks softly and the receiver signal transmitted in the receiver-side signal path is larger, the echo suppressor will strongly act on the input signal transmitted in the receiver-side signal path, and the voice of the near-end speaker may disappear.
[0008] The present invention has been made in view of the above circumstances, and an object of the present invention is to provide an echo suppression device, an echo suppression method, and an echo suppression program capable of detecting speech even when the speech is small and appropriately suppressing the echo.
[0009] -Methods for solving the problem-
[0010] In order to solve the above-mentioned problems, the echo suppression device involved in the present invention is, for example, the following echo suppression device, which is provided in a sending-side signal path, and the sending-side signal path transmits an input signal input from the microphone in a near-end terminal having a speaker and a microphone, and is characterized in that it comprises: a masking signal storage unit, which stores a basic masking signal, which is one or more masking signals generated based on a learning signal transmitted in the sending-side signal path when a sound is output from the speaker without a sound being input to the microphone; and a masking signal selection unit, which selects a basic masking signal from the receiving-side signal path whenever a basic masking signal is obtained. A sampling point of a transmitted received call signal is generated or selected from the basic masking signal in sequence based on the received call signal obtained within a given period before the sampling point is obtained; a double-end talk detection unit detects whether a double-end talk state is present based on the result of comparing the input signal with the optimal masking signal each time the optimal masking signal is generated or selected; and an echo suppressor performs echo suppression processing on the input signal in sequence when the double-end talk detection unit detects that no sound is input to the microphone and the received call signal includes sound.
[0011] According to the echo suppression device of the present invention, each time a sample point of a received speech signal transmitted through a receiver-side signal path that transmits the signal to a speaker is acquired, an optimal masking signal is sequentially generated or selected from one or more masking signals (i.e., base masking signals) generated based on a learning signal, based on the received speech signal acquired within a predetermined period prior to the sample point. Each time an optimal masking signal is selected, a double-talk state is sequentially detected based on the result of comparing the input signal with the optimal masking signal. If no speech is detected in the microphone input and the received speech signal includes speech, echo suppression processing is sequentially performed on the input signal. By varying the size of the masking signal in accordance with the size of the received speech signal, speech can be detected and echo can be appropriately suppressed even in the case of relatively small speech.
[0012] The system comprises a masking signal generator configured to generate multiple masking signals by varying the magnitude of the learning signal; a masking signal storage unit configured to store the multiple masking signals generated by the masking signal generator as the basic masking signal; and a masking signal selection unit configured to select the optimal masking signal from the basic masking signals based on the magnitude of the input signal. This accurately stores the frequency characteristics of the residual echo for each call level, and allows the magnitude of the masking signal to be varied according to the magnitude of the call signal. Furthermore, the echo suppressor's activation mode does not need to be frequently changed, ensuring stable communication.
[0013] The system comprises a masking signal generating unit that generates a single masking signal based on the learning signal; a masking signal storage unit that stores the single masking signal generated by the masking signal generating unit as the basic masking signal; and a masking signal selecting unit that generates the optimal masking signal by multiplying the basic masking signal by a coefficient based on the magnitude of the input signal. This allows accurate storage of the frequency characteristics of the residual echo for each reception level, and allows the magnitude of the masking signal to be varied according to the magnitude of the reception signal. Furthermore, there is no need to store multiple basic masking signals, thereby reducing memory usage.
[0014] The present invention includes a signal measuring unit configured to measure a first time, a time when a signal is no longer transmitted through the transmitting-side signal path, when a state in which a sound is output from the speaker but no sound is input to the microphone changes to a state in which a sound is not input to the microphone and no sound is output from the speaker. The masking signal selecting unit sequentially generates or selects the optimal masking signal using the first time as the predetermined period. Thus, the predetermined period can be determined based on the length of an echo generated by the received signal.
[0015] The system comprises a first power spectrum calculation unit that calculates an input signal power spectrum, which is a power spectrum of the input signal, and a learning power spectrum, which is a power spectrum of the learning signal, wherein the masking signal is the maximum value of each frequency band of the learning power spectrum obtained within a certain interval, and the optimal masking signal has a value for each frequency band. The double-talk detection unit detects whether a double-talk state exists based on a result obtained by comparing the value of the input signal power spectrum with the value of the optimal masking signal for each frequency band. This makes it possible to accurately detect a double-talk state.
[0016] The system includes a second power spectrum calculation unit configured to calculate a power spectrum of the received signal, namely, a received signal power spectrum. The masker signal selection unit compares the maximum value of the received signal power spectrum with the optimal masker signal for each frequency band to generate or select the optimal masker signal. This allows for the appropriate generation or selection of the optimal masker signal, taking into account the frequency characteristics of the received signal.
[0017] The double-talk detector compares the input signal power spectrum with the optimal masking signal for each frequency band. If the number of frequency bands in which the input signal power spectrum exceeds the optimal masking signal is less than a first threshold, or if the integral value of the region in which the input signal power spectrum exceeds the optimal masking signal is less than a second threshold, the detector detects that no signal is being transmitted in the receiving-side signal path. This enables accurate near-end speech detection.
[0018] To address the aforementioned issues, the echo suppression method according to the present invention is characterized, for example, by comprising the steps of: generating and storing one or more basic masking signals as masking signals based on a learning signal, the learning signal being transmitted in a speaker-side signal path that transmits a signal input from the microphone when no speech is spoken into a microphone input of a near-end terminal and a speaker of the near-end terminal is outputting sound; sequentially generating or selecting, each time a sampling point of a received speech signal transmitted in a speaker-side signal path that transmits the signal to the speaker is obtained, optimal masking signals having a magnitude corresponding to the magnitude of the input signal input from the microphone based on the received speech signal obtained within a predetermined period before the sampling point and the basic masking signals; sequentially detecting whether a double-talk state is present based on a comparison result of the input signal and the optimal masking signals when the optimal masking signals are selected; and performing echo suppression processing on the input signal to suppress an echo when it is detected that no speech is spoken into the microphone input and the received speech signal includes speech.
[0019] In order to solve the above-mentioned problems, the echo suppression program involved in the present invention is the following echo suppression program, for example, set in a sending-side signal path, which transmits a signal input from the microphone in a near-end terminal having a speaker and a microphone, and is characterized in that a computer functions as the following elements: a masking signal storage unit, which stores a basic masking signal, which is one or more masking signals generated based on a learning signal transmitted in the sending-side signal path when a sound is output from the speaker without a sound input to the microphone; a masking signal selection unit, which selects a basic masking signal each time a basic masking signal is obtained in the receiving-side signal path transmitting a signal to the speaker. The invention also provides a method for generating or selecting an optimal masking signal corresponding to the magnitude of the received signal from the basic masking signal based on the sampling point of the received signal transmitted in the path, based on the received signal obtained within a given period before the sampling point is obtained; a double-end talk detection unit, each time the optimal masking signal is selected, sequentially detecting whether a double-end talk state is present based on a result of comparing the input signal input from the microphone with the optimal masking signal; and an echo suppressor, when the double-end talk detection unit detects that no sound is input to the microphone and the received signal includes sound, sequentially performing echo suppression processing on the input signal.
[0020] -Effects of the Invention-
[0021] According to the present invention, even when the sound is small, it is possible to detect the sound and appropriately suppress the echo. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 11 is a diagram schematically showing a voice communication system 100 provided with the echo suppression apparatus 1 according to the first embodiment.
[0023] Figure 2 This is a diagram schematically showing the functional blocks of the echo suppression device 1 .
[0024] Figure 3 This is a diagram schematically showing functional blocks when the masking signal is generated in the echo suppressor 1 .
[0025] Figure 4 This is an example of a learning power spectrum at time t1.
[0026] Figure 5 The input includes Figure 4 An example of a masking signal when there are multiple learning power spectra shown in the learning power spectrum.
[0027] Figure 6 This is a diagram showing an example of two masking signals having different reception levels.
[0028] Figure 7 The diagram shows the relationship between the received signal and the input signal when there is no near-end speech. (A) shows the received signal, and (B) shows the input signal.
[0029] Figure 8 The diagram shows the relationship between the received signal and the input signal when there is no near-end speech. (A) shows the received signal, and (B) shows the input signal.
[0030] Figure 9 This is a diagram schematically showing the relationship between the maximum value of each frequency band of the received signal obtained within a given period before the sampling point and the optimal masking signal.
[0031] Figure 10 This is a diagram schematically showing the relationship between the maximum value of each frequency band of the received signal obtained within a given period before the sampling point and the optimal masking signal.
[0032] Figure 11 This is a diagram schematically showing an example of selecting an optimal masker signal based on the total power of the received signal obtained for each frequency band.
[0033] Figure 12 This is a diagram schematically showing a state in which the value of the input signal power spectrum is compared with the value of the masking signal.
[0034] Figure 13 This is a diagram schematically showing a state in which the value of the input signal power spectrum is compared with the value of the masking signal.
[0035] Figure 14This is a diagram schematically showing a state in which the value of the input signal power spectrum is compared with the value of the masking signal.
[0036] Figure 15 1 is a flowchart showing the flow of processing for sequentially reducing echoes by the echo suppression device 1 .
[0037] Figure 16 This is a diagram schematically showing the functional blocks of the echo suppression device 2 .
[0038] Figure 17 This is a diagram schematically showing a state in which the value of the input signal power spectrum when the signal level of the received signal is equal to or greater than the threshold II is compared with the value of the optimal masking signal.
[0039] Figure 18 This is a diagram schematically showing the functional blocks of the echo suppression device 3 .
[0040] Figure 19 Schematically illustrates a process of generating an optimal masker signal by the masker signal selection unit 14A.
[0041] Figure 20 This is a diagram schematically showing the functional blocks of the echo suppression device 4 .
[0042] Figure 21 This is a diagram schematically showing the functional blocks of the echo suppression device 5 .
[0043] Figure 22 This is a diagram schematically showing functional blocks when the masking signal is generated in the echo suppressor 5 .
[0044] Figure 23 1 is a flowchart showing the flow of processing for sequentially reducing echoes by the echo suppressor 5 . DETAILED DESCRIPTION
[0045] Hereinafter, embodiments of the echo suppression device according to the present invention will be described in detail with reference to the accompanying drawings. The echo suppression device is a device that suppresses echoes generated during a call in a voice communication system.
[0046] <First embodiment>
[0047] Figure 1 Schematically shows a voice communication system 100 equipped with the echo suppression apparatus 1 according to the first embodiment. The voice communication system 100 mainly includes a terminal 50 having a microphone 51 and a speaker 52 , two mobile phones 53 and 54 , a speaker amplifier 55 , and the echo suppression apparatus 1 .
[0048] Voice communication system 100 enables voice communication between a near-end speaker (user A on the near-end side) using terminal 50 (near-end terminal) and a far-end speaker (user B on the far-end side) using mobile phone 54 (far-end terminal). Speaker 52 amplifies and outputs voice signals input via mobile phone 54, while microphone 51 collects and transmits the near-end user's voice to mobile phone 54. This allows user A to conduct a hands-free conversation without having to hold mobile phone 53. Mobile phones 53 and 54 are connected via a standard telephone line.
[0049] The echo suppression device 1 is provided in a transmitting-side signal path that transmits a signal input via a microphone 51 from a terminal 50 to a mobile phone 53 .
[0050] The echo suppression device 1 can also be configured as a dedicated board installed in a communication terminal (e.g., an in-vehicle device, a conference system, or a mobile terminal) within the voice communication system 100. Furthermore, the echo suppression device 1 can also be primarily comprised of a computer system including a computing device such as a CPU (Central Processing Unit) for performing information processing, storage devices such as RAM (Random Access Memory) and ROM (Read Only Memory), and software (echo suppression program). The echo suppression program can be pre-stored in a storage medium built into a computer or other device, such as a hard disk drive (HDD), or a ROM within a microcomputer with a CPU, and installed from there into the computer. Furthermore, the echo suppression program can be temporarily or permanently stored (or stored) in a removable storage medium such as a semiconductor memory, a memory card, an optical disk, a magneto-optical disk, or a magnetic disk.
[0051] Figure 2 This is a diagram schematically showing the functional blocks of the echo suppression device 1. The echo suppression device 1 mainly includes an echo removal unit 11, a frequency analyzer (FFT unit) 12, 19, a masking signal storage unit 13, a masking signal selection unit 14, a double talk detection unit 15, an echo suppressor 16, and a restoration unit (IFFT unit) 17. Figure 2 In the figure, the upper signal path is a transmitting-side signal path that transmits an input signal from microphone 51, and the lower signal path is a receiving-side signal path that transmits a signal to speaker 52. Furthermore, the functional components of the echo suppression device 1 can be further classified into more components according to the processing content, or a single component can perform the processing of multiple components.
[0052] The echo removal unit 11 removes echoes using, for example, an adaptive filter. The echo removal unit 11 updates the filter coefficients according to a given procedure, generates a pseudo-echo signal from the signal transmitted through the receiver-side signal path, and subtracts the pseudo-echo signal from the signal transmitted through the transmitter-side signal path to remove the echo. Since adaptive filters are well known, their description will be omitted.
[0053] In addition, in this embodiment, an adaptive filter is applied to the echo removal unit 11, but other well-known echo removal technologies may also be applied to the echo removal unit 11. In addition, the echo removal unit 11 is not essential, but by using a learning signal from which a portion of the echo is removed to generate a masking signal, as will be described in detail later, even when the value of the masking signal decreases and the input signal is small, the value of the power spectrum of the input signal (hereinafter referred to as the input signal power spectrum) is likely to exceed the value of the masking signal, making it possible to more accurately detect the presence of near-end speech (user A (refer to FIG. 1 )). Figure 1 ) in this case, it is preferable to provide an echo removal unit 11.
[0054] Frequency analyzers (FFT units) 12 and 19 perform Fast Fourier Transforms (FFTs) on the signals. FFT unit 12 performs a Fast Fourier Transform on the signal transmitted through the transmitter signal path, in this case, the signal that has passed through echo removal unit 11. FFT unit 19 performs a Fast Fourier Transform on the received signal transmitted through the receiver signal path. FFT units 12 and 19 convert the function of time into a function of frequency, obtaining the result as X[i] for each frequency band i.
[0055] The mask signal storage unit 13 stores the mask signal generated by the mask signal generation unit 18 (see Figure 3 ) generates a mask signal. The generation of the mask signal will be described in detail below. Before the echo suppression device 1 performs the process of suppressing the echo, the mask signal is generated in advance.
[0056] Figure 3 1 is a diagram schematically showing functional blocks when a masking signal is generated in the echo suppressor 1. The echo suppressor 1 functionally includes a masking signal generating unit 18. The masking signal generating unit 18 mainly performs the masking signal generating process.
[0057] The masking signal generation process will be described in detail. First, after the adaptive filter learning is fully completed in the echo removal unit 11, the far-end unilateral speech (single-ended conversation) is repeatedly performed, with the speaker 52 outputting speech without near-end speech. Furthermore, the signal transmitted through the transmitting-side signal path during single-ended conversation serves as the learning signal. In the echo suppression device 1, the signal from which the echo has been removed by the echo removal unit 11 serves as the learning signal.
[0058] The learning signal is input to the FFT unit 12. The FFT unit 12 performs a Fast Fourier Transform on the learning signal and then inputs it to the masking signal generator 18. The masking signal generator 18 calculates the power spectrum of the learning signal for each fixed interval, thereby obtaining multiple learning power spectra. Here, a fixed interval is an arbitrarily determined, given time period, represented by times t1, t2, t3, etc.
[0059] The power spectrum P[i] represents the power of X[i] for each frequency element i obtained by fast Fourier transform as a function of the frequency element (see equation (1)).
[0060] [Mathematical formula 1]
[0061] P[i]=|X[i]| 2 =X[i]*X[i]···(1)
[0062] Figure 4 This is an example of a learning power spectrum at time t1. Hereinafter, the power in the power spectrum (the value on the vertical axis) is referred to as the value of the power spectrum. The horizontal axis of the power spectrum represents frequency. The masking signal generator 18 stores a plurality of learning power spectra calculated for each predetermined interval.
[0063] The masking signal generating unit 18 obtains the maximum value among the values of the plurality of learning power spectra for each frequency band, and uses the maximum value as the masking signal. Figure 5 The input includes Figure 4 An example of a masking signal when there are multiple learning power spectra is shown in FIG. Then, the masking signal generating unit 18 outputs the generated masking signal to the masking signal storing unit 13, and the masking signal storing unit 13 stores the masking signal.
[0064] In the present embodiment, the masker signal generating unit 18 generates a plurality of masker signals by changing the magnitude (receiving level) of the learning signal. Figure 6 This is a diagram showing an example of two masking signals having different reception levels. Figure 6 The solid line in the figure represents the masking signal when the reception level is high, that is, when the echo can return with a large amount of force. Figure 6 The dotted line in the figure represents the masking signal for a low reception level. In this way, the masking signal generator 18 generates multiple masking signals by varying the magnitude of the learning signal multiple times. This allows the frequency characteristics of the residual echo to be accurately stored for each reception level.
[0065] The number of masking signals generated by the masking signal generating unit 18 and stored in the masking signal storing unit 13 is not limited to two, but may be three or more. Hereinafter, the plurality of masking signals stored in the masking signal storing unit 13 are referred to as basic masking signals.
[0066] return Figure 2 The power spectrum of the received signal (hereinafter referred to as the received signal power spectrum) is sequentially input from the double-talk detector 15 to the masking signal selector 14. Upon receiving the received signal power spectrum (with sampling points obtained), the masking signal selector 14 sequentially selects a masking signal (hereinafter referred to as the optimal masking signal) corresponding to the magnitude of the received signal from the basic masking signal based on the received signal obtained within a predetermined period before the sampling points were obtained.
[0067] Here, the predetermined period before the sampling point is obtained is calculated based on the time required from the moment the received speech signal becomes 0 (no sound is output from speaker 52) until the value of the input signal reaches 0. This predetermined time varies depending on the magnitude of the received speech signal, and can be as short as tens to hundreds of milliseconds, or as long as one to two seconds.
[0068] Figure 7 、 Figure 8 5 is a diagram showing the relationship between the received speech signal and the input signal when there is no near-end speech (no speech is input to the microphone 51 ), (A) shows the received speech signal, and (B) shows the input signal. Figure 7 If the level of the received signal is low, Figure 8 Indicates that the level of the received signal is high.
[0069] Because reflections of sound within the vehicle and vibrations of speaker 52 are output as sound from speaker 52, an echo signal exists as an input signal even in the absence of near-end sound. When the level of the received signal is low, the input signal persists for approximately 100 milliseconds even if the received signal is zero. When the level of the received signal is high, the input signal persists for approximately 150 milliseconds even if the received signal is zero. Therefore, in this embodiment, the predetermined time is set to approximately 100 to 300 milliseconds.
[0070] The masker signal selection unit 14 selects the optimal masker signal based on the maximum value of the power of the received signal acquired within approximately 100 milliseconds to approximately 300 milliseconds before the sampling point of the received signal power spectrum.
[0071] Figure 9 、 Figure 10 Schematically shows the relationship between the maximum value of each frequency band of the received signal power spectrum obtained in a given period before the sampling point and the optimal masking signal. Figure 9 、 10In the figure, the solid line represents the maximum value of the received signal spectrum obtained within a given period, and the dotted line represents the basic masking signal. Here, it is assumed that three masking signals are stored as basic masking signals. The masking signal selection unit 14 compares the maximum value of the received signal power with the basic masking signal in each frequency band, and selects the masking signal closest to the received signal as the optimal masking signal in any frequency band so that the value of the masking signal is not less than the maximum value of the received signal. Figure 9 In the case shown, the masking signal with the largest value is selected (refer to Figure 9 Thick dotted line), Figure 10 In the case shown, the mask signal with an intermediate value is selected (refer to Figure 10 Thick dotted line). This makes it possible to select an optimal masking signal taking into account the frequency characteristics of the received signal.
[0072] Alternatively, the masking signal selection unit 14 may select the optimal masking signal based on the sum or average of the power of the received signal obtained within approximately 100 ms to approximately 300 ms prior to the sampling point of the received signal power spectrum, rather than the maximum value of the power of the received signal obtained within approximately 100 ms to approximately 300 ms prior to the sampling point of the received signal power spectrum.
[0073] Figure 11 Schematically shows an example of selecting an optimal masking signal based on the average value of the power of the received signal obtained for each frequency band. Figure 11 In the figure, the thin solid line is the maximum value of the power spectrum of the received signal, and the thick solid line is the maximum value of the power spectrum of the received signal (at Figure 9 The thin line in the figure is added together (summed) by frequency band, and the average value is obtained by dividing it by the frequency band. In other words, the average value has the same meaning as the sum. Figure 11 In the figure, the dotted line is the masking signal.
[0074] The masker signal selection unit 14 compares the average value of the received signal with the masker signal for each frequency band and selects the masker signal closest to the received signal as the optimal masker signal so that the masker signal does not become smaller than the average value of the received signal. Figure 11 Select the mask signal with the smallest value (refer to Figure 11 thick dotted line).
[0075] Furthermore, when selecting the optimal masker signal based on the sum of the power of the received signal calculated for each frequency band, the sum of the power of the received signal calculated for each frequency band is compared with the sum of the power of the basic masker signals, and the masker signal that is closest to the received signal is selected as the optimal masker signal, so that the masker signal does not fall below the sum of the power of the received signals. In this way, by selecting the optimal masker signal based on the sum or average of the power of the received signals, it is possible to reduce the impact of a prominent power in only one frequency band.
[0076] return Figure 2 Double-talk detector 15 calculates the input signal power spectrum and the received signal power spectrum per unit time based on the spectral waveforms input from FFT units 12 and 19, respectively. FFT unit 12 and a portion of double-talk detector 15 correspond to the first power spectrum calculation unit of the present invention, while FFT unit 19 and a portion of double-talk detector 15 correspond to the second power spectrum calculation unit of the present invention.
[0077] Furthermore, each time the masking signal selection unit 14 selects the optimal masking signal, the double-talk detection unit 15 sequentially compares the value of the input signal power spectrum with the value of the optimal masking signal selected by the masking signal selection unit 14 for each frequency band. Based on the comparison results, the double-talk detection unit 15 detects whether a double-talk state exists. The double-talk detection unit 15 performs the process of detecting whether a double-talk state exists during each unit time period during which the input signal power spectrum is calculated.
[0078] The following describes in detail how the double-talk detection unit 15 detects a double-talk state. Here, the double-talk state is a state in which both the near-end speaker (user A) and the far-end speaker (user B) are speaking.
[0079] First, the double-talk detector 15 compares the input signal power spectrum value with the optimal masking signal value for each frequency band and counts the number of frequency bands where the input signal power spectrum value exceeds the optimal masking signal value (hereinafter referred to as the "exceedance count"). The double-talk detector 15 determines whether the "exceedance count" is below a pre-set threshold value I (equivalent to the first threshold value). The threshold value I can be set to any value.
[0080] Figure 12 、 13 are diagrams schematically showing how the value of the input signal power spectrum is compared with the value of the masking signal. Figure 12 、 13 In the figure, the solid line represents the input signal power spectrum, the dotted line represents the received signal, and the dashed line represents the masking signal.
[0081] exist Figure 12In the case shown, the received signal acquired in the most recent given period is large, and the masker signal with the large value is selected as the optimal masker signal. The double-talk detector 15 detects no near-end speech because the number of excesses is 0, which is below the threshold value I (e.g., threshold value I=3).
[0082] exist Figure 13 In the case shown, the received signal obtained in the latest given period is small, and the masking signal with the small value is selected as the optimal masking signal. Figure 13 ) is above the threshold value I, so the double-end talk detection unit 15 detects near-end speech.
[0083] Furthermore, the double-talk detector 15 obtains the power spectrum of the received signal transmitted from the mobile phone 53 to the terminal 50 and determines its signal level. The power spectrum of the received signal is obtained from the received-side signal path via the FFT unit 19. The double-talk detector 15 compares the signal level of the received signal with a pre-set threshold value III. The threshold value III can be set to any value.
[0084] When the signal level of the received signal is equal to or higher than the threshold value III prepared in advance, the double-talk detection unit 15 detects the presence of far-end speech (user B (refer to Figure 1 ) of the voice), the receiving signal includes the voice.
[0085] In this way, the double-end talk detection unit 15 detects the presence or absence of near-end speech and far-end speech based on thresholds I and III to detect whether the call is a double-end talk state with both near-end and far-end speech, a single-end talk with only near-end speech, or a single-end talk with only far-end speech.
[0086] Furthermore, the method by which the double-talk detector 15 detects the presence of near-end speech is not limited to the method based on whether the number of excesses is greater than or equal to threshold I. For example, the double-talk detector 15 may determine whether the sum (integral value) of the portion of the input signal power spectrum where the values exceed the masking signal value is less than or equal to a pre-determined threshold II (equivalent to the second threshold), and detect the presence of near-end speech based on this result. Furthermore, threshold II can be set to any value.
[0087] Figure 14 is a diagram schematically showing a comparison between the value of the input signal power spectrum and the value of the optimal masking signal. Figure 14 In the figure, the solid line represents the input signal power spectrum, the dotted line represents the received signal, and the dashed line represents the optimal masking signal. Figure 14 In FIG, the portion where the value of the input signal power spectrum exceeds the value of the masking signal is shaded with oblique lines. The double-talk detection unit 15 calculates the area of the shaded portion. Figure 14In the example, since the area of the portion where the value of the input signal power spectrum exceeds the value of the masking signal is greater than the threshold value III, it is detected that the signal is being transmitted in the signal path on the transmitting side (there is near-end speech).
[0088] return Figure 2 The echo suppressor 16 performs echo suppression (strongly suppresses echoes) on the input signal passed through the FFT unit 12. The echo suppressor 16 enables echo suppression in single-ended calls where only the far-end voice is present, and disables echo suppression in other cases. Echo suppression is well known, so a detailed description thereof will be omitted.
[0089] In this embodiment, the echo suppressor 16 disables the echo suppression process and switches the echo suppression process on and off except in single-ended calls with only far-end speech. However, the echo suppression process may be switched between strong and weak. For example, the echo may be suppressed strongly in single-ended calls with only far-end speech, and weakly in other cases.
[0090] The result of detecting whether or not a double talk state exists is inputted every unit time from the double talk detector 15 to the echo suppressor 16. Therefore, the echo suppressor 16 switches between validating and invalidating the echo suppression process every unit time.
[0091] The IFFT unit 17 performs inverse FFT (IFFT) on the input signal passed through the FFT unit 12 .
[0092] Figure 15 1 is a flowchart showing the flow of a process for sequentially reducing echoes by the echo suppressor 1. This process is continuously performed at predetermined time intervals while the received signal and the input signal are input to the echo suppressor 1.
[0093] First, the echo remover 11 removes echoes from the input signal (step S11), and the double-talk detector 15 calculates the power spectrum of the echo-removed input signal (step S12). Furthermore, the double-talk detector 15 calculates the power spectrum of the received signal (step S13), and the masker signal selector 14 selects the optimal masker signal from the base masker signals based on the received signal power spectrum (step S14). Alternatively, steps S11 or S12 and S13 may be performed simultaneously.
[0094] Next, the double-talk detector 15 detects whether a double-talk state exists based on the input signal power spectrum calculated in step S12 and the received signal power spectrum calculated in step S13 (step S15). Furthermore, if a single-ended call, where only far-end speech is present, is not a double-talk state, the echo suppressor 16 performs echo suppression on the input signal power spectrum calculated in step S12 (step S16). Finally, the IFFT unit 17 converts the input signal power spectrum back into a time-axis signal (step S17).
[0095] According to this embodiment, the input signal caused by the near-end sound and the residual echo of the far-end sound have different frequency characteristics, and the frequency characteristics of the residual echo are stored as a masking signal. By comparing the frequency characteristics of the input signal with the masking signal, the double-end call state is accurately detected, and the echo suppression processing is effectively performed when the double-end call state is not in place, so that the near-end voice (the voice input from the microphone 51) is not degraded and the echo is reliably suppressed.
[0096] Furthermore, according to the present embodiment, since the magnitude of the masker signal is changed according to the magnitude of the received speech signal, even when the speech is relatively small, it is possible to detect speech and appropriately suppress echo.
[0097] For example, if only a masking signal generated assuming a large received signal is used, if the user on the microphone side (near-end speaker) speaks softly and the received signal is large, the echo suppressor will strongly affect the input signal transmitted through the received signal path, potentially eliminating the near-end speaker's voice. In contrast, in this embodiment, the magnitude of the learning signal is varied to generate multiple masking signals, and the masking signal closest to the received signal is selected as the optimal masking signal. In other words, the optimal masking signal that matches the magnitude of the likely echo is used to accurately detect a double-talk state. This allows detection of speech even in the presence of soft speech, and prevents the echo suppressor from operating more effectively than necessary.
[0098] Furthermore, for example, if the far-end speaker (User B) is at a call center, the voice of a speaker adjacent to User B may be included in the received signal. In such a situation, since the small received signal persists, a masking signal generated assuming a large received signal cannot adequately detect a double-talk state. In contrast, this embodiment accurately detects a double-talk state using an optimal masking signal tailored to the size of the received signal, making it possible to address such situations.
[0099] Furthermore, according to this embodiment, when the power spectrum of the received signal is sequentially inputted, the masking signal selection unit 14 sequentially selects the optimal masking signal from the basic masking signals based on the received signal acquired within a given period before the sampling point is acquired. This ensures stable communication without frequently changing the activation mode of the echo suppressor.
[0100] Because mobile phones 53 and 54 are connected via a standard telephone line, the volume of the sound output from speaker 52 (the volume of the received signal) fluctuates frequently depending on the communication status. If the optimal masking signal is selected solely based on the volume of the received signal at the time of the sampling point, the frequent fluctuations in the volume of the received signal will cause frequent switching of the masking signal, potentially making it difficult for the far-end speaker to hear the near-end speaker. In contrast, by selecting the optimal masking signal based on the received signal acquired within a predetermined period before the sampling point, frequent switching of the masking signal can be prevented, ensuring stable call quality.
[0101] Furthermore, even when there is no incoming signal from the receiving end, sound may be reflected within the vehicle, or sound may be output from speaker 52 due to vibration of speaker 52, for example. In such cases, if the optimal masking signal is selected solely based on the magnitude of the receiving signal at the time the sampling point is obtained, the receiving signal will be zero, and echo suppressor 16 will not function, failing to cancel the echo. In contrast, by selecting the optimal masking signal based on the receiving signal obtained within a predetermined period before the sampling point, the optimal masking signal can be selected taking into account previous conditions, thereby canceling echoes that are output from speaker 52 as sound due to reflections of sound within the vehicle, vibration of speaker 52, and the like.
[0102] Furthermore, in the embodiment of the present invention, when the masking signal selection unit 14 selects the optimal masking signal based on the received call signal acquired within a predetermined period before the sampling point of the received call signal, the predetermined period is predetermined to be approximately 100 to 300 milliseconds. However, the value of the predetermined period and the method for determining the predetermined time are not limited to this. For example, when generating the masking signal, the masking signal generation unit 18 may measure the time from when the received call signal becomes 0 to when the input signal becomes 0, and determine the predetermined time based on this measured time. In this way, the predetermined period can be determined based on the length of the echo generated by the received call signal.
[0103] Furthermore, in the embodiment of the present invention, the masking signal generating unit 18 generates a plurality of masking signals by changing the magnitude of the learning signal. However, the types of masking signals generated by the masking signal generating unit 18 are not limited thereto. For example, the masking signal generating unit 18 may generate a masking signal when only an echo signal caused by reflections of sounds inside the vehicle, vibrations of the speaker 52, etc., which are outputted as sounds from the speaker 52, is input as an input signal. In this case, after the learning of the adaptive filter is fully completed in the echo removing unit 11, the masking signal generating unit 18 transmits a signal (see FIG. 1 ) in a state where only an echo signal caused by reflections of sounds inside the vehicle, vibrations of the speaker 52, etc., which are outputted as sounds from the speaker 52, is generated in the signal path on the transmitting side. Figure 7 、 Figure 8 (B)) is used as a learning signal, and the maximum value among the values of the learning power spectrum is obtained for each frequency band, which is used as a masking signal.
[0104] The masker signal selection unit 14 then sequentially obtains the power spectra of the received and input signals. Once each sampling point is obtained, the optimal masker signal is sequentially selected from the basic masker signals based on the received and input signals obtained within a predetermined period prior to the sampling point. For example, if the received signal is zero and the input signal is low for several milliseconds, the masker signal selection unit 14 selects the optimal masker signal corresponding to a state where only echo signals, such as those caused by reflections of sounds within the vehicle or vibrations of the speaker 52, are generated and outputted as sound from the speaker 52. This effectively eliminates echo signals caused by reflections of sounds within the vehicle or vibrations of the speaker 52, which are outputted as sound from the speaker 52.
[0105] <Second embodiment>
[0106] The second embodiment detects a double-talk state for each frequency band. The following describes an echo suppression device 2 according to the second embodiment. Components identical to those of the echo suppression device 1 according to the first embodiment are denoted by the same reference numerals, and their descriptions are omitted.
[0107] Figure 16 This diagram schematically illustrates the functional blocks of the echo suppression device 2. The echo suppression device 2 mainly includes an echo removal unit 11, FFT units 12 and 19, a masking signal storage unit 13, a masking signal selection unit 14, a double-talk detector 15A, an echo suppressor 16A, an IFFT unit 17, and a masking signal generator 18 (not shown).
[0108] The double-talk detection unit 15A detects whether a double-talk state exists for each frequency band. The double-talk detection unit 15A sequentially performs a process of detecting whether a double-talk state exists for each unit time when the input signal power spectrum is calculated.
[0109] The following describes in detail the method by which double-talk detector 15A detects a double-talk state. First, double-talk detector 15A compares the power spectrum of the input signal input from FFT unit 12 with the value of the optimal masker signal selected by masker signal selector 14 for each frequency band.
[0110] Furthermore, the double-talk detector 15A acquires a call signal transmitted from the mobile phone 53 to the terminal and determines its signal level. The double-talk detector 15A compares the signal level of the call signal with a threshold value II.
[0111] Moreover, regarding the frequency band where the value of the input signal power spectrum does not exceed the value of the optimal masking signal, when the signal level of the received signal is above threshold II, the double-end talk detection unit 15A detects that it is a single-end talk with only far-end speech, not a double-end talk state.
[0112] Figure 17 Schematically shows how the value of the input signal power spectrum is compared with the value of the optimal masking signal when the signal level of the received signal is equal to or greater than the threshold II. Figure 17 In , the solid line represents the input signal power spectrum, and the dotted line represents the optimal masking signal.
[0113] exist Figure 17 In the frequency band enclosed by the solid circle, the value of the input signal power spectrum exceeds the value of the optimal masking signal. Therefore, in this frequency band, the double-talk detector 15A detects both far-end and near-end speech, i.e., a double-talk state.
[0114] In contrast, in Figure 17 In the frequency band enclosed by the dotted circle, the value of the input signal power spectrum does not exceed the value of the optimal masking signal. Therefore, in this frequency band, double-talk detector 15A detects a single-ended call with only far-end speech, i.e., a state not involving a double-talk call.
[0115] Return to Figure 16 Echo suppressor 16A performs echo suppression on the input signal passed through FFT unit 12. Echo suppressor 16A enables echo suppression in the frequency band where single-ended speech, where only far-end speech, is detected, and disables echo suppression in other frequency bands. Echo suppressor 16A switches between enabling and disabling echo suppression per unit time.
[0116] According to this embodiment, a double-talk state can be accurately detected for each frequency band, and echo suppression processing can be effectively performed for each frequency band.
[0117] <Third embodiment>
[0118] The third embodiment is a method in which a masking signal storage unit stores a single base masking signal, and a masking signal selection unit generates an optimal masking signal. The following describes an echo suppression device 3 according to the third embodiment. Components identical to those of the echo suppression devices 1 and 2 according to the first and second embodiments are denoted by the same reference numerals, and their descriptions are omitted.
[0119] Figure 18 This diagram schematically illustrates the functional blocks of the echo suppression device 3. The echo suppression device 3 mainly includes an echo removal unit 11, FFT units 12 and 19, a masking signal storage unit 13A, a masking signal selection unit 14A, a double-talk detector 15, an echo suppressor 16, an IFFT unit 17, and a masking signal generator 18 (not shown).
[0120] The masking signal generating unit 18 generates a masking signal based on the power spectrum of the learning signal calculated by the FFT unit 12, and stores the generated masking signal. The masking signal generating unit 18 generates a masking signal only when the signal of the reception side signal path is assumed to be large (see Figure 5 ), only the mask signal is stored as a basic mask signal in the mask signal storage unit 13A.
[0121] The masker signal selection unit 14A generates an optimal masker signal by multiplying the basic masker signal by a coefficient based on the maximum value of the received signal power obtained within a predetermined period before the sampling point of the received signal power spectrum.
[0122] Figure 19 Schematically shows the process of generating the optimal masking signal by the masking signal selection unit 14A. Figure 19 In FIG, the solid line represents the maximum value of the received signal spectrum obtained within a given period, and the dotted line represents the basic masking signal. The masking signal selection unit 14A compares the maximum value of the received signal power with the basic masking signal for each frequency band, multiplies the basic masking signal by a coefficient, and generates an optimal masking signal so that the value of the optimal masking signal is not less than the maximum value of the received signal in any frequency band, and the optimal masking signal is close to the maximum value of the received signal. Figure 18 In the example shown, the masker signal selection unit 14A generates an optimal masker signal by multiplying the power of each frequency band of the basic masker signal by a coefficient of 0.3. This allows the optimal masker signal to be generated in consideration of the frequency characteristics of the received speech signal.
[0123] According to this embodiment, it is not necessary to store multiple basic masker signals, and the memory used can be reduced. This embodiment is effective when the shapes of the masker signals are similar regardless of the size of the received signal.
[0124] Furthermore, in this embodiment, the masking signal selection unit 14A generates an optimal masking signal by multiplying the power of the basic masking signal in each frequency band by an arbitrary coefficient, regardless of the frequency band. However, the coefficient used to multiply the basic masking signal can also be changed for each frequency band. For example, the coefficient can be reduced as the frequency band increases. In this case, the masking signal storage unit 13A stores the equation representing the relationship between the frequency band size and the coefficient. The masking signal selection unit 14A can then determine the coefficient for each frequency band based on the coefficient at any frequency and the equation representing the relationship between the frequency band size and the coefficient. This allows the generation of an optimal masking signal that further reflects the frequency characteristics of the received signal.
[0125] <Fourth embodiment>
[0126] The fourth embodiment does not use the FFT unit 19. The echo suppressor 4 according to the fourth embodiment will be described below. Components identical to those of the echo suppressors 1 to 3 according to the first to third embodiments are denoted by the same reference numerals, and their description will be omitted.
[0127] Figure 20 This diagram schematically illustrates the functional blocks of the echo suppression device 4. The echo suppression device 4 mainly includes an echo removal unit 11, an FFT unit 12, a masking signal storage unit 13, a masking signal selection unit 14B, a double-talk detector 15, an echo suppressor 16, an IFFT unit 17, and a masking signal generation unit 18 (not shown).
[0128] The received call signal is sequentially input to the masking signal selection unit 14B. When the received call signal is sequentially input (sampling points are obtained), the masking signal selection unit 14B sequentially selects a masking signal (hereinafter referred to as an optimal masking signal) corresponding to the magnitude of the received call signal from the basic masking signal based on the received call signal obtained within a predetermined period before the sampling points are obtained.
[0129] In this embodiment, since FFT unit 19 is not used, the power of the received signal, not separated by frequency band, is input to masker signal selector 14B. Masker signal selector 14A then compares the sum of the power of the received signal input over a certain period of time with the sum of the power of the masker signal for each frequency band. Masker signal selector 14B then selects the optimal masker signal from among the basic masker signals stored in masker signal storage unit 13, the masker signal whose sum of the power of the received signal is smaller than the sum of the power of the masker signals and whose sum of the power of the masker signals is closest to the sum of the power of the received signal.
[0130] Double-talk detector 15B compares the power spectrum of the input signal received from echo remover 11 with the value of the optimal masker signal selected by masker signal selector 14C and counts the number of frequency bands in which the power spectrum of the input signal exceeds the value of the optimal masker signal (exceedance count). Double-talk detector 15B detects the absence of near-end speech if the exceedance count falls below an arbitrary threshold.
[0131] In addition, the double-talk detector 15B compares the magnitude of the received signal with a pre-set threshold. If the magnitude of the received signal is greater than the pre-set threshold, the double-talk detector 15B detects a far-end voice (user B (refer to Figure 1 ) sound), the detection signal is being transmitted in the signal path on the receiving side.
[0132] According to this embodiment, the amount of calculation required for the mask signal selection process can be reduced.
[0133] <Fifth embodiment>
[0134] The fifth embodiment does not use the FFT units 12 and 19. The echo suppressor 5 according to the fifth embodiment will be described below. Components identical to those of the echo suppressors 1 to 4 according to the first to fourth embodiments are denoted by the same reference numerals and their descriptions are omitted.
[0135] Figure 21 This is a diagram schematically showing the functional blocks of the echo suppression device 5 . Figure 22 This diagram schematically shows the functional blocks when generating a masking signal in the echo suppressor 5. The echo suppressor 5 mainly includes an echo removing unit 11, a masking signal storage unit 13B, a masking signal selecting unit 14C, a double-talk detector 15C, an echo suppressor 16B, and a masking signal generating unit 18A.
[0136] First, use Figure 22 The masking signal generation process will be described in detail. First, after the adaptive filter learning is fully completed in the echo removal unit 11, the far-end unilateral speech (single-ended conversation) is repeatedly performed, with no sound input from the microphone 51. The signal from which the echo has been removed by the echo removal unit 11 is used as the learning signal.
[0137] The power of the learning signal (learning power) calculated for each fixed interval is input to the masking signal generator 18A. The masking signal generator 18A stores the multiple learning power values input. The masking signal generator 18A obtains the maximum value of the multiple learning power values input and uses it as the masking signal. Therefore, the generated masking signal has only a single value.
[0138] In this embodiment, the masker signal generator 18A generates multiple masker signals by changing the magnitude of the learning signal (reception level) multiple times. This allows the magnitude of the residual echo to be accurately stored for each reception level.
[0139] Return to Figure 21 The mask signal storage unit 13B stores the plurality of mask signals generated by the mask signal generation unit 18A as basic mask signals.
[0140] The received signal is sequentially input to the masking signal selection unit 14C. Upon receiving the received signal power spectrum (sampling points obtained), the masking signal selection unit 14C sequentially selects a masking signal (hereinafter referred to as an optimal masking signal) corresponding to the magnitude of the received signal from the basic masking signal based on the received signal obtained within a predetermined period before the sampling points were obtained.
[0141] In this embodiment, since the FFT unit 19 is not used, the power of the received signal, not separated by frequency band, is input to the masker signal selector 14C. The masker signal selector 14C compares the sum of the received signal powers input over a certain period of time with the power of the masker signal. The masker signal selector 14C then selects, as the optimal masker signal, the masker signal whose sum of the received signal powers is smaller than the power of the masker signal and whose sum of the masker signal powers is closest to the sum of the received signal powers, from among the basic masker signals stored in the masker signal storage unit 13B.
[0142] For example, if three masking signals are stored in the masking signal storage unit 13B (a first masking signal for a call reception level of 3, a second masking signal for a call reception level of 6, and a third masking signal for a call reception level of 9), and the power of the call reception signal input to the masking signal selection unit 14C is 2, the masking signal selection unit 14C selects the first masking signal as the optimal masking signal. Furthermore, if the power of the call reception signal input to the masking signal selection unit 14C is 4, the masking signal selection unit 14C selects the second masking signal as the optimal masking signal.
[0143] The double-talk detector 15C compares the magnitude of the input signal from the echo remover 11 with the value of the optimal masker signal selected by the masker signal selector 14C, and detects near-end speech when the magnitude of the input signal is greater than the value of the optimal masker signal.
[0144] Furthermore, the double-talk detector 15C compares the magnitude of the received call signal with a preset threshold value. If the magnitude of the received call signal is equal to or greater than the preset threshold value, the double-talk detector 15C detects that a far-end voice is present.
[0145] The echo suppressor 16B enables echo suppression processing for the input signal passing through the echo removing unit 11 when the conversation is single-ended with only far-end speech rather than double-ended conversation, and disables echo suppression processing in other cases.
[0146] Figure 23 1 is a flowchart showing the flow of a process for sequentially reducing echoes by the echo suppressor 5. This process is continuously performed at predetermined time intervals while the received signal and the input signal are input to the echo suppressor 1.
[0147] First, the echo removing unit 11 removes the echo from the input signal (step S11 ), and the masker signal selecting unit 14 selects the optimal masker signal from the basic masker signals based on the power of the received signal (step S18 ).
[0148] Next, the double-talk detector 15 detects whether a double-talk state exists based on the power of the input signal from which the echo has been removed in step S11 and the power of the received signal (step S19). Furthermore, in the case of a single-talk state, in which only the far-end voice is present, the echo suppressor 16 performs echo suppression on the input signal from which the echo has been removed in step S11 (step S20).
[0149] According to this embodiment, since FFT processing and IFFT processing are not performed, the amount of calculation can be reduced.
[0150] While the embodiments of the present invention have been described in detail above with reference to the accompanying drawings, the specific configuration is not limited to these embodiments and encompasses design modifications that do not depart from the spirit of the present invention. In particular, while the embodiments utilize power represented by the square of the amplitude to generate the base masking signal, generate and select the optimal masking signal, and detect a double-talk state, these processes may also be performed based on the absolute value of the amplitude.
[0151] -Description of Reference Numerals-
[0152] 1, 2, 3, 4, 5: Echo suppression devices
[0153] 11: Echo removal unit
[0154] 12: FFT section
[0155] 13, 13A, 13B: Masking signal storage unit
[0156] 14, 14A, 14B, 14C: Mask signal selection unit
[0157] 15, 15A, 15B: Double-ended call detection unit
[0158] 16, 16A, 16B: Echo suppressor
[0159] 17: IFFT Department
[0160] 18, 18A: Masking signal generation unit
[0161] 19: FFT
[0162] 50: Terminal
[0163] 51: Microphone
[0164] 52: Speaker
[0165] 53, 54: Mobile phones
[0166] 55: Speaker amplifier
[0167] 100: Voice communication system.
Claims
1. An echo suppression device, provided in a transmitting-side signal path, the transmitting-side signal path transmitting an input signal input from a microphone, the microphone being provided in a near-end terminal having a speaker and the microphone, characterized in that: have: a masking signal generating unit that generates a masking signal based on a learning signal transmitted in the transmitting-side signal path when a speech sound is output from the speaker but no speech sound is input to the microphone, the masking signal generating unit generating a plurality of masking signals by performing a process of changing the magnitude of the learning signal a plurality of times and generating the masking signal based on the plurality of learning signals at each predetermined interval; a mask signal storage unit configured to store the plurality of mask signals generated by the mask signal generation unit as a plurality of basic mask signals; a masking signal selection unit, each time a sampling point of a received call signal is obtained, sequentially selecting an optimal masking signal corresponding to the magnitude of the received call signal from the basic masking signals based on the received call signal obtained within a predetermined period before the sampling point is obtained, wherein the masking signal selection unit selects a masking signal closest to the received call signal from the plurality of basic masking signals as the optimal masking signal, wherein the received call signal is transmitted in a receiving-side signal path that transmits a signal to the speaker; a double-talk detection unit that sequentially detects whether a double-talk state exists based on a result of comparing the input signal with the optimal masking signal each time the optimal masking signal is generated or selected; and The echo suppressor sequentially performs echo suppression processing on the input signal when the double-talk detection unit detects that no speech sound is input to the microphone and the received signal includes the speech sound.
2. An echo suppression device, provided in a transmitting-side signal path, the transmitting-side signal path transmitting an input signal input from a microphone, the microphone being provided in a near-end terminal having a speaker and the microphone, characterized in that: have: a masking signal generating unit for generating a masking signal based on a learning signal transmitted in the transmitting-side signal path when a speech sound is not input to the microphone but a sound is output from the speaker; a masking signal storage unit configured to store a masking signal generated by the masking signal generating unit as a basic masking signal; a masking signal selecting unit, each time a sampling point of a received speech signal is obtained, sequentially generating an optimal masking signal corresponding to the magnitude of the received speech signal from the basic masking signal based on the received speech signal obtained within a predetermined period before the sampling point, wherein the masking signal selecting unit generates the optimal masking signal by multiplying the basic masking signal by a coefficient based on the magnitude of the input signal, wherein the received speech signal is transmitted in a receiving-side signal path that transmits a signal to the speaker; a double-talk detection unit that sequentially detects whether a double-talk state exists based on a result of comparing the input signal with the optimal masking signal each time the optimal masking signal is generated or selected; and The echo suppressor sequentially performs echo suppression processing on the input signal when the double-talk detection unit detects that no speech sound is input to the microphone and the received signal includes the speech sound.
3. The echo suppression device according to claim 1 or 2, wherein: The echo suppression device includes: a signal measuring unit for measuring a first time, a time when a signal is no longer transmitted in the transmitting-side signal path, when a state in which no speech sound is input to the microphone and the speaker outputs a sound changes to a state in which no speech sound is input to the microphone and the speaker outputs a sound; The mask signal selection unit sequentially generates or selects the optimal mask signal using the first time as the predetermined period.
4. The echo suppression device according to claim 1 or 2, wherein: The echo suppression device includes a first power spectrum calculation unit that calculates an input signal power spectrum that is a power spectrum of the input signal and a learning power spectrum that is a power spectrum of the learning signal. The masking signal is the maximum value of each frequency band of the learning power spectrum obtained within a certain interval. The optimal masking signal has a value corresponding to each frequency band, The double-talk detection unit detects whether a double-talk state exists based on a result obtained by comparing a value of the input signal power spectrum with a value of the optimal masker signal for each frequency band.
5. The echo suppression device according to claim 4, wherein: The echo suppression device includes: a second power spectrum calculation unit for calculating a power spectrum of the received speech signal, that is, a received speech signal power spectrum; The masker signal selection unit compares the maximum value of the received signal power spectrum with the optimal masker signal for each frequency band, and generates or selects the optimal masker signal.
6. The echo suppression device according to claim 4, wherein: The double-end talk detection unit compares the value of the input signal power spectrum with the value of the optimal masking signal in each frequency band, and detects that no speech sound is input to the microphone when the number of frequency bands in which the value of the input signal power spectrum exceeds the value of the optimal masking signal is less than a first threshold, or when the integral value of the area in which the value of the input signal power spectrum exceeds the value of the optimal masking signal is less than a second threshold.
7. An echo suppression method, characterized in that: The steps include: A plurality of basic masking signals are generated and stored based on a learning signal, the size of the learning signal is changed multiple times, and a masking signal is generated based on the plurality of learning signals at each predetermined interval to generate the plurality of masking signals, and the plurality of masking signals generated are stored as the plurality of basic masking signals, wherein the learning signal is transmitted in a transmitting-side signal path that transmits a signal input from the microphone when no speech is input to the microphone of the near-end terminal but when a sound is output from a speaker of the near-end terminal; Whenever a sampling point of a received call signal is obtained, an optimal masking signal, which is a masking signal having a magnitude corresponding to the magnitude of the input signal input from the microphone, is sequentially selected based on the received call signal obtained within a given period before the sampling point and a plurality of basic masking signals, wherein a masking signal closest to the received call signal is selected as the optimal masking signal from among the basic masking signals, wherein the received call signal is transmitted in a received-side signal path for transmitting a signal to the speaker; If the optimal masking signal is selected, detecting whether a double-talk state is present based on a result of comparing the input signal with the optimal masking signal; and When it is detected that no speech sound is input to the microphone and the received speech signal includes the speech sound, an echo suppression process for suppressing an echo is performed on the input signal.
8. An echo suppression method, characterized in that: The steps include: generating and storing a basic masking signal as a masking signal based on a learning signal, the learning signal being transmitted in a transmitting-side signal path that transmits a signal input from the microphone when a speech sound is output from a speaker of the near-end terminal without a speech sound being input to the microphone of the near-end terminal; Whenever a sampling point of a received speech signal is obtained, sequentially generating an optimal masking signal as a masking signal having a magnitude corresponding to the magnitude of an input signal input from the microphone based on the received speech signal and the basic masking signal obtained within a predetermined period before the sampling point is obtained, wherein the optimal masking signal is generated by multiplying the basic masking signal by a coefficient based on the magnitude of the input signal, wherein the received speech signal is transmitted through a received speech-side signal path that transmits a signal to the speaker; If the optimal masking signal is generated, detecting whether a double-talk state exists in sequence based on a result of comparing the input signal with the optimal masking signal; and When it is detected that no speech sound is input to the microphone and the received speech signal includes the speech sound, an echo suppression process for suppressing an echo is performed on the input signal.
9. An echo suppression program product, comprising an echo suppression program, wherein the echo suppression program is applied to a transmitting-side signal path that transmits a signal input from a microphone, the microphone being provided in a near-end terminal having a speaker and the microphone, wherein: The echo suppression program causes the computer to function as the following elements: a masking signal generating unit that generates a masking signal based on a learning signal transmitted in the transmitting-side signal path when a speech sound is output from the speaker but no speech sound is input to the microphone, the masking signal generating unit generating a plurality of masking signals by performing a process of changing the magnitude of the learning signal a plurality of times and generating the masking signal based on the plurality of learning signals at each predetermined interval; a mask signal storage unit configured to store the plurality of mask signals generated by the mask signal generation unit as basic mask signals; a masking signal selection unit, each time a sampling point of a received call signal is obtained, sequentially selecting an optimal masking signal corresponding to the magnitude of the received call signal from the basic masking signals based on the received call signal obtained within a predetermined period before the sampling point is obtained, wherein the masking signal selection unit selects the masking signal closest to the received call signal from the basic masking signals as the optimal masking signal, wherein the received call signal is transmitted in a receiving-side signal path that transmits a signal to the speaker; a double-talk detection unit that sequentially detects whether a double-talk state exists based on a result of comparing an input signal input from the microphone with the optimal masking signal each time the optimal masking signal is selected; and The echo suppressor sequentially performs echo suppression processing on the input signal when the double-talk detection unit detects that no speech sound is input to the microphone and the received signal includes the speech sound.
10. An echo suppression program product, comprising an echo suppression program, wherein the echo suppression program is applied to a transmitting-side signal path that transmits a signal input from a microphone, the microphone being provided in a near-end terminal having a speaker and the microphone, wherein: The echo suppression program causes the computer to function as the following elements: a masking signal generating unit for generating a masking signal based on a learning signal transmitted in the transmitting-side signal path when a speech sound is not input to the microphone but a sound is output from the speaker; a masking signal storage unit configured to store a masking signal generated by the masking signal generating unit as a basic masking signal; a masking signal selecting unit, each time a sampling point of a received speech signal is obtained, sequentially generating an optimal masking signal corresponding to the magnitude of the received speech signal from the basic masking signal based on the received speech signal obtained within a predetermined period before the sampling point, wherein the masking signal selecting unit generates the optimal masking signal by multiplying the basic masking signal by a coefficient based on the magnitude of the input signal input from the microphone, wherein the received speech signal is transmitted in a receiving-side signal path that transmits a signal to the speaker; a double-talk detection unit that sequentially detects whether a double-talk state exists based on a result of comparing the input signal with the optimal masking signal each time the optimal masking signal is generated; and The echo suppressor sequentially performs echo suppression processing on the input signal when the double-talk detection unit detects that no speech sound is input to the microphone and the received signal includes the speech sound.
Citation Information
Patent Citations
Echo suppression apparatus, echo suppression method, and echo suppression program
JP2018201147A