Call processing device and call processing method

The call processing device addresses the complexity and cost issue of existing systems by implementing echo removal, noise suppression, and dummy noise generation to eliminate sound interruptions with a simpler and cost-effective setup.

JP7843159B2Active Publication Date: 2026-04-09DENSO TEN LTD
View PDF 9 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Existing call processing systems require complex control circuits for precise noise addition to eliminate the feeling of sound interruption, leading to increased product costs.

Method used

A call processing device that performs echo removal, noise suppression, and dummy noise generation to continuously add a dummy noise signal, eliminating the need for high-precision switching control.

Benefits of technology

The solution allows for eliminating the feeling of sound interruption with a simpler configuration, using a DSP with low responsiveness without causing unnatural sounds, and reducing noise fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007843159000001
    Figure 0007843159000001
  • Figure 0007843159000002
    Figure 0007843159000002
  • Figure 0007843159000003
    Figure 0007843159000003
Patent Text Reader

Abstract

To provide a call processor capable of removing sound interruption with a simple construction and also provide a call processing method.SOLUTION: In a communication system, a call processor 10 of a communication device 1, comprises a control part 10a. The control part includes: an echo elimination processing 12 that eliminates an echo component as a sound of a remote end talker contained in a sound signal collected by a microphone mounted onto a vehicle; an echo suppression processing 14 that suppresses a residual component of the echo component from the sound signal after the echo elimination processing; a noise suppression processing 13 that suppresses the noise component contained in the sound signal; a dummy noise generation processing 15 that generates the dummy noise signal on the basis of the noise component contained in the sound signal; and an accumulator 18 that steadily adds the dummy noise signal that is amplified by an amplifier 17 to the sound signal after the processing by the echo elimination processing 12, the echo suppression processing 14, and the noise suppression processing 13.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a call processing apparatus and a call processing method.

Background Art

[0002] Conventionally, in a technique for performing a call such as a hands-free call, for example, a technique for suppressing an echo component generated by re-inputting the voice on the far-end side output from a speaker into a microphone is known. Further, in this type of technique, a technique for intentionally adding noise has been proposed in order to eliminate the feeling of interrupted sound on the far-end side due to the operation of the echo suppression processing (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the prior art, in order to add noise at the timing when the sound cut occurs, high-precision switching control is required. For this reason, a complicated control circuit is required, and there is a possibility that the product cost increases.

[0005] The present invention has been made in view of the above, and an object thereof is to provide a call processing apparatus and a call processing method that can eliminate the feeling of interrupted sound with a simple configuration.

Means for Solving the Problems

[0006] To solve the above-mentioned problems and achieve the objective, the call processing device according to the present invention includes a control unit. The control unit performs: an echo removal process to remove the echo component, which is the voice of a far-end speaker, contained in the voice signal collected by a microphone mounted on a vehicle; an echo suppression process to suppress the residual component of the echo component from the voice signal after the echo removal process; a noise suppression process to suppress the noise component contained in the voice signal; a dummy noise generation process to generate a dummy noise signal based on the noise component contained in the voice signal; and a dummy noise signal to be continuously added to the voice signal after processing by the echo removal process, the echo suppression process and the noise suppression process. [Effects of the Invention]

[0007] According to the present invention, the feeling of sound interruption can be eliminated with a simple configuration. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 is a diagram showing an overview of the call processing method according to the embodiment. [Figure 2] Figure 2 is a block diagram showing an example of the functional configuration of a call processing device according to an embodiment. [Figure 3] Figure 3 is an explanatory diagram illustrating the audio signal output from the call processing unit to the far end. [Figure 4] Figure 4 is a flowchart showing the overall processing procedure executed by the call processing device according to the embodiment. [Figure 5] Figure 5 is a flowchart showing the processing procedure for dummy noise generation performed by the call processing device according to the embodiment. [Modes for carrying out the invention]

[0009] Hereinafter, embodiments of the call processing device and call processing method disclosed in this application will be described in detail with reference to the attached drawings. However, the present invention is not limited to the embodiments described below.

[0010] First, an overview of the call processing method according to the embodiment will be described using Figure 1. Figure 1 is a diagram showing an overview of the call processing method according to the embodiment. Figure 1 shows an example of the configuration of the call system S according to the embodiment. As shown in Figure 1, the call system S according to the embodiment includes a call device 1 and a far-end device 100. The call device 1 and the far-end device 100 are connected by a call-enabled network such as a telephone line. Note that the call device 1 and the far-end device 100 may be connected by a communication line using, for example, the Internet Protocol.

[0011] The communication device 1 is, for example, a communication device with a hands-free function, and is installed inside the vehicle. As shown in Figure 1, the communication device 1 comprises a communication processing unit 10, a microphone 20, an RF circuit 30, and a speaker 40.

[0012] The microphone 20 is mounted in the vehicle and outputs an audio signal collected from the surrounding sounds to the communication processing unit 10. The audio signal may include an audio component, which is the occupant's speech; a noise component, such as road noise; and an echo component, which is the sound from the far end output from the speaker 40.

[0013] The RF circuit 30 modulates the audio signal, which has been processed by the call processing device 10 as described later, with the carrier wave fundamental signal, and transmits it to the far-end device 100. The RF circuit 30 also demodulates the modulated signal transmitted from the far-end device 100 to extract the audio signal, including the speech from the far-end side, and outputs it to the call processing device 10.

[0014] The speaker 40 is installed inside the vehicle and acquires voice signals, including speech from the far end, from the call processing unit 10 and outputs them into the vehicle.

[0015] The call processing device 10 executes the call processing method according to the embodiment. Specifically, by executing the call processing method, the call processing device 10 performs echo removal processing, noise removal processing, echo suppression processing, and dummy noise generation processing.

[0016] The echo cancellation process is a process of removing the echo component, which is the voice of the remote speaker, included in the voice signal acquired from the microphone 20. Specifically, in the echo cancellation process, the echo component is removed from the voice signal using a pseudo echo signal generated with the voice signal on the remote side before being output from the speaker 40 as a reference signal.

[0017] The noise reduction process is a process of suppressing the noise component included in the voice signal after the echo cancellation process. Specifically, in the noise reduction process, the signal component of the frequency estimated to be noise is suppressed. Thereby, the S / N ratio with the speech being S and the noise being N is improved.

[0018] The echo suppression process suppresses the residual component of the echo component from the voice signal after the echo cancellation process. Specifically, in the echo suppression process, the signal component of the frequency estimated to be the echo component is suppressed. In other words, in the echo suppression process, the speech component, noise component, and echo component (residual component) of the occupant are collectively suppressed.

[0019] Here, in the echo suppression process, since the noise component is also suppressed, when the remote side speaks in a state where the noise component on the near end side is large, the noise fluctuates temporarily (the noise decreases) at the moment when the remote side speaks. That is, there was a risk that the remote speaker would feel a cut-off in the voice because the noise included in the voice heard by the remote speaker was temporarily interrupted immediately after speaking.

[0020] Therefore, in the present disclosure, the dummy noise signal generated by the dummy noise generation process is constantly added to the voice signal on the near end side and transmitted to the remote side. Specifically, the dummy noise generation process generates a dummy noise signal based on the noise component included in the voice signal collected by the microphone 20.

[0021] Then, the call processing device 10 adds the dummy noise signal generated by the dummy noise generation process to the voice signal after the echo suppression process and outputs it to the RF circuit 30.

[0022] In other words, since the call processing method simply involves continuously adding a dummy noise signal, it is possible to eliminate the feeling of sound dropout with a simpler configuration compared to a configuration that performs high-precision switching control.

[0023] Furthermore, because switching control itself is not performed, even if a DSP (Digital Signal Processor) with low responsiveness is used, for example, it is possible to avoid causing unnatural sounds at the far end due to the addition of dummy noise.

[0024] Furthermore, in the call processing method according to this embodiment, the amount of dummy noise signal added can be changed according to the amount of echo suppression processing, but this point will be described later.

[0025] Next, an example of the configuration of the call processing device 10 according to the embodiment will be described using Figure 2. Figure 2 is a block diagram showing an example of the functional configuration of the call processing device 10 according to the embodiment.

[0026] As shown in Figure 2, the call processing device 10 includes a control unit 10a. The control unit 10a includes an ADC 11, an echo rejection unit 12, a noise suppression unit 13, an echo suppression unit 14, a dummy noise generation unit 15, a noise suppression unit 16, an amplifier 17, an adder 18, and a DAC 19.

[0027] Here, the call processing unit 10 includes, for example, a computer having a CPU (Central Processing Unit), ROM (Read Only Memory), RAM (Random Access Memory), flash memory, input / output ports, and various circuits.

[0028] The computer's CPU functions, for example, by reading and executing a program stored in ROM, as the ADC 11, echo rejection unit 12, noise suppression unit 13, echo suppression unit 14, dummy noise generation unit 15, noise suppression unit 16, amplifier 17, adder 18, and DAC 19 of the control unit 10a.

[0029] Furthermore, at least one or all of the components of the control unit 10a, including the ADC 11, echo rejection unit 12, noise suppression unit 13, echo suppression unit 14, dummy noise generation unit 15, noise suppression unit 16, amplifier 17, adder 18, and DAC 19, can be configured using hardware such as ASICs (Application Specific Integrated Circuits), FPGAs (Field Programmable Gate Arrays), digital circuits, or analog circuits.

[0030] Furthermore, RAM and flash memory correspond to storage units not shown. RAM and flash memory can store information about various programs, etc. The call processing unit 10 may also acquire the above-mentioned programs and various information via other computers or portable recording media connected by wired or wireless networks.

[0031] The ADC11 converts the analog audio signal collected by the microphone 20 into a digital signal and outputs it to the echo removal unit 12.

[0032] The echo removal unit 12 comprises a delay unit 121, an A-FIR 122, and a calculator 123, and performs echo removal processing. The delay unit 121 delays the audio signal from the far end acquired from the RF circuit 30 and outputs it to the A-FIR 122.

[0033] Specifically, the delay unit 121 determines the amount of delay based on the positional relationship between the microphone 20 and the speaker 40 and the acoustic transfer function in the vehicle cabin. More specifically, the delay unit 121 delays the far-end audio signal that is the source of the pseudo-echo signal so that the timing of when the audio signal is input from the ADC 11 to the arithmetic unit 123 coincides with the timing of when the pseudo-echo signal is input from the A-FIR 122 to the arithmetic unit 123.

[0034] A-FIR122 is an adaptive filter that estimates the impulse response of the echo component that leaks from speaker 40 to microphone 20, and generates a pseudo-echo signal based on the estimation result. A-FIR122 outputs the generated pseudo-echo signal to arithmetic unit 123.

[0035] The arithmetic unit 123 removes the echo component from the audio signal by subtracting the pseudo-echo signal input from the A-FIR 122 from the audio signal input from the ADC 11. The arithmetic unit 123 outputs the audio signal after the echo component has been removed to the noise suppression unit 13.

[0036] The noise suppression unit 13 performs noise suppression processing to suppress noise components contained in the audio signal input from the echo removal unit 12. Specifically, the noise suppression unit 13 estimates the frequency characteristics of the noise components based on the audio signal and suppresses the noise components with the estimated frequency characteristics.

[0037] The echo suppression unit 14 performs echo suppression processing to suppress residual echo components from the audio signal after echo removal processing by the echo removal unit 12. Specifically, the echo suppression unit 14 estimates the frequency band of the echo component (residual component) from the audio signal and suppresses the echo component in the estimated frequency band. In addition to the echo component, the echo suppression processing also suppresses noise components (residual components from the noise suppression unit 13).

[0038] The dummy noise generation unit 15 generates a dummy noise signal, which is a pseudo-signal of the noise component, based on the noise component contained in the audio signal output from the ADC 11. For example, the dummy noise generation unit 15 extracts the noise component from the audio signal by suppressing components other than the noise component (echo component and speech component), and generates a dummy noise signal based on the signal of the extracted noise component.

[0039] The noise suppression unit 16 performs noise suppression processing on the dummy noise signal generated by the dummy noise generation unit 15. The noise suppression unit 16 performs noise suppression processing on the dummy noise signal, for example, when the echo suppression unit 14 does not perform echo suppression processing. In other words, the noise suppression unit 16 may change the amount of dummy noise signal added to the audio signal in accordance with the echo suppression processing.

[0040] In other words, if echo suppression processing is not performed, the control unit 10a may perform processing to reduce the amount of dummy noise signal added. This prevents the noise level that is unnecessarily high for the speaker at the far end when echo suppression processing is not performed, i.e., when the speaker at the near end is speaking, or when neither the near nor far end speakers are speaking. Furthermore, the processing to change the amount of dummy noise signal added does not need to be performed instantaneously in response to the operation of echo suppression processing, but can be done by gradually changing the amount added while continuously adding dummy noise.

[0041] The amplifier 17 amplifies the dummy noise signal input from the noise suppression unit 16 and outputs it to the adder 18. For example, the amplifier 17 changes the amount of dummy noise signal added in accordance with the echo suppression processing by the echo suppression unit 14. Alternatively, an attenuator may be used instead of an amplifier to attenuate the dummy noise signal before outputting it to the adder 18.

[0042] Specifically, when the echo suppression unit 14 performs echo suppression processing, the amplifier 17 (or attenuator) generates a dummy noise signal that is amplified (or attenuated) to an amount proportional to the amount of suppression performed by the echo suppression processing.

[0043] This makes it possible to minimize noise fluctuations (noise reduction) caused by echo suppression processing by adding an optimal amount of dummy noise signal. In other words, it eliminates the feeling of sound dropout at the far end.

[0044] The adder 18 adds the dummy noise signal amplified by the amplifier 17 to the audio signal output from the echo suppression unit 14. In other words, the adder 18 continuously adds the dummy noise signal to the audio signal after processing by echo removal, echo suppression, and noise suppression. The adder 18 outputs the audio signal with the added dummy noise signal to the RF circuit 30.

[0045] The DAC19 acquires the audio signal from the far end collected by the microphone of the far end device 100 via the RF circuit 30, converts it from a digital signal to an analog signal, and outputs it to the speaker 40.

[0046] Next, we will explain the audio signal output from the call processing device 10 to the far end using Figure 3. Figure 3 is an explanatory diagram illustrating the audio signal output from the call processing device 10 to the far end.

[0047] In Figure 3, it is assumed that the speaker at the near end is not speaking, and the audio signal output to the far end is noise only. Then, during the period from time t1 to time t2, the speaker at the far end speaks, and consequently, echo suppression processing is performed during the same period from time t1 to time t2.

[0048] Furthermore, in Figure 3, the dashed line represents the audio signal without noise suppression (NC) or dummy noise signal (DN) addition. Echo suppression significantly reduces the level from time t1 to time t2 (variation amount H1).

[0049] Therefore, when a distant speaker begins to speak, they may perceive the noise coming from the other party as suddenly becoming silent (a feeling of disconnection), which can cause anxiety that the call has been dropped, potentially hindering them from continuing to speak smoothly.

[0050] Next, in Figure 3, the solid line represents the audio signal when noise suppression (NC) and dummy noise signal (DN) are added. During periods when the speaker at the far end is not speaking (before time t1 and after time t2), the noise level is reduced by noise suppression (NC) (arrow NC), and during the period from time t1 to time t2, the noise level is increased by the addition of the dummy noise signal (arrow DN). Therefore, the decrease in level from time t1 to time t2 due to echo suppression is mitigated (variation amount H2), and because a dummy noise signal is added, there is no complete silence.

[0051] In this way, by combining noise suppression (NC) processing with the addition of a dummy noise signal, it is possible to reduce the sense of interruption experienced by the speaker at the far end and promote smoother speech.

[0052] Furthermore, since the dummy noise signal can be continuously added to the audio signal, there is no need to perform high-precision switching control in conjunction with the echo suppression process. In other words, the feeling of sound dropout can be eliminated with a simple configuration.

[0053] Next, the processing procedure of the call processing device 10 according to the embodiment will be described using Figures 4 and 5. Figure 4 is a flowchart showing the overall processing procedure of the call processing device 10 according to the embodiment.

[0054] As shown in Figure 4, the call processing device 10 first acquires the voice signal collected by the microphone 20 (step S101).

[0055] Next, the call processing device 10 performs echo removal processing to remove echo components contained in the acquired voice signal (step S102).

[0056] Next, the call processing device 10 performs noise suppression processing to suppress noise components contained in the voice signal (step S103).

[0057] Next, the call processing device 10 performs an echo suppression process to suppress any residual echo components in the audio signal after the echo removal process (step S104).

[0058] Next, the call processing device 10 performs a dummy noise generation process to generate a dummy noise signal based on the noise components contained in the voice signal (step S105).

[0059] Next, the call processing device 10 performs an additional processing to add the generated dummy noise signal to the audio signal (step S106).

[0060] Next, the call processing device 10 transfers the processed audio signal to the far-end device 100 (step S107), and terminates the process.

[0061] Next, Figure 5 is a flowchart showing the processing procedure for the dummy noise generation process performed by the call processing device 10 according to the embodiment.

[0062] As shown in Figure 5, the call processing device 10 according to the embodiment generates a dummy noise signal based on the noise component contained in the audio signal collected by the microphone 20 (step S201).

[0063] Next, the call processing device 10 performs noise suppression processing on the generated dummy noise signal (step S202).

[0064] Next, the call processing device 10 determines whether or not echo suppression processing has been performed on the voice signal (step S203). If echo suppression processing has been performed (step S203: Yes), the call processing device 10 increases the amount of dummy noise signal added in proportion to the amount of suppression in the echo suppression processing (step S204), and then terminates the processing.

[0065] On the other hand, if echo suppression processing is not performed (step S203: No), the call processing device 10 reduces the amount of dummy noise signal added (step S205) and terminates the process.

[0066] As described above, the call processing device 10 according to the embodiment includes a control unit 10a. The control unit 10a performs echo removal processing to remove the echo component, which is the voice of the far-end speaker, contained in the voice signal collected by the microphone 20 mounted on the vehicle; echo suppression processing to suppress residual echo components from the voice signal after echo removal processing; noise suppression processing to suppress noise components contained in the voice signal; dummy noise generation processing to generate a dummy noise signal based on the noise components contained in the voice signal; and continuously adds the dummy noise signal to the voice signal after processing by the echo removal processing, echo suppression processing and noise suppression processing. This makes it possible to eliminate the feeling of sound dropout with a simple configuration.

[0067] Further effects and modifications can be readily derived by those skilled in the art. Therefore, broader aspects of the present invention are not limited to the specific details and representative embodiments expressed and described above. Accordingly, various modifications are possible without departing from the spirit or scope of the overall concept of the invention as defined by the appended claims and their equivalents. [Explanation of Symbols]

[0068] 1 Telephone device 10. Call Processing Device 10a Control Unit 12 Echo removal section 13. Noise suppression section 14 Echo suppression section 15 Dummy noise generation unit 16. Noise suppression section 17 Amplifier 18 Adder 20 microphones 30 RF circuit 40 speakers 100 Far end device 121 Delay section 123 Arithmetic unit S Calling System

Claims

1. An echo removal process that removes echo components, which are the voice of a far-end speaker, from an audio signal collected by a microphone mounted on a vehicle; an echo suppression process that suppresses residual echo components from the audio signal after the echo removal process; a noise suppression process that suppresses noise components contained in the audio signal; a dummy noise generation process that generates a dummy noise signal based on the noise components contained in the audio signal; and a control unit that continuously adds the dummy noise signal to the audio signal after processing by the echo removal process, the echo suppression process, and the noise suppression process. Equipped with, The control unit, The amount of dummy noise signal added is changed depending on whether or not the echo suppression process is performed. Call processing device.

2. The control unit, If the echo suppression process is not performed, the amount of addition is reduced. The call processing device according to claim 1.

3. The control unit, When performing the echo suppression process, the dummy noise signal is added in an amount proportional to the amount of suppression in the echo suppression process. The call processing device according to claim 1 or 2.

4. A method for handling phone calls performed by a computer, An echo removal process to remove echo components, which are the voice of a far-end speaker, from an audio signal collected by a microphone mounted on a vehicle; an echo suppression process to suppress residual echo components from the audio signal after the echo removal process; a noise suppression process to suppress noise components contained in the audio signal; a dummy noise generation process to generate a dummy noise signal based on the noise components contained in the audio signal; and a control process to continuously add the dummy noise signal to the audio signal after processing by the echo removal process, the echo suppression process, and the noise suppression process. Includes, The control process described above is: The amount of dummy noise signal added is changed depending on whether or not the echo suppression process is performed. Call handling method.

Citation Information

Patent Citations

  • Center clipper

    JP1989133433A

  • Acoustic echo control system and simultaneous speech detector of the same system and simultaneous speech control method for the same system

    JP1999074822A

  • Echo canceller

    JP2000138619A

  • Hands-free telephone device, noise canceling function discriminator, and method for canceling noise

    JP2008306561A

  • Method and test signal for measuring speech intelligibility

    JP2009509396A