Voice processing device, voice processing method, program, and storage medium
The audio processing device addresses the issue of erroneous recognition in voice recognition systems by estimating and suppressing system audio echoes in microphone inputs, improving recognition accuracy and reliability.
Patent Information
- Application Number
- JP2023191953
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-22
AI Technical Summary
Echo cancellation techniques in speech recognition systems, such as smart speakers, often fail to adequately remove system audio as an echo component, leading to erroneous recognition results, especially when threshold values fluctuate.
An audio processing device and method that estimate and suppress system audio as an echo component in microphone inputs, applying a predetermined delay and adjusting the audio output based on the correlation between the microphone audio and the echo component to minimize erroneous recognition.
The proposed solution effectively suppresses erroneous recognition in voice recognition systems by accurately removing system audio echoes, even under conditions of fluctuating threshold values, thereby improving the reliability of speech recognition.
Smart Images

Figure 2025079376000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to techniques for processing audio. [Background technology]
[0002] 2. Description of the Related Art Echo cancellation techniques for removing echo components contained in speech are known in the art.
[0003] Specifically, for example, Patent Document 1 discloses a technique for removing a component corresponding to an echo from a sound including a user's speech input to a microphone and the echo input to the microphone. Patent Document 1 also discloses a technique for detecting a sound portion having a sound pressure exceeding a threshold as a user's speech section. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] JP 2009-109536 A Summary of the Invention [Problem to be solved by the invention]
[0005] For example, when echo cancellation is applied to speech input into an interactive device that uses speech recognition, such as a smart speaker, the speech emitted from the device itself may not be sufficiently removed as an echo component, resulting in an erroneous recognition result.
[0006] According to the technology disclosed in Patent Document 1, for example, in a situation where the threshold value frequently changes, the above-mentioned problem may occur. Therefore, according to the technology disclosed in Patent Document 1, a problem corresponding to the above-mentioned problem occurs.
[0007] In view of the above-mentioned problems, a main object of the present disclosure is to provide a voice processing device capable of suppressing erroneous recognition in voice recognition. [Means for solving the problem]
[0008] The invention described in the claims is an audio processing device comprising: an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component; a suppression unit that outputs a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio; a buffer unit that outputs the second microphone audio with a predetermined delay; and a control unit that decides whether to attenuate the second microphone audio delayed by the predetermined time depending on the correlation between the first microphone audio and the echo component.
[0009] The invention described in the claims is an audio processing method executed by a computer, comprising: an estimation step of estimating, as an echo component, a component corresponding to a system audio contained in a first microphone audio input from a microphone; a suppression step of outputting a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio; a buffering step of outputting the second microphone audio with a predetermined delay; and a control step of determining whether to attenuate the second microphone audio delayed by the predetermined time depending on the correlation between the first microphone audio and the echo component.
[0010] The invention described in the claims is a program executed by a computer, which causes the computer to function as an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component, a suppression unit that subtracts the echo component from the first microphone audio to output a second microphone audio in which the system audio contained in the first microphone audio is suppressed, a buffer unit that outputs the second microphone audio with a predetermined delay, and a control unit that determines whether to attenuate the second microphone audio delayed by the predetermined time depending on the correlation between the first microphone audio and the echo component. [Brief description of the drawings]
[0011] [Figure 1] FIG. 1 is a diagram showing an example of a configuration of a voice processing system according to an embodiment. [Diagram 2] FIG. 1 is a block diagram showing an example of a hardware configuration of a voice processing device according to an embodiment. [Diagram 3] FIG. 1 is a block diagram showing an example of a functional configuration of a voice processing device according to an embodiment. [Figure 4] 6 is a diagram showing an example of a change over time in a correlation coefficient used in the voice processing according to the embodiment. [Diagram 5] 4 is a flowchart showing an example of processing performed by a voice processing device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] In one preferred embodiment of the present invention, an audio processing device includes an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component, a suppression unit that outputs a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio, a buffer unit that outputs the second microphone audio with a predetermined delay, and a control unit that decides whether to attenuate the second microphone audio delayed by the predetermined time depending on the correlation between the first microphone audio and the echo component.
[0013] The above voice processing device includes an estimation unit, a suppression unit, a buffer unit, and a control unit. The estimation unit estimates a component corresponding to a system voice included in a first microphone voice input from a microphone as an echo component. The suppression unit outputs a second microphone voice in which the system voice included in the first microphone voice is suppressed by subtracting the echo component from the first microphone voice. The buffer unit outputs the second microphone voice with a predetermined delay. The control unit determines whether or not to attenuate the second microphone voice delayed by the predetermined time according to the correlation between the first microphone voice and the echo component. This makes it possible to suppress erroneous recognition in voice recognition.
[0014] In one aspect of the above-mentioned audio processing device, the control unit attenuates the second microphone audio delayed by the predetermined time when the correlation between the first microphone audio and the echo component is high, while not attenuating the second microphone audio delayed by the predetermined time when the correlation between the first microphone audio and the echo component is low.
[0015] In one aspect of the above-mentioned audio processing device, the control unit further includes a calculation unit that calculates a correlation coefficient indicating the correlation between the first microphone audio and the echo component, and when the correlation coefficient is greater than a predetermined threshold, the control unit sets the gain value applied to the second microphone audio delayed by the predetermined time to a value less than 1, and when the correlation coefficient is equal to or less than the predetermined threshold, sets the gain value applied to the second microphone audio delayed by the predetermined time to 1.
[0016] In one aspect of the above audio processing device, the control unit sets the gain value applied to the second microphone audio delayed by the predetermined time to a value greater than or equal to 0.1 and less than or equal to 0.5 when the correlation coefficient is greater than the predetermined threshold.
[0017] In another preferred embodiment of the present invention, a computer-executed voice processing method includes an estimation step of estimating a component corresponding to a system voice included in a first microphone voice input from a microphone as an echo component, a suppression step of outputting a second microphone voice in which the system voice included in the first microphone voice is suppressed by subtracting the echo component from the first microphone voice, a buffering step of outputting the second microphone voice with a predetermined delay, and a control step of determining whether or not to attenuate the second microphone voice delayed by the predetermined time according to the correlation between the first microphone voice and the echo component. This makes it possible to suppress erroneous recognition in voice recognition.
[0018] In yet another preferred embodiment of the present invention, a program executed by a computer causes the computer to function as an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component, a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound, a buffer unit that delays the second microphone sound by a predetermined time and outputs it, and a control unit that determines whether or not to attenuate the second microphone sound delayed by the predetermined time according to the correlation between the first microphone sound and the echo component. This makes it possible to suppress erroneous recognition in voice recognition. The above voice processing device can be realized by executing the above program on a computer. The above program can be stored in a storage medium and used. EXAMPLES
[0019] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings.
[0020] [System Configuration] 1 is a diagram showing a schematic configuration of a voice processing system according to an embodiment. As shown in FIG. 1, the voice processing system 1 includes a voice processing device 100, a microphone 200, a speaker 300, and a voice recognition engine 400.
[0021] The voice processing device 100 acquires processed microphone voice MVZ by performing voice processing such as echo cancellation on the microphone voice MVA input from the microphone 200, and outputs the acquired processed microphone voice MVZ to the voice recognition engine 400.
[0022] The voice recognition engine 400 recognizes the speech content included in the processed microphone voice MVZ by analyzing the processed microphone voice MVZ obtained from the voice processing device 100. The voice recognition engine 400 also outputs a recognition result NK of the speech content included in the processed microphone voice MVZ to the voice processing device 100. The recognition result NK may include, for example, text data or the like corresponding to the speech content included in the processed microphone voice MVZ.
[0023] The voice processing device 100 generates a system voice SV that includes the recognition result NK obtained from the voice recognition engine 400 as the speech content, and outputs the generated voice to the speaker 300.
[0024] [Audio output device] The voice processing device 100 can provide various information to the user through dialogue with the user. The voice processing device 100 may be incorporated, for example, into a navigation device installed in a vehicle and providing route guidance to a set destination. The voice processing device 100 may also be incorporated, for example, into a mobile terminal such as a smartphone carried by a user.
[0025] (Hardware configuration) 2 is a block diagram showing an example of a hardware configuration of a voice processing device according to an embodiment. As shown in FIG. 2, the voice processing device 100 includes an interface (IF) 111, a processor 112, a memory 113, and a recording medium 114.
[0026] The IF 111 inputs and outputs data to and from an external device. For example, the microphone voice MVA and the recognition result NK are input to the voice processing device 100 through the IF 111. In addition, for example, the system voice SV is output to the speaker 300 through the IF 111.
[0027] The processor 112 is a computer such as a CPU (Central Processing Unit) and executes a program prepared in advance to control the entire sound processing device 100. The processor 112 performs, for example, a process of removing the system sound SV included in the microphone sound MVA as an echo component.
[0028] The memory 113 is composed of a read only memory (ROM), a random access memory (RAM), etc. The memory 113 is also used as a working memory while the processor 112 is executing various processes.
[0029] The recording medium 114 is a non-volatile and non-temporary recording medium such as a disk-shaped recording medium or a semiconductor memory, and has a configuration that is detachable from the voice processing device 100. The recording medium 114 records various programs executed by the processor 112. When the voice processing device 100 executes various processes, the programs recorded in the recording medium 114 are loaded into the memory 113 and executed by the processor 112.
[0030] (Functional configuration) 3 is a block diagram showing an example of a functional configuration of a voice processing device according to an embodiment. As shown in FIG. 3, the voice processing device 100 includes a preprocessing unit 11, an adaptive signal processing unit 12, a buffer unit 13, an attenuation processing unit 14, a noise removal unit 15, a system voice generating unit 16, and a correlation coefficient calculation unit 17.
[0031] The preprocessing unit 11 acquires the microphone audio MVB by performing preprocessing on the microphone audio MVA input to the microphone 200. The preprocessing may include, for example, sampling, normalization, and filtering using a band-pass filter. The preprocessing unit 11 also outputs the microphone audio MVB to the adaptive signal processing unit 12 and the correlation coefficient calculation unit 17.
[0032] The adaptive signal processing unit 12 has a function capable of suppressing the system audio SV included in the microphone audio MVB by performing processing using, for example, an NLMS (Normalized Least Mean Square) algorithm. The adaptive signal processing unit 12 outputs the audio in which the system audio SV included in the microphone audio MVB is suppressed as the microphone audio MVC to the buffer unit 13. The adaptive signal processing unit 12 also performs processing related to estimation of the echo component ECC based on the microphone audio MVC output to the buffer unit 13 and the system audio SV generated by the system audio generation unit 16. The adaptive signal processing unit 12 also has an echo suppression unit 12A and an echo component estimation unit 12B.
[0033] The echo suppression unit 12A has a function as a suppression unit. The echo suppression unit 12A also subtracts the echo component ECC from the microphone audio MVB to obtain the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed. The echo suppression unit 12A also outputs the microphone audio MVC to the echo component estimation unit 12B and the buffer unit 13.
[0034] The echo component estimation unit 12B has a function as an estimation unit. Moreover, the echo component estimation unit 12B performs a process of estimating a component corresponding to the system audio SV included in the microphone audio MVB as an echo component ECC based on the microphone audio MVC output to the buffer unit 13 and the system audio SV generated by the system audio generation unit 16. Moreover, the echo component estimation unit 12B performs, for example, a process using an FIR (Finite Impulse Response) filter as a process related to the estimation of the echo component ECC. Moreover, the echo component estimation unit 12B has a parameter setting unit 12P.
[0035] The parameter setting unit 12P sets parameters for processing related to estimation of the echo component ECC based on the microphone audio MVC and the correlation coefficient CK. Specifically, the parameter setting unit 12P sets, for example, the filter coefficient FC of the FIR filter as a coefficient that changes according to the microphone audio MVC. The details of the correlation coefficient CK will be described later.
[0036] According to the above-described process, the echo component estimation unit 12B can estimate the echo component ECC by applying an FIR filter in which a filter coefficient FC is set to the system sound SV.
[0037] The buffer unit 13 outputs the microphone audio MVC delayed by a predetermined time BD as the microphone audio MVT to the attenuation processing unit 14. The predetermined time BD is desirably set, for example, as a time corresponding to a time lag until the calculation result of the correlation coefficient CK according to the current speech state is obtained. Specifically, the predetermined time BD is desirably set, for example, as a time corresponding to any of the elements that may be the cause of the time lag described above, such as the sampling rate of the microphone audio MVA in the preprocessing of the preprocessing unit 11 and the calculation speed of the correlation coefficient CK in the correlation coefficient calculation unit 17.
[0038] The attenuation processing unit 14 determines whether to attenuate the microphone audio MVT based on the correlation coefficient CK. In addition, the attenuation processing unit 14 acquires audio equivalent to the microphone audio MVT or audio obtained by attenuating the microphone audio MVT as the microphone audio MVD according to the above-mentioned determination, and outputs the acquired microphone audio MVD to the noise removal unit 15. In addition, the attenuation processing unit 14 has an attenuation control unit 14S.
[0039] The attenuation control unit 14S has a function as a control unit. Also, the attenuation control unit 14S sets the gain value GV applied to the microphone voice MVT as a value that changes according to the correlation coefficient CK calculated by the correlation coefficient calculation unit 17. Specifically, for example, when the correlation coefficient CK is greater than a predetermined threshold value THK, the attenuation control unit 14S sets the gain value GV to a value that satisfies the relationship of 0 < GN < 1. Also, for example, when the correlation coefficient CK is less than or equal to the predetermined threshold value THK, the attenuation control unit 14S sets the gain value GV to 1. In this embodiment, it is desirable that the attenuation control unit 14S sets the gain value GV when the correlation coefficient CK is greater than the predetermined threshold value THK to a value belonging to the range of 0.1 or more and 0.5 or less. Also, in this embodiment, for example, when the correlation coefficient CK is calculated as a value belonging to the range of 0 or more and 1 or less, it is desirable that the threshold value THK is set to 0.8.
[0040] According to the processing as described above, the attenuation processing unit 14 can set the gain value GV applied to the microphone voice MVT to a value of 1 or less than 1 according to the correlation coefficient CK. Also, the attenuation processing unit 14 can obtain the microphone voice MVD by applying the gain value GV set as described above to the microphone voice MVT, and output the obtained microphone voice MVD to the noise removal unit 15.
[0041] The noise removal unit 15 obtains the processed microphone voice MVZ by performing processing related to noise removal on the microphone voice MVD, and outputs the obtained processed microphone voice MVZ to the speech recognition engine 400.
[0042] The system voice generation unit 16 generates a system voice SV including the recognition result NK obtained from the speech recognition engine 400 as the utterance content, and outputs the generated voice to the echo component estimation unit 12B and the speaker 300.
[0043] The correlation coefficient calculation unit 17 has a function as a calculation unit. Moreover, the correlation coefficient calculation unit 17 calculates a value belonging to the range of 0 or more and 1 or less as a correlation coefficient CK indicating the correlation between the microphone audio MVB and the echo component ECC. Specifically, for example, when the correlation between the microphone audio MVB and the echo component ECC is high, the correlation coefficient calculation unit 17 calculates a value of 1 or close to 1 as the correlation coefficient CK. Moreover, for example, when the correlation between the microphone audio MVB and the echo component ECC is low, the correlation coefficient calculation unit 17 calculates a value of 0 or close to 0 as the correlation coefficient CK. Moreover, the correlation coefficient calculation unit 17 outputs the correlation coefficient CK to the attenuation processing unit 14. According to such a process, the attenuation control unit 14S can attenuate the microphone audio MVT when the correlation between the microphone audio MVB and the echo component ECC is high. Moreover, according to the above-mentioned process, the attenuation control unit 14S can prevent the microphone audio MVT from being attenuated when the correlation between the microphone audio MVB and the echo component ECC is low. Furthermore, according to the above-described processing, the attenuation processing unit 14 can determine whether or not to attenuate the microphone audio MVT depending on the correlation between the microphone audio MVB and the echo component ECC.
[0044] [Examples of voice processing] Next, a specific example of the voice processing in this embodiment will be described. FIG. 4 is a diagram showing an example of a temporal change in the correlation coefficient used in the voice processing according to the embodiment. In this specific example, the period during which the voice processing device 100 continues to speak is represented as a period PSN, and the period during which the user speaks in the period PSN is represented as a period PSU. In other words, in this specific example, a case will be described in which the system voice SV continues to be output from the speaker 300 in the period PSN, and the user voice UV corresponding to the user's speech is generated in the period PSU included in the period PSN. In addition, in this specific example, the start time of the period PSN is represented as time T1, the start time of the period PSU is represented as time T2, the end time of the period PSU is represented as time T3, and the end time of the period PSN is represented as time T4.
[0045] During a period PA from time T1 to time T2, a sound including the system sound SV but not including the user sound UV is input to the sound processing device 100 as the microphone sound MVA. Accordingly, the correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK during the period PA (see FIG. 4). Furthermore, the attenuation control unit 14S sets the gain value GV to be applied to the microphone sound MVT to a value larger than 0 and smaller than 1 according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17. It is preferable that the attenuation control unit 14S sets the gain value GV during the period PA to a value that falls within the range of 0.1 or more and 0.5 or less.
[0046] According to the above-described processing, the attenuation processing unit 14 can sufficiently attenuate the system sound SV included in the microphone sound MVT during the period PA.
[0047] In a period PSU from time T2 to time T3, a mixed voice including the system voice SV and the user voice UV is input to the voice processing device 100 as the microphone voice MVA.
[0048] The correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK for the period from time T2 until the predetermined time BD has elapsed (see FIG. 4). Moreover, the correlation coefficient calculation unit 17 calculates a value equal to or smaller than the threshold value THK as the correlation coefficient CK at time T21 when the predetermined time BD has elapsed since time T2 (see FIG. 4). Moreover, the correlation coefficient calculation unit 17 calculates a value equal to or smaller than the threshold value THK as the correlation coefficient CK for the period from time T21 to time T3 (see FIG. 4). Moreover, the correlation coefficient calculation unit 17 calculates a value equal to or smaller than the threshold value THK as the correlation coefficient CK for the period from time T3 until time T31 when the predetermined time BD has elapsed (see FIG. 4).
[0049] The attenuation control unit 14S sets the gain value GV to be applied to the microphone voice MVT for the period from time T2 until a predetermined time BD has elapsed, to a value greater than 0 and less than 1, in accordance with the calculation result of the correlation coefficient CK for that period. Also, the attenuation control unit 14S sets the gain value GV to be applied to the microphone voice MVT for the period from time T21 to time T31, to 1, in accordance with the calculation result of the correlation coefficient CK for that period.
[0050] According to the above-described process, the attenuation processing unit 14 can determine whether or not to attenuate the microphone voice MVT while considering the time lag of a predetermined time BD that occurs from the time T2 when the speech state changes until the calculation result of the correlation coefficient CK at the time T2 is obtained. Specifically, according to the above-described process, for example, when the microphone voice MVC at the time T2 is input to the attenuation processing unit 14 as the microphone voice MVT at the time T21, the attenuation control unit 14S can prevent the microphone voice MVT from being attenuated according to the calculation result of the correlation coefficient CK at the time T21. Therefore, according to the above-described process, the attenuation control unit 14S can prevent the user voice UV included in the microphone voice MVT from being attenuated during the period from the time T21 to the time T31 when the microphone voice MVC of the period PSU is input as the microphone voice MVT.
[0051] In a period PB from time T3 to time T4, a sound that includes the system sound SV but does not include the user sound UV is input to the sound processing device 100 as the microphone sound MVA.
[0052] The correlation coefficient calculation unit 17 calculates a value larger than a threshold value THK as the correlation coefficient CK for the period from time T31 to time T4 (see FIG. 4). Furthermore, the attenuation control unit 14S sets the gain value GV to be applied to the microphone voice MVT for the period from time T31 to time T4 to a value larger than 0 and smaller than 1, depending on the calculation result of the correlation coefficient CK for the period from time T31 to time T4.
[0053] According to the above-described process, the attenuation processing unit 14 can sufficiently attenuate the system sound SV included in the microphone sound MVT in the period from time T31 to time T4 in the period PB.
[0054] [Processing flow] Next, a description will be given of a flow of processing performed by the voice processing device 100. Fig. 5 is a flowchart showing an example of processing performed by the voice processing device. It should be noted that the voice processing device 100 repeats the processing of Fig. 5 during an operation period corresponding to a period from when the power is turned on to when it is turned off, for example.
[0055] First, the audio processing device 100 acquires the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed by subtracting the echo component ECC from the microphone audio MVB (step S11).
[0056] Next, the sound processing device 100 calculates a correlation coefficient CK indicating the correlation between the microphone sound MVB that is the processing target of step S11 and the echo component ECC used in the processing of step S11 (step S12).
[0057] Next, the sound processing device 100 determines whether the correlation coefficient CK calculated in step S12 is greater than a threshold value THK (step S13).
[0058] If the correlation coefficient CK is greater than the threshold THK (step S13: YES), the voice processing device 100 sets the gain value GV to a value greater than 0 and less than 1 (step S14). If the correlation coefficient CK is equal to or less than the threshold THK (step S13: NO), the voice processing device 100 sets the gain value GV to 1 (step S15).
[0059] The audio processing device 100 applies the gain value GV set in step S14 or step S15 to the microphone audio MVT obtained by delaying the microphone audio MVC obtained by the process of step S11 by a predetermined time BD (step S16).
[0060] As described above, according to this embodiment, even if the user and the voice processing device 100 speak at the same time, a voice that can identify the contents of the user's speech can be output to the voice recognition engine 400. Therefore, according to this embodiment, it is possible to suppress erroneous recognition in voice recognition.
[0061] As described above, according to this embodiment, even if the voice (system voice) of the voice processing device 100 is input to the system through the microphone 200 during a period when the user is not speaking, the speech content included in the voice can be prevented from being recognized by the voice recognition engine 400. Therefore, according to this embodiment, it is possible to suppress erroneous recognition in voice recognition.
[0062] In the above-described embodiment, the program can be stored using various types of non-transitory computer readable media and supplied to a control unit, which is a computer. The non-transitory computer readable media includes various types of tangible storage media. Examples of the non-transitory computer readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROM (Read Only Memory), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, and RAM (Random Access Memory).
[0063] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. In other words, the present invention naturally includes various modifications and corrections that a person skilled in the art could make in accordance with the entire disclosure, including the claims, and the technical ideas. In addition, the disclosures of the above cited patent documents and the like are incorporated herein by reference. [Explanation of symbols]
[0064] 11 Pretreatment section 12 Adaptive signal processing section 12A Echo suppression section 12B Echo component estimation section 12P Parameter setting section 13 Buffer section 14 Attenuation processing section 14S Damping control section 15 Noise Reduction Section 16 System audio generation section 17 Correlation coefficient calculation section
Claims
1. an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; a buffer unit that delays the second microphone voice by a predetermined time and outputs the delayed voice; a control unit that determines whether to attenuate the second microphone sound delayed by the predetermined time in accordance with a correlation between the first microphone sound and the echo component; 13. An audio processing device comprising:
2. 2. The audio processing device according to claim 1, wherein the control unit attenuates the second microphone audio delayed by the predetermined time when the correlation between the first microphone audio and the echo component is high, and does not attenuate the second microphone audio delayed by the predetermined time when the correlation between the first microphone audio and the echo component is low.
3. A calculation unit is further provided to calculate a correlation coefficient indicating a correlation between the first microphone voice and the echo component, 2. The audio processing device of claim 1, wherein the control unit sets a gain value applied to the second microphone audio delayed by the predetermined time to a value less than 1 when the correlation coefficient is greater than a predetermined threshold, and sets a gain value applied to the second microphone audio delayed by the predetermined time to 1 when the correlation coefficient is equal to or less than the predetermined threshold.
4. The audio processing device according to claim 3 , wherein the control unit sets a gain value applied to the second microphone audio delayed by the predetermined time to a value greater than or equal to 0.1 and less than or equal to 0.5 when the correlation coefficient is greater than the predetermined threshold value.
5. 1. A computer implemented method for audio processing, comprising: an estimation step of estimating a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression step of outputting a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; a buffering step of delaying the second microphone voice by a predetermined time and outputting the delayed voice; a control step of determining whether or not to attenuate the second microphone sound delayed by the predetermined time in accordance with a correlation between the first microphone sound and the echo component; 13. A method for processing audio comprising the steps of:
6. A program executed by a computer, an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; a buffer unit that delays the second microphone voice by a predetermined time and outputs the delayed voice; A program that causes the computer to function as a control unit that determines whether or not to attenuate the second microphone sound delayed by the predetermined time in accordance with the correlation between the first microphone sound and the echo component.
7. A storage medium storing the program according to claim 6.
Citation Information
Patent Citations
Voice recognition system and voice recognizer
JP2009109536A