Voice processing device, voice processing method, program, and storage medium

The audio processing device addresses the issue of erroneous voice recognition by estimating and suppressing echo components in microphone inputs and adjusting attenuation periods based on correlation, resulting in improved recognition accuracy.

JP2025079373APending Publication Date: 2025-05-22PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023191950
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Echo cancellation techniques in speech recognition systems, such as smart speakers, often fail to adequately remove system audio as an echo component, leading to erroneous recognition results, especially when threshold values fluctuate.

Method used

An audio processing device comprising an estimation unit to identify echo components from system audio in microphone inputs, a suppression unit to subtract these echo components from the microphone audio, and a control unit to set a non-attenuation period based on the correlation between microphone audio and echo components, ensuring accurate voice recognition.

Benefits of technology

The proposed solution effectively suppresses erroneous recognition in voice recognition systems by accurately removing echo components and adjusting attenuation periods according to the correlation between microphone audio and echo components.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079373000001_ABST
    Figure 2025079373000001_ABST
Patent Text Reader

Abstract

To provide a voice processing device, a voice processing method, and a program which can suppress misrecognition in voice recognition.SOLUTION: A voice processing device includes an estimation section, a suppression section, and a control section. The estimation section estimates a component corresponding to a system voice which is included in a first microphone voice input from a microphone as an echo component. The suppression section subtracts the echo component from the first microphone voice, so as to output a second microphone voice in which the system voice included in the first microphone voice is suppressed. The control section sets a non-attenuation period equivalent to a period during which the second microphone voice is not attenuated according to the correlation between the first microphone voice and the echo component.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for processing audio. [Background technology]

[0002] 2. Description of the Related Art Echo cancellation techniques for removing echo components contained in speech are known in the art.

[0003] Specifically, for example, Patent Document 1 discloses a technique for removing a component corresponding to an echo from a sound including a user's speech input to a microphone and the echo input to the microphone. Patent Document 1 also discloses a technique for detecting a sound portion having a sound pressure exceeding a threshold as a user's speech section. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2009-109536 A Summary of the Invention [Problem to be solved by the invention]

[0005] For example, when echo cancellation is applied to speech input into an interactive device that uses speech recognition, such as a smart speaker, the speech emitted from the device itself may not be sufficiently removed as an echo component, resulting in an erroneous recognition result.

[0006] According to the technology disclosed in Patent Document 1, for example, in a situation where the threshold value frequently changes, the above-mentioned problem may occur. Therefore, according to the technology disclosed in Patent Document 1, a problem corresponding to the above-mentioned problem occurs.

[0007] In view of the above-mentioned problems, a main object of the present disclosure is to provide a voice processing device capable of suppressing erroneous recognition in voice recognition. [Means for solving the problem]

[0008] The invention described in the claims is an audio processing device comprising: an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component; a suppression unit that outputs a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio; and a control unit that sets a non-attenuation period corresponding to a period during which the second microphone audio is not attenuated in accordance with the correlation between the first microphone audio and the echo component.

[0009] The invention described in the claims is an audio processing method executed by a computer, comprising: an estimation step of estimating a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component; a suppression step of outputting a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio; and a control step of setting a non-attenuation period corresponding to a period during which the second microphone audio is not attenuated in accordance with the correlation between the first microphone audio and the echo component.

[0010] The invention described in the claims is a program executed by a computer, which causes the computer to function as an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component, a suppression unit that subtracts the echo component from the first microphone audio to output a second microphone audio in which the system audio contained in the first microphone audio is suppressed, and a control unit that sets a non-attenuation period corresponding to a period during which the second microphone audio is not attenuated depending on the correlation between the first microphone audio and the echo component. [Brief description of the drawings]

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

[0012] In one preferred embodiment of the present invention, an audio processing device includes an estimation unit that estimates a component corresponding to a system audio contained in a first microphone audio input from a microphone as an echo component, a suppression unit that outputs a second microphone audio in which the system audio contained in the first microphone audio is suppressed by subtracting the echo component from the first microphone audio, and a control unit that sets a non-attenuation period corresponding to a period during which the second microphone audio is not attenuated in accordance with the correlation between the first microphone audio and the echo component.

[0013] The above voice processing device includes an estimation unit, a suppression unit, and a control unit. The estimation unit estimates a component corresponding to a system voice included in a first microphone voice input from a microphone as an echo component. The suppression unit subtracts the echo component from the first microphone voice to output a second microphone voice in which the system voice included in the first microphone voice is suppressed. The control unit sets a non-attenuation period corresponding to a period during which the second microphone voice is not attenuated according to the correlation between the first microphone voice and the echo component. This makes it possible to suppress erroneous recognition in voice recognition.

[0014] In one aspect of the above audio processing device, the control unit sets the non-attenuation period to be the period from the time when the correlation between the first microphone audio and the echo component transitions from a high state to a low state until a predetermined time has elapsed.

[0015] In one embodiment of the above-mentioned audio processing device, the device further includes a calculation unit that calculates a correlation coefficient indicating the correlation between the first microphone audio and the echo component, and the control unit sets the non-attenuation period to be the period from the time when the correlation coefficient transitions from a value greater than a predetermined threshold to a value equal to or less than the predetermined threshold until a predetermined time has elapsed.

[0016] In one aspect of the above voice processing device, the control unit sets the specified time to the time required to speak a voice command to activate a voice recognition engine that recognizes the speech content contained in the second microphone voice.

[0017] In another preferred embodiment of the present invention, a computer-executed voice processing method includes an estimation step of estimating a component corresponding to a system voice included in a first microphone voice input from a microphone as an echo component, a suppression step of outputting a second microphone voice in which the system voice included in the first microphone voice is suppressed by subtracting the echo component from the first microphone voice, and a control step of setting a non-attenuation period corresponding to a period during which the second microphone voice is not attenuated according to the correlation between the first microphone voice and the echo component. This makes it possible to suppress erroneous recognition in voice recognition.

[0018] In yet another preferred embodiment of the present invention, a program executed by a computer causes the computer to function as an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component, a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound, and a control unit that sets a non-attenuation period corresponding to a period during which the second microphone sound is not attenuated according to the correlation between the first microphone sound and the echo component. This makes it possible to suppress erroneous recognition in voice recognition. The above voice processing device can be realized by executing the above program on a computer. The above program can be stored in a storage medium and used. EXAMPLES

[0019] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings.

[0020] [System Configuration] 1 is a diagram showing a schematic configuration of a voice processing system according to an embodiment. As shown in FIG. 1, the voice processing system 1 includes a voice processing device 100, a microphone 200, a speaker 300, and a voice recognition engine 400.

[0021] The voice processing device 100 acquires processed microphone voice MVZ by performing voice processing such as echo cancellation on the microphone voice MVA input from the microphone 200, and outputs the acquired processed microphone voice MVZ to the voice recognition engine 400.

[0022] The voice recognition engine 400 recognizes the speech content included in the processed microphone voice MVZ by analyzing the processed microphone voice MVZ obtained from the voice processing device 100. The voice recognition engine 400 also outputs a recognition result NK of the speech content included in the processed microphone voice MVZ to the voice processing device 100. The recognition result NK may include, for example, text data or the like corresponding to the speech content included in the processed microphone voice MVZ.

[0023] The voice processing device 100 generates a system voice SV that includes the recognition result NK obtained from the voice recognition engine 400 as the speech content, and outputs the generated voice to the speaker 300.

[0024] [Audio output device] The voice processing device 100 can provide various information to the user through dialogue with the user. The voice processing device 100 may be incorporated, for example, into a navigation device installed in a vehicle and providing route guidance to a set destination. The voice processing device 100 may also be incorporated, for example, into a mobile terminal such as a smartphone carried by a user.

[0025] (Hardware configuration) 2 is a block diagram showing an example of a hardware configuration of a voice processing device according to an embodiment. As shown in FIG. 2, the voice processing device 100 includes an interface (IF) 111, a processor 112, a memory 113, and a recording medium 114.

[0026] The IF 111 inputs and outputs data to and from an external device. For example, the microphone voice MVA and the recognition result NK are input to the voice processing device 100 through the IF 111. In addition, for example, the system voice SV is output to the speaker 300 through the IF 111.

[0027] The processor 112 is a computer such as a CPU (Central Processing Unit) and executes a program prepared in advance to control the entire sound processing device 100. The processor 112 performs, for example, a process of removing the system sound SV included in the microphone sound MVA as an echo component.

[0028] The memory 113 is composed of a read only memory (ROM), a random access memory (RAM), etc. The memory 113 is also used as a working memory while the processor 112 is executing various processes.

[0029] The recording medium 114 is a non-volatile and non-temporary recording medium such as a disk-shaped recording medium or a semiconductor memory, and has a configuration that is detachable from the voice processing device 100. The recording medium 114 records various programs executed by the processor 112. When the voice processing device 100 executes various processes, the programs recorded in the recording medium 114 are loaded into the memory 113 and executed by the processor 112.

[0030] (Functional configuration) 3 is a block diagram showing an example of a functional configuration of a voice processing device according to an embodiment. As shown in FIG. 3, the voice processing device 100 includes a preprocessing unit 11, an adaptive signal processing unit 12, an attenuation processing unit 14, a noise reduction unit 15, a system voice generating unit 16, and a correlation coefficient calculation unit 17.

[0031] The preprocessing unit 11 acquires the microphone audio MVB by performing preprocessing on the microphone audio MVA input to the microphone 200. The preprocessing may include, for example, sampling, normalization, and filtering using a band-pass filter. The preprocessing unit 11 also outputs the microphone audio MVB to the adaptive signal processing unit 12 and the correlation coefficient calculation unit 17.

[0032] The adaptive signal processing unit 12 has a function capable of suppressing the system audio SV included in the microphone audio MVB by performing processing using, for example, an NLMS (Normalized Least Mean Square) algorithm. The adaptive signal processing unit 12 outputs the audio in which the system audio SV included in the microphone audio MVB is suppressed as the microphone audio MVC to the attenuation processing unit 14. The adaptive signal processing unit 12 also performs processing related to estimation of the echo component ECC based on the microphone audio MVC output to the attenuation processing unit 14 and the system audio SV generated by the system audio generating unit 16. The adaptive signal processing unit 12 also has an echo suppression unit 12A and an echo component estimation unit 12B.

[0033] The echo suppression unit 12A has a function as a suppression unit. The echo suppression unit 12A also subtracts the echo component ECC from the microphone audio MVB to obtain the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed. The echo suppression unit 12A also outputs the microphone audio MVC to the echo component estimation unit 12B and the attenuation processing unit 14.

[0034] The echo component estimation unit 12B has a function as an estimation unit. Moreover, the echo component estimation unit 12B performs a process of estimating a component corresponding to the system audio SV included in the microphone audio MVB as an echo component ECC based on the microphone audio MVC output to the attenuation processing unit 14 and the system audio SV generated by the system audio generation unit 16. Moreover, the echo component estimation unit 12B performs, for example, a process using an FIR (Finite Impulse Response) filter as a process related to the estimation of the echo component ECC. Moreover, the echo component estimation unit 12B has a parameter setting unit 12P.

[0035] The parameter setting unit 12P sets parameters for processing related to estimation of the echo component ECC based on the microphone audio MVC and the correlation coefficient CK. Specifically, the parameter setting unit 12P sets, for example, the filter coefficient FC of the FIR filter as a coefficient that changes according to the microphone audio MVC. The details of the correlation coefficient CK will be described later.

[0036] According to the above-described processing, the echo component estimation unit 12B can estimate the echo component ECC by applying an FIR filter in which a filter coefficient FC is set to the system sound SV.

[0037] The attenuation processing unit 14 sets a non-attenuation period NDP corresponding to a period during which the microphone audio MVC is not attenuated and an attenuation period PDP corresponding to a period during which the microphone audio MVC is attenuated based on the correlation coefficient CK. In addition, the attenuation processing unit 14 acquires audio equivalent to the microphone audio MVC as the microphone audio MVD during the non-attenuation period NDP set as described above, and outputs the acquired microphone audio MVD to the noise elimination unit 15. In addition, the attenuation processing unit 14 acquires audio obtained by attenuating the microphone audio MVC as the microphone audio MVD during the attenuation period PDP set as described above, and outputs the acquired microphone audio MVD to the noise elimination unit 15. In addition, the attenuation processing unit 14 has an attenuation control unit 14S.

[0038] The attenuation control unit 14S has a function as a control unit. The attenuation control unit 14S sets a non-attenuation period NDP and a attenuation period PDP based on the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17. Specifically, the attenuation control unit 14S sets, for example, a period from the timing when the correlation coefficient CK transitions from a value larger than a threshold value THK to a value equal to or smaller than the threshold value THK until a predetermined time TJN has elapsed as the non-attenuation period NDP. The attenuation control unit 14S also sets a gain value GV applied to the microphone voice MVC in the non-attenuation period NDP to 1. The attenuation control unit 14S also sets, for example, a time required to utter a voice command for starting the voice recognition engine 400 as the predetermined time TJN. The attenuation control unit 14S also sets, for example, a period during which the correlation coefficient CK is greater than the threshold value THK as the attenuation period PDP, in a period excluding the non-attenuation period NDP. The attenuation control unit 14S also sets a gain value GV applied to the microphone voice MVC in the attenuation period PDP to a value greater than 0 and less than 1.

[0039] In this embodiment, the "voice command" can be rephrased as a "wake-up word." In this embodiment, it is preferable that the attenuation control unit 14S sets the predetermined time TJN to 0.8 seconds, for example. The predetermined time TJN may be a fixed value, or may be a variable value that changes depending on the type of voice command, etc. In this embodiment, it is preferable that the threshold value THK is set to 0.8 when the correlation coefficient CK is calculated as a value that falls within the range of 0 or more and 1 or less.

[0040] According to this embodiment, the attenuation control unit 14S sets a non-attenuation period corresponding to a period during which the second microphone sound is not attenuated, in accordance with the correlation between the first microphone sound and the echo component.

[0041] The noise removal unit 15 performs processing related to noise removal on the microphone voice MVD to obtain a processed microphone voice MVZ, and outputs the obtained processed microphone voice MVZ to the voice recognition engine 400.

[0042] The system voice generating unit 16 generates a system voice SV including the recognition result NK obtained from the voice recognition engine 400 as the speech content, and outputs the generated voice to the echo component estimating unit 12B and the speaker 300.

[0043] The correlation coefficient calculation unit 17 has a function as a calculation unit. The correlation coefficient calculation unit 17 calculates a value in the range of 0 to 1 as a correlation coefficient CK indicating the correlation between the microphone audio MVB and the echo component ECC. Specifically, for example, when the correlation between the microphone audio MVB and the echo component ECC is high, the correlation coefficient calculation unit 17 calculates a value of 1 or close to 1 as the correlation coefficient CK. For example, when the correlation between the microphone audio MVB and the echo component ECC is low, the correlation coefficient calculation unit 17 calculates a value of 0 or close to 0 as the correlation coefficient CK. The correlation coefficient calculation unit 17 outputs the correlation coefficient CK to the attenuation processing unit 14. According to such processing, the attenuation control unit 14S can set a non-attenuation period NDP corresponding to a period during which the microphone audio MVC is not attenuated according to the correlation between the microphone audio MVB and the echo component ECC. In addition, according to the above-mentioned processing, the attenuation control unit 14S can set the period from the time when the correlation between the microphone audio MVB and the echo component ECC transitions from a high state to a low state until a predetermined time TJN has elapsed as the non-attenuation period NPD.

[0044] [Examples of voice processing] Next, a specific example of the voice processing in this embodiment will be described. FIG. 4 is a diagram showing an example of a temporal change in the correlation coefficient used in the voice processing according to the embodiment. In this specific example, the period during which the voice processing device 100 continues to speak is represented as a period PSN, and the period during which the user speaks in the period PSN is represented as a period PSU. In other words, in this specific example, a case will be described in which the system voice SV continues to be output from the speaker 300 in the period PSN, and the user voice UV corresponding to the user's speech is generated in the period PSU included in the period PSN. In addition, in this specific example, the start time of the period PSN is represented as time T1, the start time of the period PSU is represented as time T2, the end time of the period PSU is represented as time T3, and the end time of the period PSN is represented as time T4.

[0045] During a period PA from time T1 to time T2, a sound including the system sound SV but not including the user sound UV is input to the sound processing device 100 as the microphone sound MVA. Accordingly, the correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK in the period PA (see FIG. 4). Furthermore, the attenuation control unit 14S sets the period PA to the attenuation period PDP according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17. Furthermore, the attenuation control unit 14S sets the gain value GV applied to the microphone sound MVC in the period PA to a value larger than 0 and smaller than 1.

[0046] According to the above-described process, the attenuation processing unit 14 can sufficiently attenuate the system sound SV included in the microphone sound MVC during the period PA.

[0047] In a period PSU from time T2 to time T3, a mixed sound including the system sound SV and the user sound UV is input to the sound processing device 100 as the microphone sound MVA. Accordingly, the correlation coefficient calculation unit 17 calculates a value equal to or less than the threshold value THK as the correlation coefficient CK at time T2 (see FIG. 4). In other words, the correlation coefficient CK transitions from a value greater than the threshold value THK to a value equal to or less than the threshold value THK at time T2. In addition, the correlation coefficient calculation unit 17 calculates a correlation coefficient CK according to the correlation between the microphone sound MVB and the echo component ECC in the period PSU (see FIG. 4). In addition, the attenuation control unit 14S sets a period from time T2 until a predetermined time TJN has elapsed as a non-attenuation period NDP according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17. Note that, in this specific example, for simplicity, a case will be described in which the attenuation control unit 14S sets the non-attenuation period NDP to coincide with the period PSU. Moreover, the attenuation control unit 14S sets the gain value GV applied to the microphone audio MVC in the period PSU to 1.

[0048] According to the processing described above, even if a situation occurs in which a correlation coefficient CK greater than the threshold value THK is temporarily calculated in a period PSU due to reasons such as the volume of the user voice UV included in the microphone audio MVA being unstable (see FIG. 4), the attenuation processing unit 14 can prevent the user voice UV included in the microphone audio MVC of the period PSU from being attenuated.

[0049] During a period PB from time T3 to time T4, a sound including the system sound SV but not including the user sound UV is input to the sound processing device 100 as the microphone sound MVA. Accordingly, the correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK in the period PB (see FIG. 4). Furthermore, the attenuation control unit 14S sets the period PB to the attenuation period PDP according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17. Furthermore, the attenuation control unit 14S sets the gain value GV applied to the microphone sound MVC in the period PB to a value larger than 0 and smaller than 1.

[0050] According to the above-described process, the attenuation processing unit 14 can sufficiently attenuate the system sound SV included in the microphone sound MVC during the period PB.

[0051] In addition, for a period in which audio including user audio UV but not system audio SV is input as microphone audio MVA, the attenuation control unit 14S obtains a value of 0 or close to 0 as the calculation result of the correlation coefficient CK, and sets the period in question as a non-attenuation period NDP.

[0052] [Processing flow] Next, a description will be given of a flow of processing performed by the voice processing device 100. Fig. 5 is a flowchart showing an example of processing performed by the voice processing device. It should be noted that the voice processing device 100 repeats the processing of Fig. 5 during an operation period corresponding to a period from when the power is turned on to when it is turned off, for example.

[0053] First, the audio processing device 100 acquires the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed by subtracting the echo component ECC from the microphone audio MVB (step S11).

[0054] Next, the sound processing device 100 calculates a correlation coefficient CK indicating the correlation between the microphone sound MVB that is the processing target of step S11 and the echo component ECC used in the processing of step S11 (step S12).

[0055] Next, the sound processing device 100 determines whether or not the timing at which the calculation result of the correlation coefficient CK in step S12 is obtained belongs to a non-decay period (step S13).

[0056] If the timing at which the calculation result of the correlation coefficient CK in step S12 is obtained does not belong to the non-decay period NDP (step S13: NO), the audio processing device 100 performs the process of step S14 described below. If the timing at which the calculation result of the correlation coefficient CK in step S12 is obtained belongs to the non-decay period NDP (step S13: YES), the audio processing device 100 ends a series of processes while maintaining the non-decay period NDP. In this embodiment, it is preferable that the attenuation control unit 14S performs the process of step S13.

[0057] Based on the calculation result of the correlation coefficient CK in step S12, the sound processing device 100 determines whether or not the correlation coefficient CK has transitioned from a value larger than the threshold value THK to a value equal to or smaller than the threshold value THK (step S14).

[0058] When the correlation coefficient CK transitions from a value larger than the threshold THK to a value equal to or smaller than the threshold THK (step S14: YES), the audio processing device 100 sets the period from the timing when the calculation result of the correlation coefficient CK is obtained until a predetermined time TJN has elapsed as a non-decay period NDP (step S15), and then ends the series of processes. When the correlation coefficient CK maintains a value larger than the threshold THK (step S14: NO), the audio processing device 100 sets the period until the next calculation result of the correlation coefficient CK is obtained as a decay period PDP (step S16), and then ends the series of processes. In this embodiment, it is preferable that the attenuation control unit 14S performs the processes of steps S14 to S16.

[0059] As described above, according to this embodiment, even if the user and the voice processing device 100 speak at the same time, a voice that can identify the contents of the user's speech can be output to the voice recognition engine 400. Therefore, according to this embodiment, it is possible to suppress erroneous recognition in voice recognition.

[0060] As described above, according to this embodiment, even if the voice (system voice) of the voice processing device 100 is input to the system through the microphone 200 during a period when the user is not speaking, the speech content included in the voice can be prevented from being recognized by the voice recognition engine 400. Therefore, according to this embodiment, it is possible to suppress erroneous recognition in voice recognition.

[0061] In the above-described embodiment, the program can be stored using various types of non-transitory computer readable media and supplied to a control unit, which is a computer. The non-transitory computer readable media includes various types of tangible storage media. Examples of the non-transitory computer readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROM (Read Only Memory), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, and RAM (Random Access Memory).

[0062] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. In other words, the present invention naturally includes various modifications and corrections that a person skilled in the art could make in accordance with the entire disclosure, including the claims, and the technical ideas. In addition, the disclosures of the above cited patent documents and the like are incorporated herein by reference. [Explanation of symbols]

[0063] 11 Pretreatment section 12 Adaptive signal processing section 12A Echo suppression section 12B Echo component estimation section 12P Parameter setting section 14 Attenuation processing section 14S Damping control section 15 Noise Reduction Section 16 System audio generation section 17 Correlation coefficient calculation section

Claims

1. an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; a control unit that sets a non-attenuation period corresponding to a period during which the second microphone voice is not attenuated in accordance with a correlation between the first microphone voice and the echo component; 13. An audio processing device comprising:

2. The audio processing device according to claim 1 , wherein the control unit sets the non-decay period to a period from when the correlation between the first microphone audio and the echo component transitions from a high state to a low state until a predetermined time has elapsed.

3. A calculation unit is further provided to calculate a correlation coefficient indicating a correlation between the first microphone voice and the echo component, The audio processing device according to claim 1 , wherein the control unit sets, as the non-decay period, a period from when the correlation coefficient transitions from a value greater than a predetermined threshold to a value equal to or less than the predetermined threshold until a predetermined time has elapsed.

4. The voice processing device according to claim 2 , wherein the control unit sets, as the predetermined time, a time required for uttering a voice command for activating a voice recognition engine that recognizes a speech content included in the second microphone voice.

5. 1. A computer implemented method for audio processing, comprising: an estimation step of estimating a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression step of outputting a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; a control step of setting a non-attenuation period corresponding to a period during which the second microphone sound is not attenuated in accordance with a correlation between the first microphone sound and the echo component; 13. A method for processing audio comprising the steps of:

6. A program executed by a computer, an estimation unit that estimates a component corresponding to a system sound included in a first microphone sound input from a microphone as an echo component; a suppression unit that outputs a second microphone sound in which the system sound included in the first microphone sound is suppressed by subtracting the echo component from the first microphone sound; and A program that causes the computer to function as a control unit that sets a non-attenuation period corresponding to a period during which the second microphone sound is not attenuated, according to the correlation between the first microphone sound and the echo component.

7. A storage medium storing the program according to claim 6.

Citation Information

Patent Citations

  • Voice recognition system and voice recognizer

    JP2009109536A