Voice processing device, voice processing method, program, and storage medium

The audio processing device enhances voice recognition accuracy in voice recognition systems by estimating and suppressing echo components based on correlation, addressing the issue of degraded voice quality and recognition accuracy.

JP2025079377APending Publication Date: 2025-05-22PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023191954
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-22

AI Technical Summary

Technical Problem

Echo cancellation techniques in voice recognition systems, such as smart speakers, often degrade voice quality, leading to decreased recognition accuracy.

Method used

An audio processing device with an estimation unit that identifies echo components in microphone audio and a setting unit that adjusts the step size for echo estimation based on the correlation between microphone audio and echo components.

Benefits of technology

Improves voice recognition accuracy by effectively suppressing echo components while maintaining voice quality, even during simultaneous user and device speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079377000001_ABST
    Figure 2025079377000001_ABST
Patent Text Reader

Abstract

To provide a voice processing device, a voice processing method, and a program which can achieve improved recognition accuracy of voice recognition.SOLUTION: A voice processing device includes an estimation section and a setting section. The estimation section estimates a component corresponding to a system voice which is included in a microphone voice input in a microphone as an echo component. The setting section sets a step size related to the estimation of the echo component to a different size according to the correlation between the microphone voice and the echo component.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to techniques for processing audio. [Background technology]

[0002] 2. Description of the Related Art Echo cancellation techniques for removing echo components contained in speech are known in the art.

[0003] Specifically, for example, Patent Document 1 discloses a technique for removing a component corresponding to an echo from a voice including a user's speech input to a microphone and the echo input to the microphone. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] JP 2009-109536 A Summary of the Invention [Problem to be solved by the invention]

[0005] For example, when echo cancellation is applied to voice input into an interactive device that uses voice recognition, such as a smart speaker, a problem occurs in which the recognition accuracy of the voice recognition decreases due to degradation of the voice.

[0006] In contrast, Patent Document 1 does not particularly disclose a method for solving the above-mentioned problems. Therefore, the technology disclosed in Patent Document 1 causes problems corresponding to the above-mentioned problems.

[0007] In view of the above-mentioned problems, a main object of the present disclosure is to provide a voice processing device capable of improving the recognition accuracy of voice recognition. [Means for solving the problem]

[0008] The invention described in the claims is an audio processing device having an estimation unit that estimates a component corresponding to system audio contained in microphone audio input to a microphone as an echo component, and a setting unit that sets a step size related to the estimation of the echo component to different sizes depending on the correlation between the microphone audio and the echo component.

[0009] The invention described in the claims is an audio processing method executed by a computer, comprising an estimation step of estimating a component corresponding to a system audio contained in a microphone audio input to a microphone as an echo component, and a setting step of setting a step size related to the estimation of the echo component to a different size depending on the correlation between the microphone audio and the echo component.

[0010] The invention described in the claims is a program executed by a computer, which causes the computer to function as an estimation unit that estimates a component corresponding to system audio contained in microphone audio input to a microphone as an echo component, and a setting unit that sets a step size related to the estimation of the echo component to different sizes depending on the correlation between the microphone audio and the echo component. [Brief description of the drawings]

[0011] [Figure 1] FIG. 1 is a diagram showing an example of a configuration of a voice processing system according to an embodiment. [Diagram 2] FIG. 1 is a block diagram showing an example of a hardware configuration of a voice processing device according to an embodiment. [Diagram 3] FIG. 1 is a block diagram showing an example of a functional configuration of a voice processing device according to an embodiment. [Figure 4] 6 is a diagram showing an example of a change over time in a correlation coefficient used in the voice processing according to the embodiment. [Diagram 5] 4 is a flowchart showing an example of processing performed by a voice processing device. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0012] In one preferred embodiment of the present invention, an audio processing device has an estimation unit that estimates a component corresponding to a system audio contained in a microphone audio input to a microphone as an echo component, and a setting unit that sets a step size related to the estimation of the echo component to different sizes depending on the correlation between the microphone audio and the echo component.

[0013] The voice processing device includes an estimation unit and a setting unit. The estimation unit estimates a component corresponding to a system voice included in a microphone voice input to a microphone as an echo component. The setting unit sets a step size related to the estimation of the echo component to a different size according to the correlation between the microphone voice and the echo component. This makes it possible to improve the recognition accuracy of voice recognition.

[0014] In one aspect of the above audio processing device, the setting unit sets the step size to a relatively large first size when the correlation between the microphone audio and the echo component is high, and sets the step size to a relatively small second size when the correlation between the microphone audio and the echo component is low.

[0015] In one aspect of the above audio processing device, the device further includes a calculation unit that calculates a correlation coefficient indicating the correlation between the microphone audio and the echo component, and the setting unit sets the step size to the first size when the correlation coefficient is greater than a predetermined threshold, and sets the step size to the second size when the correlation coefficient is equal to or less than the predetermined threshold.

[0016] In one aspect of the above-mentioned audio processing device, the device further includes a suppression unit that suppresses the system audio contained in the microphone audio by subtracting the echo component from the microphone audio, and an attenuation unit that attenuates the microphone audio including the suppressed system audio.

[0017] In one aspect of the above sound processing device, the setting unit sets, as the step size, a variable width of a filter coefficient in a filter used to estimate the echo component.

[0018] In one aspect of the above sound processing device, the setting unit sets the step size related to the estimation of the echo component to one step size among a plurality of step sizes.

[0019] In another preferred embodiment of the present invention, a computer-executed voice processing method includes an estimation step of estimating a component corresponding to a system voice included in a microphone voice input to a microphone as an echo component, and a setting step of setting a step size related to the estimation of the echo component to a different size according to the correlation between the microphone voice and the echo component, thereby improving the recognition accuracy of voice recognition.

[0020] In yet another preferred embodiment of the present invention, a program executed by a computer causes the computer to function as an estimation unit that estimates a component corresponding to a system voice included in a microphone voice input to a microphone as an echo component, and a setting unit that sets a step size related to the estimation of the echo component to a different size according to the correlation between the microphone voice and the echo component, thereby improving the recognition accuracy of voice recognition. The above-mentioned voice processing device can be realized by executing the above-mentioned program on a computer. The above-mentioned program can be stored in a storage medium for use. EXAMPLES

[0021] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings.

[0022] [System Configuration] 1 is a diagram showing a schematic configuration of a voice processing system according to an embodiment. As shown in FIG. 1, the voice processing system 1 includes a voice processing device 100, a microphone 200, a speaker 300, and a voice recognition engine 400.

[0023] The voice processing device 100 acquires processed microphone voice MVZ by performing voice processing such as echo cancellation on the microphone voice MVA input from the microphone 200, and outputs the acquired processed microphone voice MVZ to the voice recognition engine 400.

[0024] The voice recognition engine 400 recognizes the speech content included in the processed microphone voice MVZ by analyzing the processed microphone voice MVZ obtained from the voice processing device 100. The voice recognition engine 400 also outputs a recognition result NK of the speech content included in the processed microphone voice MVZ to the voice processing device 100. The recognition result NK may include, for example, text data or the like corresponding to the speech content included in the processed microphone voice MVZ.

[0025] The voice processing device 100 generates a system voice SV that includes the recognition result NK obtained from the voice recognition engine 400 as the speech content, and outputs the generated voice to the speaker 300.

[0026] [Audio output device] The voice processing device 100 can provide various information to the user through dialogue with the user. The voice processing device 100 may be incorporated, for example, into a navigation device installed in a vehicle and providing route guidance to a set destination. The voice processing device 100 may also be incorporated, for example, into a mobile terminal such as a smartphone carried by a user.

[0027] (Hardware configuration) 2 is a block diagram showing an example of a hardware configuration of a voice processing device according to an embodiment. As shown in FIG. 2, the voice processing device 100 includes an interface (IF) 111, a processor 112, a memory 113, and a recording medium 114.

[0028] The IF 111 inputs and outputs data to and from an external device. For example, the microphone voice MVA and the recognition result NK are input to the voice processing device 100 through the IF 111. In addition, for example, the system voice SV is output to the speaker 300 through the IF 111.

[0029] The processor 112 is a computer such as a CPU (Central Processing Unit) and executes a program prepared in advance to control the entire sound processing device 100. The processor 112 performs, for example, a process of removing the system sound SV included in the microphone sound MVA as an echo component.

[0030] The memory 113 is composed of a read only memory (ROM), a random access memory (RAM), etc. The memory 113 is also used as a working memory while the processor 112 is executing various processes.

[0031] The recording medium 114 is a non-volatile and non-temporary recording medium such as a disk-shaped recording medium or a semiconductor memory, and has a configuration that is detachable from the voice processing device 100. The recording medium 114 records various programs executed by the processor 112. When the voice processing device 100 executes various processes, the programs recorded in the recording medium 114 are loaded into the memory 113 and executed by the processor 112.

[0032] (Functional configuration) 3 is a block diagram showing an example of a functional configuration of a voice processing device according to an embodiment. As shown in FIG. 3, the voice processing device 100 includes a preprocessing unit 11, an adaptive signal processing unit 12, an attenuation processing unit 14, a noise reduction unit 15, a system voice generating unit 16, and a correlation coefficient calculation unit 17.

[0033] The preprocessing unit 11 acquires the microphone audio MVB by performing preprocessing on the microphone audio MVA input to the microphone 200. The preprocessing may include, for example, sampling, normalization, and filtering using a band-pass filter. The preprocessing unit 11 also outputs the microphone audio MVB to the adaptive signal processing unit 12 and the correlation coefficient calculation unit 17.

[0034] The adaptive signal processing unit 12 has a function capable of suppressing the system audio SV included in the microphone audio MVB by performing processing using, for example, an NLMS (Normalized Least Mean Square) algorithm. The adaptive signal processing unit 12 outputs the audio in which the system audio SV included in the microphone audio MVB is suppressed as the microphone audio MVC to the attenuation processing unit 14. The adaptive signal processing unit 12 also performs processing related to estimation of the echo component ECC based on the microphone audio MVC output to the attenuation processing unit 14 and the system audio SV generated by the system audio generating unit 16. The adaptive signal processing unit 12 also has an echo suppression unit 12A and an echo component estimation unit 12B.

[0035] The echo suppression unit 12A has a function as a suppression unit. The echo suppression unit 12A also subtracts the echo component ECC from the microphone audio MVB to obtain the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed. The echo suppression unit 12A also outputs the microphone audio MVC to the echo component estimation unit 12B and the attenuation processing unit 14.

[0036] The echo component estimation unit 12B has a function as an estimation unit. Moreover, the echo component estimation unit 12B performs a process of estimating a component corresponding to the system audio SV included in the microphone audio MVB as an echo component ECC based on the microphone audio MVC output to the attenuation processing unit 14 and the system audio SV generated by the system audio generation unit 16. Moreover, the echo component estimation unit 12B performs, for example, a process using an FIR (Finite Impulse Response) filter as a process related to the estimation of the echo component ECC. Moreover, the echo component estimation unit 12B has a parameter setting unit 12P.

[0037] The parameter setting unit 12P sets parameters for processing related to estimation of the echo component ECC based on the microphone audio MVC and the correlation coefficient CK. Specifically, the parameter setting unit 12P sets, for example, the filter coefficient FC of the FIR filter as a coefficient that changes according to the microphone audio MVC. The details of the correlation coefficient CK will be described later.

[0038] The parameter setting unit 12P has a function as a setting unit. The parameter setting unit 12P sets a step size related to the estimation of the echo component ECC as a size that changes according to the correlation coefficient CK calculated by the correlation coefficient calculation unit 17. The step size can be expressed as, for example, a variable width of a filter coefficient FC in an FIR filter used to estimate the echo component ECC. In such a case, the parameter setting unit 12P can set the variable width of the filter coefficient FC in the FIR filter as the step size related to the estimation of the echo component ECC.

[0039] Specifically, for example, when the correlation coefficient CK is greater than the threshold THK, the parameter setting unit 12P sets the step size related to the estimation of the echo component ECC to a relatively large step size SS1. Also, for example, when the correlation coefficient CK is equal to or less than the threshold THK, the parameter setting unit 12P sets the step size related to the estimation of the echo component ECC to a relatively small step size SS2. According to such processing, the parameter setting unit 12P can set the step sizes SS1 and SS2 as sizes that satisfy the relationship SS1>SS2. Also, according to the above-mentioned processing, when the correlation coefficient CK is greater than the threshold THK, the echo component estimation unit 12B can estimate the echo component ECC while sharply changing the filter coefficient FC. Also, according to the above-mentioned processing, when the correlation coefficient CK is equal to or less than the threshold THK, the echo component estimation unit 12B can estimate the echo component ECC while finely changing the filter coefficient FC according to the microphone voice MVC. In this embodiment, for example, when the correlation coefficient CK is calculated as a value that falls within the range of 0 or more and 1 or less, the threshold value THK is desirably set to 0.8.

[0040] According to the above-described process, the echo component estimation unit 12B can estimate the echo component ECC by applying an FIR filter in which a filter coefficient FC is set to the system sound SV.

[0041] According to this embodiment, the parameter setting unit 12P may set the step size related to the estimation of the echo component ECC to any one of three or more step sizes. Specifically, for example, when the correlation coefficient CK is greater than a threshold value THK and the magnitude of the external noise included in the microphone voice MVC is equal to or less than a predetermined level, the parameter setting unit 12P may set the step size related to the estimation of the echo component ECC to a step size SSA greater than the step size SS1. Also, for example, when the correlation coefficient CK is equal to or less than a threshold value THK and the magnitude of the external noise included in the microphone voice MVC is greater than a predetermined level, the parameter setting unit 12P may set the step size related to the estimation of the echo component ECC to a step size SSB smaller than the step size SS2. Note that the above-mentioned external noise may include sounds unnecessary for voice recognition other than the system voice SV, such as music and environmental sounds. According to this embodiment, the parameter setting unit 12P may set the step size related to the estimation of the echo component ECC to one of a plurality of step sizes.

[0042] The attenuation processing unit 14 has a function as an attenuation unit. The attenuation processing unit 14 also acquires a microphone audio MVD by performing a process of attenuating the microphone audio MVC. Specifically, the attenuation processing unit 14 acquires the microphone audio MVD by performing a process of applying a gain GN set as a value larger than 0 and smaller than 1 to the microphone audio MVC. The attenuation processing unit 14 also outputs the microphone audio MVD to the noise removal unit 15.

[0043] The noise removal unit 15 performs processing related to noise removal on the microphone voice MVD to obtain a processed microphone voice MVZ, and outputs the obtained processed microphone voice MVZ to the voice recognition engine 400.

[0044] The system voice generating unit 16 generates a system voice SV including the recognition result NK obtained from the voice recognition engine 400 as the speech content, and outputs the generated voice to the echo component estimating unit 12B and the speaker 300.

[0045] The correlation coefficient calculation unit 17 has a function as a calculation unit. The correlation coefficient calculation unit 17 calculates a value in the range of 0 to 1 as a correlation coefficient CK indicating the correlation between the microphone audio MVB and the echo component ECC. Specifically, for example, when the correlation between the microphone audio MVB and the echo component ECC is high, the correlation coefficient calculation unit 17 calculates a value of 1 or close to 1 as the correlation coefficient CK. For example, when the correlation between the microphone audio MVB and the echo component ECC is low, the correlation coefficient calculation unit 17 calculates a value of 0 or close to 0 as the correlation coefficient CK. The correlation coefficient calculation unit 17 outputs the correlation coefficient CK to the echo component estimation unit 12B. According to such a process, the parameter setting unit 12P can set the step size related to the estimation of the echo component ECC to a different size depending on the correlation between the microphone audio MVB and the echo component ECC.

[0046] [Examples of voice processing] Next, a specific example of the voice processing in this embodiment will be described. FIG. 4 is a diagram showing an example of a temporal change in the correlation coefficient used in the voice processing according to the embodiment. In this specific example, the period during which the voice processing device 100 continues to speak is represented as a period PSN, and the period during which the user speaks in the period PSN is represented as a period PSU. In other words, in this specific example, a case will be described in which the system voice SV continues to be output from the speaker 300 in the period PSN, and the user voice UV corresponding to the user's speech is generated in the period PSU included in the period PSN. In addition, in this specific example, the start time of the period PSN is represented as time T1, the start time of the period PSU is represented as time T2, the end time of the period PSU is represented as time T3, and the end time of the period PSN is represented as time T4.

[0047] During a period PA from time T1 to time T2, a sound including the system sound SV but not including the user sound UV is input as the microphone sound MVA to the sound processing device 100. Accordingly, the correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK during the period PA (see FIG. 4). In addition, the parameter setting unit 12P sets the step size related to the estimation of the echo component ECC to the step size SS1 according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17.

[0048] According to the above-described processing, the adaptive signal processor 12 can quickly suppress the system sound SV included in the microphone sound MVB during the period PA.

[0049] In a period PSU from time T2 to time T3, a mixed voice including the system voice SV and the user voice UV is input as the microphone voice MVA to the voice processing device 100. Accordingly, the correlation coefficient calculation unit 17 calculates a value equal to or smaller than a threshold value THK as the correlation coefficient CK in the period PSU (see FIG. 4). In addition, the parameter setting unit 12P sets the step size related to the estimation of the echo component ECC to a step size SS2 according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17.

[0050] According to the above-described processing, the adaptive signal processing unit 12 can suppress the system audio SV included in the microphone audio MVB during the period PSU while preventing deterioration of the user audio UV included in the microphone audio MVB.

[0051] During a period PB from time T3 to time T4, a voice including the system voice SV but not including the user voice UV is input as the microphone voice MVA to the voice processing device 100. Accordingly, the correlation coefficient calculation unit 17 calculates a value larger than the threshold value THK as the correlation coefficient CK during the period PB (see FIG. 4). In addition, the parameter setting unit 12P sets the step size related to the estimation of the echo component ECC to the step size SS1 according to the calculation result of the correlation coefficient CK by the correlation coefficient calculation unit 17.

[0052] According to the above-described processing, the adaptive signal processor 12 can quickly suppress the system sound SV included in the microphone sound MVB during the period PB.

[0053] When a voice that includes the user voice UV but does not include the system voice SV is input as the microphone voice MVA, the adaptive signal processor 12 may perform the same processing as that performed in the above-mentioned period PSU.

[0054] [Processing flow] Next, a description will be given of a flow of processing performed by the voice processing device 100. Fig. 5 is a flowchart showing an example of processing performed by the voice processing device. It should be noted that the voice processing device 100 repeats the processing of Fig. 5 during an operation period corresponding to a period from when the power is turned on to when it is turned off, for example.

[0055] First, the audio processing device 100 acquires the microphone audio MVC in which the system audio SV included in the microphone audio MVB is suppressed by subtracting the echo component ECC from the microphone audio MVB (step S11).

[0056] Next, the sound processing device 100 calculates a correlation coefficient CK indicating the correlation between the microphone sound MVB that is the processing target of step S11 and the echo component ECC used in the processing of step S11 (step S12).

[0057] Next, the sound processing device 100 determines whether the correlation coefficient CK calculated in step S12 is greater than a threshold value THK (step S13).

[0058] When the correlation coefficient CK is greater than the threshold THK (step S13: YES), the audio processing device 100 sets the step size for estimating the echo component ECC to a relatively large step size SS1 (step S14).When the correlation coefficient CK is equal to or smaller than the threshold THK (step S13: NO), the audio processing device 100 sets the step size for estimating the echo component ECC to a relatively small step size SS2 (step S15).

[0059] The voice processing device 100 performs processing for estimating the echo component ECC using the step size SS1 set in step S14 or the step size SS2 set in step S15 based on the microphone voice MVC obtained in step S11 and the system voice SV generated according to the recognition result NK (step S16).

[0060] As described above, according to this embodiment, even if the user and the voice processing device 100 speak at the same time, a voice that can identify the contents of the user's speech can be output to the voice recognition engine 400. Therefore, according to this embodiment, the recognition accuracy of the voice recognition can be improved.

[0061] As described above, according to this embodiment, even if the voice (system voice) of the voice processing device 100 is input to the system through the microphone 200 during a period when the user is not speaking, the speech content included in the voice can be prevented from being recognized by the voice recognition engine 400. Therefore, according to this embodiment, the recognition accuracy of voice recognition can be improved.

[0062] In the above-described embodiment, the program can be stored using various types of non-transitory computer readable media and supplied to a control unit, which is a computer. The non-transitory computer readable media includes various types of tangible storage media. Examples of the non-transitory computer readable media include magnetic storage media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical storage media (e.g., magneto-optical disks), CD-ROM (Read Only Memory), CD-R, CD-R / W, and semiconductor memory (e.g., mask ROM, PROM (Programmable ROM), EPROM (Erasable PROM), flash ROM, and RAM (Random Access Memory).

[0063] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above-mentioned embodiments. Various modifications that can be understood by a person skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. In other words, the present invention naturally includes various modifications and corrections that a person skilled in the art could make in accordance with the entire disclosure, including the claims, and the technical ideas. In addition, the disclosures of the above cited patent documents and the like are incorporated herein by reference. [Explanation of symbols]

[0064] 11 Pretreatment section 12 Adaptive signal processing section 12A Echo suppression section 12B Echo component estimation section 12P Parameter setting section 14 Attenuation processing section 15 Noise Reduction Section 16 System audio generation section 17 Correlation coefficient calculation section

Claims

1. an estimation unit that estimates a component corresponding to a system sound included in a microphone sound input to a microphone as an echo component; a setting unit that sets a step size related to the estimation of the echo component to a different size depending on a correlation between the microphone voice and the echo component; 13. An audio processing device comprising:

2. 2. The audio processing device according to claim 1, wherein the setting unit sets the step size to a relatively large first size when the correlation between the microphone audio and the echo component is high, and sets the step size to a relatively small second size when the correlation between the microphone audio and the echo component is low.

3. A calculation unit is further provided for calculating a correlation coefficient indicating a correlation between the microphone voice and the echo component, 3. The audio processing device according to claim 2, wherein the setting unit sets the step size to the first size when the correlation coefficient is greater than a predetermined threshold, and sets the step size to the second size when the correlation coefficient is equal to or less than the predetermined threshold.

4. a suppression unit that suppresses the system sound included in the microphone sound by subtracting the echo component from the microphone sound; The audio processing device according to claim 1 , further comprising: an attenuation unit that attenuates the microphone audio including the suppressed system audio.

5. The audio processing device according to claim 1 , wherein the setting unit sets, as the step size, a variable width of a filter coefficient in a filter used to estimate the echo component.

6. The audio processing device according to claim 1 , wherein the setting unit sets the step size related to the estimation of the echo component to one of a plurality of step sizes.

7. 1. A computer implemented method for audio processing, comprising: an estimation step of estimating a component corresponding to a system sound included in a microphone sound input to a microphone as an echo component; a setting step of setting a step size related to the estimation of the echo component to a different size according to a correlation between the microphone voice and the echo component; 13. A method for processing audio comprising the steps of:

8. A program executed by a computer, an estimation unit that estimates a component corresponding to a system sound included in a microphone sound input to a microphone as an echo component; A program that causes the computer to function as a setting unit that sets a step size related to the estimation of the echo component to a different size depending on the correlation between the microphone voice and the echo component.

9. A storage medium storing the program according to claim 8.

Citation Information

Patent Citations

  • Voice recognition system and voice recognizer

    JP2009109536A