Voice processing device, voice processing method, voice processing program, and voice processing system

The audio processing apparatus addresses the issue of false detection in audio recognition by determining noise levels from a speaker and outputting appropriate signals to the audio recognition unit, thereby enhancing recognition accuracy.

JP7685813B2Active Publication Date: 2025-05-30PANASONIC AUTOMOTIVE SYST CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022015324
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-03
Publication Date
2025-05-30
Estimated Expiration
2042-02-03

AI Technical Summary

Technical Problem

Existing audio processing systems face challenges in suppressing false detection of audio recognition due to noise components like residual echo that cannot be completely removed by echo cancellers.

Method used

An audio processing apparatus with an audio acquisition unit, a determination unit, an audio processing unit, and a switching unit. The apparatus acquires audio signals from a microphone, determines if the reference signal from a speaker exceeds a threshold, and outputs a removal signal or a replacement signal (comfort noise or mute) to the audio recognition unit based on this determination.

Benefits of technology

The solution effectively suppresses false detection in audio recognition by managing noise levels and preventing recognition during high noise conditions, thereby improving the accuracy of voice commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007685813000001
    Figure 0007685813000001
  • Figure 0007685813000002
    Figure 0007685813000002
  • Figure 0007685813000003
    Figure 0007685813000003
Patent Text Reader

Abstract

To suppress erroneous detection of voice recognition.SOLUTION: A voice processing unit 10 includes: a voice acquisition part 20; a determination part 22; a voice processing part 24; and a switching part 26. The voice acquisition part 20 acquires a voice signal from a microphone MC for collecting voice in space. The determination part 22 determines whether or not a level of a reference signal which is a reproduction signal reproduced from a speaker SP for emitting sound in the space is equal to or higher than a threshold value. The voice processing part 24 outputs a removal signal obtained by removing a voice component of the reference signal from the voice signal, to a voice recognition part 40 as an output signal. . When the level of the reference signal is determined to be equal to or higher than the threshold value, the switching part 26 outputs a replacement signal which is at least one of comfort noise and a mute signal, to the voice recognition part 40 as the output signal instead of the removal signal.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an audio processing apparatus, an audio processing method, an audio processing program, and an audio processing system.

Background Art

[0002] There is known an audio processing system that processes an audio recognition command based on audio uttered by a speaker. For example, audio picked up by a microphone is recognized by a first audio recognition unit, and audio output from a speaker is recognized by a second audio recognition unit. And when an audio recognition command is included in the audio recognized by the second audio recognition unit, a configuration is disclosed in which the recognition by the first audio recognition unit is stopped (see, for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, in the prior art, when noise components such as residual echo components that cannot be completely removed by an echo canceller are included in the audio picked up by the microphone, false detection of audio recognition may occur. That is, in the prior art, it may be difficult to suppress false detection of audio recognition.

[0005] An object of the present disclosure is to provide an audio processing apparatus, an audio processing method, an audio processing program, and an audio processing system that can suppress false detection of audio recognition.

Means for Solving the Problems

[0006] An audio processing apparatus according to one aspect of the present disclosure includes an audio acquisition unit, a determination unit, an audio processing unit, and a switching unit. The audio acquisition unit acquires an audio signal from a microphone that picks up the audio in the space. The determination unit determines whether the level of a reference signal, which is a reproduction signal reproduced from a speaker that outputs sound in the space, is equal to or higher than a threshold value. The audio processing unit outputs, as an output signal, a removal signal obtained by removing the audio component of the reference signal from the audio signal to an audio recognition unit. When it is determined that the level of the reference signal is equal to or higher than the threshold value, the switching unit outputs, as the output signal, a replacement signal that is at least one of comfort noise and a mute signal to the audio recognition unit instead of the removal signal.

Effect of the Invention

[0007] According to the present disclosure, false detection in audio recognition can be suppressed.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Mode for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings as appropriate. However, overly detailed descriptions may be omitted. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter described in the claims.

[0010] FIG. 1 is a diagram showing an example of the schematic configuration of the audio processing system 1 of the present embodiment.

[0011] The voice processing system 1 is a system for recognizing voices in a space. In this embodiment, a case where the space is the space inside the vehicle cabin of the vehicle 2 will be described as an example. Also, in this embodiment, a form in which the voice processing system 1 is mounted on the vehicle 2 will be described as an example. Note that the space is not limited to the inside of the vehicle cabin of the vehicle 2.

[0012] The voice processing system 1 includes a microphone MC, a speaker SP, a voice processing device 10, a sound source device 30, a voice recognition unit 40, an electronic device 50, and a display 60. The microphone MC, the speaker SP, the voice recognition unit 40, and the display 60 are communicably connected to the voice processing device 10. The voice processing system 1 may have a configuration including at least the microphone MC, the speaker SP, the voice processing device 10, and the voice recognition unit 40.

[0013] The microphone MC picks up the voice in the space. In this embodiment, the microphone MC picks up at least the voice in the space inside the vehicle cabin of the vehicle 2. In this embodiment, a form in which the microphone MC is provided near the driver's seat, which is the seat of the driver hm1 of the vehicle 2, will be described as an example. For this reason, in this embodiment, the microphone MC picks up the voice including at least the voice component spoken by the driver hm1.

[0014] The vehicle 2 may have a configuration in which a plurality of microphones MC are provided. In this case, it is preferable that these plurality of microphones MC are arranged at different positions in the vehicle cabin of the vehicle 2. Specifically, for example, microphones MC may be respectively arranged near the seats of the driver hm1, the passengers hm2, hm3, and hm3 of the vehicle 2. In this embodiment, a form in which one microphone MC is provided in the vehicle 2 will be described as an example.

[0015] The microphone MC may be either a directional microphone or an omnidirectional microphone. The microphone MC may be either a small MEMS (Micro Electro Mechanical Systems) microphone or an ECM (Electret Condenser Microphone). The microphone MC may also be a microphone capable of beamforming. For example, the microphone MC may be a microphone array that has directivity in a specific direction and can pick up sound in the direction of pointing.

[0016] The microphone MC outputs the audio signal of the picked-up sound to the audio processing device 10. The audio processing device 10 is provided in association with the microphone MC. Therefore, when the audio processing system 1 is configured with a plurality of microphones MC, the audio processing system 1 may be configured with a plurality of audio processing devices 10 corresponding to each of the plurality of microphones MC. In the present embodiment, a form in which the audio processing system 1 includes one microphone MC and one audio processing device 10 communicably connected to the microphone MC will be described as an example.

[0017] The speaker SP outputs sound to the same space as the space where the sound is picked up by the microphone MC. In the present embodiment, the speaker SP outputs sound to at least the space inside the vehicle compartment of the vehicle 2.

[0018] In the present embodiment, a form in which four speakers SP, namely speakers SP1 to SP4, are arranged inside the vehicle compartment of the vehicle 2 will be described as an example. Note that the audio processing system 1 may be configured with at least one speaker SP, and the number and arrangement position of the speakers SP are not limited. In the present embodiment, a form in which speakers SP1, SP2, SP3, and SP4 are arranged near the seats of the driver hm1, the passengers hm2, hm3, and hm3 inside the vehicle compartment of the vehicle 2 will be described as an example. When these speakers SP1 to SP4 are collectively described, they will be simply referred to as the speaker SP.

[0019] Speaker SP is electrically connected to the sound source device 30. Speaker SP outputs the sound represented by the reproduction signal received from the sound source device 30. The reproduction signal is a signal output from the sound source device 30 to Speaker SP. Speaker SP outputs the sound corresponding to the reproduction signal received from the sound source device 30. Specifically, Speaker SP outputs the sound with a volume corresponding to the level of the reproduction signal received from the sound source device 30. That is, in the present embodiment, the level means the level of the signal, and specifically, it means the magnitude of the sound represented by the signal.

[0020] The sound source device 30 is, for example, a radio receiver, a television broadcast device, an audio device, etc. The radio receiver receives a radio broadcast signal, generates a reproduction signal from the received radio broadcast signal, and outputs it to Speaker SP. In this case, the reproduction signal is, for example, a radio audio signal of radio voice. The television broadcast device receives a television broadcast signal, generates a reproduction signal from the received television broadcast signal, and outputs it to Speaker SP. In this case, the reproduction signal is, for example, a television audio signal of television voice. The audio device outputs a reproduction signal such as an audio signal recorded in a memory or the like to Speaker SP. In this case, the reproduction signal is, for example, an audio signal, etc.

[0021] In the present embodiment, the sound source device 30 generates a 4-channel reproduction signal in order to use four speakers SP (Speaker SP1 to Speaker SP4), and outputs it to each of the four speakers SP as a reference signal. Specifically, the sound source device 30 outputs reference signal 1, which is a reproduction signal, to Speaker SP1, outputs reference signal 2, which is a reproduction signal, to Speaker SP2, outputs reference signal 3, which is a reproduction signal, to Speaker SP3, and outputs reference signal 4, which is a reproduction signal, to Speaker SP4. These reference signals 1 to reference signals 4 are reproduction signals output to each of the plurality of speakers SP. When collectively explaining reference signals 1 to reference signals 4, they are simply referred to as reference signals.

[0022] The voice processing device 10 outputs an output signal based on the voice signal received from the microphone MC and the reference signal which is the reproduction signal reproduced from the speaker SP, to the voice recognition unit 40. Details of the voice processing device 10 will be described later.

[0023] The voice recognition unit 40 recognizes the voice represented by the output signal received from the voice processing device 10, and outputs a signal representing the voice recognition result to the electronic device 50. For example, the voice recognition unit 40 recognizes the voice command represented by the output signal and outputs it to the electronic device 50. The voice command is a signal for causing the electronic device 50 to execute various processes. The voice command may be referred to as a voice recognition command, keyword, wake-up word, etc.

[0024] The electronic device 50 executes a process according to the voice command which is a signal representing the voice recognition result received from the voice recognition unit 40. For example, the electronic device 50 executes a process of opening and closing a window, a process related to driving the vehicle 2, a process of changing the temperature of the air conditioner, a process of changing the volume of the audio device, etc. based on the voice command. The electronic device 50 is, for example, a car navigation device, an air conditioner, a panel meter, a television, a mobile terminal, a driving device for driving each part of the vehicle 2, etc.

[0025] The display 60 is a display device for displaying various information. The display 60 is, for example, various displays provided in the vehicle 2, a head-up display, a display of a car navigation system, a multi-information display provided in the meter of the vehicle 2, a center display capable of receiving audio operations, etc. In the present embodiment, information is displayed on the display 60 by the voice processing device 10 described later. Note that the display 60 may function as an example of the electronic device 50.

[0026] The voice processing device 10 will be described in detail. First, an example of the hardware configuration of the voice processing device 10 will be described.

[0027] Figure 2 is a hardware configuration diagram of an example of the voice processing device 10.

[0028] In the voice processing device 10, a CPU (Central Processing Unit) 11A, a ROM (Read Only Memory) 11B, a RAM 11C, an I / F 11D, etc. are mutually connected by a bus 11E, and it has a hardware configuration using a normal computer.

[0029] The CPU 11A is an arithmetic unit that controls the voice processing device 10 of the present embodiment. The ROM 11B stores programs and the like for realizing various processes by the CPU 11A. The RAM 11C stores data necessary for various processes by the CPU 11A. The I / F 11D is an interface for transmitting and receiving data.

[0030] The program for executing the information processing executed by the voice processing device 10 of the present embodiment is provided by being pre - incorporated in the ROM 11B or the like. Note that the program executed by the voice processing device 10 of the present embodiment may be configured to be recorded and provided on a computer - readable recording medium such as a CD - ROM, a flexible disk (FD), a CD - R, a DVD (Digital Versatile Disk) in a form installable or executable on the voice processing device 10.

[0031] Next, the configuration of the voice processing device 10 will be described in detail.

[0032] Figure 3 is a block diagram showing an example of the configuration of the voice processing device 10. For the purpose of explanation in Figure 3, in addition to the voice processing device 10, a microphone MC, a sound source device 30, a voice recognition unit 40, an electronic device 50, and a display 60 are shown.

[0033] The voice processing device 10 includes a voice acquisition unit 20, a determination unit 22, a voice processing unit 24, a switching unit 26, a generation unit 28, and an output control unit 29.

[0034] Part or all of the voice acquisition unit 20, determination unit 22, voice processing unit 24, switching unit 26, generation unit 28, and output control unit 29 may be realized, for example, by causing a processing device such as the CPU 11A to execute a program, that is, by software, or by hardware such as an IC (Integrated Circuit), or by a combination of software and hardware. Further, at least one of the voice acquisition unit 20, determination unit 22, voice processing unit 24, switching unit 26, generation unit 28, and output control unit 29 may be configured to be mounted on an external information processing device communicably connected to the voice processing device 10 via a network or the like.

[0035] The voice acquisition unit 20 acquires a voice signal from the microphone MC. The voice acquisition unit 20 outputs the acquired voice signal to the voice processing unit 24.

[0036] The determination unit 22 determines whether or not the level of a reference signal, which is a reproduction signal reproduced from the speaker SP, is equal to or greater than a threshold value. The level of the reference signal represents the loudness of the sound represented by the reproduction signal, which is the reference signal. As described above, the speaker SP outputs a sound with a volume corresponding to the level of the reproduction signal received from the sound source device 30. Therefore, the louder the level of the reference signal, which is the reproduction signal, the louder the volume of the sound output from the speaker SP.

[0037] The threshold value may be set in advance to be equal to or less than the level of the reproduction signal when the distortion starts to occur in the sound output from the speaker SP in response to the reproduction signal as the level of the reproduction signal is gradually increased, and to be a value close to that level. Further, the threshold value may be a value that coincides with the level of the reproduction signal when the distortion starts to occur in the sound output from the speaker SP in response to the reproduction signal as the level of the reproduction signal is gradually increased. The distortion of the sound output from the speaker SP may also be referred to as sound cracking.

[0038] For example, the determination unit 22 determines a threshold value that satisfies the above conditions for each of the plurality of speakers SP1 to SP4.

[0039] Then, the determination unit 22 determines whether at least one of the levels of the reference signals 1 to 4 received from each of the plurality of speakers SP1 to SP4 is equal to or higher than the threshold value corresponding to each of the speakers SP1 to SP4.

[0040] Further, the determination unit 22 may set, as a threshold value common to the plurality of speakers SP1 to SP4, the minimum value, average value, or maximum value of the threshold values that satisfy the above conditions for each of the plurality of speakers SP1 to SP4. Then, the determination unit 22 may determine whether at least one of the levels of the reference signals 1 to 4 received from each of the plurality of speakers SP1 to SP4 is equal to or higher than the threshold value set as the common threshold value.

[0041] In the present embodiment, as an example, a mode will be described in which the determination unit 22 determines whether at least one of the levels of the reference signals 1 to 4 received from each of the plurality of speakers SP1 to SP4 is equal to or higher than the threshold value corresponding to each of the speakers SP1 to SP4.

[0042] Note that the threshold value corresponding to each of the plurality of speakers SP1 to SP4 may be stored in advance in the memory or the like of the determination unit 22. Further, the threshold value corresponding to each of the plurality of speakers SP1 to SP4 may be appropriately changed within a range that satisfies the above conditions by an operation instruction or the like by the user according to the type and installation position of the speaker SP provided in the audio processing system 1.

[0043] The audio processing unit 24 generates a removal signal obtained by removing the audio component of the reference signal from the audio signal received from the audio acquisition unit 20.

[0044] The audio processing unit 24 removes the audio component of the reference signal, which is a reproduction signal, included in the audio signal received from the audio acquisition unit 20. The audio processing unit 24 may remove the audio component of the reference signal included in the audio signal by using at least one of a known echo canceller and a crosstalk canceller.

[0045] For example, the voice processing unit 24 includes an adaptive filter F, an adaptive filter control unit 24A, and a subtraction unit 24B.

[0046] The adaptive filter F is a filter having a function of changing the characteristics of a reference signal. In the present embodiment, the adaptive filter F includes adaptive filters F1 to F4. The number of adaptive filters F is appropriately set based on the number of input reference signals and the like.

[0047] The adaptive filter control unit 24A sets the filter coefficients of each of the adaptive filters F1 to F4 by a known method according to the cancellation signal output from the subtraction unit 24B. The adaptive filters F1 to F4 output, as subtraction signals, the passed signals based on each of the received reference signals 1 to 4 and the set filter coefficients to the subtraction unit 24B. For this reason, a subtraction signal, which is a signal obtained by adding together the passed signals based on each of the reference signals 1 to 4 and the set filter coefficients, output from each of the adaptive filters F1 to F4 is output to the subtraction unit 24B.

[0048] The subtraction unit 24B executes a cancellation process of removing the voice component of the reference signal from the voice signal by subtracting the subtraction signal from the voice signal received from the voice acquisition unit 20. The subtraction unit 24B outputs the cancellation signal obtained by the cancellation process, that is, the cancellation signal obtained by removing the voice component of the reference signal from the voice signal, to the adaptive filter control unit 24A and the switching unit 26.

[0049] When it is determined that the level of the reference signal is equal to or higher than the threshold value, the switching unit 26 outputs, as an output signal, a replacement signal that is at least one of a comfort noise and a mute signal to the voice recognition unit 40 instead of the cancellation signal received from the voice processing unit 24.

[0050] Specifically, when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold value, the switching unit 26 switches to output the replacement signal received from the generation unit 28 to the voice recognition unit 40 instead of the cancellation signal received from the voice processing unit 24.

[0051] The generation unit 28 generates a replacement signal that is at least one of a comfort noise and a mute signal, and outputs it to the switching unit 26. The mute signal is a signal with a sound level of "0". In other words, the mute signal is a signal representing a silent state, a muted state, or no signal (MUTE).

[0052] When the generation unit 28 generates comfort noise as a replacement signal, it is preferable to generate comfort noise at a level corresponding to the noise level included in the voice signal at the timing immediately before being determined to be equal to or higher than the threshold value by the determination unit 22. For example, the voice acquisition unit 20 outputs the voice signal acquired from the microphone MC to the voice processing unit 24 and the generation unit 28. The generation unit 28 identifies, by a known method, the noise level included in the voice signal at the timing immediately before being determined to be equal to or higher than the threshold value by the determination unit 22 in the voice signal received from the voice acquisition unit 20. Then, the generation unit 28 generates comfort noise at a level corresponding to the identified noise level. For example, the generation unit 28 generates comfort noise at the same level as the identified noise level, that is, comfort noise representing the same volume level.

[0053] By generating comfort noise at a level corresponding to the noise level included in the voice signal at the timing immediately before being determined to be equal to or higher than the threshold value as a replacement signal, the generation unit 28 suppresses a rapid change in the level of the output signal output to the voice recognition unit 40. For example, when the acoustic environment of the space varies according to a change in the driving environment of the vehicle 2 or the like, comfort noise at a level corresponding to the variation in the acoustic environment of the space is output to the voice recognition unit 40 as a replacement signal. Therefore, when the output signal output to the voice recognition unit 40 switches from the replacement signal to the removal signal or from the removal signal to the replacement signal, a rapid change in the level of the output signal is suppressed. Therefore, it is possible to suppress a decrease in the voice recognition performance of the voice recognition unit 40 due to a rapid change in the level of the output signal.

[0054] Further, the generation unit 28 may generate a replacement signal including both comfort noise and a mute signal, and output it to the switching unit 26. For example, the generation unit 28 generates a replacement signal in which comfort noise and the mute signal are alternately arranged. In this case, it is preferable that the generation unit 28 generates an output signal with its level adjusted so that the level when the comfort noise and the mute signal switch gradually changes.

[0055] Note that the generation unit 28 may always generate a replacement signal, but preferably generates a replacement signal and outputs it to the switching unit 26 when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold. Then, when the determination unit 22 determines that the level of the reference signal is less than the threshold, the generation unit 28 may stop the generation process of the replacement signal.

[0056] When the determination unit 22 determines that the level of the reference signal is less than the threshold, by stopping the generation process of the replacement signal by the generation unit 28, it is possible to reduce the processing calculation amount of the audio processing apparatus 10.

[0057] When the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold, the switching unit 26 outputs the replacement signal received from the generation unit 28 as an output signal to the speech recognition unit 40 instead of the removal signal received from the audio processing unit 24. Therefore, when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold, a replacement signal is output to the speech recognition unit 40 instead of the removal signal.

[0058] Note that the switching unit 26 may output the replacement signal as an output signal to the speech recognition unit 40 instead of the removal signal during the period when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold. Then, during the period when the determination unit 22 determines that the level of the reference signal is less than the threshold, the switching unit 26 may output the removal signal received from the audio processing unit 24 as an output signal to the speech recognition unit 40.

[0059] In this case, during the period when the level of the reference signal is equal to or higher than the threshold value, the replacement signal is output as the output signal to the voice recognition unit 40. Also, during the period when the level of the reference signal is less than the threshold value, the removal signal is output as the output signal to the voice recognition unit 40.

[0060] Further, when it is determined that the level of the reference signal is equal to or higher than the threshold value, the switching unit 26 may output the replacement signal as the output signal instead of the removal signal, and continuously output it to the voice recognition unit 40 for a predetermined first time.

[0061] The first time may be determined in advance. For example, for the first time, when the output signal output to the voice recognition unit 40 repeatedly switches between the removal signal and the replacement signal in a short time, resulting in a performance degradation of the voice recognition unit 40, a time longer than the continuous output time of the replacement signal to the voice recognition unit 40 may be determined. Also, for example, for the first time, a value that is equal to or longer than the average utterance period required for uttering one voice command and less than the average utterance period when two voice commands are uttered continuously may be determined. Further, the first time may be appropriately changeable according to an operation instruction by the user or the like.

[0062] In this case, the replacement signal is output as the output signal to the voice recognition unit 40 continuously for at least the first time from the timing when the level of the reference signal becomes equal to or higher than the threshold value. Then, after the elapse of the first time, the removal signal is output as the output signal to the voice recognition unit 40.

[0063] Further, when it is determined by the determination unit 22 that the level of the reference signal continues to be equal to or higher than the threshold value for a predetermined second time or longer, the switching unit 26 may output the replacement signal as the output signal instead of the removal signal to the voice recognition unit 40.

[0064] The second time may be determined in advance. For example, for the second time, when the output signal output to the voice recognition unit 40 repeatedly switches between a removal signal and a replacement signal in a short time, resulting in a performance degradation of the voice recognition unit 40, a time longer than the continuous output time of the removal signal or the replacement signal to the voice recognition unit 40 may be determined. Also, for example, for the second time, a value that is equal to or longer than the average utterance period required for uttering one voice command and less than the average utterance period when two voice commands are uttered continuously may be determined. Further, the second time may be appropriately changeable according to an operation instruction by the user or the like.

[0065] In this case, when the state where the level of the reference signal is equal to or higher than the threshold continues for the second time, the replacement signal is output to the voice recognition unit 40 as the output signal. And when the level of the reference signal is less than the threshold or the continuous time of the state where the level is equal to or higher than the threshold is less than the second time, the removal signal is output to the voice recognition unit 40 as the output signal.

[0066] Note that the voice processing unit 24 may always perform a removal process of removing the voice component of the reference signal from the voice signal, but when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold, the removal process may be stopped. For example, when the determination unit 22 determines that the level of the reference signal is equal to or higher than the threshold, the determination unit 22 controls the voice processing unit 24 to stop the removal process.

[0067] When it is determined that the level of the reference signal is equal to or higher than the threshold, by stopping the removal process by the voice processing unit 24, it is possible to reduce the processing calculation amount of the voice processing apparatus 10.

[0068] When it is determined that the level of the reference signal is equal to or higher than the threshold, the output control unit 29 outputs information indicating that voice recognition is stopped. The output control unit 29 outputs, for example, information indicating that voice recognition is stopped to the display 60.

[0069] As described above, when the level of the reference signal is equal to or higher than the threshold value, a replacement signal is output as an output signal to the voice recognition unit 40. Since the replacement signal is at least one of the comfort noise and the mute signal, the voice recognition unit 40 does not perform voice recognition during the period when the replacement signal is received. For this reason, for example, in a situation where a sound with a volume corresponding to a reproduction signal having a level equal to or higher than the threshold value is output by the speaker SP in the space inside the passenger compartment of the vehicle 2, even if the driver hm1 or the like utters a voice command or the like, the voice recognition unit 40 does not perform voice recognition. Therefore, when it is determined that the level of the reference signal, which is the reproduction signal, is equal to or higher than the threshold value, the output control unit 29 outputs information indicating that the voice recognition is stopped, so that the status of the voice recognition of the voice recognition unit 40 can be easily presented to the user.

[0070] Note that the output target of the information by the output control unit 29 is not limited to the display 60. For example, the output control unit 29 may transmit information indicating that the voice recognition is stopped to an information processing device such as a mobile terminal managed by the pre-registered driver hm1. Further, the output control unit 29 may output information indicating that the voice recognition is stopped from the speaker SP. In this case, the level of the reproduction signal of the information indicating that the voice recognition is stopped may be set to a level lower than the above threshold value.

[0071] Next, an example of the flow of information processing executed by the voice processing apparatus 10 of the present embodiment will be described.

[0072] FIG. 4 is a flowchart showing an example of the flow of information processing executed by the voice processing apparatus 10 of the present embodiment.

[0073] The voice acquisition unit 20 acquires a voice signal from the microphone MC (step S100).

[0074] The determination unit 22 determines whether or not the level of the reference signal, which is the reproduction signal reproduced from the speaker SP, is equal to or higher than the threshold value (step S102). When it is determined that the level of the reference signal is equal to or higher than the threshold value (step S102: Yes), the process proceeds to step S104.

[0075] In step S104, the determination unit 22 controls the audio processing unit 24 to stop the removal process. By the process of step S104, the audio processing unit 24 stops the removal process.

[0076] The generation unit 28 generates a replacement signal that is at least one of comfort noise and a mute signal, and outputs it to the switching unit 26 (step S106).

[0077] The switching unit 26 outputs the replacement signal generated by the generation unit 28 to the speech recognition unit 40 as an output signal (step S108). Since the replacement signal is at least one of comfort noise and a mute signal, the replacement signal does not include a voice command. Therefore, during the period when the replacement signal is being received, the speech recognition unit 40 is in a state where it does not recognize voice commands.

[0078] The output control unit 29 outputs information indicating that speech recognition is stopped to the display 60 (step S110).

[0079] Next, the voice processing device 10 determines whether to end the process (step S112). For example, the voice processing device 10 makes the determination in step S112 by determining whether the power supply to the voice processing device 10 has been instructed to be cut off by an operation instruction from the user or the like. If an affirmative determination is made in step S112 (step S112: Yes), the voice processing device 10 ends this routine. If the voice processing device 10 makes a negative determination in step S112 (step S112: No), the process returns to step S100 above.

[0080] On the other hand, in step S102 above, if it is determined that the level of the reference signal, which is the reproduction signal reproduced from the speaker SP, is less than the threshold value (step S102: No), the process proceeds to step S114.

[0081] In step S114, the audio processing unit 24 executes a removal process to generate a removal signal obtained by removing the audio component of the reference signal from the audio signal received from the audio acquisition unit 20. If the removal process by the audio processing unit 24 has been stopped by the process of step S104, the determination unit 22 controls the audio processing unit 24 to cancel the stop of the removal process, and then the audio processing unit 24 may execute the removal process of step S114.

[0082] The switching unit 26 outputs the removal signal generated by the audio processing unit 24 to the speech recognition unit 40 as an output signal (step S116). Since the removal signal is a signal obtained by removing the reproduction signal, which is the reference signal, from the audio signal, the removal signal may include an audio command. Therefore, during the period when the removal signal is received as the output signal, the speech recognition unit 40 is in a state where it can recognize the audio command. Then, the process proceeds to step S112 above.

[0083] As described above, the audio processing apparatus 10 of the present embodiment includes an audio acquisition unit 20, a determination unit 22, an audio processing unit 24, and a switching unit 26. The audio acquisition unit 20 acquires an audio signal from a microphone MC that picks up the audio in the space. The determination unit 22 determines whether or not the level of the reference signal, which is the reproduction signal output from the speaker SP that outputs sound in the space, is equal to or higher than a threshold value. The audio processing unit 24 outputs, as an output signal, a removal signal obtained by removing the audio component of the reference signal from the audio signal to the speech recognition unit 40. When it is determined that the level of the reference signal is equal to or higher than the threshold value, the switching unit 26 outputs, as an output signal to the speech recognition unit 40, a replacement signal that is at least one of a comfort noise and a mute signal instead of the removal signal.

[0084] Here, in the prior art, the voice collected by the microphone is recognized by the first voice recognition unit, the voice output from the speaker is recognized by the second voice recognition unit, and when the voice recognized by the second voice recognition unit includes a voice recognition command, the recognition by the first voice recognition unit is stopped. However, in the prior art, when the voice collected by the microphone includes noise components such as residual echo components that cannot be completely removed by an echo canceller or the like, misdetection of voice recognition may occur. That is, in the prior art, it may be difficult to suppress misdetection of voice recognition. Also, in the prior art, depending on the performance of the second voice recognition unit and the like, misdetection may occur in the voice recognition by the first voice recognition unit.

[0085] On the other hand, in the voice processing apparatus 10 of the present embodiment, when it is determined that the level of the reference signal, which is the reproduction signal, is equal to or higher than the threshold value, instead of the removal signal obtained by removing the voice component of the reference signal from the voice signal acquired from the microphone MC, a replacement signal that is at least one of comfort noise and a mute signal is output to the voice recognition unit 40 as the output signal. Since the replacement signal is at least one of comfort noise and a mute signal, the replacement signal does not include a voice command. Therefore, during the period when the replacement signal is being received, the voice recognition unit 40 is in a state where it does not perform recognition of voice commands.

[0086] Therefore, in the voice processing apparatus 10 of the present embodiment, for example, even in a voice environment where the level of the reproduction signal reproduced from the speaker SP is large and components that cannot be completely cancelled by the removal process remain in the voice signal collected by the microphone MC, misdetection of voice recognition caused by the reproduction signal can be suppressed.

[0087] Therefore, the voice processing apparatus 10 of the present embodiment can suppress misdetection of voice recognition.

[0088] Also, in the voice processing apparatus 10 of the present embodiment, the determination unit 22 determines whether or not the level of the reproduction signal reproduced from the speaker SP is equal to or greater than a threshold value, rather than the level of the voice signal acquired from the microphone MC. Therefore, in the voice processing apparatus 10 of the present embodiment, regardless of the magnitude of the level of the voice spoken by the user, when the level of the reproduction signal is less than the threshold value, a removal signal including the voice component of the user picked up by the microphone MC can be output to the voice recognition unit 40 as a voice recognition target. Thus, in addition to the above effects, the voice processing apparatus 10 of the present embodiment can efficiently recognize a voice signal including a voice command or the like spoken by the user.

[0089] Further, in the voice processing system 1 of the present embodiment, since voice recognition by the voice recognition unit 40 is not performed on the reproduction signal of the speaker SP, in addition to the above effects, it is possible to reduce the processing calculation amount of the voice processing system 1. Also, in the present embodiment, since voice recognition is not performed on the reproduction signal, regardless of the voice recognition accuracy of the voice recognition unit 40, false detection of voice recognition can be suppressed.

[0090] Note that, in the present embodiment, the voice processing system 1 has been described by taking as an example the form mounted on the vehicle 2. However, the voice processing system 1 may be configured to be arranged in any space to be voice-processed, and is not limited to the form mounted on the vehicle 2.

[0091] Note that the above describes the embodiments, but the above embodiments are presented as examples and are not intended to limit the scope of the invention. The above novel embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. The above embodiments are included in the scope or gist of the invention and are included in the invention described in the claims and the equivalent scope thereof.

Explanation of Reference Numerals

[0092] 1 Voice processing system 10 Voice processing apparatus 20 Voice acquisition unit 22 Determination unit 24 Voice processing unit 26 Switching unit 28 Generation unit 40 Voice recognition unit 50 Electronic device 60 Display MC Microphone SP Speaker

Claims

1. An audio acquisition unit that acquires an audio signal from a microphone that picks up the audio in the space; A determination unit that determines whether the level of a reference signal, which is a reproduction signal reproduced from a speaker that outputs sound to the space, is equal to or greater than a threshold value; An audio processing unit that outputs, as an output signal, a removal signal obtained by removing the audio component of the reference signal from the audio signal to an audio recognition unit; A switching unit that, when it is determined that the level of the reference signal is equal to or greater than the threshold value, outputs, as the output signal, a replacement signal that is at least one of a comfort noise and a mute signal to the audio recognition unit instead of the removal signal; An audio processing apparatus comprising:

2. The determination unit: When it is determined that the level of the reference signal is equal to or greater than the threshold value, controls the audio processing unit so as to stop a removal process of removing the audio component of the reference signal from the audio signal. The audio processing apparatus according to claim 1. The audio processing apparatus according to claim 1.

3. Further comprising a generation unit that generates the replacement signal, The generation unit: Generates the replacement signal, which is the comfort noise, according to the noise level included in the audio signal immediately before it is determined to be equal to or greater than the threshold value. The audio processing apparatus according to claim 1 or claim 2. The audio processing apparatus according to claim 1 or claim 2.

4. The switching unit: During a period in which it is determined that the level of the reference signal is equal to or greater than the threshold value, outputs, as the output signal, the replacement signal to the audio recognition unit instead of the removal signal. The audio processing apparatus according to any one of claims 1 to 3. The audio processing apparatus according to any one of claims 1 to 3.

5. The switching unit: When it is determined that the level of the reference signal is equal to or greater than the threshold value, outputs, as the output signal, the replacement signal to the audio recognition unit instead of the removal signal and continuously outputs it to the audio recognition unit for a predetermined first time. The audio processing apparatus according to any one of claims 1 to 3. The audio processing apparatus according to any one of claims 1 to 3.

6. The switching unit: When it is determined that the level of the reference signal continuously remains equal to or greater than the threshold value for a predetermined second time or more, outputs, as the output signal, the replacement signal to the audio recognition unit instead of the removal signal. The audio processing apparatus according to any one of claims 1 to 3.

7. An output control unit that outputs information indicating that audio recognition is stopped when it is determined that the level of the reference signal is equal to or greater than the threshold value. The audio processing apparatus according to any one of claims 1 to 6. The audio processing apparatus according to any one of claims 1 to 6.

8. An audio processing method executed by an audio processing apparatus, comprising: A step of acquiring an audio signal from a microphone that picks up the audio in the space; A step of determining whether the level of a reference signal, which is a reproduction signal reproduced from a speaker that outputs sound to the space, is equal to or higher than a threshold value; A step of outputting, as an output signal, a removal signal obtained by removing the voice component of the reference signal from the voice signal to a voice recognition unit; When it is determined that the level of the reference signal is equal to or higher than the threshold value, a step of outputting, as the output signal, a replacement signal that is at least one of comfort noise and a mute signal to the voice recognition unit instead of the removal signal; A voice processing method including the above.

9. A step of obtaining a voice signal from a microphone that picks up the voice in the space; A step of determining whether the level of a reference signal, which is a reproduction signal reproduced from a speaker that outputs sound to the space, is equal to or higher than a threshold value; A step of outputting, as an output signal, a removal signal obtained by removing the voice component of the reference signal from the voice signal to a voice recognition unit; When it is determined that the level of the reference signal is equal to or higher than the threshold value, a step of outputting, as the output signal, a replacement signal that is at least one of comfort noise and a mute signal to the voice recognition unit instead of the removal signal; A voice processing program for causing a computer to execute the above.

10. A voice processing system including a voice processing device, a microphone that picks up the voice in the space, a speaker that outputs sound to the space, and a voice recognition unit that recognizes voice, wherein the voice processing device includes a voice acquisition unit that acquires a voice signal from the microphone, a determination unit that determines whether the level of a reference signal, which is a reproduction signal reproduced from the speaker, is equal to or higher than a threshold value, a voice processing unit that outputs, as an output signal, a removal signal obtained by removing the voice component of the reference signal from the voice signal to the voice recognition unit, and a switching unit that, when it is determined that the level of the reference signal is equal to or higher than the threshold value, outputs, as the output signal, a replacement signal that is at least one of comfort noise and a mute signal to the voice recognition unit instead of the removal signal. A voice processing system comprising the above.

Citation Information

Patent Citations

  • Waste straw cutter

    JP1987025920A

  • On-vehicle speech recognition device

    JP2009181025A

  • Video system and method for controlling the same

    JP2013142903A

  • Sound recognition device

    JP2019176431A

  • Sound processing system and sound processing device

    JP2021057807A